Enhancement of intelligent speech or dialogue
A machine learning network in audio systems automatically enhances dialogue clarity by detecting speech components and adjusting equalization settings, addressing the challenge of inconsistent sound quality in multimedia content.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BOSE CORP
- Filing Date
- 2023-07-19
- Publication Date
- 2026-07-22
AI Technical Summary
Existing audio systems struggle to automatically adjust equalization settings to enhance dialogue clarity in multimedia content, leading to inconsistent sound quality due to manual adjustments being cumbersome and delayed user reactions, and traditional equalization profiles often negatively affect non-dialogue content.
Implementing a machine learning network to detect speech components in audio signals and transition to a speech equalization mode, enhancing dialogue clarity by adjusting frequency bands and dynamic range compression without user input, while maintaining quality for other audio components.
Improves dialogue clarity by dynamically adjusting equalization based on detected speech content, reducing misapplication of dialogue modes to non-dialogue content, and enhancing overall audio performance with minimal user input.
Smart Images

Figure 0007893963000001 
Figure 0007893963000002 
Figure 0007893963000003
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure generally relate to audio signal processing.
Background Art
[0002] In recording and playback, an equalizer may perform equalization to adjust the magnitudes of different frequency bands in an audio signal. For example, the equalizer may use filters that adjust bass and treble to improve the listening experience. The equalization may be dynamically adjusted in real time by a user or may include one or more preset profiles for different genres of audio input (e.g., jazz, classical, pop, etc.).
[0003] In a home theater, speech or dialogue may be made less understandable than intended due to different surround sound configurations and / or different recording or compression profiles. A user may prefer to use an equalization profile to enhance the sound components of speech or dialogue (e.g., by increasing the magnitude of the relevant frequency bands).
Summary of the Invention
[0004] All examples and features mentioned herein can be combined in any technically possible way.
[0005] Aspects of the present disclosure provide a method for processing and generating an audio signal. The method includes detecting a speech component in an original set of audio signals using a machine learning network trained based on a mixed category of audio content, the speech component consisting of sound elements that convey linguistic meaning. The method includes enhancing the detected speech component by transitioning from an original equalization mode to a speech equalization mode that enables the user to better understand the linguistic meaning therein. The method further includes outputting the original set of audio signals in the original equalization mode if no speech component is detected.
[0006] In some embodiments, detecting speech components using a trained machine learning network is based on calculating the root mean square of the energy level of each category in a mixed category of audio content. In some cases, the trained machine learning network includes a deep learning model that estimates the energy level of each category in a mixed category of audio content, where the mixed category of audio content includes speech components, musical components, and singing components. In some cases, detecting speech components involves determining the ratio of the energy level of the speech component to the overall energy level that exceeds a threshold.
[0007] In some cases, speech components include sounds whose meaning is conveyed based on linguistic features. Musical components include sounds that lack linguistic features. Singing components include a mixture of sounds that simultaneously contain components of linguistic expression and components of musical expression. In some cases, a trained machine learning network may be trained to identify each category and threshold of the mixed categories of audio content based on a known database of film content.
[0008] In some cases, detecting speech components involves processing the ongoing audio signal in the time preceding a transition or output operation.
[0009] In some embodiments, enhancing speech components includes gradually fading from the original equalization mode to a speech equalization mode. For example, the speech equalization mode includes at least one of increasing the magnitude or contrast of speech-related frequency bands, decreasing the magnitude of non-speech-related frequency bands or signal channels, or changing the equalization settings or dynamic range compression settings on non-speech-related frequency bands or signal channels in order to improve the clarity of speech components.
[0010] In one embodiment, enhancing the detected speech components is performed in a first device, and outputting the original set of audio signals is performed in a second device. For example, the first and second devices are paired in a short-range wireless communication network. In one instance, the method further includes extracting the detected speech components and playing the extracted speech components in a third device. In one instance, the first, second, and third devices are configured to generate mixed surround sound.
[0011] In one embodiment, outputting the original set of audio signals includes determining the loss of speech components.
[0012] Aspects of this disclosure provide an apparatus for processing and generating audio signals. The apparatus includes a memory and a processor coupled to the memory. The processor and memory are configured to detect speech components in an original set of audio signals using a trained machine learning network based on mixed categories of audio content, where speech components consist of sound elements that carry linguistic meaning. The processor and memory are configured to enhance the detected speech components by transitioning from the original equalization mode to a speech equalization mode that enables the user to better understand the linguistic meaning therein. The processor and memory are further configured to output the original set of audio signals in the original equalization mode if no speech components are detected.
[0013] In one embodiment, the processor and memory are configured to enhance the detected speech components and output the original set of audio signals in a second device. In one instance, the device includes a soundbar configured to output surround sound, and the second device includes noise-canceling headphones.
[0014] In one embodiment, the processor and memory are configured to enhance the speech component by gradually fading from the original equalization mode to the speech equalization mode.
[0015] In one embodiment, the processor and memory are configured to extract detected speech components and to play the extracted speech components on a third device. In some cases, the first, second, and third devices are configured to generate mixed surround sound.
[0016] The embodiment provides a method for processing an audio signal, the method comprising: analyzing the content of an audio signal before playback of the content to determine whether one or more predefined conditions are met to indicate that the content contains utterances during playback of the audio signal; automatically applying a first playback equalizer configured to enhance utterances in the content to the audio signal in response to determining that one or more predefined conditions are met; and applying to the audio signal, i) no playback equalizer is applied, or ii) a second playback equalizer different from the first playback equalizer, in response to determining that one or more predefined conditions are not met.
[0017] In one embodiment, the analysis is performed using a trained machine learning model. In one embodiment, the trained machine learning model comprises a deep learning model that estimates the energy level of an audio signal. In one embodiment, the energy level of the audio signal includes the energy levels of any combination of speech, musical components of the audio signal, and singing components of the audio signal.
[0018] In one embodiment, the analysis includes analyzing metadata associated with the content, where the metadata indicates that the content contains utterances. In another embodiment, the analysis includes analyzing the audio track of an audio signal, where one or more predefined conditions include the audio track exceeding a threshold.
[0019] In one embodiment, the audio signal comprises different channels, and the analysis includes analyzing the different channels. In another embodiment, the analysis of different channels includes comparing the correlated content between two of the different channels.
[0020] In this embodiment, the analysis includes analyzing the central channel of the audio signal.
[0021] In one embodiment, automatically applying a first playback equalization configured to enhance speech in the content to an audio signal includes transitioning from either i) no playback equalization or ii) a second playback equalization to the first playback equalization.
[0022] In one embodiment, automatically applying a first playback equalization configured to enhance speech within content to an audio signal includes increasing the volume of speech within content relative to other content within the audio signal.
[0023] In one embodiment, automatically applying a first playback equalization configured to enhance speech within the content to an audio signal includes reducing the volume of non-speech content within the audio signal. In another embodiment, automatically applying the first playback equalization further includes increasing the volume of speech within the content.
[0024] In one embodiment, the second reproduction equalization comprises at least one of low-frequency enhancement or music reproduction enhancement.
[0025] In one embodiment, the first regeneration equalization or at least one of one or more predefined conditions is configurable by the user.
[0026] In some embodiments, the method further includes analyzing the sounds in the environment in which the audio signal is to be reproduced in order to help determine whether the first reproduction equalization should be applied to the audio signal.
[0027] Aspects provide an apparatus for audio signal processing comprising a memory and a processor coupled to the memory, wherein the processor and the memory are configured to analyze the content of an audio signal during playback of the audio signal prior to playback of the content to determine whether one or more predefined conditions are met to indicate that the content includes speech, and in response to determining that one or more predefined conditions are met, automatically apply a first playback equalization configured to enhance speech in the content to the audio signal, and in response to determining that one or more predefined conditions are not met, be configured to either i) not apply playback equalization or ii) apply a second playback equalization different from the first playback equalization to the audio signal.
[0028] In aspects, the processor and the memory are configured to detect using a trained machine learning model. In aspects, the audio signal comprises different channels, and the memory and the processor are configured to detect by analyzing the different channels, which includes comparing correlated content between two of the different channels.
[0029] In aspects, the processor and the memory are configured to detect by analyzing a center channel of the audio signal.
[0030] In aspects, the processor and the memory are configured to automatically apply to the audio signal a first playback equalization configured to enhance speech in the content by transitioning from either i) no playback equalization or ii) a second playback equalization to the first playback equalization.
[0031] In aspects, the second playback equalization comprises at least one of low frequency enhancement or music playback enhancement. In aspects, at least one of the first playback equalization or one or more predefined conditions is configurable by a user.
[0032] In one embodiment, the processor and memory are further configured to analyze the sounds in the environment in which the audio signal is to be reproduced in order to help determine whether to apply a first reproduction equalization to the audio signal.
[0033] The embodiment provides a non-temporary computer-readable medium for storing instructions, which, when executed by a device for processing and generating audio signals, causes the device to analyze the content of an audio signal before playback of the content, determine whether one or more predefined conditions are met to indicate that the content includes utterances, automatically apply a first playback equalization configured to enhance utterances in the content to the audio signal in response to determining that one or more predefined conditions are met, and, in response to determining that one or more predefined conditions are not met, i) not apply playback equalization, or ii) apply a second playback equalization different from the first playback equalization to the audio signal.
[0034] Two or more features described in this disclosure, including those described in the summary section of the present invention, may be combined to form implementations not specifically described herein.
[0035] Details of one or more implementation configurations are described in the attached drawings and the following description. Other features, purposes, and advantages will become apparent from this description and drawings, as well as from the "Claims."
[0036] Similar figures represent similar elements, and "utterance" and "dialogue" can be used interchangeably. [Brief explanation of the drawing]
[0037] [Figure 1] Examples of systems in which aspects of this disclosure may be implemented are shown. [Figure 2] Block diagrams of machine learning networks and related components according to several aspects of this disclosure are shown. [Figure 3A]This disclosure illustrates exemplary determination of different sound categories in an audio signal, according to several aspects of this disclosure. [Figure 3B] This disclosure illustrates exemplary determination of different sound categories in an audio signal, according to several aspects of this disclosure. [Figure 4] The present disclosure illustrates exemplary equalization processes for speech according to several aspects of this disclosure. [Figure 5] This disclosure provides exemplary processes for generating training datasets according to several aspects of this disclosure. [Figure 6] This flowchart illustrates exemplary operations for automatic speech enhancement according to several aspects of the present disclosure. [Figure 7] This flowchart illustrates exemplary operations for automatic speech enhancement according to several aspects of the present disclosure. [Modes for carrying out the invention]
[0038] Users often react with considerable delay to audio content such as movies, which often includes various types of audio signals (e.g., dialogue, music, singing), so manual equalization adjustments by the user may be too slow. In addition, manually selecting speech-based equalization profiles is cumbersome and may discourage users. Therefore, there is a need for methods to automatically update equalization profiles (without additional user input) taking speech content into account, as well as devices and systems configured to implement these methods.
[0039] This disclosure provides processes, methods, systems, and devices for intelligently detecting speech or dialogue content in an audio signal (for example, by implementing machine learning and artificial intelligence) and for smoothly transitioning to speech equalization modes that enable a user to better understand the linguistic meaning of the speech content in the audio signal. For example, aspects of this disclosure provide methods for processing and generating audio signals. In one example, the method uses a machine learning network trained on mixed categories of audio content to detect speech components in an original set of audio signals. Regardless of how the speech components are detected, the detected speech components can be enhanced by transitioning from the original equalization mode to a speech equalization mode that enables a user to better understand the linguistic meaning within them.
[0040] Multimedia played in home theaters or entertainment centers often includes different types of movies or television programs. These different types of multimedia may include sound reproduction that benefits from different equalization settings. For example, documentaries may include substantial linguistic narration. Linguistic narration can benefit from equalization settings for speech more than, for example, concerts that include substantial singing or music. Traditionally, equalization settings can be adjusted on high-fidelity (hi-fi) equipment by applying filters or amplification across different frequency bands from low to high. Hi-fi equipment may allow manual adjustment or recall of pre-configured equalization settings (e.g., pre-configured for jazz, classical, pop, concerts, etc.). However, such manual adjustment or recall can be inconvenient, especially when the user is cycling through television channels or when a movie contains different scenes with different primary sound elements (e.g., dialogue, music, sound effects, etc.).
[0041] For example, a user can manually select a pre-configured equalization profile that enhances speech, and then recognize that the media includes a musical background. The sound quality of the music may be negatively affected by the selected equalization profile. Similarly, if a user selects a pre-configured equalization profile that enhances music or other sound effects, or in situations with complex and loud surround sound, the sound quality for dialogue may be negatively affected or become less understandable.
[0042] Aspects of the Disclosure overcome such challenges by detecting sounds of different categories and identifying whether or when dialogue or utterances are the primary content being played. In some aspects, machine learning models are used to detect or identify different sound categories. For ease of understanding, aspects of the Disclosure implement smooth transitions to utterance enhancement equalization profiles so that users can better understand the meaning therein. When utterance content ends, aspects of the Disclosure provide that playback smoothly transitions to the original equalization profile or another preferred equalization profile.
[0043] Dialogue often suffers from a lack of clarity in A4V (audio-for-video) content. To improve dialogue clarity, sound products (e.g., speakers, soundbars, headphones, etc.) may include pre-configured equalization modes that enhance the clarity of dialogue or utterances. When dialogue equalization mode is turned on, non-dialogue audio quality may be negatively affected. However, due to interlacing with other non-dialogue content, users may not turn on dialogue mode in a timely or actual manner at the appropriate time (e.g., on multiple different occasions within the same playback session).
[0044] This disclosure provides the benefits of improved (e.g., better sounding) dialogue equalization modes. In general, automatic equalization of audio content to enhance utterances within content, as variously described herein, allows equalization to be applied when utterances are detected, thereby making the application of utterance enhancement equalization dynamic based on the content itself, as opposed to not statically turning equalization on or off. This enables a variety of benefits, such as more aggressive / desirable utterance enhancement and / or more aggressive / desirable equalization of non-utterance content (e.g., enhancing bass in action scenes in movies, or applying music enhancement equalization when music is playing). In other words, dialogue enhancement as variously described herein may only be applied when one or more predefined conditions are met to indicate that there is known or a high degree of confidence that there are actual utterances in the audio content. One exemplary method for determining whether the audio being played contains utterances (or whether it is likely to contain utterances, which can be determined by preprocessing the audio data immediately before the audio data is played) is to use an algorithm such as a trained machine learning model. By implementing machine learning to detect actual utterances / dialogues, dialogue tuning can be implemented as needed (e.g., taking background sounds and ambient noise into account). By using machine learning and artificial intelligence in the dialogue recognition model, aspects of this disclosure significantly reduce the misapplication of dialogue modes to non-dialogue modes and result in improved audio performance for music or action scenes. Automated aspects of implementing this disclosure require minimal user input and improve overall dialogue clarity.
[0045] In one embodiment, a deep learning model estimates multiple energy levels in media audio content in real time. These energy levels may include speech energy, musical energy, and singing energy. The system may also calculate the ratio of different sound categories to the total energy in the content signal in order to identify the dominant sound category.
[0046] Deep learning models (e.g., artificial intelligence) can be trained so that sound category determination is not pre-programmed but determined based on machine learning. In one example, data is synthesized to emulate A4V content in order to train a deep learning model. A4V content can be emulated by mixing separate tracks of music, speech, sound effects, and singing. Using these energy levels, control logic is applied to seamlessly fade or transition in and out of dialogue modes.
[0047] In addition to classifying content types, deep learning models can also be trained for content understanding (e.g., what content is, in addition to which category it belongs to). For example, in addition to estimating the energy level of music or singing, the sound of the instruments used can also be estimated. In some cases, music or video genre classifiers may also be included in deep learning models to determine whether a sound could belong to hip hop or an action movie.
[0048] Based on content detection, aspects of this disclosure can improve control logic for playback devices. For example, in addition to the output of a deep learning model, system states such as volume level and subwoofer availability may also be output. Ambient noise monitoring can also be implemented to improve control logic (e.g., for active noise cancellation).
[0049] The detected and / or measured information can be used to calculate an equalization profile of the dialogue mode. Therefore, a comprehensive machine learning model for estimating the sampling or clarity of utterances can be applied.
[0050] In one embodiment, a deep learning model is trained to calculate a level of confidence about a determined content category. For example, the deep learning model may output a measure of confidence that the content is, for example, speech or non-speech. The confidence parameter may help control the manner in which equalization is modified for the benefit of the user. For example, when the model indicates a sudden high probability or level of confidence that the detected sound is speech, the ballistic time constant when switching equalization may occur rapidly in order to switch to speech mode as quickly as possible. Additionally or alternatively, in one embodiment, the level of confidence of the determined content category influences how much equalization modification is implemented. For example, if the maximum change in channel center is 4 dB and the model indicates a high level of confidence that the content category is speech, the control logic may add a full 4 dB for the output by the playback device. If the maximum change in channel center is 4 dB and the model indicates a low level of confidence that the content category is speech, the control logic may apply less than a full 4 dB for the output by the playback device. In this way, the adjustment amount is based at least in part on the model's confidence level.
[0051] In some cases, to improve sound quality, speech intelligibility, and specialization, deep learning models detect or recognize sources in an audio signal and isolate or extract certain sources or types of content for a specific output channel. For example, speech content may be extracted from the rest of a soundtrack and played back through a pair of synchronized open audio headphones. In this way, surround sound is maintained, and the user can clearly understand the speech content.
[0052] In some cases, aspects of this disclosure may apply to a variety of different ambient speakers, including car audio systems, portable speakers, headphones, earphones, and the like, as described below.
[0053] Figure 1 shows an example of a system 100 in which an aspect of the present disclosure is put into practice. As shown, the system 100 includes one or more sound processing and playback devices 110 (e.g., wireless audio devices such as a soundbar or smart speaker) communicably coupled to a source device 120 (e.g., a computing device or user device such as a smartphone or tablet computer). One or more partner devices 112 (e.g., portable speakers, headsets, etc.) may be available to accept pairing requests from the sound processing and playback devices 110 or the source device 120. The sound processing and playback devices 110 may be paired with the source device 120 and may receive content data (including audio signals) from the source device 120. The sound processing and playback devices 110 may also receive content data directly from the network 130. The partner devices 112 may be battery-powered portable devices suitable for mobile or privacy applications.
[0054] According to aspects of this disclosure, a sound processing and playback device 110 may receive the original set of audio signals from at least one of a source device 120, the network 130, or a cloud 140 (via the network 130). The content of the audio signals is analyzed before playback to determine whether one or more predefined conditions are met indicating that the content contains utterances. The analysis can be performed using a variety of different techniques, such as using a trained machine learning model as described herein, analyzing metadata associated with the content (e.g., if the metadata indicates that the content contains utterances), analyzing the audio track of the audio signal (if there is one), analyzing different channels of the audio signal, and / or other techniques that can be understood based on this disclosure. Techniques that utilize metadata associated with the audio signal and / or its content can analyze the metadata immediately before playback to determine whether the metadata indicates that the audio content contains utterances or otherwise meets a predefined condition, and if so, speech / dialogue enhancement equalization (as described herein in various ways) may be automatically applied to assist in utterance clarity. Metadata can be text related to audio content (such as closed caption data or other subtitles), genre data (such as indicating that the content is a podcast or talk show), and / or other data that helps determine whether speech is present in the content. If the audio signal or associated content includes an audio track, the energy levels of the audio track can be analyzed to determine whether the audio content includes speech. If the audio signal includes different channels, they can be analyzed and / or compared to determine whether speech is likely to occur.For example, such analysis may include comparing correlated content between two channels (such as correlated content between the left and right channels of a stereo audio signal) and / or analyzing the central channel (e.g., in a 5.0, 5.1, 7.0, or 7.1 audio signal), since the central channel is typically where the majority of dialogue from movies and television takes place. For example, central channel analysis may include determining when central channel playback exceeds a threshold (e.g., a nominal threshold or a threshold relative to other channels) to determine that the content is likely to contain speech, and therefore speech reinforcement equalization should be automatically applied. Numerous different techniques will become apparent in light of this disclosure.
[0055] In response to determining that one or more predefined conditions are met, a first repetition equalization configured to enhance utterances within the content is automatically applied. As described herein, "automatically" may mean without user input. Thus, in response to determining that one or more predefined conditions are met, a first repetition equalization configured to enhance utterances within the content is applied without user input.
[0056] In some embodiments, applying a first reproduction equalization to an audio signal includes a transition from either no reproduction equalization or a second reproduction equalization to the first reproduction equalization. The transition includes a stepwise change from either no reproduction equalization or a second reproduction equalization to the first reproduction equalization. In some embodiments, the second reproduction equalization includes low-frequency enhancement and / or music reproduction enhancement.
[0057] In some embodiments, applying a first reproduction equalization to an audio signal involves increasing the volume of speech within the content relative to other content in the audio signal. In some embodiments, speech content is extracted using a machine learning algorithm. Additionally or alternatively, in some embodiments, speech content is taken from at least one of the correlated content between two channels or the central channel (as described above). In some embodiments, speech content is taken from the speech component of an audio signal.
[0058] In one embodiment, applying a first reproduction equalization to an audio signal includes reducing the volume of non-speaking content within the audio signal. In another embodiment, in addition to reducing the volume of non-speaking content, the volume of speech within the content is increased.
[0059] In one embodiment, the second reproduction equalization includes enhancing the low-frequency components of the content and / or enhancing the music reproduction of the content.
[0060] In some embodiments, the sound processing and playback device 110 may analyze the original set of audio signals and detect speech components using a trained machine learning network (e.g., the deep learning model 260 in Figure 2). The machine learning network is trained to identify speech components based on mixed categories of audio content such as dialogue, music, singing, and other sound categories (e.g., musical instruments or digital sound effects). Speech components may include any sound elements that convey linguistic meaning.
[0061] When a speech component is detected, the sound processing and playback device 110 can enhance the detected speech component by transitioning from the original equalization mode to a speech (or dialogue) equalization mode. The speech equalization mode can enable the user to better understand the linguistic meaning of the speech component. For example, the speech equalization mode can enhance the frequency spectrum related to dialogue and / or suppress the non-speech frequency spectrum. When the sound processing and playback device 110 does not detect a speech component, or when a detected speech component is lost or interrupted in the incoming audio signal, the sound processing and playback device 110 can output the original set of audio signals in the original equalization mode. Thus, the sound processing and playback device 110 intelligently applies the speech equalization mode only when necessary to achieve speech enhancement and minimize adverse effects on non-speech audio content.
[0062] In some scenarios, the playback device 110 provides different volume settings for different frequency bands. Dynamic equalization can adjust the overall system frequency response as a function of the detected mode and can provide a loudness compensation function that can emphasize high-frequency content without emphasizing lower-frequency content in order to improve clarity related to volume settings.
[0063] In this embodiment, the predicted sound pressure level (SPL) of the entire system is monitored and maintained when the playback device 110 switches between equalization modes. By monitoring the SPL, similar SPLs are maintained between the two equalization modes, which reduces the perceived volume changes during non-interactive content.
[0064] The sound processing and playback device 110 may further include hardware and circuitry, including, but not limited to, a processor / processing system and memory configured to implement one or more sound management capabilities or other capabilities, including noise cancellation circuits (not shown) and / or noise masking circuits (not shown), body movement detection devices / sensors and circuits (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetometers, etc.), geolocation circuits, and other sound processing circuits.
[0065] In one embodiment, the sound processing and playback device 110 is wirelessly connected to a source device 120 or a partner device 112 using one or more wireless communication methods, including but not limited to Bluetooth, Wi-Fi, BLE (Bluetooth Low Energy), and other RF-based techniques. In one embodiment, the sound processing and playback device 110 includes a transceiver that transmits and receives data via one or more antennas to exchange audio data and other information with the source device 120.
[0066] In one embodiment, the sound processing and playback device 110 includes a communication circuit capable of transmitting and receiving audio data and other information from a source device 120. The sound processing and playback device 110 also includes an arriving audio buffer, such as a render buffer, which buffers at least a portion of the arriving audio signal (e.g., audio packets) to allow time for retransmission of any missing or lost data packets from the source device 120. For example, when the sound processing and playback device 110 receives a Bluetooth transmission from the source device 120, the communication circuit typically buffers at least a portion of the arriving audio data in the render buffer before the audio is actually rendered and output as audio to at least one of the transducers (e.g., audio speakers) of the sound processing and playback device 110. This is done to ensure that even if there is an RF collision that causes audio packets to be lost during transmission, there is time for the lost audio packets to be retransmitted by the source device 120 before they need to be rendered by the sound processing and playback device 110 for output by one or more acoustic transducers of the sound processing and playback device 110.
[0067] An example of the partner device 112 is shown as noise-canceling headphones, but the techniques described herein apply to other wireless audio devices, such as wearable audio devices, including any audio output device that fits around, over, inside, or near the ear (including open-ear audio devices worn on the user's head or shoulders), or to any other part of the user's body, such as the head or neck. The partner device 112 can take any form, wearable or otherwise, including standalone devices (including automotive speaker systems), stationary devices (including portable devices such as battery-powered portable speakers), headphones, earphones, earpieces, headsets, goggles, headbands, earphones, armbands, sports headphones, neckbands, or glasses with integrated speakers.
[0068] In one embodiment, the sound processing and playback device 110 is connected to the source device 120 using a wired connection, with or without a corresponding wireless connection. The source device 120 may be a smartphone, tablet computer, laptop computer, digital camera, or other user device connected to the sound processing and playback device 110. As shown, the source device 120 may be connected to a network 130 (e.g., the Internet) and may have access to one or more services on the network. As shown, these services may include one or more cloud services 140.
[0069] In one embodiment, the source device 120 can access a cloud server in the cloud 140 on the network 130 using a mobile web browser or a local software application or “app” running on the source device 120. In one embodiment, the software application or “app” is a local application that is locally installed and run on the source device 120. In one embodiment, a cloud server accessible on the cloud 140 includes one or more cloud applications running on the cloud server. The cloud application can be accessed and run by the source device 120. For example, the cloud application can generate a web page that is rendered by a mobile web browser on the source device 120. In one embodiment, a mobile software application installed on the source device 120 or a cloud application installed on the cloud server can be used individually or in combination to implement techniques for low-latency Bluetooth communication between the source device 120 and the sound processing and playback device 110 according to embodiments of this disclosure. In one embodiment, examples of local software applications and cloud applications include game applications, audio AR applications, and / or game applications with audio AR capabilities. The source device 120 can receive signals (e.g., data and control) from the sound processing and playback device 110 and transmit signals to the sound processing and playback device 110.
[0070] An exemplary sound processing and playback device 110 or partner device 112 may include components described below (not shown in Figure 1). For example, each sound processing and playback device 110 or partner device 112 may include one or more processors, memory modules, communication modules, and / or input interfaces for receiving user input. Each sound processing and playback device 110 or partner device 112 may include one or more electroacoustic transducers (or speakers) for outputting audio. The sound processing and playback device 110 also includes a user input interface. The user input interface may include a plurality of preset indicators, which may be hardware buttons. The preset indicators can provide the user with easy one-press access to entities assigned to those buttons. The assigned entities can be associated with different digital audio sources so that a single sound processing and playback device 110 can provide one-press access to a variety of different digital audio sources.
[0071] The sound processing and playback device 110 or partner device 112 may essentially include an acoustic driver or speaker for converting audio signals into acoustic energy via audio hardware. The sound processing and playback device 110 also includes a network interface, at least one processor, audio hardware, a power supply for powering various components of the sound processing and playback device 110, and memory. In one embodiment, the processor, network interface, power supply, and memory are interconnected using various buses, and some of the components may be mounted on a common motherboard or in other ways as needed. In some cases, the sound processing and playback device 110 or partner device 112 may include a housing that accommodates an optional graphical interface (e.g., an OLED display) that can provide the user with information about the music currently being played ("Now Playing").
[0072] The network interface can provide communication between the sound processing and playback device 110 and other electronic user devices such as the source device 120 and partner device 112 via one or more communication protocols, such as Bluetooth® Classic protocol and Bluetooth® Low Energy protocol. Generally, the network interface provides either a wireless network interface or a wired interface (optional), or both. The wireless interface allows the sound processing and playback device 110 to communicate wirelessly with other devices according to a wireless communication protocol such as IEEE 802.11. The wired interface provides network interface functionality via a wired (e.g., Ethernet) connection for reliability and high transfer speed, for example, when the sound processing and playback device 110 is not being worn by a user.
[0073] In some embodiments, the network interface includes a network media processor to support Apple AirPlay® and / or Apple AirPlay® 2. For example, when a user connects an AirPlay® or Apple AirPlay® 2-compatible device, such as an iPhone or iPad, to a LAN, the user can then stream music to an audio playback device connected to the network via Apple AirPlay® or Apple AirPlay® 2. In particular, the audio playback device can support audio streaming via the UPnP protocol of AirPlay®, Apple AirPlay® 2, and / or DLNA, all integrated into a single device.
[0074] All other digital audio received as part of a network packet is passed straight from the network media processor to the processor via a USB bridge (not shown), then to a decoder, DSP, and finally can be played back (rendered) via an electroacoustic transducer.
[0075] The network interface may further include a Bluetooth circuit for Bluetooth applications (e.g., for wireless communication with Bluetooth-enabled audio sources such as smartphones or tablets) or other Bluetooth-enabled speaker packages. In some embodiments, the Bluetooth circuit may be the primary network interface due to energy constraints. For example, the network interface may use the Bluetooth circuit only for mobile applications when the sound processing and playback device 110 or partner device 112 adopts any wearable form. For example, BLE technology may be used in the sound processing and playback device 110 or partner device 112 to extend battery life, reduce package weight, and provide high-quality performance without other backup or alternative network interfaces.
[0076] In one embodiment, the network interface supports communication with other devices using multiple communication protocols simultaneously. For example, the sound processing and playback device 110 can support Wi-Fi / Bluetooth coexistence and support simultaneous communication using both Wi-Fi and Bluetooth protocols at the same time. For example, the sound processing and playback device 110 can receive an audio stream from a smartphone using Bluetooth and then simultaneously redistribute the audio stream to one or more other devices over Wi-Fi. In one embodiment, the network interface may include only one RF chain capable of communicating using only one communication method (e.g., Wi-Fi or Bluetooth) at a time. In this context, the network interface can simultaneously support Wi-Fi and Bluetooth communication by, for example, time-division multiplexing (TDM) between Wi-Fi and Bluetooth using a single RF chain.
[0077] Streamed data can be passed from the network interface to the processor. The processor can execute instructions (e.g., for performing digital signal processing, decoding, and equalization functions, among others) including instructions stored in memory. The processor may be implemented as a chipset of chips including multiple separate analog and digital processors. The processor can provide coordination of other components of the audio sound processing and playback device 110, such as control of the user interface.
[0078] The memory can store software / firmware related to protocols and their versions used by the sound processing and playback device 110 or partner device 112 to communicate with other networked devices, including the source device 120. For example, the software / firmware manages how the sound processing and playback device 110 communicates with other devices for synchronous audio playback. In one embodiment, the software / firmware includes lower-level frame protocols related to control path management and audio path management. Protocols related to control path management generally include protocols used to exchange messages between speakers. Protocols related to audio path management generally include protocols used for clock synchronization, audio distribution / frame synchronization, audio decoder / time matching, and playback of audio streams. In one embodiment, the memory can also store various codecs supported by the speaker package for audio playback of their respective media formats. In one embodiment, the software / firmware stored in the memory may be accessible and executable by the processor for synchronous audio playback with other networked speaker packages.
[0079] In some embodiments, the protocol stored in memory may include, for example, BLE conforming to Bluetooth Core Specification Version 5.2 (BT5.2). The sound processing and playback device 110 or partner device 112, and various components therein, are provided herein to fully comply with or perform embodiments of the protocol and associated specifications. For example, BT5.2 includes an enhanced attribute protocol (EATT) that supports concurrent transactions. A new L2CAP mode is defined to support EATT. Thus, the sound processing and playback device 110 includes sufficient hardware and software components to support the specifications and operating modes of BT5.2, even if not explicitly shown or discussed herein. For example, the sound processing and playback device 110 may utilize the LE isochronous channel specified in BT5.2.
[0080] The processor may provide the processed digital audio signal to audio hardware that includes one or more digital-to-analog (D / A) converters for converting the digital audio signal to an analog audio signal. The audio hardware may also include one or more amplifiers that provide the amplified analog audio signal to an electroacoustic transducer for sound output. In addition, the audio hardware may include circuitry for processing analog input signals to provide a digital audio signal for sharing with other devices, such as other speaker packages for synchronous digital audio output.
[0081] The memory may include, for example, non-temporary memory such as flash memory and / or non-volatile random-access memory (NVRAM). In some embodiments, instructions (e.g., software) are stored on an information carrier. When the instructions are executed by one or more processing devices (e.g., processors), they perform one or more processes, such as those described elsewhere in this specification. Instructions may also be stored in one or more storage devices, such as one or more computer-readable or machine-readable media (e.g., memory, or memory on a processor). Instructions may include instructions for performing decoding (i.e., including an audio codec for a software module to decode a digital audio stream), as well as digital signal processing and equalization. In some embodiments, the memory and processor may cooperate with a sound processing and playback device 110 or a microphone on a source device 120 in data acquisition and real-time processing.
[0082] Exemplary intelligent dialogue or speech enhancement Aspects of this disclosure provide techniques for intelligently detecting and enhancing detected speech components in an audio signal. For example, an audio device may detect speech components in the original set of audio signals using any number of methods. One example is a machine learning network trained on mixed categories of audio content. Speech components include sound elements that convey linguistic meaning. The audio device may enhance detected speech components by transitioning from the original equalization mode to a speech equalization mode that enables the user to better understand the linguistic meaning of the utterances. If no speech components are detected (e.g., before or after detection of speech components), the audio device may output the original set of audio signals in the original equalization mode. To correctly detect speech components (e.g., as opposed to singing components), the machine learning network may be trained to recognize what constitutes speech without user intervention. An exemplary intelligent dialogue or speech enhancement according to this disclosure is provided in Figure 2.
[0083] Figure 2 is a block diagram 200 showing the relationship between audio signals, a training dataset, and processing components according to an aspect of this disclosure. As shown, the original set of audio signals 210, such as those received by the sound processing and playback device 110, is provided to the machine learning network 220. For example, the sound processing and playback device 110 may be coupled to the machine learning network 220 via the network 130 in a communicative manner.
[0084] The machine learning network 220 may analyze the original set of audio signals 210 using a deep learning model 260 coupled with the machine learning network 220. Figure 2 shows the deep learning model 260 as separate from the machine learning network 220, but in some cases, the deep learning model 260 may be integrated with the machine learning network 220. Figure 2 shows one deep learning model 260 coupled with the machine learning network 220, but in some cases, two or more different deep learning models may be coupled or integrated with the machine learning network 220. In some cases, the machine learning network 220 or its interface (e.g., a graphical user interface such as an application on an operating system) may be installed on a source device 120, which may be a smartphone.
[0085] The deep learning model 260 can utilize various machine learning techniques based on artificial neural networks. For example, the deep learning model 260 may include deep learning architectures such as deep neural networks, deep belief networks, deep reinforcement learning, recurrent neural networks, and convolutional neural networks. Similar to speech recognition, the deep learning model 260 can identify sound elements that contain linguistic meaning. On the other hand, the deep learning model 260 is trained to distinguish between sound elements that primarily represent tonality or musical elements other than linguistic meaning in context and sound elements that primarily represent linguistic meaning.
[0086] For example, deep learning model 260 is trained to distinguish speech components from musical or singing components. Speech components may include sounds that convey meaning based on linguistic features. Musical components may include sounds that lack linguistic features. Singing components may include a mixture of sounds that simultaneously contain components of linguistic expression and components of musical expression.
[0087] The deep learning model 260 is trained to identify sounds that do not contain musical expressions or lack linguistic features, based on the training dataset 250. For example, the training dataset 250 may include various types of singing, music, and dialogue. The deep learning model 260 may be supervised, semi-supervised, or unsupervised to learn whether a sound pattern belongs to one of three categories. For example, the training dataset 250 may include various samples such as music, opera, rap music, choruses, conversations, dialogues, and speech.
[0088] The machine learning network 220 can then use the deep learning model 260 to calculate the root mean square of the energy level of each category in the mixed category of audio content. For example, the deep learning model 260 may estimate the energy level of each category in the mixed category of audio content. The mixed category of audio content may include speech components, musical components, and singing components, which may correspond to the samples contained in the training dataset 250. Figure 5 and the corresponding description below provide further details of the training dataset 250.
[0089] Speech components can be defined or detected by determining the ratio of the energy level of the speech component to the overall energy level that exceeds a threshold (see example in Figure 3). In some cases, a trained machine learning network 220 is trained to identify each category and threshold of mixed categories of audio content based on a known database of movie content (e.g., samples in the training dataset 250).
[0090] The machine learning network 220 may include a speech component enhancement module 230. When a speech component is detected, the machine learning network 220 performs calculations in speech equalization mode and provides a calculation output 240.
[0091] The speech component enhancement module 230 can implement and be trained to improve speech equalization modes. For example, the speech component enhancement module 230 may increase the magnitude or contrast of speech-related frequency bands to improve the clarity of speech components. In some cases, the speech component enhancement module 230 may decrease the magnitude of non-speech-related frequency bands or signal channels. In some cases, the speech component enhancement module 230 may change or update equalization settings or dynamic range compression settings on non-speech-related frequency bands or signal channels. The speech component enhancement module may combine two or more of these operations. An example of a speech equalization mode is provided in Figure 4 and described below.
[0092] Exemplary dialogue or utterance detection Figures 3A and 3B illustrate exemplary determination of different sound categories in an audio signal according to several aspects of the present disclosure. Figure 3A shows a first instance 310 of a video clip 305, and Figure 3B shows a second instance 320. In the first instance, a character in the video clip is speaking. The window includes a category indication frame 312 that monitors the category of sound currently detected (e.g., singing, music, and speech). In the first instance 310, the window highlights frame 314 indicating the speech component, and in the second instance 320, the window highlights frame 316 indicating the singing component. Such indications may be useful in supervised learning so that a trainer can verify whether the machine learning network 220 has correctly classified the sound content.
[0093] Figures 3A and 3B further illustrate the data processing windows 320 within each video window. The data processing window 320 plots the calculated model outputs, such as energy levels, for each sound category. For example, in the data processing window 320 of the first instance 310, the root mean square of the energy level of the speech component 322 is dominant in the output. The root mean squares of the energy levels of the music component 324 and the singing component 326 are low. In some cases, the ratio of the energy level of the speech component to the overall energy level may be compared to a threshold level (learned based on the dataset). The speech component is detected when the ratio exceeds the threshold. Thus, in the first instance 310, the machine learning network 220 detects the speech component and indicates detection by highlighting the frame 314 corresponding to singing.
[0094] In the data processing window of the second instance 320, the root mean square of the energy level of the singing component 326 is dominant in the output, while the root mean square of the energy levels of the musical component 324 and the speech component 322 are low. Therefore, in the second instance 320, the machine learning network 220 detects the singing component and indicates the detection by highlighting the frame 316 corresponding to the singing.
[0095] Figures 3A and 3B illustrate examples of energy determination for three sound categories, although different sets of categories may be used. For example, the machine learning network 220 may be further trained to identify noise that does not belong to any of the speech, music, or singing categories. According to aspects of this disclosure, detecting and identifying noise may enable the machine learning network 220 to receive and process captured ambient noise for active noise cancellation. Noise categories may also improve accuracy in detecting other sound categories.
[0096] Exemplary dialogue or speech processing Upon detecting speech components, the machine learning network 220 may enhance the detected speech components by transitioning from the original equalization mode to a speech equalization mode. Figure 4 shows an example of such equalization processing. The original equalization mode 410 (or alternatively, not an equalization mode) is transitioned (via process 432) to speech equalization mode 420 when speech components are detected and are dominant in the content (e.g., excluding background speech noise). As shown, in the original equalization mode 410, three profiles are used to tune the center channel, left / right channels, and left surround / right surround channels. Before transitioning to speech equalization mode 420, the bass and treble portions of all channels have amplification that is not close to zero in order to provide a rich surround sound. Dialogue or speech in such an equalization mode may be difficult to understand, given other sound components.
[0097] Process 432 implements a transition that minimizes the perceptibility of the change from the original equalization mode 410 (which may also not be an equalization mode) to the speech equalization mode 420. For example, frequency band changes, channel changes, or various other parameters can be used to smoothly transition between the two equalization modes 410 and 420 at different rates or times, over different durations, to seamlessly blend them. In this way, the user does not have to perceive the change. Process 432 may be performed by the speech component enhancement module 230 shown in Figure 2.
[0098] In the second equalization mode 420, the base frequency and high frequency spectra of all three channels (center, lateral, and peripheral) can be tuned to values close to zero, as shown, thus allowing the spectral range corresponding to the speech frequency to be more distinct from other non-speech sound components. Consequently, speech is more understandable when reproduced in speech equalization mode 420 than in the original equalization mode 410.
[0099] Exemplary training datasets for machine learning Figure 5 illustrates an exemplary process for generating a training dataset, such as the training dataset 250 in Figure 2, according to several aspects of this disclosure. For example, the training dataset may be used by a deep learning model based on a convolutional recurrent neural network (CRNN). The CRNN receives an input created by mixing several clean sound sources with some noise to generate a synthetic example. In some cases, the input may take the form of a Log Mel Spectrogram. As illustrated in Figure 2, the deep learning model may provide an output about the energy levels of different categories of sounds. The output may include several frames having a total duration equal to the input duration.
[0100] In some cases, the dataset for the deep learning model may be configured by the user to update or mix the parameters of the dataset and / or the parameters for the deep learning model. The training process may include closed-loop feedback (e.g., supervised), where the model under training receives evaluations based on both sound content detection and equalization (e.g., utterance and non-utterance modes) output evaluations. For example, as shown in Figure 5, four source datasets are used to generate a composite example 570. The utterance dataset 510, the music dataset 520, the singing dataset 530, and the noise dataset 540 are mixed in different ways to form a first sound category 550 and a second sound category 560. The first sound category 550 includes all datasets 510-540 and emulates video audio. The second sound category 560 includes the singing dataset 530 and the noise dataset 540, emulating concert audio, etc. Synthesis example 570 can select a first category and one of the second categories 550 or 560 to train a deep learning model.
[0101] Methods and processes for enhancing intelligent dialogue Figure 6 is a flowchart illustrating exemplary operations 600 that may be performed by a target device to establish wireless communication with other devices. For example, exemplary operations 600 may be performed by the sound processing and playback device 110 shown in Figure 1.
[0102] An exemplary operation 600 is performed in 602 during playback by analyzing the content of the audio signal before playback of the content to determine whether one or more predefined conditions are met to indicate that the content contains speech.
[0103] In 604, in response to determining that one or more predefined conditions are met, a first playback equalization configured to enhance speech in the content is automatically applied to the audio signal. Automatically applying the first playback equalization may mean applying the first playback equalization without further user input.
[0104] In 606, in response to determining that one or more predefined conditions are not met, the audio signal is either i) not subjected to reproduction equalization, or ii) subjected to a second reproduction equalization different from the first reproduction equalization.
[0105] In particular, in some embodiments, the content is analyzed immediately before and during playback of the audio signal. This analysis may be performed 20 milliseconds to 2 seconds before audio playback. In some embodiments, the analysis is more likely to be performed in the range of 0.2 to 1 second. In some embodiments, the analysis is performed 0.5 seconds before playback. By analyzing the content before and during playback, embodiments of this disclosure differ from existing techniques that do not have the real-time or near real-time processing challenges associated with the disclosed techniques.
[0106] In one embodiment, the analysis is performed using a trained machine learning model, as described above. In another embodiment, metadata associated with the content is analyzed, including whether the content contains utterances. In yet another embodiment, the audio track of an audio signal is analyzed when one or more predefined conditions exceed a threshold.
[0107] In this approach, different channels of an audio signal are analyzed. For example, content is correlated between two channels.
[0108] If an audio signal contains different channels, they may be analyzed and / or compared to determine whether speech is likely to occur. For example, such analysis may include comparing correlated content between two channels (such as correlated content between the left and right channels of a stereo audio signal) and / or analyzing the central channel (e.g., in a 5.0, 5.1, 7.0, or 7.1 audio signal), since the central channel is typically where the majority of dialogue from movies and television takes place. For example, central channel analysis may include determining when central channel playback exceeds a threshold (e.g., a nominal threshold or a threshold relative to other channels) to determine that the content is likely to contain speech, and therefore speech reinforcement equalization should be automatically applied.
[0109] Depending on the embodiment, applying a first reproduction equalization to an audio signal includes a transition from either i) no reproduction equalization or ii) a second reproduction equalization to the first reproduction equalization. The transition may include a stepwise change from either i) no reproduction equalization or ii) a second reproduction equalization to the first reproduction equalization.
[0110] In some embodiments, applying a first reproduction equalization to an audio signal involves increasing the volume of speech within the content relative to other content in the audio signal. In some embodiments, speech content is extracted using a machine learning algorithm. As described herein, speech content is taken from i) correlated content between two channels, or ii) at least one of the central channels. In some embodiments, speech content is taken from the speech component of an audio signal.
[0111] In one embodiment, applying a first reproduction equalization to an audio signal includes reducing the volume of non-speaking content within the audio signal. Furthermore, in another embodiment, the volume of speech within the content is increased.
[0112] In this embodiment, the second reproduction equalization includes low-frequency enhancement, music reproduction enhancement, or a combination of both.
[0113] In one embodiment, the sounds in the environment in which the audio signal is to be reproduced are analyzed to help determine whether to apply a first reproduction equalization to the audio signal.
[0114] In some embodiments, predefined conditions are configurable by the user. Additionally or alternatively, in some embodiments, the first regeneration equalization may be configurable by the user.
[0115] Figure 7 is a flowchart illustrating exemplary operations 700 that may be performed by a target device to establish wireless communication with other devices. For example, exemplary operations 700 may be performed by the sound processing and playback device 110 shown in Figure 1.
[0116] An exemplary operation 700 begins in 702 by detecting speech components in the original set of audio signals using a machine learning network trained on mixed categories of audio content, where speech components consist of sound elements that convey linguistic meaning.
[0117] In 704, the detected speech components are enhanced by transitioning from the original equalization mode to a speech equalization mode that allows the user to better understand the linguistic meaning within them.
[0118] In 706, the original set of audio signals is output in the original equalized mode without detection of the speech component.
[0119] During operation, the performance of 704 and 706 depends on the instance of content within the audio signal, which may change from time to time. Therefore, the enhancement performed in 704 is dynamic and can be done automatically for detected speech components.
[0120] In some embodiments, speech components may be detected using a trained machine learning network. For example, detection may be based on calculating the root mean square of the energy level of each category in a mixed category of audio content. In some cases, the trained machine learning network may include a deep learning model that estimates the energy level of each category in a mixed category of audio content. The mixed category of audio content may include speech components, musical components, and singing components.
[0121] In some cases, speech components may be detected by determining the ratio of the energy level of the speech component to the overall energy level that exceeds a threshold. For example, speech components may include sounds whose meaning is conveyed based on linguistic features. Musical components may include sounds that lack linguistic features. Singing components may include a mixture of sounds that simultaneously contain components of linguistic expression and components of musical expression. In some cases, a trained machine learning network may be trained to identify each category and threshold of a mixed category of audio content based on a known database of film content.
[0122] In some cases, speech components may be detected by processing the audio signal in progress at an advanced time prior to a transition or output operation. For example, the processing or detection operation may be performed in real time or near real time with the minimum delay permitted by the computing power of the handling device. In some cases, the processing device and the sound output device may be separate and independent of each other.
[0123] In some embodiments, enhancing speech components may include gradually fading from the original equalization mode to a speech equalization mode. In this way, speech enhancement equalization may be engaged in or initiated without the user's awareness. Similarly, playback may include smoothly fading from the speech equalization mode to the original equalization mode without the user's awareness. In some cases, the speech equalization mode may include at least one of increasing the magnitude or contrast of speech-related frequency bands, decreasing the magnitude of non-speech-related frequency bands or signal channels, or changing the equalization settings or dynamic range compression settings on non-speech-related frequency bands or signal channels in order to improve the clarity of speech components.
[0124] In one embodiment, enhancing detected speech components may be performed in a first device, and outputting the original set of audio signals may be performed in a second device. For example, the first device may include a soundbar configured to output surround sound (e.g., sound processing and playback device 110 in Figure 1), and the second device may include noise-canceling headphones (e.g., partner device 112 in Figure 1). Thus, when the soundbar receives an audio signal, it detects speech components and performs calculations (e.g., supported by machine learning and neural network calculations as described above) to automatically apply speech mode equalization. The output of speech mode enhancement may be performed by the soundbar, the noise-canceling headphones, or both. The first and second devices may be paired in a short-range wireless communication network.
[0125] In some cases, the first device may be noise-canceling headphones, and the second device may be a soundbar configured to output surround sound. The first and second devices may each include different devices such as a smartphone or other types of wearable electronic devices. Thus, various configurations based on different devices can be configured.
[0126] In some embodiments, the detected speech components may be extracted and played separately in a third device (e.g., another short-range paired speaker or noise-canceling headphones, such as the second partner device 112 in Figure 1). In some cases, the extracted speech components may be used to improve the equalization profile, for example, by identifying several frequency spectra for processing in the equalizer. In some cases, the first, second, and third devices may be configured to produce a mixed surround sound. For example, each of the different devices may have an equalization profile for each category of sound, such as speech, singing, and background music. In some cases, one or more of the devices may include a microphone or noise sensor for noise cancellation. The devices may be paired with a microphone for measuring ambient noise for cancellation.
[0127] In some embodiments, the original set of audio signals may be played back or output without speech content enhancement based on a decision to remove speech components. For example, in a multimedia clip containing occasional speech or dialogue content, speech enhancement processing may be applied only to portions where speech content is detected.
[0128] In some embodiments, the disclosed method is applicable to wireless earphones, earhooks, or ear-to-ear devices. For example, a host such as a mobile phone may be connected to a bud (e.g., the right bud) via Bluetooth®, and the right bud further connects to a left bud using a Bluetooth® link or other wireless technology such as NFMI or NFEMI. The left bud is initially time-synchronized with the right bud. Audio frames (compressed to mono) are transmitted from the left bud, having a timestamp (synchronized with the timestamp of the right bud), as described in the technique above. The right bud transmits these encoded mono frames along with its own frames. The right bud does not wait for audio frames from the left bud having the same timestamp. Instead, the right bud transmits any frames that are available and ready to be transmitted with appropriate packing. It is the responsibility of the receiving application in the host to assemble the packets using the timestamp and channel number. Depending on how it is configured, the receiving application may choose to merge the decoded mono channels of one bud with the decoded mono channels of the other bud into a stereo track based on the timestamp contained in the header of the received encoded frame. This disclosure allows the right bud to simply transfer audio frames from the left bud without decoding the frames. This helps to conserve battery power in truly wireless audio devices.
[0129] In some embodiments, the techniques described herein may be used to determine contextual information about a source device and / or the user of the source device. For example, the techniques may be used to help determine aspects of the user's environment (e.g., noisy place, quiet place, indoors, outdoors, on an airplane, inside a car, etc.) and / or activity (e.g., commuting, walking, running, sitting, driving, flying, etc.). In some such embodiments, sensor data received from the source device may be processed in the target device to determine such contextual information and provide the user with a new or enhanced experience. For example, this could enable, to name a few, customization of playlists or audio content, noise cancellation adjustment, and / or other setting adjustments (e.g., audio equalizer settings, volume settings, notification settings, etc.). Since the source device (e.g., headphones or earphones) typically has limited resources (e.g., memory and / or processing resources), using the techniques described herein to offload data processing from the source device's sensors to the target device, while having a system for synchronizing sensor data in the target device, offers a variety of applications. In some embodiments, the techniques disclosed herein enable a user device to automatically identify an optimized or most preferred configuration or setting for synchronized audio capture operation.
[0130] In some embodiments, the techniques described herein can be used for a wide range of audio / video applications. For example, the techniques can be used for capturing stereo or surround sound audio from a source device to be synchronized on the target device with video captured from the same source device, another source device, and / or target device. For instance, the techniques can be used to synchronize stereo or surround sound audio captured by microphones on a pair of headphones with video captured from a camera on or connected to the headphones, a separate camera, and / or a smartphone camera, where the smartphone (in this example, the target device) performs the audio-video synchronization. This can enable real-time playback of stereo or surround sound audio with video (e.g., for live streaming) and capture of recorded video with stereo or surround sound audio (e.g., for posting to a social media platform or news platform). In addition, the techniques described herein can enable wirelessly captured audio for audio or video messages without interrupting the user's music or audio playback. Thus, the techniques described herein enable the ability to produce immersive and / or noise-free audio for video using a wireless configuration. Furthermore, as can be understood from this disclosure, the described technique enables schemes that were previously only achievable using wired configurations, and thus the described technique frees users from the undesirable and unpleasant experience of being connected by one or more wires.
[0131] While the descriptions of the embodiments of this disclosure are provided above for illustrative purposes, it should be noted that the embodiments of this disclosure are not intended to be limited to any of the embodiments disclosed. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described.
[0132] The above refers to the embodiments presented in this disclosure. However, the scope of this disclosure is not limited to the specific embodiments described. Embodiments of this disclosure may take the form of entirely hardware embodiments, entirely software embodiments (including firmware, resident software, microcode, etc.), or embodiments that combine software embodiments and hardware embodiments, which in this specification may generally be referred to as “components,” “circuits,” “modules,” or “systems.” Furthermore, embodiments of this disclosure may take the form of computer program products embodied in one or more computer-readable media having computer-readable program code embodied thereon.
[0133] Any combination of one or more computer-readable media can be used. A computer-readable media may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of computer-readable storage media include electrical connections with one or more wires, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM, or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the current context, a computer-readable storage medium may be any tangible medium that contains or can store programs.
[0134] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of assumed implementations for various systems, methods, and computer program products. In this regard, each block in a flowchart or block diagram may correspond to a module, segment, or portion of an instruction set containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions described in a block may occur in the order shown in the diagram. For example, two consecutively shown blocks may actually be executed substantially simultaneously, or, in some cases, the blocks may be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, as well as combinations of blocks in a block diagram and / or flowchart, can be implemented in a dedicated hardware-based system that performs a specified function or operates a combination of dedicated hardware and computer instructions. [Explanation of symbols]
[0135] 100 Systems 110 Playback devices 112 Partner Devices 120 Source Devices 130 Networks 140 Cloud 200 Block Diagram 210 Audio signal 220 Machine Learning Networks 230 Speech component enhancement module 240 Calculation Output 250 training datasets 260 Deep Learning Models 305 video clips 310 First instance 312 Category Indicator Frame 314 frames 316 frames 320 Second instance 322 speech components 324 Musical Components 326 Singing component 410 Equalization Mode 420 Second equalization mode 510 utterance dataset 520 Music Datasets 530 Singing Datasets 540 Noise Datasets 550 First Sound Category 560 Second Sound Category 600 operations 700 operations
Claims
1. A method for processing audio signals, Determining whether, during playback of a portion of an audio signal, before playback of a subsequent portion of the audio signal, the content of the subsequent portion of the audio signal is analyzed to determine whether one or more predefined conditions are met to indicate that the content includes speech, wherein the analysis includes detecting speech based on estimating the respective energy levels of the speech component, musical component, and singing component of the subsequent portion of the audio signal. In response to determining that one or more of the above predefined conditions are met, a first playback equalization configured to enhance the utterances in the content is automatically applied to the subsequent portion of the audio signal, A method comprising, in response to determining that one or more of the above predefined conditions are not met, i) not applying a playback equalization to the subsequent portion of the audio signal, or ii) applying a second playback equalization different from the first playback equalization.
2. The method according to claim 1, wherein the analysis is performed using a trained machine learning model.
3. The method according to claim 2, wherein the trained machine learning model comprises a deep learning model for estimating the respective energy levels of the subsequent portions of the audio signal.
4. The method according to claim 1, wherein the analysis further comprises analyzing metadata associated with the content, the metadata indicating that the content includes an utterance.
5. The method according to claim 1, wherein the analysis further comprises analyzing the audio track of the content.
6. The aforementioned content has different channels, The method according to claim 1, wherein the analysis further comprises analyzing the different channels.
7. The method according to claim 1, wherein the analysis further comprises analyzing the central channel of the content.
8. The method according to claim 1, wherein automatically applying the first playback equalization configured to enhance the utterance in the content to the subsequent portion of the audio signal includes a transition from either i) no playback equalization or ii) the second playback equalization to the first playback equalization.
9. The method according to claim 1, wherein the first playback equalization configured to enhance the utterance in the content is automatically applied to the subsequent portion of the audio signal, which includes increasing the volume of the utterance in the content relative to other content in the subsequent portion of the audio signal.
10. The method according to claim 1, wherein the first playback equalization configured to enhance the speech in the content is automatically applied to the subsequent portion of the audio signal, the volume of non-speech content in the subsequent portion of the audio signal is reduced.
11. The method according to claim 10, further comprising increasing the volume of the utterance in the content.
12. The method according to claim 1, wherein the second regeneration equalization comprises at least low-frequency enhancement.
13. The method according to claim 1, wherein the first regeneration equalization or at least one of the one or more predefined conditions is configurable by the user.
14. The method according to claim 1, further comprising analyzing the sound in the environment in which the subsequent portion of the audio signal is to be reproduced.
15. A device for audio signal processing, Memory and The system comprises a processor coupled to the memory, and the processor and the memory are During playback of a portion of an audio signal, the content of the subsequent portion of the audio signal is analyzed before playback of the subsequent portion of the audio signal to determine whether one or more predefined conditions are met to indicate that the content includes speech, and the analysis includes detecting speech based on estimating the respective energy levels of the speech component, musical component, and singing component of the subsequent portion of the audio signal. In response to determining that one or more of the above predefined conditions are met, a first playback equalization configured to enhance the utterances in the content is automatically applied to the subsequent portion of the audio signal. An apparatus configured to, in response to determining that one or more of the above predefined conditions are not met, to either i) not apply a playback equalization to the subsequent portion of the audio signal, or ii) apply a second playback equalization different from the first playback equalization.
16. The apparatus according to claim 15, wherein the processor and the memory are further configured to analyze the content using a trained machine learning model.
17. The aforementioned content has different channels, The apparatus according to claim 15, wherein the memory and the processor are further configured to analyze the content by analyzing the different channels.
18. The apparatus according to claim 15, wherein the processor and the memory are further configured to analyze the content by analyzing the central channel of the content.
19. The apparatus according to claim 15, wherein the processor and the memory are configured to automatically apply the first playback equalization, which is configured to enhance the utterance in the content by transitioning from either i) no playback equalization or ii) the second playback equalization to the first playback equalization, to the subsequent portion of the audio signal.
20. The apparatus according to claim 15, wherein the second regeneration equalization comprises at least low-frequency enhancement.
21. The apparatus according to claim 15, wherein the first regeneration equalization or at least one of the one or more predefined conditions is configurable by the user.
22. The apparatus according to claim 15, wherein the processor and the memory are further configured to analyze the sound in the environment in which the subsequent portion of the audio signal is to be reproduced.
23. A non-temporary computer-readable medium for storing instructions, wherein when an instruction is executed by a device for processing and generating audio signals, the device... During playback of a portion of an audio signal, before playback of the subsequent portion of the audio signal, the content of the subsequent portion of the audio signal is analyzed to determine whether one or more predefined conditions are met to indicate that the content includes speech, and the analysis includes detecting speech based on estimating the respective energy levels of the speech component, musical component, and singing component of the subsequent portion of the audio signal. In response to determining that one or more of the above predefined conditions are met, a first playback equalization configured to enhance the utterances in the content is automatically applied to the subsequent portion of the audio signal. A non-temporary computer-readable medium that, in response to determining that one or more of the aforementioned predefined conditions are not met, either i) does not apply playback equalization to the subsequent portion of the audio signal, or ii) applies a second playback equalization different from the first playback equalization.