Systems and methods for enhancing audio in various environments

By creating personalized audio profiles for users, the problem of unclear dialogue when playing media on mobile devices in noisy environments is solved, and audio optimization to improve dialogue clarity in noisy environments is achieved, avoiding unnecessary noise cancellation and low cost.

CN115362499BActive Publication Date: 2025-08-12DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180026166.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-14
Filing Date
2021-04-02
Publication Date
2025-08-12
Estimated Expiration
2041-04-02

AI Technical Summary

Technical Problem

In noisy environments, it is difficult to clearly hear the actor's conversation when playing media content with conversations on mobile devices, and existing noise reduction devices are expensive and may eliminate ambient noise that users do not want to eliminate.

Method used

Create personalized audio profiles for users, store and apply these settings to optimize audio playback by identifying ambient noise, setting up conversation enhancements, and graphics equalizers, including creating and using profiles on mobile devices, adjusting audio to enhance conversation clarity.

Benefits of technology

Improve the intelligibility of conversations in noisy environments, provide a clear audio experience, while avoiding the ambient noise that users don't want to eliminate, which is cheaper.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115362499B_ABST
    Figure CN115362499B_ABST
Patent Text Reader

Abstract

A novel method and system for creating and using user profiles for dialogue enhancement and sound equalizer adjustments to compensate for various ambient sound scenarios. To create profiles when no ambient sound is present, synthetic / pre-recorded ambient noise can be mixed with the media to simulate noisy conditions.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to International Patent Application No. PCT / CN 2020 / 083083, filed on April 2, 2020; U.S. Provisional Patent Application No. 63 / 014,502, filed on April 23, 2020; and U.S. Provisional Patent Application No. 63 / 125,132, filed on December 14, 2020, which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates to improvements to the audio playback of media content. In particular, the present disclosure relates to setting and applying optimal noise compensation for the audio of media content played in various environments, particularly on mobile devices. Background Art

[0004] Audio playback for media with dialogue (movies, TV shows, etc.) is typically created for enjoyment in relatively quiet environments, such as at home or in a theater. However, it's becoming increasingly common for people to consume this content on the go, using their mobile devices. This becomes a problem because it can be difficult to hear what the actors are saying when there's excessive ambient noise (traffic, crowds, etc.) or due to audio quality limitations of the mobile hardware or the type of audio playback device being used (headphones, etc.).

[0005] A common solution is to use noise-canceling headphones / earbuds. However, this can be an expensive solution and has the disadvantage of eliminating ambient noise (car horns, sirens, loud warnings, etc.) that the user may want to hear. Summary of the Invention

[0006] Various audio processing systems and methods are disclosed herein. Some of these systems and methods may involve creating and using audio adjustment profiles that are customized for a user and specific to different environmental conditions.

[0007] According to a first aspect, a method for configuring a user of a mobile device for use in ambient noise is described, the method comprising: receiving a situation identification of the ambient noise from the user; receiving a noise level of the ambient noise from the user; receiving a dialog boost level of the ambient noise at the noise level from the user; receiving a graphic equalizer setting for the ambient noise at the noise level from the user; playing sample audio for the user from the mobile device when the user sets the dialog boost level and the graphic equalizer setting; and storing the dialog boost level and the graphic equalizer setting for the situation identification using the noise level in a profile on the mobile device, wherein the device is configured to play audio media using the dialog boost level and the graphic equalizer setting when the user selects the profile.

[0008] According to a second aspect, a method for adjusting audio of a mobile device for a user is described, the method comprising: receiving a profile selection from the user, wherein the profile selection is at least related to an ambient noise condition; receiving a noise level of the ambient noise condition from the user; retrieving a dialogue enhancement level and a graphic equalizer setting from a memory on the mobile device; and adjusting the level of the audio using the dialogue enhancement level and the graphic equalizer setting.

[0009] Some or all of the methods described herein may be performed by one or more devices according to instructions (e.g., software) stored on one or more non-transient media. Such non-transient media may include memory devices as described herein, including but not limited to random access memory (RAM) devices, read-only memory (ROM) devices, etc. Accordingly, the various innovative aspects of the subject matter described in this disclosure may be implemented in a non-transient medium on which software is stored. For example, the software may be performed by one or more components of a control system such as those disclosed herein. The software may, for example, include instructions for executing one or more of the methods disclosed herein.

[0010] At least some aspects of the present disclosure may be implemented via one or more devices. For example, one or more devices may be configured to at least partially perform the methods disclosed herein. In some embodiments, the device may include an interface system and a control system. The interface system may include one or more network interfaces, one or more interfaces between the control system and a memory system, one or more interfaces between the control system and another device, and / or one or more external device interfaces.

[0011] The details of one or more embodiments of the subject matter described in this specification are set forth in the following drawings and description. Additional features, aspects, and advantages will become apparent from the description, drawings, and claims. Note that the relative dimensions of the following figures may not be drawn to scale. Like reference numerals and names in the various figures generally indicate similar elements, but different reference numerals between different figures do not necessarily designate different elements. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 An example flow chart for configuring case-specific audio settings is illustrated.

[0013] Figure 2 An example flow diagram for utilizing situation-specific audio settings is illustrated.

[0014] Figure 3 An example flow chart for configuring situation-specific audio settings, including synthesizing ambient noise, is illustrated.

[0015] Figure 4 Illustrated is an example comparison of unadjusted audio and adjusted audio obtained through subjective testing experiments.

[0016] Figure 5A and 5B Illustrated are example frequency response curves for dialogue enhancement applied to one type of codec. Figure 5A The response curve of the output speech segment is shown. Figure 5B The response curves of the output non-speech segments are shown.

[0017] Figure 6A and 6B The diagram shows the application of Figure 5A and Figure 5B Example frequency response curves for dialogue enhancement for different types of codecs. Figure 6A The response curve of the output speech segment is shown. Figure 6B The response curves of the output non-speech segments are shown.

[0018] Figure 7 An example graphical user interface for use with the methods herein is shown.

[0019] Figure 8 An example hardware / software configuration for use with the methods herein is shown. DETAILED DESCRIPTION

[0020] This article describes a solution to the problem of providing intelligible speech in media playback (audio or audio / visual) in noisy environments (ambient noise) by creating and using dialogue enhancement and equalizer settings in profiles for specific users in specific noise levels and types (environment types).

[0021] As used herein, the term "mobile device" refers to a device that is capable of audio playback and can be carried by a user and used in multiple locations. Examples include mobile phones, laptops, tablets, mobile game systems, wearable devices, small media players, etc.

[0022] As used herein, the term "environmental condition," "case," or "case identification" refers to a class of noisy locations / environments that may or may not interfere with the enjoyment of listening to audio media on a mobile device. Examples include home (e.g., "default"), outdoors in densely populated areas (e.g., walking), on public transportation, in noisy indoor environments (e.g., airports), and other situations.

[0023] The term "dialog boost" refers to the application of general sound amplification to the speech component of audio while amplifying non-speech components negligibly. For example, dialog boost can be implemented as an algorithm that continuously monitors the audio being played, detects the presence of dialogue, and dynamically applies processing to improve the intelligibility of the spoken portions of the audio content. In some embodiments, dialog boost analyzes features from the audio signal and applies a pattern recognition system to detect the presence of dialogue at any time. When speech is detected, the speech spectrum is altered as necessary to emphasize the speech content and make it more concise for the listener to hear.

[0024] The term "equalization" or "graphic equalizer" or "GED" refers to frequency-based amplitude adjustments of audio. In a real GED, the amplitude setting would be set by a slider whose position corresponds to the frequency range it controls, but GED in this context also refers to the specific settings a graphic equalizer may have, giving a particular frequency response curve.

[0025] As used herein, the terms "media" or "content" refer to anything with audio content. This could be music, movies, videos, video games, phone conversations, alarms, and the like. In particular, the systems and methods herein are useful for media that has a combination of speech and non-speech components, but these systems and methods can be applied to any media.

[0026] Figure 1A flowchart for creating profiles for different environmental conditions (scenarios) is shown. A user selects a start setting 110 from the device user interface (UI), and a sample playback sample is played for the user 140. This sample can be selected by the system or by the user. The user selects the volume 141 at which it will be played. The user can then automatically or manually (user-selected) experience different ambient noise scenarios 142 (default, walking, public transportation, airport, etc.). The system can cycle through all scenarios or only a selected subset that includes only one selected scenario. If the selected scenario is not the default scenario, the user enters an estimated ambient noise level 125 for their current situation, as well as a dialogue enhancement level 130 and a graphic equalizer (GEQ) setting 135, which combine to provide the user with their subjectively optimal listening experience. These settings 125 and 130 can be applied multiple times in any order (not necessarily the order shown in the figure) and are set based on the playback sample 140 of audio with a speech component. Once the dialogue enhancement 130 and GEQ settings 135 are set according to the user's preferences, they are stored in a profile database / storage 145 for future use. The system can then determine whether all applicable situations have been set up 115. If so, the setup ends 150. If not, the system can move on to the next situation 142 and repeat the setup process for that situation. In some embodiments, the saved setup profiles are also indexed based on the injected noise level 125.

[0027] In some embodiments, each of the dialogue enhancement level and / or GEQ settings is a value from a short list of possible settings, for example, "3" on a scale of 0 to 5. In some embodiments, the settings are real values associated with the setting, such as +10 dB (e.g., at a specific frequency range).

[0028] Figure 2An example flow chart for using profiles created according to the methods described herein is shown. A user launches their media 205 and selects a situation profile 210 that best describes their current situation. If the profiles are indexed based on ambient noise level, an ambient noise level can also be selected. The system then retrieves 210 a profile from a database / memory 215 that matches the selected situation (and, if applicable, the selected noise level). The system then determines whether playback is in a mobility situation 220 (i.e., a situation requiring dialogue enhancement and GEQ adjustment). This determination can be based on user input, device identification, location data, or other means. If the system determines that this is not a mobility situation, normal playback / mixing 240 occurs from the media being played 245, regardless of the ambient noise 250 present at the user's location. This continues until the situation profile is changed 255, at which point a new profile is retrieved 210 and the process begins again. In some embodiments, a new mobility state check is performed at or before the situation switch 255, and the process repeats only when a mobility situation exists. If a mobility context is found 220 , dialogue enhancement 230 and GEQ adjustment 235 are applied to the mix 240 to adjust the media playback 245 to provide intelligible dialogue despite the presence of ambient noise 250 .

[0029] Figure 3 An example flow chart for creating a profile is shown, including using synthesized ambient noise to virtually mix ambient noise with media (as opposed to actually mixing actual ambient noise with media, such as Figure 1 and Figure 2 ). This system is similar to Figure 1 The system is similar to the system described above, except that a check 310 is performed to see whether the user is creating a profile in a noisy location or pre-setting a situation from a relatively noise-free environment (e.g., home). This check can be determined by asking the user or by determining through location services that the mobile device is at "home." In some embodiments, the system always assumes that the user is in a relatively noise-free environment. If the user is not in that location, a situation (ambient noise condition) is selected 320 by the user or the system. The ambient noise for that situation is synthesized 330. In some embodiments, this synthesis can be or be based on pre-recorded noise stored in a database / memory 340. This noise is added to the playback sample 350, and the user can set the noise level 360 to the level they want to experience, thereby adjusting the simulated noise 330. The dialogue enhancement level 370 and GEQ level 380 can be set in the same manner as performing a live setup. The settings are then saved to the database / memory 390 for future use. In some embodiments, the recorded ambient noise is taken from a surround sound source and rendered in a binaural format.

[0030] Figure 4Examples of the perceptual differences the system can make and how performance can be evaluated using comparisons to a reference condition are shown. It can be seen that dialogue enhancement outperforms the reference condition without enhancement and increases the intelligibility of dialogue, making it beneficial for users to tune into media where understanding dialogue is important. For example, Figure 4 It shows that level "2" of dialogue enhancement (DE) 415 exhibits a high level of user preference 405 and subjective intelligibility 410, and thus is likely to be the preferred setting for most users.

[0031] Figure 5A and Figure 5B as well as Figure 6A and Figure 6B An example graph of conversation enhancement is shown. Figure 5A and Figure 5B A graph showing different dialogue enhancement settings for the speech component of a media. As shown, different settings show different curves. In contrast, Figure 5B and Figure 6B The same graphs are shown for different dialogue enhancement settings but for the non-speech component of the media, where the differences between the curves for the different settings are negligible (ie, dialogue enhancement does not enhance the non-speech component). Figure 5A and Figure 5B The level of dialogue enhancement indicated is higher than Figure 6A and Figure 6B Smaller intervals between levels. Depending on how noisy the environment is, different curves can be used: the louder the environment, the more aggressive the curve can be, with more dialogue enhancement on the overall playback content. Figure 5A , where the level of dialogue enhancement indicates that for the speech component of the audio, there is a stronger enhancement in the lower frequencies 505 than in the higher frequencies 510. Figure 6A shows that there is a stronger enhancement in the higher frequencies 610 compared to the lower frequencies 605. In both cases, Figure 5B and Figure 6B It shows that the non-speech components have negligible enhancement at all frequencies.

[0032] Figure 7An example UI (specifically, a graphical user interface, GUI, in this case) for setting a profile is shown. On mobile device 700, inputs for settings can be presented in a simplified form for ease of use. Noise level control 710 can be presented as a finite number of noise levels (e.g., 0 to 5), starting at 0 for no noise and increasing in value as the noise level increases in uniform increments (in actual dB or perceptual steps). Dialogue enhancement setting 720 can be presented as a graphical slider from no enhancement to maximum enhancement. Similarly, GEQ setting 730 can be simplified to a single range of values (shown here as a slider) for selecting a preset GEQ setting (e.g., corresponding to "loud," "flat," "deep," etc.). Conditions 740 can be displayed as icons (with or without text). For example, "Default" can be displayed as a house, "Walking" can be displayed as a person, "Public Transportation" can be displayed as a train or bus, and "Indoor Locations" can be displayed as an airplane (to indicate an airport). Other conditions and icons can be used so that the icons provide users with a quick reference to the conditions they represent.

[0033] Figure 8 An example mobile device architecture for implementing the features and processes described herein according to an embodiment is shown. Architecture 800 can be implemented in any electronic device, including but not limited to desktop computers, consumer audio / video (AV) equipment, radio broadcasting equipment, mobile devices (e.g., smartphones, tablets, laptops, wearable devices). In the example embodiment shown, architecture 800 is for a smartphone and includes (multiple) processors 801, peripheral device interfaces 802, audio subsystems 803, loudspeakers 804, microphones 805, sensors 806 (e.g., accelerometers, gyroscopes, barometers, magnetometers, cameras), location processors 807 (e.g., GNSS receivers), wireless communication subsystems 808 (e.g., Wi-Fi, Bluetooth, cellular), and (multiple) I / O subsystems 809, including touch controllers 810 and other input controllers 811, touch surfaces 812 and other input / control devices 813. The memory interface 814 is coupled to the processor 801, the peripheral device interface 802, and the memory 815 (e.g., flash memory, RAM, ROM). The memory 815 stores computer program instructions and data, including but not limited to: operating system instructions 816, communication instructions 817, GUI instructions 818, sensor processing instructions 819, telephony instructions 820, electronic messaging instructions 821, web browsing instructions 822, audio processing instructions 823, GNSS / navigation instructions 824, and application / data 825. The audio processing instructions 823 include instructions for performing the audio processing described herein. Other architectures with more or fewer components may also be used to implement the disclosed embodiments.

[0034] The system can be provided as a service driven from a remote server, can be provided as a standalone program on the device, can be integrated into a media player application, or can be included as part of the operating system as part of its sound settings.

[0035] A number of embodiments of the present disclosure have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the present disclosure. Therefore, other embodiments are also within the scope of the appended claims.

[0036] As described herein, embodiments of the present invention may therefore be directed to one or more of the following enumerated exemplary embodiments. Accordingly, the present invention may be implemented in any form described herein, including but not limited to the following enumerated exemplary embodiments (EEE) that describe the structure, features, and functions of some portions of the present invention:

[0037] EEE1. A method for configuring a user of a mobile device for use in ambient noise, the method comprising: receiving a situation identification of the ambient noise from the user; receiving a noise level of the ambient noise from the user; receiving a conversation enhancement level for the ambient noise at the noise level from the user; receiving a graphic equalizer setting for the ambient noise at the noise level from the user; playing sample audio for the user from the mobile device when the user sets the conversation enhancement level and the graphic equalizer setting; and storing the conversation enhancement level and the graphic equalizer setting for the situation identification at the noise level in a configuration file on the mobile device, wherein the device is configured to play audio media using the conversation enhancement level and the graphic equalizer setting when the user selects the configuration file.

[0038] EEE2. The method of EEE1 further comprises: simulating the ambient noise at the noise level; and mixing the simulated ambient noise with the sample audio before playing the sample audio.

[0039] EEE3. The method of EEE2, wherein the simulation comprises retrieving stored pre-recorded environmental noise from a memory.

[0040] EEE4. The method of EEE3, wherein the stored pre-recorded ambient noise is in binaural format.

[0041] EEE5. The method of any one of EEE1 to EEE4, further comprising: presenting graphical user interface controls on the mobile device for setting the situation identification, the noise level, the dialogue enhancement level, and the graphic equalizer setting.

[0042] EEE6. The method of any one of EEE1 to EEE5, wherein the profile corresponds to both the situation identification and the noise level.

[0043] EEE7. A method for adjusting the audio of a mobile device for a user, the method comprising: receiving a profile selection from the user, wherein the profile selection is at least related to an ambient noise condition; receiving a noise level of the ambient noise condition from the user; obtaining a dialogue enhancement level and a graphic equalizer setting from a memory on the mobile device; and adjusting the level of the audio using the dialogue enhancement level and the graphic equalizer setting.

[0044] EEE8. The method of EEE7 further comprises presenting, on the mobile device, a graphical user interface control for selecting a profile corresponding to the ambient noise condition.

[0045] EEE9. A device configured to execute at least one of the methods described in EEE1 to EEE8 in software or firmware.

[0046] EEE10. A non-transitory computer-readable medium, which, when read by a computer, instructs the computer to execute at least one of the methods described in EEE1 to EEE8.

[0047] EEE11. The device of EEE9, wherein the device is a telephone.

[0048] EEE12. The device of EEE9, wherein the device is at least one of the following: a mobile phone, a laptop computer, a tablet computer, a mobile gaming system, a wearable device, and a small media player. EEE13.

[0049] EEE13. The device of any one of EEE9, EEE11 or EEE12, wherein the software or firmware is part of an operating system of the device. EEE13.

[0050] EEE14. The device of any one of EEE9, EEE11 or EEE12, wherein the software or firmware runs a standalone program on the device.

[0051] EEE15. The method as described in any one of EEE1 to EEE8, wherein the method is performed by an operating system of a mobile device.

[0052] The present disclosure relates to certain embodiments for the purpose of describing some innovative aspects described herein and examples of contexts in which these innovative aspects can be implemented. However, the teachings of this article can be applied in various different ways. In addition, the described embodiments can be implemented in various hardware, software, firmware, etc. For example, various aspects of the present application can be at least partially embodied in a device, a system including more than one device, a method, a computer program product, etc. Therefore, various aspects of the present application can take the form of hardware embodiments, software embodiments (including firmware, resident software, microcode, etc.), and / or embodiments combining both software and hardware aspects. These embodiments can be referred to as "circuits", "modules", "devices", "devices" or "engines" in this article. Some aspects of the present application can take the form of a computer program product implemented in one or more non-transient media, and the non-transient media has a computer-readable program code implemented thereon. Such non-transient media can, for example, include a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. Thus, the teachings of the present disclosure are not intended to be limited to the embodiments shown in the drawings and / or described herein, but rather have broad applicability.

Claims

1. A method for configuring a user of a mobile device for use in ambient noise, the method comprising: receiving, from the user, identification of a situation in which the environmental noise is received; receiving a noise level of the ambient noise from the user; receiving from the user a conversation enhancement level of the ambient noise at the noise level; receiving from the user a graphic equalizer setting of the ambient noise at the noise level; playing sample audio for the user from the mobile device while the user sets the dialogue enhancement level and the graphic equalizer setting; as well as storing the dialogue enhancement level and the graphic equalizer setting for the situation identification of the noise level in a profile on the mobile device, wherein the mobile device is configured to play audio media using the dialogue enhancement level and the graphic equalizer setting when the user selects the profile, wherein the method further comprises: simulating the ambient noise at the noise level; and The simulated ambient noise is mixed with the sample audio before playing the sample audio.

2. The method according to claim 1, wherein The simulation includes retrieving stored pre-recorded ambient noise from a memory.

3. The method according to claim 2, wherein: The stored pre-recorded ambient noise is in binaural format.

4. The method according to any one of claims 1 to 3, further comprising: Graphical user interface controls for setting the situation identification, the noise level, the dialogue enhancement level, and the graphic equalizer setting are presented on the mobile device.

5. The method according to any one of claims 1 to 3, wherein The profile corresponds to both the situation identification and the noise level.

6. The method according to any one of claims 1 to 3, wherein The method is executed by an operating system of a mobile device.

7. A method for adjusting audio of a mobile device for a user, the method comprising: receiving a profile selection from the user, wherein the profile selection is related to at least an ambient noise condition, wherein the profile selection indicates a profile stored on the mobile device when performing the method according to any one of claims 1-6; receiving a noise level of the ambient noise condition from the user; Retrieving the graphic equalizer settings and the dialogue enhancement level of the selected profile from memory on the mobile device; and The level of the audio is adjusted using the dialogue enhancement level and the graphic equalizer settings.

8. The method of claim 7, further comprising presenting, on the mobile device, a graphical user interface control for selecting a profile corresponding to the ambient noise conditions.

9. The method according to any one of claims 7 to 8, wherein The method is executed by an operating system of a mobile device.

10. A mobile device comprising: processor; and A memory storing audio processing instructions, the audio processing instructions comprising instructions for performing the processing of the method according to any one of claims 1 to 9.

11. The mobile device according to claim 10, wherein: The mobile device is a phone.

12. The mobile device according to claim 10, wherein: The mobile device is at least one of: a cell phone, a laptop computer, a tablet computer, a mobile gaming system, a wearable device, and a small media player.

13. The mobile device according to any one of claims 10 to 12, wherein: The audio processing instructions are included in software or firmware that is part of an operating system of the mobile device.

14. The mobile device according to any one of claims 10 to 12, wherein: The audio processing instructions are included in software or firmware that runs as a standalone program on the mobile device. 15 . A non-transitory computer-readable medium, which, when read by a computer, instructs the computer to execute the method according to claim 1 .

Citation Information

Patent Citations

  • Speech Intelligibility

    US20110125494A1

  • Systems and methods for enhancing targeted audibility

    US20150281853A1