Customized voice processing based on user-specific and hardware-specific voice information.

The method addresses the challenge of tailoring audio settings to individual listeners by using user-specific and device-specific information to generate customized audio signals, ensuring consistent and personalized audio experiences across different environments.

JP7862465B2Active Publication Date: 2026-05-19HARMAN INT IND INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HARMAN INT IND INC
Filing Date
2024-04-25
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing audio environments often fail to tailor audio settings to the personal preferences of individual listeners, requiring repetitive customization and are impractical for shared environments, especially for users with hearing impairments.

Method used

A method for processing audio signals that accesses user-specific and device-specific audio processing information to generate a customized audio signal, enabling personalized audio experiences across different environments and devices, including cloud-based and device-based processing.

Benefits of technology

Enables personalized audio experiences that automatically adapt to individual preferences and hearing needs, providing consistent audio quality across various environments without the need for repeated customization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007862465000001
    Figure 0007862465000001
  • Figure 0007862465000002
    Figure 0007862465000002
  • Figure 0007862465000003
    Figure 0007862465000003
Patent Text Reader

Abstract

To provide customized audio processing based on user-specific audio information and hardware-specific audio information.SOLUTION: A method of audio signal processing comprises: accessing user-specific audio processing information for a particular user; determining identity information of an audio device for producing sound output from an audio signal; based on the identity information of the audio device, accessing device-specific audio processing information for the audio device; generating a customized audio-processing procedure for the audio signal based on the user-specific audio processing information and the device-specific audio processing information; and generating a customized audio signal by processing the audio signal with the customized audio-processing procedure.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority benefit of U.S. Provisional Patent Application No. 62 / 788,677, filed on January 4, 2019. The subject matter of this related application is hereby incorporated by reference in its entirety.

[0002] Embodiments of the present disclosure generally relate to audio devices, and more specifically, to customized audio processing based on user - specific audio information and hardware - specific audio information.

Background Art

[0003] In the field of audio entertainment, a listener's audio experience can be affected by various aspects of the current audio environment (such as a room, a vehicle, and a headset device, etc.). For example, the settings of the bass and treble ranges, the volume balance between speakers, and other characteristics of the audio environment can either degrade or enhance the listener's audio experience depending on whether such characteristics match the listener's personal audio preferences. Thus, when the audio environment conflicts with the listener's personal audio preferences (e.g., the bass is too loud), even if the listener selects a favorite audio, the audio experience may decline.

[0004] By customizing individual audio products such as in-car audio systems, wireless headphones, and home entertainment systems, the audio environment can be tailored to the listener's personal audio preferences for that environment. For example, the performance of an audio system in a particular room may be optimized by room equalization, which can compensate for problems caused by the interaction of sounds generated within the room itself and / or further take into account the listener's audio preferences. In another example, a listener can set up equalization, volume, and other settings in the audio system in a vehicle, thereby optimizing the resulting audio environment for that listener in that vehicle. As a result, that particular listener has an optimal in-room listening experience tailored to their personal audio preferences and the acoustic characteristics of the audio environment.

[0005] One drawback of customizing audio environments is that such customizations are generally not tailored to the current listener, but rather relate to the unique audio environment customized by the last listener who performed the customization. Therefore, when a new listener enters a room or vehicle with an audio environment customized by a previous listener, the customization set up by that previous listener is applied as the default setting. As a result, the customization process must be repeated whenever a different listener enters an optimized audio environment, which can be time-consuming and frustrating for the new listener. Furthermore, it may be impractical or impossible to achieve certain personal audio preferences every time a user enters an optimized audio environment. For example, gain adjustments can be employed in an audio environment to compensate for a particular listener's hearing impairment profile, but conducting hearing tests every time a listener recustomizes the audio environment is not very practical. Consequently, such gain adjustments are generally impractical in audio environments shared with other listeners, nor can they be conveniently applied to different audio environments.

[0006] Considering the above, it seems that more effective techniques for performing customized audio processing in audio environments would be useful. [Overview of the project] [Means for solving the problem]

[0007] Various embodiments describe methods for processing audio signals, which include accessing user-specific audio processing information relating to a specific user; determining identification information for an audio device for generating sound output from an audio signal; accessing device-specific audio processing information relating to an audio device based on the identification information for the audio device; generating a customized audio processing procedure relating to an audio signal based on the user-specific audio processing information and the device-specific audio processing information; and generating a customized audio signal by processing the audio signal with the customized audio processing procedure.

[0008] At least one technical advantage of the disclosed technology over the prior art is that it enables the personalization of the listener's voice experience regardless of the current voice environment. Specifically, the listener's personal preferences and / or hearing impairment profile can be automatically applied to any voice environment, while also understanding the voice characteristics of the voice environment, so that the listener does not need to recustomize the voice system in each voice environment. A further advantage is that the personalized voice experience can be implemented in voice environments including high-performance voice devices that perform some or all of the voice signal processing necessary to produce the personalized voice experience, or "low-performance" voice devices that do not perform any voice signal processing. These technical advantages represent one or more technical improvements over the prior art approach.

[0009] To allow for a more detailed understanding of the features listed above in one or more embodiments, a more specific description of one or more embodiments summarized above may be provided by reference to a specific embodiment, some of which are shown in the accompanying drawings. However, it should be noted that the accompanying drawings only show general embodiments and are not intended to limit their scope in any way, but rather to encompass other embodiments with respect to the scope of various embodiments. The present invention provides, for example, the following: (Item 1) A method for processing audio signals, wherein the method is: Accessing user-specific speech processing information about a particular user, In order to generate sound output from an audio signal, the identification information of the audio device is determined, Based on the identification information of the voice device, access device-specific voice processing information relating to the voice device, Based on the user-specific voice processing information and the device-specific voice processing information, a customized voice processing procedure is generated for the voice signal. By processing the aforementioned audio signal with the customized audio processing procedure, a customized audio signal is generated. The method, including the method described above. (Item 2) The method of the item, further comprising causing the audio device to generate an audio output from the customized audio signal. (Item 3) The method of any of the above items, wherein causing the audio device to generate sound output from the customized audio signal includes transmitting the customized audio signal to the audio device via a wireless connection. (Item 4) The method according to any of the above items, wherein processing the audio signal with the customized audio processing procedure is performed using a processor located outside the audio device. (Item 5) The method according to any of the above items, wherein processing the audio signal with the customized audio processing procedure is performed using a processor included in the audio device. (Item 6) Accessing user-specific speech processing information relating to the aforementioned specific user means Determining the identification information of the aforementioned specific user, Based on the identification information of the aforementioned specific user, the user-specific voice processing information is retrieved from the cloud-based repository. A method of any of the above items, including: (Item 7) Accessing user-specific speech processing information relating to the aforementioned specific user means Determining the identification information of the aforementioned specific user, Based on the identification information of the specific user, the user-specific speech processing information is read from a computing device configured to generate the customized speech processing procedure. A method of any of the above items, including: (Item 8) The method according to any of the above items, wherein generating the customized speech processing procedure includes generating a synthesized speech equalization curve from information contained in at least one of the user-specific speech processing information or the device-specific speech processing information. (Item 9) The method according to any of the above items, wherein generating the synthesized sound equalization curve includes combining all sound equalization curves included in the user-specific speech processing information or the device-specific speech processing information. (Item 10) Generating the customized audio signal using the customized audio processing procedure is The process involves modifying the aforementioned audio signal using the synthesized sound equalization curve to generate a modified audio signal. Performing a gain correction operation indicated in at least one of the user-specific audio information or the device-specific audio information of the modified audio signal, A method of any of the above items, including: (Item 11) A non-transient computer-readable medium, which, when executed by a processor, the processor, Accessing user-specific speech processing information about a particular user, In order to generate sound output from an audio signal, the identification information of the audio device is determined, Based on the identification information of the voice device, access device-specific voice processing information relating to the voice device, Based on the user-specific voice processing information and the device-specific voice processing information, a customized voice processing procedure is generated for the voice signal. By processing the aforementioned audio signal with the customized audio processing procedure, a customized audio signal is generated. A non-transient computer-readable medium that stores instructions for performing steps such as the above. (Item 12) A non-transient computer-readable medium according to any of the above items, wherein generating the customized voice processing procedure for the voice signal based on the user-specific voice processing information and the device-specific voice processing information further includes generating the customized voice processing procedure for the voice signal based on environment-specific information. (Item 13) The method further comprises determining the environment-specific information based on at least one of the identification information of the voice device and the identification information of a particular user, according to any of the above items, for a non-transient computer-readable medium. (Item 14) Accessing user-specific speech processing information relating to the aforementioned specific user means Receiving user input indicating a unique equalization profile, accessing the unique equalization profile; A non-transitory computer-readable medium according to any of the above items, including . (Item 15) Generating the customized voice processing procedure includes generating the customized voice processing procedure based on the unique equalization profile. A non-transitory computer-readable medium according to any of the above items. (Item 16) The method further includes generating the unique equalization profile based on a personalization test performed by the specific user. A non-transitory computer-readable medium according to any of the above items. (Item 17) Accessing the user-specific voice processing information regarding the specific user includes: determining the identification information of the specific user; reading the user-specific voice processing information from a cloud-based repository based on the identification information of the specific user; A non-transitory computer-readable medium according to any of the above items, including . (Item 18) Accessing the user-specific voice processing information regarding the specific user includes: determining the identification information of the specific user; reading the user-specific voice processing information from a computing device configured to generate the customized voice processing procedure based on the identification information of the specific user; A non-transitory computer-readable medium according to any of the above items, including . (Item 19) Generating the customized voice processing procedure includes generating a synthetic sound equalization curve from information included in at least one of the user-specific voice processing information or the device-specific voice processing information. A non-transitory computer-readable medium according to any of the above items. (Item 20) It is a system, Memory for storing instructions, A processor coupled to the memory, when executing the instruction, Accessing user-specific speech processing information about a particular user, In order to generate sound output from an audio signal, the identification information of the audio device is determined, Based on the identification information of the voice device, access device-specific voice processing information relating to the voice device, Based on the user-specific voice processing information and the device-specific voice processing information, a customized voice processing procedure is generated for the voice signal. By processing the aforementioned audio signal with the customized audio processing procedure, a customized audio signal is generated. The processor is configured to perform steps such as the following: The system including the above. (Summary) A method for processing an audio signal, the method comprising: accessing user-specific audio processing information relating to a specific user; determining identification information for an audio device for generating sound output from an audio signal; accessing device-specific audio processing information relating to an audio device based on the identification information for the audio device; generating a customized audio processing procedure relating to an audio signal based on the user-specific audio processing information and the device-specific audio processing information; and generating a customized audio signal by processing the audio signal with the customized audio processing procedure. [Brief explanation of the drawing]

[0010] [Figure 1] This is a schematic diagram showing a personalized voice system configured to implement one or more aspects of the present disclosure. [Figure 2] This is a flowchart of method steps for generating user-specific information to personalize the voice experience, according to various embodiments of the present disclosure. [Figure 3] This is a flowchart of method steps for generating a customized audio signal according to various embodiments of the present disclosure. [Figure 4] This is a schematic diagram illustrating an individualized voice system according to various embodiments of the present disclosure. [Figure 5] This is a conceptual block diagram of a computing system configured to implement one or more aspects of various embodiments.

[0011] For clarity, the same reference numeral is used to designate the same element that is common to multiple figures, where appropriate. It is conceivable that features of one embodiment may be incorporated into other embodiments without being further enumerated. [Modes for carrying out the invention]

[0012] Embodiments described herein provide users with device-based and / or cloud-based personalized audio experiences in various audio environments, such as at home, in a car, and / or while traveling (e.g., using headphones). The personalized audio experience is optimized for a specific user's listening preferences and hearing impairments through individual sound and audio experience adjustments. When a user transitions from listening to audio content in one audio environment (e.g., using headphones) to another audio environment (e.g., using an in-car audio system), the individual listening preferences and / or hearing impairment settings associated with the user are implemented in each audio environment. Thus, the embodiments generate a customized audio experience for a specific user, continuously following that user from one audio environment to another. As a result, the user's audio experience remains substantially the same, even though different audio devices in different audio environments are providing the audio content. In various embodiments, combinations of mobile computing devices, software applications (or “apps”), and / or cloud services bring personalized audio experiences to a wide variety of devices and environments. One such embodiment is described below in conjunction with Figure 1.

[0013] Figure 1 is a schematic diagram showing a personalized voice system 100 configured to implement one or more aspects of the present disclosure. The personalized voice system 100 includes, but is not limited to, one or more voice environments 110, a user profile database 120, a device profile database 130, and a mobile computing device 140. The personalized voice system 100 is configured to provide a personalized voice experience to a particular user, regardless of which unique voice environment 110 is currently providing the voice experience to the user. In some embodiments, the voice content relating to the voice experience is stored locally within the mobile computing device 140, and in other embodiments, such voice content is provided by a streaming service 104 implemented on a cloud infrastructure 105. The cloud infrastructure 105 may be any technically feasible internet-based computing system, such as a distributed computing system and / or a cloud-based storage system.

[0014] Each of the one or more audio environments 110 is configured to play audio content related to a specific user. For example, the audio environment 110 may include, but is not limited to, one or more audio environments 101 in a car (or other vehicle), headphones 102, and smart speakers 103. In the embodiment shown in Figure 1, the audio environment 110 plays audio content received from a mobile computing device 140, for example, via a wireless connection (e.g., Bluetooth® and / or WiFi®) and / or a wired connection. As a result, the audio environment 110 may include any audio device that can receive audio content directly from a mobile computing device 140, such as a “low-end” speaker in a home, a stereo system in a car, or a conventional pair of headphones. Furthermore, in the embodiment shown in Figure 1, the audio environment 110 does not rely on the ability to process audio signals within the cloud-based infrastructure 105, or the ability to receive audio content or other information from entities implemented within the cloud-based infrastructure 105.

[0015] Each of the one or more audio environments 110 includes one or more speakers 107, and in some embodiments, includes one or more speakers 107 and one or more sensors 108. The speaker(s) 107 are audio output devices configured to generate sound output based on customized audio signals received from a mobile computing device 140. The sensor(s) 108 are configured to acquire biometric data (e.g., heart rate, skin conductance, and / or equivalent) from the user and transmit signals associated with the biometric data to the mobile computing device 140. The biometric data acquired by the sensor(s) 108 can then be processed by a control algorithm 145 running on the mobile computing device 140 to determine one or more individual audio preferences of a particular user. In various embodiments, the sensor(s) 108 may include any type of image sensor, electrical sensor, biosensor, etc., capable of acquiring biometric data, including, but not limited to, cameras, electrodes, microphones, etc.

[0016] The user profile database 120 stores user-specific and device-specific information that enables a similar personalized voice experience occurring in any of the voice environments 110 for a particular user. As shown, the user profile database 120 is implemented within the cloud-based infrastructure 105 and is therefore available for access by the mobile computing device 140 whenever the mobile computing device 140 has an internet connection. Such internet connection can be via a cellular connection, a Wi-Fi® connection, and / or a wired connection. The user-specific and device-specific information stored in the user profile database 120 may include one or more user preference equalization (equalization) profiles 121, one or more environment equalization (equalization) profiles 122, and a hearing impairment compensation profile 123. In some embodiments, the information associated with a particular user and stored in the user profile database 120 is also stored locally within the mobile computing device 140 associated with that particular user. In this embodiment, user preference profiles 121, environmental equalization profiles 122, and / or hearing impairment compensation profiles 123 are stored in the local user profile database 143 of the mobile computing device 140.

[0017] A user preference profile(s) 121 includes user-specific information that is employed to produce a personalized audio experience in one of the audio environments 110 for a particular user. In some embodiments, the user preference profile(s) 121 includes acoustic filters and / or equalization curves associated with a particular user. Generally, when employed as part of a customized audio processing procedure on an audio signal by an audio processing application 146 of a mobile computing device 140, the acoustic filters or equalization curves adjust the amplitude of the audio signal at a particular frequency. Thus, audio content selected by a particular user and played in one of the audio environments 110 is modified to match that user's individual listening preferences. Alternatively, or in addition, in some embodiments, the user preference profile(s) 121 includes signal processing preferred by other users, such as dynamic range compression, dynamic expansion, audio limiting, and / or spatial processing of the audio signal. In this embodiment, when selected by the user, the signal processing preferred by the user can also be employed by an audio processing application 146 that modifies the audio content when it is played back in one of the audio environments 110.

[0018] In some embodiments, the user preference profile(s) 121 includes one or more user preference-based equalization curves that reflect the voice equalization preferred by a particular user associated with the user profile database 120. In such embodiments, the user preference-based equalization curve may be a pre-configured equalization curve selected by the user during the setup of their preferred listening settings. Alternatively, or in addition, in such embodiments, the user preference-based equalization curve may be a pre-configured equalization curve associated with a different user, such as a preference-based equalization curve associated with a known musician or celebrity. Alternatively, or in addition, in such embodiments, the user preference-based equalization curve may be an equalization curve that includes one or more individual amplitude adjustments made by the user during the setup of their preferred listening settings. Alternatively, or in addition, in such embodiments, the user preference-based equalization curve may include head-related transfer function (HRTF) information specific to a particular user. When such user preference-based equalization curve is employed by the speech processing application 146 as part of a customized speech processing procedure, it can enable an immersive and / or three-dimensional speech experience for a specific user associated with the user preference equalization curve.

[0019] In some embodiments, each user preference profile 121 may be associated with a specific category or multiple categories of music playback, a specific time or multiple times of the day, a specific set of biofeedback (which may indicate atmosphere) received from the user via one or more sensors 108, etc. Thus, different user preference profiles 121 can be employed to produce different personalized audio environments for the same user. For example, different user preference equalization curves can be employed to produce a personalized audio environment for the user based on user selection using the user interface of the mobile computing device 140.

[0020] The environmental equalization profile(s) 122 includes location-specific information that is employed to produce a personalized audio experience in any of the audio environments 110 for a particular user. In some embodiments, the environmental equalization profile(s) 122 includes acoustic filters and / or equalization curves configured for the specific audio environment 110 and / or specific locations within the specific audio environment 110, respectively.

[0021] In some embodiments, one of the environmental equalization profiles 122 is configured to provide equalization compensation for problems arising from sounds and / or surface interactions within the specific sound environment 110. For example, the user's audio experience can be improved with respect to a specific seat position or location within a room when such environmental equalization profile 122 is employed by an audio processing application 146 as part of a customized audio processing procedure. With respect to a fixed environment such as the interior of a specific vehicle with known loudspeaker types and locations, such environmental equalization profile 122 can be determined without user interaction and optionally provided to the user as a pre-set corrective equalization. Alternatively, or in addition, such pre-set environmental equalization profile 122 can further be modified by the user during user sound preference testing or setup operations with respect to the personalized audio system 100. With respect to other environments, such as specific locations within a particular room, the environmental equalization profile 122 can be determined by user interaction-based tests, such as user sound preference tests, conducted at specific locations within a particular room via a speaker 107 (e.g., a smart speaker 103), a sensor 108, and a mobile computing device 140. In some embodiments, the user sound preference test can be performed using a control application 145, a voice processing application 146, or any other suitable software application running on the mobile computing device 140.

[0022] The hearing impairment compensation profile 123 includes user-specific information that can be employed to compensate for hearing impairments associated with a particular user. According to various embodiments, such hearing impairment compensation may be a component of a personalized speech experience for a user associated with the user profile database 120. Generally, the hearing impairment compensation profile 123 includes one or more gain compression curves selected to compensate for hearing impairments detected in the user profile database 120, or otherwise, hearing impairments associated with a user associated with the user profile database 120. In some embodiments, such gain compression curves may enable multiband compression, where different parts of the frequency spectrum of the speech signal undergo different levels of gain compression. Gain compression can increase low-level sounds below a threshold level without producing high-level sounds that become uncomfortably loud. As a result, gain compression is employed to compensate for hearing impairments of a particular user, and such gain compression is performed using one or more gain compression curves included in the hearing impairment compensation profile 123.

[0023] In some embodiments, a particular user's hearing impairment is determined based on demographic information collected from the user, for example, using a questionnaire delivered to the user using an appropriate software application running on a mobile computing device 140. In such embodiments, the questionnaire may be delivered to the user during the setup operation of the personalized voice system 100. In other embodiments, such hearing impairment is determined based on one or more hearing tests performed using one or more speakers 107, one or more sensors 108, and the mobile computing device 140. In either case, based on such hearing impairment, hearing impairment in a certain frequency band is determined, and an appropriate hearing impairment compensation profile 123 is selected. For example, an intrinsic gain compression curve can be selected or configured for a user based on demographic information and / or hearing test information collected from the user. The intrinsic gain compression curve is then included in the hearing impairment compensation profile 123 for that user and can be employed by a voice processing application 146 as part of a customized voice processing procedure to produce a personalized voice experience for the user. As a result, a personalized voice experience, including hearing compensation, can be provided to the user in any of the voice environments 110.

[0024] Figure 2 is a flowchart of method steps for generating user-specific information to personalize the audio experience, according to various embodiments of this disclosure. The user-specific information generated by the method steps may include one or more user preference profiles 121, an environmental equalization profile 122, and / or a hearing impairment compensation profile 123. Although the method steps are described in relation to the system in Figure 1, those skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of various embodiments.

[0025] As shown, method 200 begins in step 201, where a suitable software application launched on a mobile computing device 140, such as a control application 145, receives user input to initiate a hearing impairment test for the current user.

[0026] In step 202, the software application selects a specific hearing impairment test to perform. Each such hearing impairment test can determine hearing impairment compensation information associated with the user. For example, in some embodiments, a particular hearing impairment test can be specific to a different sound environment 110 and / or a specific user. Thus, in such embodiments, different hearing impairment tests are selectable for the user depending on the current sound environment 110. Furthermore, in some embodiments, various types of hearing impairment tests are selectable, such as hearing impairment tests based on demographic information and hearing impairment tests based on quantifying hearing loss across multiple frequency bands.

[0027] In step 203, the software application performs the hearing impairment test selected in step 202. For example, in some instances, user demographic information may be collected to determine what hearing impairment compensation is likely to be beneficial to the user. Alternatively or further, in some embodiments, the hearing impairment test is performed via the software application, one or more speakers 107 placed in the current sound environment 110, and one or more sensors 108 placed in the current sound environment 110. In such embodiments, the user's hearing impairment is quantifiable for each of several frequency bands, and the results of such tests are included in the hearing impairment compensation profile 123 for the user.

[0028] In step 204, the software application determines whether there are any remaining hearing impairment tests to be performed on the user in the current audio environment 110. For example, in some embodiments, the software application prompts the user with a list of hearing impairment tests that have not yet been performed by the user. If this is done, method 200 returns to step 202 to select another hearing impairment test to be performed; otherwise, method 200 proceeds to step 205.

[0029] In step 205, the software application receives user input to initiate a personalization test of the current user and / or voice environment 110.

[0030] In step 206, the software application selects specific personalization tests to perform. For example, in some embodiments, possible personalization tests include, but are not limited to, a personal equalization preference test to determine a specific user preference profile 121 for the user, an environmental equalization test to determine a specific environmental equalization profile 122 for a particular voice environment 110 identified by the user, and an HRTF test to determine a specific HRTF for the user.

[0031] In step 207, the software application performs the personalization test selected in step 206. For example, in instances where a personal equalization preference test is performed, pre-configured acoustic filters or other acoustic profiles may be demonstrated to the user through the current audio environment 110 so that the user can select the pre-configured acoustic profile that provides the best audio experience. During such a personalization test, the software application may display an acoustic pre-configuration ranking screen containing one or more pre-configured acoustic filter responses. The user also listens to test tones processed sequentially by each of the pre-configured acoustic filter responses and ranks the pre-configured acoustic filter responses based on personal preferences. In some embodiments, as so employed, the pre-configured acoustic filters are based on data related to the user. For example, the software application may search historical data related to demographic data associated with or entered by the user in order to select one or more pre-configured acoustic filters that users within a demographic range have previously ranked highly. Alternatively or further, in some embodiments, such a personalization test includes a “sight test” type test that relies on A / B choices made by the user. Such visual acuity tests can rapidly narrow down the selection based on A / B comparison listening tests. Alternatively, or further, in some embodiments, such personalization tests provide individual editing of the intrinsic frequency band levels of selected pre-set acoustic filter responses.

[0032] In instances where an environmental equalization test is performed, pre-configured acoustic filters can be demonstrated to the user through the current audio environment 110, allowing the user to select a highly-ranked pre-configured acoustic filter to provide the best audio experience for a specific audio environment 110 indicated by the user. During such a personalization test, the software application may display an acoustic pre-configuration ranking screen containing one or more pre-configured acoustic filter responses, and may also perform sequential or A / B testing of various pre-configured acoustic filters. Alternatively, or further, in some embodiments, such an environmental equalization test allows the user to individually edit the specific frequency band levels of the selected pre-configured acoustic filter response. For example, in some embodiments, different sliders are displayed for selecting the desired gain for each frequency band.

[0033] In instances where HRTF testing is performed, a unique HRTF value for a user is determined based on user characteristics that influence sound localization in the audio environment 110, such as the user's anthropometric features. This unique HRTF value for a user is also included in the user profile database 120 as a user preference profile 121 that can be used to process the audio signal. When the audio output based on the HRTF-processed audio signal is played back in the audio environment 110, the user's hearing generally interprets the audio output as originating from all directions, rather than from individual audio devices placed within the audio environment 110.

[0034] In step 208, the software application determines whether there are any remaining personalization tests to be performed for the user in the current voice environment 110. For example, in some embodiments, the software application prompts the user with a list of personalization tests that have not yet been performed by the user. If this is done, method 200 returns to step 206 to select another personalization test to be performed; otherwise, method 200 proceeds to step 209.

[0035] In step 209, the software application includes user-specific and / or environment-specific information determined by the personalization tests described above in the user profile database 120.

[0036] Returning to Figure 1, the device profile database 130 includes a number of device-specific equalization curves 131, each associated with a specific audio device, such as a specific manufacturer and model of headphones, an in-car audio system, or a smart speaker manufacturer and model. Furthermore, each device-specific equalization curve 131 is configured to modify the audio signal before it is played back by the associated audio device, in which case the audio signal is modified to compensate for the non-ideal frequency response of the audio device. In some embodiments, an ideal audio system produces an audio output that has little to no distortion of the input signal on which the audio output is based. That is, an ideal audio system operates with a constant and flat frequency response over the system's operating frequency (e.g., 20Hz to 20kHz). Furthermore, in an ideal audio system, the audio output is delayed by exactly the same amount of time at all operating frequencies of the system. In practice, any given audio system will have a different frequency response that deviates from the above-described frequency response of an ideal audio system. Furthermore, many speakers have a rough, non-flat frequency response that includes peaks and dips at certain frequencies and / or excessively emphasizes the response at certain frequencies. Generally, speakers with a non-flat frequency response produce audio output that is audible to most users and universally disliked, with added resonance or coloration. As a result, even when considerable effort and resources are directed towards capturing a particular performance with a high-quality recording, the frequency response of the playback device can significantly degrade the user experience when listening to the recording.

[0037] In some embodiments, each device-specific equalization curve 131 is constructed by benchmarking or other performance quantification tests of specific audio devices, such as headphone devices, smart speakers, in-car audio system speakers, and conventional speakers. The device-specific equalization curves 131 are also stored in a device profile database 130 and made available to the audio processing application 146 of the mobile computing device 140. Thus, according to various embodiments, when a specific audio device is detected by the audio processing application 146, the appropriate device-specific equalization curve 131 can be incorporated into a customized audio processing procedure for the audio signal by the audio processing application 146. As a result, the personalized audio experience generated from the audio signal for a particular user by the customized audio processing procedure may include compensation for the non-ideal frequency response of the audio device providing the personalized audio experience.

[0038] The mobile computing device 140 can be any mobile computing device that can be configured to implement at least one aspect of the disclosure described herein, including smartphones, electronic tablets, and laptop computers. Generally, the mobile computing device 140 can be any type of device capable of running an application program that includes instructions associated with a control application 145 and / or a voice processing application 146, but is not limited to this. In some embodiments, the mobile computing device 140 is further configured to store a local user profile database 143 which may include one or more user preference profiles 121, environmental equalization profiles 122, and / or hearing impairment compensation profiles 123. Alternatively or further, in some embodiments, the mobile computing device 140 is further configured to store audio content 144, such as a digital recording of audio content.

[0039] The control application 145 is configured to communicate between the mobile computing device 140 and the user profile database 120, the device profile database 130, and the voice environment 110. In some embodiments, the control application 145 is also configured to present the user with a (not shown) user interface to enable user sound preference tests, hearing tests, and / or setup operations for the personalized voice system 100. In some embodiments, the control application 145 is further configured to produce customized voice processing procedures for voice signals based on user-specific voice processing information and device-specific voice processing information. For example, the user-specific voice processing information may include one or more user preference profiles 121 and / or hearing impairment compensation profiles 123, and the device-specific voice processing information may include one or more environment equalization profiles 122 and / or device-specific equalization curves 131.

[0040] In some embodiments, the control application 145 generates customized speech processing procedures by generating a synthesized equalization curve 141 and / or a synthesized gain curve 142 for one or more specific listening scenarios. Generally, each specific listening scenario is a unique combination of user and listening environment 110. Therefore, for a particular user, the control application 145 is configured to generate different synthesized equalization curves 141 and / or synthesized nonlinear processing 142 for each listening environment 110 in which the user is expected to have an individualized speech experience. For example, when the user is in a specific automotive speech environment 101 (such as a specific seat in a specific make and model of vehicle), the control application 145 generates a synthesized equalization curve 141 based on some or all applicable equalization curves. In such instances, examples of applicable equalization curves include, but are not limited to, one or more user-associated and applicable user preference profiles 121, one or more environmental equalization profiles 122 applicable to the unique in-car audio environment 101 in which the user is located, one or more device-specific equalization curves 131 applicable to the unique in-car audio environment 101, and a hearing impairment compensation profile 123.

[0041] In some embodiments, the control application 145 generates a composite equalization curve 141 for a specific listening scenario by combining the behavior of all applicable equalization profiles into a single sound equalization curve. Thus, in a customized speech processing procedure performed by the speech processing application 146, the speech signal can be modified by the composite equalization curve 141 instead of being processed sequentially by multiple equalization profiles. In some embodiments, the control application 145 also generates a composite nonlinear processing 142 for a specific listening scenario by combining the behavior of all applicable nonlinear processing portions of the user preference profile 121 and / or hearing impairment compensation profile 123 into a single composite nonlinear processing 142. For example, such nonlinear processing may include, but is not limited to, one or more gain compensation operations included in the hearing impairment compensation profile 123, one or more dynamic range compression operations included in the user preference profile 121, and one or more speech limiting operations included in the user preference profile 121.

[0042] In some embodiments, when a control application 145 generates a synthetic equalization curve 141 for a particular listening scenario, the synthetic equalization curve is stored in the local user profile database 143 and / or the user profile database 120 for further use. Similarly, in such embodiments, when a control application 145 generates a synthetic nonlinear process 142 for a particular listening scenario, the synthetic nonlinear process 142 is also stored in the local user profile database 143 and / or the user profile database 120 for further use.

[0043] In some embodiments, each specific listening scenario is a unique combination of the user, the listening environment 110, and a user-selected user preference profile 121 from the user profile database 120. In such embodiments, the user-selected user preference profile 121 may be an equalization curve associated with a well-known musician or celebrity, an equalization curve associated with the user with a specific activity (e.g., playing video games, exercising, driving, etc.), or an equalization curve associated with the user with a specific category of music or playlist. Thus, in such embodiments, the control application 145 is configured to generate different composite equalization curves 141 for a specific combination of the user, the listening environment 110, and the user-selected user preference profile 121. By selecting a suitable user preference profile 121, the user can tailor a personalized audio experience to both a specific audio environment 110 and a user preference profile 121.

[0044] The audio processing application 146 is configured to generate a customized audio signal by processing the initial audio signal with a customized audio processing procedure generated by the control application 145. More specifically, the audio processing application 146 generates a customized audio signal by modifying the initial audio signal with a composite equalization curve 141, and in some embodiments, with a composite nonlinear processing 142. One such embodiment is described later in conjunction with Figure 3.

[0045] Figure 3 is a flowchart of method steps for generating a customized audio signal according to various embodiments of the present disclosure. Although the method steps are described with respect to the systems of Figures 1 and 2, those skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of various embodiments.

[0046] As shown, method 300 begins in step 301, where the voice processing application 146 identifies the current user. For example, the voice processing application 146 can determine the user's identification information based on user login, user information entered by the user, etc.

[0047] In step 302, the speech processing application 146 accesses user-specific speech information, such as one or more user preference equalization curves 151, hearing impairment compensation profiles 123, and / or other user-specific listening processing information that enables customized speech processing procedures to generate a personalized speech experience for the user. In some embodiments, the speech processing application 146 accesses some or all of the user-specific speech information in the user profile database 120. Alternatively or further, in some embodiments, the speech processing application 146 accesses some or all of the user-specific speech information in the local user profile database 143.

[0048] In step 303, the voice processing application 146 identifies the voice device(s) or(s) included in the current voice environment. For example, in some embodiments, the control application 145 determines identification information about the voice device(s) in the current voice environment 110 based on information entered by the user and communicates this identification information to the voice processing application 146. In other embodiments, the control application 145 determines or receives identification information about the voice device(s) by directly querying each voice device. For example, in one such embodiment, the control application 145 receives the Media Access Control (MAC) address and / or model number, etc., via a wireless connection with the voice device.

[0049] In step 304, the voice processing application 146 accesses device-specific voice information (such as one or more device-specific equalization curves 131) that enables customized voice processing procedures to generate a personalized voice experience for the user. In some embodiments, the voice processing application 146 accesses some or all of the device-specific voice information in the user profile database 120, and in some embodiments, the voice processing application 146 accesses some or all of the device-specific voice information in the local user profile database 143.

[0050] In step 305, the voice processing application 146 determines whether voice environment-specific voice processing information is applicable. For example, based on the identification information of the voice device(s) determined in step 303, the control application 145 may determine that the current voice environment 110 includes a voice system associated with a specific cabin or smart speaker that is associated with a specific room or other location where the user performed the environment equalization test. If not, method 300 proceeds to step 307; otherwise, method 300 proceeds to step 306.

[0051] In step 306, the voice processing application 146 accesses environment-specific voice information (such as one or more environment-specific equalization profiles 122) that enables customized voice processing procedures to generate a personalized voice experience for the user. In some embodiments, the voice processing application 146 accesses some or all of the environment-specific voice information in the user profile database 120, and in some embodiments, the voice processing application 146 accesses some or all of the environment-specific voice information in the local user profile database 143.

[0052] In step 307, the speech processing application 146 generates a customized speech processing procedure based on the speech information accessed in steps 302, 304, and 306. Specifically, the speech processing application 146 generates the customized speech processing procedure by generating a synthesized equalization curve 141 and / or synthesized nonlinear processing 142 for the current listening scenario. As described above, the current listening scenario can be based on a combination of the current user, the current listening environment 110, and, in some embodiments, a user preference profile 121 selected by the user and / or hearing impairment compensation profile 123.

[0053] In step 308, the audio processing application 146 modifies the audio signal using the customized audio processing procedure generated in step 307. In some embodiments, the audio signal is generated from audio content 144 stored locally on the mobile computing device 140. In other embodiments, the audio signal is generated from audio content received from the streaming service 104.

[0054] According to various embodiments, the modification of the audio signal by a customized audio processing procedure occurs in two stages. First, the audio signal is processed using a composite equalization curve 141 to generate a modified audio signal. Then, a gain modification operation is performed on the modified audio signal to generate a customized audio signal that produces an individualized audio experience for the user when played back in a suitable audio environment 110. It should be noted that the multiple equalization or filtering operations combined to form the composite equalization curve 141 are performed on the audio signal in a single operation, rather than sequentially. As a result, the noise level in the audio signal does not increase, which can occur when one equalization operation reduces the level in a particular frequency band and a subsequent equalization operation amplifies the level in that frequency band. Similarly, clipping can also be prevented or reduced, as it can occur when one equalization operation amplifies the level of the audio signal in a particular frequency band beyond a limit and a subsequent equalization operation reduces the level in that frequency band.

[0055] In the embodiment shown in Figure 1, a combination of a mobile computing device 140, one or more software applications running on the mobile computing device 140, and a cloud-based service delivers personalized voice experiences to various voice environments 110. In other embodiments, one or more voice devices in various voice environments communicate directly with the cloud-based service to enable personalized voice experiences in each of the various voice environments. In such embodiments, the mobile computing device can provide a user interface and / or a voice system control interface, but does not act as a processing engine for generating and / or performing customized voice processing procedures on voice signals. Instead, some or all of the customized voice processing procedures are performed in the cloud-based service, and some or all of the voice processing using the customized voice processing procedures is performed locally on the smart devices included in the voice environment. One such embodiment is described later in conjunction with Figure 4.

[0056] Figure 4 is a schematic diagram showing a personalized voice system 400 configured to implement one or more aspects of the present disclosure. The personalized voice system 400 includes, but is not limited to, one or more voice environments 410, including at least one programmable voice device 440, a user profile database 120, a device profile database 130, and a mobile computing device 440. The personalized voice system 400 is configured to provide a personalized voice experience to a particular user, regardless of which unique voice environment 410 is currently providing the user with a voice experience. The personalized voice system 400 operates similarly to the personalized voice system 100, except that a control application 445 running on the cloud infrastructure 105 generates customized voice processing procedures to modify the voice signal for playback in a particular voice environment. Furthermore, the voice signal processing using the customized voice processing procedures is performed in one or more programmable voice devices 440 associated with a unique voice environment. Therefore, the control application 445 generates a composite equalization curve similar to the composite equalization curve 141 in Figure 1, and / or a composite gain curve similar to the composite nonlinear processing 142 in Figure 1.

[0057] In some embodiments, customized speech processing procedures are implemented in the personalized speech system 400 by programming them into an internal speech processor 446 of the programmable speech device 440. In such embodiments, the speech processing associated with the customized speech processing procedures is performed by the internal speech processor 446, which can be a programmable digital signal processor (DSP) or other processor. The speech signal (e.g., from a streaming service 104 or based on audio content 144) is modified by the internal speech processor 446 using the customized speech processing procedures to generate the customized speech signal 444. The personalized speech experience is generated for the user in the speech environment 410 when a speaker 408 included in or associated with the programmable speech device 440 generates a sound output 449 based on the customized speech signal 444. Thus, in the embodiment shown in Figure 4, the speech signal (e.g., from a streaming service 104 or based on audio content 144) is processed by the internal speech processor 445 using the customized speech processing procedures, rather than by a processor outside the speech device included in the speech environment 410.

[0058] Figure 5 is a conceptual block diagram of a computing system 500 configured to implement one or more aspects of various embodiments. The computing system 500 may be any type of device capable of executing an application program including a control application 145, a voice processing application 146, and / or instructions associated with the control application 445, but is not limited to these. For example, the computing system 500 may be an electronic tablet, a smartphone, a laptop computer, an infotainment system incorporated in a vehicle, a home entertainment system, etc. Alternatively, the computing system 500 may be implemented as a standalone chip such as a microprocessor, or as part of a more comprehensive solution implemented as an application-specific integrated circuit (ASIC) and a system-on-a-chip (SoC), etc. It should be noted that the computing systems described herein are for illustrative purposes only, and any other technically feasible configurations are within the scope of the invention.

[0059] As shown, the computing system 500 includes, but is not limited to, a processor 550, an input / output (I / O) device interface 560 coupled to an input / output device 580, memory 510, storage 530, and an interconnect (bus) 540 connecting a network interface 570. The processor 550 may be any suitable processor implemented as a combination of different processing units, such as a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), any type of processing unit, or a CPU configured to operate in conjunction with a digital signal processor (DSP). For example, in some embodiments, the processor 550 includes a CPU and a DSP. Generally, the processor 550 may be any technically feasible hardware unit capable of processing data and / or executing instructions to facilitate the operation of the computing system 500 in Figure 5, as described herein. Furthermore, in the context of this disclosure, the computing elements shown in the computing device 500 may correspond to a physical computing system (e.g., a system in a data center) or to a virtual computing instance running in a computing cloud.

[0060] The input / output device 580 may include not only devices capable of providing inputs such as a keyboard, mouse, touchscreen, and microphone 581, but also devices capable of providing outputs such as a loudspeaker 582 and a display screen. The display screen may be a computer monitor, a video display screen, a display device incorporated into a handheld device, or any other technically feasible display screen. A particular instance of the loudspeaker 582 may include one or more loudspeakers that are elements of an audio system, such as the personalized audio system 100 in Figure 1 or the personalized audio system 400 in Figure 4.

[0061] The input / output device 580 may include additional devices capable of both receiving input and providing output, such as a touchscreen and a Universal Serial Bus (USB) port. Such an input / output device 580 may be configured to receive various types of input from the end user of the computing device 500 and to provide various types of output to the end user of the computing device 500, such as a displayed digital image or digital video. In some embodiments, one or more of the input / output devices 580 are configured to connect the computing device 500 to the communication network 505.

[0062] The input / output interface 560 enables communication between the input / output device 580 and the processor 550. The input / output interface typically includes the necessary logic for interpreting the address corresponding to the input / output device 580 generated by the processor 550. The input / output interface 560 may also be configured to perform a handshake between the processor 550 and the input / output device 580 and / or to generate interrupts associated with the input / output device 580. The input / output interface 560 may be implemented as any technically feasible CPU, ASIC, FPGA, or any other type of processing unit or device.

[0063] The network interface 570 is a computer hardware component that connects the processor 550 to the communication network 505. The network interface 570 may be implemented in the computing device 500 as a standalone card, processor, or other hardware device. In embodiments where the communication network 505 includes a WiFi® network or WPAN, the network interface 570 includes a suitable wireless transceiver. Alternatively or further, the network interface 570 may consist of cellular communication capabilities, satellite phone communication capabilities, wireless WAN communication capabilities, or other types of communication capabilities that enable communication with the communication network 505 and other computing devices 500 located outside the computing system 500.

[0064] The memory 510 may include a random access memory (RAM) module, a flash memory unit, or other types of memory units, or a combination thereof. The processor 550, the input / output device interface 560, and the network interface 570 are configured to read and write data to and from the memory 510. The memory 510 includes various software programs executable by the processor 550, and application data associated with the above software programs, including control application 145, voice processing application 145, and / or control application 445.

[0065] The storage 530 may include non-transient computer-readable media such as non-volatile storage devices. In some embodiments, the storage 530 includes a user profile database 120, a device profile database 130, and / or a local user profile database 143.

[0066] In short, various embodiments describe systems and technologies for providing users with device-based and / or cloud-based personalized audio experiences in various audio environments, where the personalized audio experience is optimized for a particular user's listening preferences and hearing impairments through individual sound and audio experience adjustments. In some embodiments, the customized audio processing procedure arises based on user-specific information, audio device-specific information, and environment-specific information. When the customized audio processing procedure is employed to modify the audio signal before playback, the user can have a personalized audio experience tailored to their listening preferences.

[0067] At least one technical advantage of the disclosed technology over the prior art is that it enables the personalization of the listener's voice experience regardless of the current voice environment. Specifically, the listener's personal preferences and / or hearing impairment profile can be automatically applied to any voice environment, while also understanding the voice characteristics of the voice environment, so that the listener does not need to recustomize the voice system in each voice environment. A further advantage is that the personalized voice experience can be implemented in voice environments including high-performance voice devices that perform some or all of the voice signal processing necessary to produce the personalized voice experience, or "low-performance" voice devices that do not perform any voice signal processing. These technical advantages represent one or more technical improvements over the prior art approach.

[0068] 1. In some embodiments, a method for processing an audio signal includes accessing user-specific audio processing information relating to a specific user; determining identification information for an audio device for generating sound output from an audio signal; accessing device-specific audio processing information relating to the audio device based on the identification information of the audio device; generating a customized audio processing procedure relating to the audio signal based on the user-specific audio processing information and the device-specific audio processing information; and generating a customized audio signal by processing the audio signal with the customized audio processing procedure.

[0069] 2. The method according to Clause 1, further comprising causing the audio device to generate an audio output from the customized audio signal.

[0070] 3. The method of Clause 1 or 2, wherein causing the audio device to generate an audio output from the customized audio signal includes transmitting the customized audio signal to the audio device via a wireless connection.

[0071] 4. The method according to any one of the clauses 1 to 3, wherein the processing of the audio signal with the customized audio processing procedure is performed using a processor located outside the audio device.

[0072] 5. The method according to any one of the clauses 1 to 4, wherein the processing of the audio signal with the customized audio processing procedure is performed using a processor included in the audio device.

[0073] 6. Accessing user-specific speech processing information relating to a specific user is the method of any one of the methods in Clauses 1 to 5, which includes determining the identification information of the specific user and retrieving the user-specific speech processing information from a cloud-based repository based on the identification information of the specific user.

[0074] 7. The method according to any one of the clauses 1 to 6, wherein accessing user-specific speech processing information relating to a particular user includes determining the identification information of the particular user and reading the user-specific speech processing information from a computing device configured to generate the customized speech processing procedure based on the identification information of the particular user.

[0075] 8. The method according to any one of the claims 1 to 7, wherein generating the customized speech processing procedure includes generating a synthesized speech equalization curve from information contained in at least one of the user-specific speech processing information or the device-specific speech processing information.

[0076] 9. The method according to any one of the claims 1 to 8, wherein generating the synthesized sound equalization curve includes combining all sound equalization curves included in the user-specific speech processing information or the device-specific speech processing information.

[0077] 10. The method according to any one of the claims 1 to 9, wherein generating the customized audio signal in the customized audio processing procedure includes generating a modified audio signal by modifying the audio signal with the synthesized sound equalization curve, and performing a gain modification operation indicated in at least one of the user-specific audio information or the device-specific audio information of the modified audio signal.

[0078] 11. In some embodiments, a non-transient computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform the following steps: access user-specific speech processing information relating to a particular user; determine identification information for a speech device for generating sound output from a speech signal; access device-specific speech processing information relating to the speech device based on the identification information of the speech device; generate a customized speech processing procedure relating to the speech signal based on the user-specific speech processing information and the device-specific speech processing information; and generate a customized speech signal by processing the speech signal with the customized speech processing procedure.

[0079] 12. A non-transient computer-readable medium according to Clause 11, wherein generating the customized speech processing procedure for the speech signal based on the user-specific speech processing information and the device-specific speech processing information further includes generating the customized speech processing procedure for the speech signal based on environment-specific information.

[0080] 13. The method described above further comprises determining the environment-specific information based on at least one of the identification information of the voice device and the identification information of the specific user, as described in Clause 11 or 12.

[0081] 14. Accessing user-specific speech processing information relating to the specific user includes receiving user input indicating a unique equalization profile and accessing the unique equalization profile, as described in any of Clauses 11 to 13, for any non-transient computer-readable medium.

[0082] 15. Generating the customized audio processing procedure, which includes generating the customized audio processing procedure based on the unique equalization profile, in any of the non-transient computer-readable media described in any of Clauses 11 to 14.

[0083] 16. The method described above further comprises generating the unique equalization profile based on a personalization test performed by the specific user, in a non-transient computer-readable medium as described in any of Clauses 11 to 15.

[0084] 17. Accessing user-specific speech processing information relating to a particular user includes determining the identification information of the particular user and retrieving the user-specific speech processing information from a cloud-based repository based on the identification information of the particular user, as described in any of Clauses 11 to 16, for any non-transient computer-readable medium.

[0085] 18. Accessing user-specific speech processing information relating to a particular user includes determining the identification information of the particular user and reading the user-specific speech processing information from a computing device configured to generate the customized speech processing procedure, as described in any of Clauses 11 to 17.

[0086] 19. A non-transient computer-readable medium according to any one of Clauses 11 to 18, wherein generating the customized speech processing procedure includes generating a synthesized speech equalization curve from information contained in at least one of the user-specific speech processing information or the device-specific speech processing information.

[0087] 20. In some embodiments, the system includes a memory for storing instructions and a processor coupled to the memory, and when executing the instructions, the system is configured to perform the following steps: access user-specific voice processing information relating to a particular user; determine identification information for a voice device for generating sound output from a voice signal; access device-specific voice processing information relating to the voice device based on the identification information of the voice device; generate a customized voice processing procedure relating to the voice signal based on the user-specific voice processing information and the device-specific voice processing information; and generate a customized voice signal by processing the voice signal with the customized voice processing procedure.

[0088] Any and all combinations of the elements of the claims listed in any of the “Claims” in any form and / or any elements described herein are within the conceivable scope of the present invention and protection.

[0089] The descriptions of various embodiments are presented for illustrative purposes but are not intended to be comprehensive or to limit oneself to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the embodiments described.

[0090] Aspects of this embodiment may be embodied as a system, method, or computer program product. Accordingly, aspects of this disclosure may take the form of hardware embodiments as a whole, software embodiments as a whole (including firmware, resident software, microcode, etc.), or embodiments that combine software and hardware embodiments, all of which may generally be referred to herein as “modules” or “systems.” In addition, any hardware and / or software technology, process, function, component, engine, module, or system described herein may be implemented as a circuit or set of circuits. Furthermore, aspects of this disclosure may take the form of a computer program product embodied in at least one computer-readable medium (having computer-readable program code embodied thereon).

[0091] Any combination of at least one computer-readable medium may be used. A computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples (non-exclusive list) of computer-readable storage media would include electrical connections having at least one wire, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable program read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In relation to this document, a computer-readable storage medium may be any tangible medium that can contain or store programs used by or in connection with an instruction execution system, apparatus, or device.

[0092] Aspects of the present disclosure are described above with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products in accordance with embodiments of the present disclosure. It will be understood that each block in a flowchart and / or block diagram, and combinations of blocks in a flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a dedicated computer, or other programmable data processing device that produces mechanical operation, thereby enabling the instructions executed by the processor of the computer or other programmable data processing device to perform the functions / actions specified in the blocks or combinations of blocks in the flowchart and / or block diagram. Such processor may be, but is not limited to, a general-purpose processor, a dedicated processor, an application-specific processor, or a field-programmable processor or field-programmable processor gate array.

[0093] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible embodiments of the systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, section, or portion of code, which contains at least one executable instruction for performing a specified logical function(s). It should also be noted that in some alternative embodiments, the functions noted in a block may occur in a different order than those noted in the figure. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or blocks may sometimes be executed in reverse order depending on the related functions. It should also be noted that each block in a block diagram and / or flowchart, as well as combinations of multiple blocks in a block diagram and / or flowchart, may be performed by a dedicated hardware-based system that performs a specified function or action, or a combination of dedicated hardware and dedicated computer instructions.

[0094] While the foregoing applies to embodiments of the present disclosure, other embodiments and further embodiments of the present disclosure may be conceived without departing from its basic scope, the scope of which is determined by the following "Claims".

Claims

1. A method for processing an audio signal, wherein the method is performed by a processor, The processor performs one or more tests to determine user-specific speech processing information relating to a particular user, and the performance of one or more tests to determine user-specific speech processing information is: The aforementioned hearing impairment test for the specific user, or At least one personalization test, wherein the at least one personalization test includes at least one of a personal equalization preference test, an environmental equalization test, and an HRTF test. This includes doing at least one of the following: The processor accesses device-specific audio processing information related to the audio device, The processor generates a synthesized sound equalization curve for the audio signal based on the user-specific audio processing information and the device-specific audio processing information. The processor generates nonlinear processing on the audio signal based on the user-specific audio processing information, The processor generates a customized speech processing procedure for the speech signal based on the synthesized sound equalization curve and the nonlinear processing. The processor generates a customized audio signal by processing the audio signal with the customized audio processing procedure, Methods that include...

2. The method according to claim 1, further comprising the processor causing the audio device to generate an audio output from the customized audio signal.

3. The method according to claim 1, wherein the processor performs one or more tests to determine user-specific speech processing information, the processor performs the hearing impairment test for the specific user.

4. The method according to claim 1, further comprising the processor selecting a first test from one or more tests based on the voice environment of the particular user.

5. The method according to claim 1, further comprising the processor selecting a first test from one or more tests based on the demographic information of the particular user.

6. The processor performs the hearing impairment test, The processor detects at least one hearing impairment of the specific user, The processor quantifies the at least one hearing impairment of the particular user for at least one frequency band, The method according to claim 3, further comprising:

7. The method according to claim 3, wherein the processor performing the hearing impairment test further includes the processor generating a hearing impairment compensation profile, and the user-specific speech processing information includes the hearing impairment compensation profile.

8. The method according to claim 1, wherein the processor generating the synthesized sound equalization curve includes the processor combining at least one sound equalization curve included in the user-specific speech processing information or the device-specific speech processing information.

9. The processor generates the customized audio signal using the customized audio processing procedure. The processor generates a modified audio signal by modifying the audio signal with the synthesized sound equalization curve. The processor performs a gain correction operation indicated in at least one of the user-specific audio processing information or the device-specific audio processing information of the modified audio signal. The method according to claim 1, including the method described in claim 1.

10. The method according to claim 1, wherein the processor generates the customized speech processing procedure for the speech signal based on the synthesized sound equalization curve and the nonlinear processing, further comprising the processor generating the customized speech processing procedure for the speech signal based on environment-specific information.

11. A non-transient computer-readable medium, which, when executed by a processor, the processor, Performing one or more tests to determine user-specific speech processing information relating to a particular user, and performing the one or more tests to determine user-specific speech processing information, The aforementioned hearing impairment test for the specific user, or At least one personalization test, wherein the at least one personalization test includes at least one of a personal equalization preference test, an environmental equalization test, and an HRTF test. This includes doing at least one of the following: Accessing device-specific audio processing information related to audio devices, Based on the user-specific voice processing information and the device-specific voice processing information, a synthesized sound equalization curve for the voice signal is generated. Based on the user-specific voice processing information, a nonlinear process is generated with respect to the voice signal. Based on the synthesized sound equalization curve and the nonlinear processing, a customized speech processing procedure is generated for the speech signal. By processing the aforementioned audio signal with the customized audio processing procedure, a customized audio signal is generated. A non-transient, computer-readable medium that stores instructions for performing steps such as the above.

12. The non-transient computer-readable medium according to claim 11, wherein generating the customized speech processing procedure for the speech signal based on the synthesized sound equalization curve and the nonlinear processing further comprises generating the customized speech processing procedure for the speech signal based on environment-specific information.

13. The non-transient computer-readable medium according to claim 12, wherein the step further comprises determining the environment-specific information based on at least one of the device-specific voice processing information relating to the voice device and the identification information of a particular user.

14. The non-transient computer-readable medium according to claim 11, further comprising the step of selecting a first test from one or more tests based on the voice environment of the particular user or demographic information of the particular user.

15. It is a system, Memory for storing instructions, A processor coupled to the memory, when executing the instruction, Performing one or more tests to determine user-specific speech processing information relating to a particular user, and performing the one or more tests to determine user-specific speech processing information, The aforementioned hearing impairment test for the specific user, or At least one personalization test, wherein the at least one personalization test includes at least one of a personal equalization preference test, an environmental equalization test, and an HRTF test. This includes doing at least one of the following: Accessing device-specific audio processing information related to audio devices, Based on the user-specific voice processing information and the device-specific voice processing information, a synthesized sound equalization curve for the voice signal is generated. Based on the user-specific voice processing information, a nonlinear process is generated with respect to the voice signal. Based on the synthesized sound equalization curve and the nonlinear processing, a customized speech processing procedure is generated for the speech signal. By processing the aforementioned audio signal with the customized audio processing procedure, a customized audio signal is generated. A processor configured to perform steps such as, A system that includes this.