Method and system for generating personalized audio stream for user

The method and system address the limitations of existing audio codecs by generating personalized audio streams based on user-specific audiograms and real-time factors, enhancing audio quality and 3D experience.

WO2026063693A1PCT designated stage Publication Date: 2026-03-26SAMSUNG ELECTRONICS CO LTD
5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing audio codecs fail to provide personalized audio experiences due to varying human hearing sensitivities, leading to unpleasant audio perception, loss of 3D spatial experience, and issues like noise artifacts, limited dynamic range, and high-bitrate dependency, without effectively utilizing user hearing profiles.

Method used

A method and system that generates a personalized audio stream by transmitting audio tones, determining user-specific audiograms, and applying audio masking parameters based on real-time physiological and environmental factors to create a customized audio experience.

Benefits of technology

Enhances user audio experience by dynamically adapting audio compression to individual hearing levels, improving 3D audio perception, and reducing noise artifacts, while optimizing bitrates for better audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025014435_26032026_PF_FP_ABST
    Figure KR2025014435_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method and a system for generating a personalized audio stream for a user. The method comprises: (a) transmitting, one or more audio tones to one or more audio devices corresponding to each ear of the user; (b) receiving, from the user, response based on transmitted audio tones; (c) generating an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user; (d) determining audio masking parameter(s) based at least on the audiogram; (e) receiving an input audio stream; and (f) generating the personalized audio stream for the user by applying the audio masking parameter(s) to the received input audio stream.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR GENERATING PERSONALIZED AUDIO STREAM FOR USER

[0001] Embodiments of the present disclosure generally relate to audio codec technologies and devices. More particularly, embodiments of the present disclosure relate to a method and a system for generating a personalized audio stream for a user.

[0002] The following description of the related art is intended to provide background information pertaining to the field of the disclosure. This section may include certain aspects of the art that may be related to various features of the present disclosure. However, it should be appreciated that this section is used only to enhance the understanding of the reader with respect to the present disclosure, and not as admissions of the prior art.

[0003] Codec is a key element when it comes to the creation and usage of media. A Codec allows the ingesting of sound and video into electronic form. Audio codecs are essential tools in the digital audio world, functioning as the technology that allows hardware or software to encode and then decode audio. That is, audio codecs are responsible for compressing audio files during encoding and decompressing them during playback. Also, they play an important role in determining the quality and size of audio files. Exemplarily summarizing, audio codecs are used for the process of encoding source audio captured by a microphone in digital form for transmission to other participants in calls, audio streaming, audio broadcasts as well as shrinking media files for users to store or send over the internet. The operation of an audio codec is based on encoding an input audio stream by the audio codec and providing a compressed file, then decoding the compressed file and providing for a media player to play the audio file.

[0004] Further, the sensitivity of the human ear is not the same for all frequencies. The sensitivity of human ear varies depending on the frequency it hears. For example, the sensitivity may be higher for certain range of frequencies, say, 500 Hz to 5 kHz, and may be lesser for other frequencies. Also, there is a loudness threshold level for the human ear, below which the human ear may not be able to hear the sound. For example, an audio tone higher than 5dB may be louder than an audio tone of 1dB, before the human ear can perceive it. Accordingly, a human ear may have a higher sensitivity for a specific frequency of sound and may hear such frequency of the sound with higher sensitivity which may appear as being louder than other frequencies. Also, different users may have different levels of hearing sensitivity. Due to this varied sensitivity, different humans may perceive different loudness for the same frequency, for which there exists a need for a solution which addresses an impact on the audio experience due the varied sensitivity of different ears of different humans and help improve user experience.

[0005] Thus, to compensate for the varying sensitivity of human ears for various frequencies, a compression technology, that is generally known as ‘perceptual audio coding’, was developed. ‘Perceptual’ means relating to the way people interpret and understand what they hear, see or notice. Perceptual encoding is a compression technique where the output audio bitstream (that is the audio bitstream that is fed to the human ears after compression and decompression), even though sounding exactly as the input audio bitstream (or close to the original), is not a replica (i.e., exact replica) of the input audio bitstream. This is based on a generalized audiogram of a normal human ear (where, an audiogram is a graph that shows the audible threshold for standardized frequencies as measured by an audiometer, and shows the pattern of the hearing loss of a human ear). Also, a normal human ear as used above refers to a human ear that does not suffer from any deficiency related to hearing, and functions properly as per the medical standards.

[0006] However, each ear of each person may have varying degrees of sensitivity to each audio frequency or range of frequencies. There is generally a level of disparity introduced in the hearing level for both of the ears for different users. Thus, a certain frequency may make one person feel happy, but it may not be healthy for others. Also, audio content being the same in both the ears often makes it unpleasant for a user to listen to, and decreases user experience. Accordingly, there remains a need for a disparity in the interaural level difference (ILD) between the input to both the ears to enable better 3D audio perception. Further, there are multiple other problems associated with the currently available solutions related to audio codecs, for example, limited dynamic range (due to encoding and decoding process, codecs sometimes fail to preserve the full range of sound, resulting in loss of detail, stereo separation and the peak audio clarity), common noise issue (certain noise artifacts may be generated after decoding encoded audio data packets, this generally happens due to compatibility issues or high-resource dependency of the codec), high-bit rate dependency (achieving high-quality audio often requires high-bitrates, which is sometimes impractical for streaming or storage purpose), audio drift (drift refers to synchronization issues that may occur over time, causing audio to become out of sync of other elements such as video). Furthermore, there are multiple imperfections in the human ear as they are susceptible to different hearing parameters (loudness, frequency) differently. Also, due to the noises present in the environment of the user, the ear of the user might not hear certain useful things properly.

[0007] Also, the sensitivity of human ears may vary due to various reasons such as health issues such as high and low blood pressure, diabetes, etc. Along with this, there may be other issues such as, but not limited to, compression artifacts being caused at lower bitrates for audio file which can cause significant harm to the audio experience of the user, loss of high frequency details and transient sounds. And improper judgement of human hearing levels leading to decreased user audio experience. Currently used audio codecs do not take advantage of a possibility of codec's expanded compression and usability, which reduces user experience with audio. The absence of some frequencies that the user could have heard, and which were essential to improve the auditory experience is limited in current solutions. Since the hearing levels of each users vary, with a generic audio codec, the user may find the amplitude reaching to them is unpleasant. Also, the difference in hearing level of both the ears may lead to loss of 3-dimensional (3D) spatial and localization experience.

[0008] One of the known solutions developed for addressing the issues related to audio codecs talks about modifying an audio signal using custom psychoacoustic methods, for encoding the audio signal. In the known solution, masking and hearing thresholds are obtained from the user’s hearing profile and applied to the frequency components of the audio sample. In this solution, the user’s imperceptible audio signal data is then disregarded and transformed audio sample is encoded. However, the existing solutions are unable to provide enhanced personalized compression based on hearing level / ability of the user. Also, the existing solutions are unable to provide enhanced 3D audio experience for the user based on audio streaming and hearing response characteristics. Furthermore, the currently used highly universal audio codecs for audio reduce user experience with audio due to limited compression and usability and do not completely utilize user hearing profile. The final audio reaching the users includes fixed amplitude or frequency that make the audio unpleasant to users based on hearing response / level. Thus, there exists a need for a technical solution that can overcome at least the above-mentioned technical limitations of the existing solutions. That is, there is a need of dynamically customizing audio codec based on hearing level and making it unique for each user. More specifically there is a need in the art to provide a method and system for generating a personalized audio stream for a user.

[0009] In an embodiment of the disclosure, a method for generating a personalized audio stream for a user is provided. The method may include transmitting one or more audio tones to one or more audio devices corresponding to each ear of the user. The method may include receiving, from the user, a response based on the transmitted one or more audio tones. The method may include generating an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user. The method may include determining one or more audio masking parameters based at least on the audiogram. The method may include receiving an input audio stream. The method may include generating the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream.

[0010] In an embodiment of the disclosure, the method may further comprise generating the one or more audio tones at a plurality of loudness levels for a frequency among a pre-defined set of frequencies

[0011] In an embodiment of the disclosure, the response may be one of an audible response and an inaudible response.

[0012] In an embodiment of the disclosure, the threshold loudness level may indicate a minimum loudness level from the set of loudness levels at which the one or more audio tones are audible at said each ear.

[0013] In an embodiment of the disclosure, generating the audiogram may include normalizing the audiogram based on at least one of a set of physiological parameters of the user and a set of environmental parameters related to the user.

[0014] In an embodiment of the disclosure, generating the personalized audio stream for the user may include: (a) receiving the input audio stream at the one or more audio devices, (b) determining an amplification factor for the one or more audio devices based on the audiogram, wherein the amplification factor is related to the received input audio stream, and (c) applying the one or more audio masking parameters based on at least one of a set of real-time physiological parameters of the user, a set of real-time environmental parameters related to the user, the determined amplification factor or a threshold loudness level corresponding to one or more frequencies, wherein the one or more audio masking parameters may include a set of parameters related to a current auditory profile information of the user.

[0015] In an embodiment of the disclosure, the amplification factor may be determined based on a difference between a loudness level of the input audio stream received at a particular frequency and a threshold loudness level corresponding to the particular frequency.

[0016] In an embodiment of the disclosure, the input audio stream may include one or more sub-streams, and the personalized audio stream may be generated based on an implementation of the one or more audio masking parameters on the one or more sub-streams.

[0017] In an embodiment of the disclosure, the method may include: (a) identifying at least one of one or more prominent streams and one or more non-prominent streams from the one or more sub-streams, and (b) applying the one or more audio masking parameters to at least the one or more prominent streams.

[0018] In an embodiment of the disclosure, the one or more sub-streams may be identified as at least one of the one or more prominent streams and the one or more non-prominent streams based on one or more parameters of the received input audio stream.

[0019] In an embodiment of the disclosure, the personalized audio stream may be generated based on at least one of an audiogram shifting technique, a frequency shifting technique, a bitwise modification technique, or a spatial audio interaural level difference (ILD) modification technique to provide a three-dimensional (3D) audio effect.

[0020] In an embodiment of the disclosure, the method may include generating a compressed audio file for the personalized audio stream based on a bitwise encoding technique.

[0021] In an embodiment of the disclosure, the one or more audio tones may be transmitted to one or more audio devices at one or more sets of loudness levels and one or more frequencies, and wherein each of the one or more sets of loudness levels may correspond to a frequency among the one or more frequencies.

[0022] In an embodiment of the disclosure, the threshold loudness level corresponding to said each ear of the user may be determined at each frequency among the one or more frequencies based on the response.

[0023] In an embodiment of the disclosure, a system for generating a personalized audio stream for a user is provided. The system may include at least one processing unit, and at least one memory unit connected to the at least one processing unit. The at least one processing unit may be configured to transmit one or more audio tones to one or more audio devices corresponding to each ear of the user. The at least one processing unit may be configured to receive, from the user, a response based on the transmitted one or more audio tones. The at least one processing unit may be configured to generate an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user. The at least one processing unit may be configured to determine one or more audio masking parameters based at least on the audiogram. The at least one processing unit may be configured to receive an input audio stream. The at least one processing unit may be configured to generate the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream.

[0024] In an embodiment of the disclosure, a non-transitory computer readable storage medium storing instructions for generating a personalized audio stream for a user is provided. The instructions may include executable code which, when executed by one or more units of a system, causes a processing unit of the system to transmit one or more audio tones to one or more audio devices corresponding to each ear of the user. The executable code, when executed, may cause the processing unit to receive, from the user, a response based on the transmitted one or more audio tones. The executable code, when executed, may cause the processing unit to generate an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user. The executable code, when executed, may cause the processing unit to determine one or more audio masking parameters based at least on the audiogram. The executable code, when executed, may cause the processing unit to receive an input audio stream. The executable code, when executed, may cause the processing unit to generate the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream to generate the personalized audio stream for the user.

[0025] The accompanying drawings, which are incorporated herein, and constitute a part of this disclosure, illustrate exemplary embodiments of the disclosed methods and systems in which like reference numerals refer to the same parts throughout the different drawings. Components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present disclosure. Also, the embodiments shown in the figures are not to be construed as limiting the disclosure, but the possible variants of the method and system according to the disclosure are illustrated herein to highlight the advantages of the disclosure. It will be appreciated by those skilled in the art that disclosure of such drawings includes disclosure of electrical components or circuitry commonly used to implement such components.

[0026] Fig. 1 illustrates a block diagram of a system for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure.

[0027] Fig. 2 illustrates a block diagram of a system comprising some exemplary modules / units for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure.

[0028] Fig. 3 illustrates a flowchart of a method for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure.

[0029] Fig. 4 illustrates a flowchart of a hearing test, in accordance with an embodiment of the disclosure.

[0030] Fig. 5 illustrates a process of determining a threshold loudness level for each frequency in accordance with an embodiment of the disclosure.

[0031] Fig. 6 illustrates a block diagram of a compression module, in accordance with an embodiment of the disclosure.

[0032] FIG. 7 illustrates a method for generating personalized audio stream for the user, in accordance with an embodiment of the disclosure.

[0033] FIG. 8 illustrates a method for generating personalized audio stream for the user, in accordance with an embodiment of the disclosure.

[0034] The foregoing shall be more apparent from the following more detailed description of the disclosure.

[0035] In the following description, for the purposes of explanation, various specific details are set forth in order to provide a thorough understanding of embodiments of the present disclosure. It will be apparent, however, that embodiments of the present disclosure may be practiced without these specific details. Several features described hereafter may each be used independently of one another or with any combination of other features. An individual feature may not address any of the problems discussed above or might address only some of the problems discussed above.

[0036] The ensuing description provides exemplary embodiments only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing an exemplary embodiment. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the disclosure as set forth.

[0037] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, it will be understood by one of ordinary skill in the art that the embodiments may be practiced without these specific details. For example, circuits, systems, processes, and other components may be shown as components in block diagram form in order not to obscure the embodiments in unnecessary detail.

[0038] Also, it is noted that individual embodiments may be described as a process which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be re-arranged. A process is terminated when its operations are completed but could have additional steps not included in a figure.

[0039] The word “exemplary” and / or “demonstrative” is used herein to mean serving as an example, instance, or illustration. For the avoidance of doubt, the subject matter disclosed herein is not limited by such examples. In addition, any aspect or design described herein as “exemplary” and / or “demonstrative” is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor is it meant to preclude equivalent exemplary structures and techniques known to those of ordinary skill in the art. Furthermore, to the extent that the terms “includes,” “has,” “contains,” and other similar words are used in either the detailed description or the claims, such terms are intended to be inclusive-in a manner similar to the term “comprising” as an open transition word-without precluding any additional or other elements.

[0040] As used herein, “a user equipment”, “a user device”, “a smart-user-device”, “a smart-device”, “an electronic device”, “an audio output device” may be any electrical, electronic and / or computing device or equipment, capable of implementing the features of the present disclosure. The user equipment / device may include, but is not limited to, a wearable device or any other computing device which is capable of implementing the features of the present disclosure. Also, the user device may contain at least one input means configured to receive an input from unit(s) which are required to implement the features of the present disclosure.

[0041] As used herein, “storage unit” or “memory unit” refers to a machine or computer-readable medium including any mechanism for storing information in a form readable by a computer or similar machine. For example, a computer-readable medium includes read-only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory devices or other types of machine-accessible storage media. The storage unit stores at least the data that may be required by one or more units of the system to perform their respective functions.

[0042] All modules, units, components used herein, unless explicitly excluded herein, may be software modules or hardware processors, the processors being a general-purpose processor, a special purpose processor, a conventional processor, a digital signal processor (DSP), a plurality of microprocessors, one or more microprocessors in association with a DSP core, a controller, a microcontroller, Application Specific Integrated Circuits (ASIC), Field Programmable Gate Array circuits (FPGA), any other type of integrated circuits, etc. The processor may perform signal coding data processing, input / output processing, and / or any other functionality that enables the working of the system according to the present disclosure. More specifically, the processor or processing unit is a hardware processor.

[0043] In an embodiment of the disclosure, a method and a system for generating the personalized audio stream for a user based on a generated audiogram of the user is provided.

[0044] In an embodiment of the disclosure, a method and a system for normalization of a generated audiogram of a user based on the current blood pressure of the user and environment noise condition around the user is provided.

[0045] In an embodiment of the disclosure, a method and a system that is able to determine a masking equation for the user which contains auditory profile information of the user is provided.

[0046] In an embodiment of the disclosure, a method and a system that is able to perform spatial audio interaural level difference (ILD) modification for 3D spatialization and localization is provided.

[0047] As discussed in the background art section, existing technologies related to generating an audio stream for a user have many limitations, such as the audiogram for a user is not normalized based on the user’s fluctuating blood pressure or the environment in which the user is present. Also, in the currently known solutions, the difficulty for localization of sound in 3D spatial audio environment of the user is also ignored.

[0048] In order to overcome at least some of the limitations of the prior known solutions which may arise due to various reasons including inefficiency due to varying hearing loudness threshold levels for the users for varying frequencies, along with other limitations, the solution comprises producing multiple tones for different set of frequencies at different loudness levels and receiving hearing response from a user. Further the solution comprises selecting threshold loudness level for each of the frequency set based on current user parameters (such as, health parameters) and user environment.

[0049] Further, an audiogram of loudness thresholds corresponding to different frequency levels of the user is generated and is normalized based on the current user parameters and environment noise condition around the user. Based on generated audiogram, a masking equation is determined which contains current auditory profile information of the user. This masking equation is implemented over audio input signal to generate a perceptualized (masked) signal, and a personalized compressed file is created by sending the generated perceptualized signal to an encoder.

[0050] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily carry out the present disclosure.

[0051] Fig. 1 illustrates a block diagram of a system for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure.

[0052] Referring to Fig. 1, an exemplary block diagram of a system 100 for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure is shown. The system 100 may comprise at least one processing unit 102, and at least one memory unit 104 (or as used herein, storage unit 104). Also, all of the components / units of the system 100 may be assumed to be connected to each other unless indicated below.

[0053] As shown in Fig. 1, all units shown within the system 100 should also be assumed to be connected to each other. Also, in Fig. 1 only a few units are shown, however, the system 100 may comprise multiple such units, or the system 100 may comprise any such numbers of said units, as required to implement the features of the present disclosure. Further, in an implementation, the system 100 may be connected to or may reside in a user device which may be an audio output device, such as an earphone or earbuds.

[0054] In an embodiment of the disclosure, the processing unit 102 may be configured to transmit one or more audio tones to one or more audio devices corresponding to each ear of the user. The one or more audio tones may be transmitted to one or more audio devices at one or more sets of loudness levels and one or more frequencies.

[0055] Also, each of the one or more sets of loudness levels may correspond to a frequency among the one or more frequencies. Therefore, in an implementation the processing unit 102 may be configured to transmit, at a set of loudness levels, to the one or more audio devices corresponding to each ear of the user, the one or more audio tones related to the one or more frequencies. This means that one or more audio tones at a plurality of sets of loudness levels, each set of loudness levels corresponding to a frequency among a set of frequencies, may be transmitted to one or more audio devices corresponding to each ear of the user (that is for example, a left earbud and right earbud).

[0056] In an embodiment of the disclosure, the processing unit 102 may be configured to receive, from the user, a response based on the transmitted one or more audio tones. In an implementation, the response may be one of an audible response and an inaudible response. Further, the processing unit 102 may be configured to determine a threshold loudness level corresponding to said each ear of the user at each frequency from the one or more frequencies based on the response.

[0057] In an embodiment of the disclosure, a minimum level of loudness at which a user’s ear (i.e., left ear or right ear) can hear a sound at a particular frequency (i.e., the minimum level of loudness at which a user responds as ‘audible’ sound) may be referred to and determined as a threshold loudness level corresponding to that frequency for the user’s ear. In other words, the threshold loudness level may indicate a minimum loudness level from the set of loudness levels at which the one or more audio tones are audible at said each ear.

[0058] In an embodiment of the disclosure, the processing unit 102 may be configured to generate an audiogram for the user based at least on the determined threshold loudness level. In an implementation, generating the audiogram may comprise normalizing the audiogram based on at least one of a set of physiological parameters of the user (such as, but not limited to, blood pressure levels of the user) or a set of environmental parameters related to the user (such as, but not limited to, environmental noise around the user).

[0059] In an embodiment of the disclosure, the processing unit 102 may be configured to determine one or more audio masking parameters based at least on the audiogram. After determining one or more audio masking parameters, the processing unit 102 may be configured to receive an input audio stream.

[0060] In an embodiment of the disclosure, the processing unit 102 may be configured to implement the one or more audio masking parameters on the received input audio stream to generate the personalized audio stream for the user. The processing unit 102 may be configured to generate the personalized audio stream for the user by applying one or more audio masking parameters to the received input audio stream.

[0061] In an embodiment of the disclosure, the personalized audio stream for the user, may be generated as a compressed audio file based on a bitwise encoding technique. For implementing the one or more audio masking parameters, the processing unit 102 may be configured to receive the input audio stream at the one or more audio devices. In an embodiment of the disclosure, the processing unit 102 may be configured to determine an amplification factor for the one or more audio devices based on the audiogram. The amplification factor may be related to the received input audio stream.

[0062] In an embodiment of the disclosure, the processing unit 102 may be configured to determine the one or more audio masking parameters based on at least one of a set of real-time physiological parameters of the user, a set of real-time environmental parameters related to the user, the determined amplification factor or a threshold loudness level of the one or more frequencies. The one or more audio masking parameters may comprise a set of parameters related to a current auditory profile information of the user. The amplification factor may be based on a difference between a loudness level of the input audio stream received at a particular frequency and a threshold loudness level corresponding to the particular frequency.

[0063] In an embodiment of the disclosure, the input audio stream may comprise one or more sub-streams, and the personalized audio stream may be generated based on an implementation of the one or more audio masking parameters on the one or more sub-streams. To generate the personalized audio stream, the processing unit 102 may be configured to identify at least one of one or more prominent streams or one or more non-prominent streams from the one or more sub-streams, and apply the one or more audio masking parameters to at least the one or more prominent streams. These one or more sub-streams may be identified as at least one of the one or more prominent streams, and the one or more non-prominent streams may be determined based on one or more parameters of the received input audio stream.

[0064] In an embodiment of the disclosure, the personalized audio stream may be generated based on at least one of an audiogram shifting technique, a frequency shifting technique, a bitwise modification technique, or a spatial audio interaural level difference (ILD) modification technique to provide a three-dimensional (3D) audio effect.

[0065] Fig. 2 illustrates a block diagram of a system comprising some exemplary modules / units for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure.

[0066] Referring to Fig. 2, a block diagram of the system 200 may comprise some exemplary modules for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure. The system 200 may comprise the exemplary modules / units / engines to implement one or more features of the disclosure. These exemplary modules as shown in Fig. 2, in an implementation, may be implemented by the processing unit 102 of the system 100 as shown in Fig. 1. As shown in Fig. 2, the system 200 may comprise an auditory test module (ATM) 202, a hearing profile generation module (HPGM) 204, a prominent stream determination module (PSDM) 206, an audio stream modification module (ASMM) 208, and a compression module (CM) 210. Each of these units / modules / engines are explained in detail in the forthcoming description.

[0067] Different users have different hearing ability. Also, different users have different listening experience hearing the same content. Thus, a hearing test helps to understand the hearing level of each ear of user quantitively.

[0068] In an embodiment of the disclosure, the auditory test module (ATM) 202 of the system 200 may conduct the hearing test. In an example, the one or more audio devices may comprise a left side earbud and a right side earbud of a pair of earbuds in possession of the user. The hearing test may be conducted by plugging both the earbuds and feeding the input to a device, such as, but not limited to, a smartphone where the processing takes place. Details related to the procedures followed by the ATM 202 for conducting the hearing test are disclosed with reference to Fig. 4 in this disclosure. The ATM 202 may be configured to transmit one or more audio tones to one or more audio devices corresponding to each ear of the user, and receive, from the user, a response based on the transmitted one or more audio tones.

[0069] Also, environmental noise effects the hearing test results conducted on the user. The audiogram generated needs to be modified in accordance with the environment in which the user is present during the hearing test. Here, noise may be construed as a low-level hiss or buzz that intrudes on audio output. Detection of environment noise in real time during a hearing test conducted with audio output device (e.g., earbuds) may involve the use of a built-in microphones present in the audio output device. Thus, a real-time frequency analysis may be performed, and the characteristics of room noise may be identified and isolated. This is done in a noisy environment which typically differs from the test signals used in hearing tests.

[0070] For this purpose, the audio environment is continuously sampled using the microphones to detect any sound while the hearing test is going on. In case the user is not able to hear the test tones properly, due to presence of disturbing sounds in the background, the test result is not ideal and an audiogram shifting may be performed for normalization of the effect of environmental noise on hearing test.

[0071] Further, Blood Pressure detection is essential for the accuracy of auditory test result and several other conditions, as the Blood Pressure fluctuations can affect the person’s auditory sensitivity and response during the test. Blood Pressure may be continuously monitored with the help of a wearable device (such as, but not limited to, a smartwatch) that will synchronize the data throughout the course of the test. The auditory result may be analyzed and co-related with the user’s blood pressure to get insights for the user’s hearing ability and sensitivity. This is performed using the equations 1, 2, 3, and 4 as given below.

[0072] [Equation 1]

[0073]

[0074] Normal Blood Flow may be derived using Equation 1.

[0075] [Equation 2]

[0076]

[0077] User’s Current blood flow (High / Low) may be derived using Equation 2.

[0078] [Equation 3]

[0079]

[0080] Then, blood flow may be derived using Equation 3.

[0081] [Equation 4]

[0082]

[0083] Hearing deviated threshold may be derived using Equation 4.

[0084] In an embodiment of the disclosure, the Pnormalmay refer to a blood flow of the user in a normal condition during which the user may not face any disturbances during the generation of the audiogram of the user. The Pcurrentmay refer to a blood flow of the user during the period for which the user is listening or consuming the content. The variable R may refer to a factor of change which may be proportional to the change in the blood pressure during the normal condition and the consumption of the content.

[0085] In an embodiment of the disclosure, the hearing profile generation module (HPGM) 204 may be configured to generate an audiogram for the user based on a threshold loudness level corresponding to said each ear of the user. An audiogram may be a graph of hearing test results that shows a person’s hearing thresholds for different frequencies and volumes. The test may be conducted on both of the ears individually as both ears may have the different hearing capabilities for an individual. The results of the left and the right ear tests are plotted on the graph of loudness vs. frequency. This may be used for the purpose of generating the hearing profile of the individual user.

[0086] In an embodiment of the disclosure, the array of pairs of loudness levels may be stored in the database in the memory unit 104 containing the details of the hearing threshold at different frequencies for both of the ears. The array may contain sets of (L,R), that are the loudness threshold for different frequencies for left and right ears respectively. For each set of the pair (L,R), for different frequencies, a graph is plotted and the plots for each of the frequencies is connected to generate the audiogram.

[0087] In an embodiment of the disclosure, the generated audiogram may be normalized based on the current environment of the user. For example, the original audiogram may be obtained due to the hearing test conducted. Thereafter, the audiogram normalized due to the environmental parameters such as environment noise, and current Blood Pressure of the user may be obtained as normalized audiogram.

[0088] For this purpose, a normalized threshold for each frequency may be obtained and finally the obtained threshold corresponding to each frequency may be plotted. Say, the noise level detected is ‘N’, the estimated abnormal Blood pressure is ‘BP’, and the threshold loudness level at frequency (f) obtained from audiogram is ‘L’. In this case, the equation 5 may be used as given below.

[0089] [Equation 5]

[0090]

[0091] where a = blood pressure coefficient, and

[0092] b = noise level coefficient

[0093] Here, a = (highest Blood pressure - threshold blood pressure) / 120

[0094] b = (Maximum Noise - Minimum noise) / (Maximum Noise)

[0095] The threshold blood pressure may refer to a blood pressure of the user which may be determined during the hearing test of the user.

[0096] In an embodiment of the disclosure, the HPGM 204 may be configured to identify the audible range and non-audible range for a particular user. The audible range of the particular user may be the audio range where the tone coming to the user is audible properly. The non-audible range of the particular user may be the audio range where the tone coming to the user is not audible. In an example, the audiogram equation of a user may be given in equation 6 below.

[0097] [Equation 6]

[0098]

[0099] where

[0100] Lf: threshold loudness at frequency 'f'

[0101] an: coefficient

[0102] f : frequency at which the hearing threshold has to be determined

[0103] Further, from above, after putting the values of f and L in equation (6),

[0104] if (Lf- anfn+ an-1fn-1+ an-2fn-2+ an-3fn-3+ ... + af) > 0

[0105] then, the audio tone lies in the audible range, and

[0106] if (Lf - anfn+ an-1fn-1+ an-2fn-2+ an-3fn-3+ ... + af) < 0

[0107] then, the audio tone lies in the inaudible range.

[0108] Since different amplitude needs to be boosted differently in order to enable the user to differently perceive various frequencies, the HPGM 204 may be configured to generate a masking equation for the purpose of masking the audio stream based on the audiogram of the user. For this purpose, the HPGM 204 may determine one or more audio masking parameters based at least on the audiogram. The masking equation may be a function of 4 audio masking parameters, namely, calculated blood pressure factor (denoted by “αA”), calculated noise factor (denoted by “βA”), calculated amplification factor, and threshold loudness level at particular frequency (denoted by “γA”), as given below in equation 7.

[0109] [Equation 7]

[0110]

[0111] The calculated blood pressure factor may depend on a change of the blood pressure during the usage condition as related to hearing test condition. Similarly, the calculated noise factor may be dependent on the change of the blood pressure.

[0112] In an embodiment of the disclosure, for calculating the amplification factor, a window of an audio tone may be taken. The minimum amplitude (amin) and maximum amplitude (amax) in the audio tone may be determined. The difference (Diff) of current loudness to threshold of the particular tone may be calculated. Here, Diff = (loudness threshold of audiogram - amplitude from the audio source). The amplification factor may be calculated based on a difference between a loudness level of the input audio stream received at a particular frequency and a threshold loudness level corresponding to the particular frequency, using equation 8 as given below.

[0113] [Equation 8]

[0114]

[0115] In an implementation, corresponding to the selected frequency, iteration of frequency is done to the left and right until the slope of the audiogram is similar.

[0116] In an embodiment of the disclosure, the input audio stream may comprise one or more sub-streams, and the personalized audio stream may be generated based on an implementation of the one or more audio masking parameters on the one or more sub-streams. In an example, implementing the one or more audio masking parameters on the received input audio stream to generate the personalized audio stream for the user may comprise: (a) receiving the input audio stream at the one or more audio devices, (b) determining an amplification factor for the one or more audio devices based on the audiogram, wherein the amplification factor is related to the received input audio stream, and (c) implementing the one or more audio masking parameters based on at least one of a set of real-time physiological parameters of the user, a set of real-time environmental parameters related to the user, the determined amplification factor or a threshold loudness level corresponding to the one or more frequencies. The one or more audio masking parameters may comprise a set of parameters related to a current auditory profile information of the user.

[0117] In an embodiment of the disclosure, the one or more sub-streams may be identified as at least one of the one or more prominent streams or the one or more non-prominent streams based on one or more parameters of the received input audio stream. The prominent stream determination module (PSDM) 206 may be configured to determine the prominent and the non-prominent streams present in input audio file. That is, to generate the personalized audio stream, at least one of one or more prominent streams or one or more non-prominent streams from the one or more sub-streams may be identified, and the one or more audio masking parameters may be applied to at least the one or more prominent streams. Here, a prominent stream may refer to a significant or dominant component within an audio signal, i.e., a stream that carries essential information or the main focus of the audio content. Also here, a non-prominent stream may refer to a sub-stream of an audio file that does not contain the primary or main elements of the audio content. These streams may be generally referred to as background noise that contain distracting elements and deflect attention of listener from the prominent streams.

[0118] In an embodiment of the disclosure, by eliminating the non-prominent or distracting sub-streams, overall space of the file can be reduced, and the clarity of audio track can be increased. The procedure for eliminating the non-prominent sub-streams may comprise: (a) determination of multiple sub-streams from an input audio file, (b) classification of sub-streams as one of a background sub-stream and foreground sub-stream, (c) scoring of the sub-streams, (d) classification of sub-streams as one of a prominent sub-stream and non-prominent sub-stream, (e) eliminating the non-prominent sub-streams, and (f) outputting the prominent sub-streams. It would be pertinent to mention that the determination / classification of sub-streams into background and foreground sub-streams is based on features of audio that are extracted using techniques such as, but not limited to, Fast Fourier Transform (FFT) and Wavelet, to detect the frequency information of the signal.

[0119] The separated audio files may be then labeled according to the pitch shifts, variation in tone, change in loudness, zero crossing rate (frequency of change) and relevance to the surrounding. For example, an input audio file contains an Audio Recording of a person singing a room, with surrounding people clapping and cheering. In this, the singer’s voice may be recognized by clear, consistent pitch and melody with a unique texture. The loudness may vary depending on the performance. Based on the extracted features in the input audio file, sub-streams may be identified.

[0120] For example, sub-stream 1 contains voice of the singer, sub-stream 2 contains audience clapping, sub-stream 3 contains audience talking, sub-stream 4 contains background static, sub-stream 5 contains other disturbing elements such as chair moving sound, step sounds, etc. The sub-streams may then be classified as foreground or background by the PSDM 206 based on their loudness level as table 1 below.

[0121]

[0122] In an embodiment of the disclosure, the PSDM 206 may analyse the identified sub-streams and generates a priority score for each sub-stream, based on the user’s personalized audiogram. Here, each audio sub-stream may be assigned a personalized score based on how well it matches the user’s profile. The score may be derived from how closely the sub-stream aligns with the user’s hearing abilities, preferences and relevance of the sub-stream to the original audio stream. The PSDM 206 may classify the scored sub-streams as a prominent sub-stream or a non-prominent sub-stream based on the score of the sub-stream. This classification may be based on a comparison with a pre-defined threshold score. The PSDM 206 may provide the prominent sub-streams as an output audio file which is of a reduced size as compared to the input audio file.

[0123] In an embodiment of the disclosure, the audio stream modification module (ASMM) 208 may be configured to shift the normalized audiogram based on current blood pressure and noise level detected from the microphone of the audio output device, such as, earbuds. This shift in the audiogram (that is, the audiogram shifting technique) may be based on the input of various factor values (such as current Blood Pressure factor, current noise factor, etc.) as noted in the masking equation, i.e., equation 7 as noted above in this disclosure.

[0124] In an embodiment of the disclosure, the ASMM 208 may be configured to modify the audio based on a frequency shifting technique, a bitwise modification technique, and a spatial audio interaural level difference (ILD) modification technique to provide a three-dimensional (3D) audio effect to generate the personalized audio stream.

[0125] Thus, for modifying the audio based on frequency analysis, the audio at a particular timestamp may be taken into consideration. For application of the compression technique, that is, for modifying the audio (sub-streams) based on frequency analysis, the number of bits for the louder stream can be reduced, or the number of bits for the less loud streams can be reduced or suppressed completely. This may be done in case the frequency of other sub-streams at same timestamps lies in a specific range (f-f1) to (f+f1) and it also lies in the auditory range. Here, f1 may be determined using equation 9 below.

[0126] [Equation 9]

[0127]

[0128] Referring to equation 9, α is the tunable parameter which can be tuned for the amount of compression required. The masking factor may refer to a ratio of a range of frequency for a similar slope of the audiogram and a total range of frequency.

[0129] Also, the number of final bits for louder signal may be determined using equation 10 below.

[0130] [Equation 10]

[0131]

[0132] In an embodiment of the disclosure, the ASMM 208 may be configured to modify the audio based on local frequency tone. For modifying the audio based on local frequency tone, bitwise modification technique may be applied. The audio stream may be taken as input and the corresponding frequency and loudness may be checked. Further, the ASMM 208 may check if the tone lies in the audible or inaudible range. If the tone lies in the audible range, then bit compression technique is applied, which involves reducing the number of bits. For example, the final number of bits (FNB) may be produced as: FNB = b / masking ratio, where ‘b’ is the initial number of bits in the audio stream. Also, the masking ratio may be given as (range of frequency for similar slope / total range of frequency). Also, if the tone lies in the audible range, then a similar input stream may be passed in the output.

[0133] In an embodiment of the disclosure, the ASMM 208 may be configured to modify the audio based on a spatial audio interaural level difference (ILD) modification technique to provide a three-dimensional (3D) audio effect to generate the personalized audio stream. ILD may be responsible for the localization and depth perception in 3D audio experience. To achieve the localization in 3D audio perception, same audio may be sent through the one or more audio devices (that is, for example, both the earbuds) at different level. Now, if the stream frequency lies in the auditory range, then number of bits sent < number of actual bits (both the ears), and the number of bits sent = (No. of actual bits / masking score). Further, if the stream frequency lies in the non-auditory range, then the number of bits sent = number of actual bits. Thus, for a level in left ear (Left audio stream) ‘L’, a level in Right ear (Right audio stream) ‘R’, frequency of the audio stream ‘f’, and if the shifting coefficient at frequency f is ‘C’, then the level sent from left ear is ‘L+C’, and the level sent from right ear is ‘R+C’. By this, compression and enhanced quality may be achieved for 3D audio.

[0134] In an embodiment of the disclosure, the compression module (CM) 210 may be used for representing the audio signal with the minimum number of bits achieving transparent signal reproduction. An exemplary block diagram of CM 210 is shown in Fig. 5. A compressed audio file for the personalized audio stream based on a bitwise encoding technique may be generated by the CM 210. Further details related to the CM 210 may be discussed with reference to Figure 5 in this disclosure.

[0135] Fig. 3 illustrates a flowchart of a method for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure.

[0136] Referring to Fig. 3, a flowchart of a method 300 for generating a personalized audio stream for a user, in accordance with an embodiment of the disclosure is shown. In an example, the method 300 may be performed by the system 100. In another example, the method 300 may be performed by the components of system 200 and therefore, the steps of Fig. 3 may be construed in conjunction with the explanation of system 200 as explained with reference to Fig. 2 above. Also, as shown in Fig. 3, the method 300 starts at step 302, and goes to step 304. It may be construed that the method 300 is triggered at step 302 where a user may access an audio output device for listening to an audio stream.

[0137] At step 304, the method may comprise transmitting one or more audio tones to one or more audio devices corresponding to each ear of the user. In an example, the one or more audio tones may be transmitted to one or more audio devices at one or more sets of loudness levels and one or more frequencies. Each of the one or more sets of loudness levels may correspond to a frequency among the one or more frequencies. This means that one or more audio tones may be transmitted to one or more audio devices corresponding to each ear of the user (that is for example, a left earbud and right earbud).

[0138] At step 306, the method may comprise receiving, from the user, a response based on the transmitted one or more audio tones. In an example, the response may be one of an audible response and an inaudible response. In an example, the threshold loudness level corresponding to said each ear of the user may be determined at each frequency among the one or more frequencies based on the response received from the user.

[0139] At step 308, the method may comprise generating an audiogram for the user based on a threshold loudness level corresponding to said each ear of the user. In a non-limiting example, a minimum level of loudness at which a user’s ear (i.e., left ear or right ear) can hear a sound at a particular frequency (i.e., the minimum level of loudness at which a user responds as ‘audible’ sound) may be referred to and determined as a threshold loudness level corresponding to that frequency for the user’s ear. In other words, the threshold loudness level may indicate a minimum loudness level from the set of loudness levels at which the one or more audio tones are audible at said each ear. Also, in an example, generating the audiogram may comprise normalizing the audiogram based on at least one of a set of physiological parameters of the user (such as, but not limited to, blood pressure levels of the user) and a set of environmental parameters related to the user (such as, but not limited to, environmental noise around the user).

[0140] At step 310, the method may comprise determining one or more audio masking parameters based at least on the audiogram.

[0141] After determining one or more audio masking parameters, at step 312, the method may comprise receiving an input audio stream.

[0142] At step 314, the method may comprise applying the one or more audio masking parameters to the received input audio stream to generate the personalized audio stream for the user. This personalized audio stream for the user, in an example, may be generated as a compressed audio file based on a bitwise encoding technique.

[0143] For implementing the one or more audio masking parameters, in an example, the method may comprise receiving the input audio stream at the one or more audio devices. The method may comprise determining an amplification factor for the one or more audio devices based on the audiogram. The amplification factor may be related to the received input audio stream. The method may comprise determining the one or more audio masking parameters based on at least one of a set of real-time physiological parameters of the user, a set of real-time environmental parameters related to the user, the determined amplification factor or a threshold loudness level of the one or more frequencies. The one or more audio masking parameters may comprise a set of parameters related to a current auditory profile information of the user.

[0144] In an embodiment of the disclosure, the amplification factor may be determined based on a difference between a loudness level of the input audio stream received at a particular frequency and a threshold loudness level corresponding to the particular frequency. Also, in an example, the input audio stream may comprise one or more sub-streams, and the personalized audio stream may be generated based on an implementation of the one or more audio masking parameters on the one or more sub-streams.

[0145] In an embodiment of the disclosure, for generating the personalized audio stream, the method may comprise identifying at least one of one or more prominent streams, or one or more non-prominent streams from the one or more sub-streams and applying the one or more audio masking parameters to at least the one or more prominent streams. These one or more sub-streams may be identified as at least one of the one or more prominent streams, or the one or more non-prominent streams based on one or more parameters of the received input audio stream. Also, in an example, the personalized audio stream may be generated based on at least one of an audiogram shifting technique, a frequency shifting technique, a bitwise modification technique, or a spatial audio interaural level difference (ILD) modification technique to provide a three-dimensional (3D) audio effect.

[0146] Fig. 4 illustrates a flowchart of a hearing test, in accordance with an embodiment of the disclosure.

[0147] Referring to Fig. 4, which illustrates a flowchart of a hearing test, in accordance with an embodiment of the disclosure. In an example, as shown in block 402 of Fig. 4, a channel may be selected for testing the one or more audio devices. For example, the one or more audio devices may comprise a pair of earbuds, i.e., a left side earbud and a right side earbud, then a suitable channel is selected for testing either of these earbuds.

[0148] In an embodiment of the disclosure, a set of frequencies may be available, corresponding to which test is conducted, as shown in block 404. As shown in block 406, corresponding to each frequency, a set of loudness may be available. As shown in block 408, a searching algorithm (such as, but not limited to, binary search) may be applied on the set of frequencies. This may be a re-iterative procedure. Each iteration of search depends on the input of the user. For the above re-iterative procedure, an audio tone at a frequency and loudness level is generated for user as shown in block 410.

[0149] In an embodiment of the disclosure, as shown in block 414, it may be checked whether the user has heard the tone or not. Also, for this re-iterative procedure, a message corresponding to each tone (frequency (f), Loudness (L)) may be sent to the device (for e.g., smartphone) to produce tone. The response may be one of an audible response and an inaudible response. If the user input, i.e., user response, as shown in block 412, indicates that the user has heard the audio tone (audible response), then the minimum loudness level may be determined, at which the user hears the audio tone at the given frequency. This minimum loudness level may be the threshold loudness level as it indicates a minimum loudness level from the set of loudness levels at which the one or more audio tones are audible at an ear of the user. Based on this threshold loudness level for each frequency for the user, an audiogram corresponding to each ear of the user may be generated for the user.

[0150] In an embodiment of the disclosure, the next frequency may be taken for testing. Also, if the user input indicates that the user has not heard the audio tone (inaudible response), then the user may be presented with another loudness level (say for example, a higher loudness level) for that frequency. And the process may be repeated / re-iterated for all frequencies in the set of frequencies, and for the one or more audio devices, that is, the left side earbud and the right side earbud according to the above example. In a non-limiting example, the set of frequencies may comprise 250Hz, 500Hz, 750Hz, 1kHz, 1.5kHz, 2kHz, 3kHz, 4kHz, 6kHz, 8kHz. Also, in this example, say, the loudness levels vary from 0dB HL ~ 80dB HL (Hearing-loss Level) with 5dB step. For example, in a first iteration for a particular frequency, 40dB may be selected according to the binary search algorithm. In the second iteration, 20dB and 60dB may be selected. In the third iteration, 10dB, 30dB, 50dB and 70dB may be selected, and so on according to the standard binary search algorithm as generally known in the art. An embodiment of the binary search algorithm will be described with reference to Fig. 5 below.

[0151] Fig. 5 illustrates a process of determining a threshold loudness level for each frequency in accordance with an embodiment of the disclosure.

[0152] Referring to Fig. 5, a search algorithm according to an embodiment of the disclosure is performed to determine a threshold loudness level of a user for a specific frequency.

[0153] In an embodiment of the disclosure, the hearing test may be conducted for a plurality of frequency sets (e.g., 250 Hz, 500 Hz, 750 Hz, 1 kHz, 1.5 kHz, 2 kHz, 3 kHz, 4 kHz, 6 kHz, and 8 kHz) with respect to each of the left and right channels. For each frequency, a loudness range from 0 dB to 80 dB may be set and subdivided into steps of 5 dB to 10 dB.

[0154] The search algorithm sequentially applies a plurality of loudness levels for each frequency and receives an input from the user indicating whether the corresponding tone has been perceived (Yes / No). Here, the algorithm may be implemented using a binary search method, and the search range of loudness is reduced iteratively depending on the user’s response.

[0155] For example, when a tone is output at 40 dB for a specific frequency and the user does not perceive the tone, the loudness level is adjusted to a higher level (e.g., 60 dB HL). When the user perceives the tone at 60 dB, a lower level (e.g., 50 dB HL) is then selected. Depending on the subsequent user responses, the final hearing threshold loudness level for the corresponding frequency is determined.

[0156] In the example of Fig. 5, 40 dB is selected in the first iteration, 60 dB in the second iteration, 50 dB in the third iteration, 55 dB in the fourth iteration, and 60 dB / 55 dB in the fifth iteration, whereby the minimum loudness value at which the user can perceive the corresponding frequency is ultimately determined.

[0157] Accordingly, the search algorithm of the present disclosure efficiently performs loudness searching according to the user’s audibility for each frequency, thereby enabling generation of a personalized audiogram of the user.

[0158] Fig. 6 illustrates a block diagram of a compression module, in accordance with an embodiment of the disclosure.

[0159] Referring to Fig. 6, which illustrates a block diagram of a compression module 210, in accordance with an embodiment of the disclosure. As shown in Fig. 6, the CM 210 may comprise a bit stream encoder 602 (or bitstream encoding module 602), a bit stream decoder 604 (or bitstream decoding module 604), a quantization module 606, and an inverse quantization module 608. The CM 210 may take the generated personalized audiogram and modified audio streams as inputs.

[0160] In an embodiment of the disclosure, in the quantization module 606, a masker and a set of masking thresholds may be determined for performing quantization, which is the process of mapping a large set of input values into a smaller set. The masker and the set of masking threshold may be used in the process of the quantization in the decoder part of the codec. Pertinently, the masking may refer to a phenomenon of masking some frequencies blocking the perception of other frequencies. Masking identifies regions in the audio spectrum where certain frequency(ies) are less perceptible due to presence of louder sounds nearby. In the identified regions, fewer bits are allocated for encoding, as the user is less sensitive to changes in these areas. Masking can also take place when the masker and signal sounds are not played simultaneously. This is called temporal masking. This may optimize the allocation of bits, focusing more on preserving the sounds more prominent and sensitive to the user.

[0161] In an embodiment of the disclosure, the audio file may be compressed by reducing the number of bits allocated to each prominent sub-stream of audio based upon the masker and the masking thresholds. Further, the compressed audio may be encoded by the bitstream encoder 602 into a digital format suitable for storage in the memory unit 104 and / or transmission to the bit stream decoder 604. The bit stream encoder 602 may take the quantized data and encodes it into a compact and binary format suitable for storage and transmission.

[0162] In an example, a variable length coding technique, which involves representing the most frequent codes using fewest bits and less frequent codes using more bits, for encoding may be used. In this technique, in a quantization block, low or middle (or high) frequency sub-bands are allocated 2 bits preferentially, and many sub-bands are allocated with 2 bits. This may be done using a bit table.

[0163] In an embodiment of the disclosure, the probability distribution of each symbol generated in the quantization block may be generally uneven. Further, the symbol (frequency) may be replaced with bits, and the encoded bit stream may be produced. Further, the bit stream decoder 604 may read the compressed audio file and extract the encoded data. Further, the output from the bit stream decoder 604 may be subject to inverse quantization by the inverse quantization module 608, which converts the compressed, encoded data into its original numeric form.

[0164] The bit stream decoder 604 may receive the transmitted bit stream and decodes it to get quantized frequency sub bands. Using the same bit table (as used for encoding), the bit-stream is decoded, and the probability distribution of the audio file is found. Once the probability distribution is found, the sub-band is created back to its original state. Further, inverse quantization may be performed by the inverse quantization module 608, which maps the discrete values back to an approximate of the original continuous values. In this, each quantized value may be assigned a representative value within the original range based on a quantized table. Using the quantized table, the audio file may be reconstructed back from the sub-bands information that was decoded.

[0165] FIG. 7 illustrates a method for generating personalized audio stream for the user, in accordance with an embodiment of the disclosure.

[0166] Referring to Fig. 7, a method for generating personalized audio stream for the user, in accordance with an embodiment of the disclosure is shown. The user may access a music platform for listening to an audio stream via an audio output device. The music platform may store favorite albums of the user. The audio stream may provide one or more audio tones which may be transmitted to one or more audio devices which may correspond to each ear of the user. Then based on the generated audiogram which may be generated based on a threshold loudness level corresponding to the each ear of the user, one or more audio masking parameters may be generated.

[0167] In an embodiment of the disclosure, the audiogram may be normalized based on a profile of the user. The profile of the user may be determined based on the set of physiological parameters of the user (such as, but not limited to, blood pressure levels of the user) and the set of environmental parameters may be related to the user (such as, but not limited to, environmental noise around the user).

[0168] In an embodiment of the disclosure, the one or more audio masking parameters may be implemented on received input audio stream to generate the personalized audio stream for the user. Also, the received input stream after the implementation of the one or more masking parameters may be encoded by the encoder to generate the personalized audio stream. The personalized audio stream may be then provided to the output device which may then be used by the user for listening. It may be noted that the personalized audio stream (or a Personalized Audio Codec) may tailor the streaming experience of the user to individual preferences by adjusting parameters like compression levels, equalization, or dynamic range based on a user’s hearing sensitivity.

[0169] The personalized audio stream may help in suppression of harsh sounds, boosting inaudible frequencies, removing distracting sub-streams, and also requires low bandwidth for streaming. Further, the implementation of the disclosure may encode the streamed audio metadata more accurately according to user’s preference while optimizing the playback quality for minimal data usage, using limited processing power and low bandwidth. Also, people with hearing loss often have difficulty with dynamic ranges in music, where soft sounds may be inaudible and loud sounds may be overwhelming, in such conditions, the implementation of the generated personalized audio stream as provided by various implementations of the present disclosure may provide the dynamic range, making all frequencies suitable for a particular user to enjoy.

[0170] FIG. 8 illustrates a method for generating personalized audio stream for the user, in accordance with an embodiment of the disclosure.

[0171] Referring to Fig. 8, another illustration of a method for generating personalized audio stream for the user, in accordance with an embodiment of the disclosure is shown. When the user accesses an audio output device for listening to an audio stream, an audio file may be shared which may be in a pulse-code modulation (PCM) Format. This original audio file for the audio stream may have a size, say 4.5 MB. The original audio file may comprise one or more audio tones which may be shared to one or more audio devices corresponding to each ear of the user.

[0172] In an embodiment of the disclosure, the one or more audio masking parameters may be implemented on the received input audio stream to generate the personalized audio stream for the user. This personalized audio stream for the user, in an example, may be generated as a compressed audio file based on a bitwise encoding technique. In such cases of compressed audio file, there may be variable compression levels for the personalized audio stream which may provide content based compression based on adaptive encoding. The file size after the compression may be then reduced to say 2.9MB while performing the content based compression adaptive encoding.

[0173] In an embodiment of the disclosure, the resulting compressed audio file after the compression is done may comprise the personalized audio stream having a compressed size of, say 2.9 MB. Further, it may be noted that by identifying frequencies that a listener with hearing loss cannot hear well or at all, the codec may apply more aggressive compression techniques to these less critical frequencies, and may result in reducing overall data size. Depending on the type of content being stored (such as speech, music, etc.), the codec may also be modified to apply different compression techniques. For example, speech content in a music file can be compressed more aggressively without loss of intelligibility, leading to even smaller file size.

[0174] In an embodiment of the disclosure, a method for generating a personalized audio stream for a user is provided. The method may include transmitting one or more audio tones to one or more audio devices corresponding to each ear of the user. The method may include receiving, from the user, a response based on the transmitted one or more audio tones. The method may include generating an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user. The method may include determining one or more audio masking parameters based at least on the audiogram. The method may include receiving an input audio stream. The method may include generating the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream.

[0175] In an embodiment of the disclosure, the method may further comprise generating the one or more audio tones at a plurality of loudness levels for a frequency among a pre-defined set of frequencies

[0176] In an embodiment of the disclosure, the response may be one of an audible response and an inaudible response.

[0177] In an embodiment of the disclosure, the threshold loudness level may indicate a minimum loudness level from the set of loudness levels at which the one or more audio tones are audible at said each ear.

[0178] In an embodiment of the disclosure, generating the audiogram may include normalizing the audiogram based on at least one of a set of physiological parameters of the user and a set of environmental parameters related to the user.

[0179] In an embodiment of the disclosure, generating the personalized audio stream for the user may include: (a) receiving the input audio stream at the one or more audio devices, (b) determining an amplification factor for the one or more audio devices based on the audiogram, wherein the amplification factor is related to the received input audio stream, and (c) applying the one or more audio masking parameters based on at least one of a set of real-time physiological parameters of the user, a set of real-time environmental parameters related to the user, the determined amplification factor or a threshold loudness level corresponding to one or more frequencies, wherein the one or more audio masking parameters may include a set of parameters related to a current auditory profile information of the user.

[0180] In an embodiment of the disclosure, the amplification factor may be determined based on a difference between a loudness level of the input audio stream received at a particular frequency and a threshold loudness level corresponding to the particular frequency.

[0181] In an embodiment of the disclosure, the input audio stream may include one or more sub-streams, and the personalized audio stream may be generated based on an implementation of the one or more audio masking parameters on the one or more sub-streams.

[0182] In an embodiment of the disclosure, the method may include: (a) identifying at least one of one or more prominent streams and one or more non-prominent streams from the one or more sub-streams, and (b) applying the one or more audio masking parameters to at least the one or more prominent streams.

[0183] In an embodiment of the disclosure, the one or more sub-streams may be identified as at least one of the one or more prominent streams and the one or more non-prominent streams based on one or more parameters of the received input audio stream.

[0184] In an embodiment of the disclosure, the personalized audio stream may be generated based on at least one of an audiogram shifting technique, a frequency shifting technique, a bitwise modification technique, or a spatial audio interaural level difference (ILD) modification technique to provide a three-dimensional (3D) audio effect.

[0185] In an embodiment of the disclosure, the method may include generating a compressed audio file for the personalized audio stream based on a bitwise encoding technique.

[0186] In an embodiment of the disclosure, the one or more audio tones may be transmitted to one or more audio devices at one or more sets of loudness levels and one or more frequencies, and wherein each of the one or more sets of loudness levels may correspond to a frequency among the one or more frequencies.

[0187] In an embodiment of the disclosure, the threshold loudness level corresponding to said each ear of the user may be determined at each frequency among the one or more frequencies based on the response.

[0188] In an embodiment of the disclosure, a system for generating a personalized audio stream for a user is provided. The system may include at least one processing unit, and at least one memory unit connected to the at least one processing unit. The at least one processing unit may be configured to transmit one or more audio tones to one or more audio devices corresponding to each ear of the user. The at least one processing unit may be configured to receive, from the user, a response based on the transmitted one or more audio tones. The at least one processing unit may be configured to generate an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user. The at least one processing unit may be configured to determine one or more audio masking parameters based at least on the audiogram. The at least one processing unit may be configured to receive an input audio stream. The at least one processing unit may be configured to generate the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream.

[0189] In an embodiment of the disclosure, a non-transitory computer readable storage medium storing instructions for generating a personalized audio stream for a user is provided. The instructions may include executable code which, when executed by one or more units of a system 100, causes a processing unit 102 of the system 100 to transmit one or more audio tones to one or more audio devices corresponding to each ear of the user. The executable code, when executed, may cause the processing unit 102 to receive, from the user, a response based on the transmitted one or more audio tones. The executable code, when executed, may cause the processing unit 102 to generate an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user. The executable code, when executed, may cause the processing unit 102 to determine one or more audio masking parameters based at least on the audiogram. The executable code, when executed, may cause the processing unit 102 to receive an input audio stream. The executable code, when executed, may cause the processing unit 102 to generate the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream to generate the personalized audio stream for the user.

[0190] Thus, the disclosure provides a novel solution for generating the personalized audio stream for a user based on a generated audiogram of the user. The present disclosure provides a solution for normalization of a generated audiogram of a user based on the current blood pressure of the user and environment noise condition around the user. Further, the present solution enables determining a masking equation for the user which contains auditory profile information of the user. Further, implementing the features as disclosed above in this disclosure, spatial audio interaural level difference (ILD) modification for 3D spatialization and localization can be performed.

[0191] While considerable emphasis has been placed herein on the preferred embodiments, it will be appreciated that many embodiments can be made and that many changes can be made in the preferred embodiments without departing from the principles of the invention. These and other changes in the preferred embodiments of the invention will be apparent to those skilled in the art from the disclosure herein, whereby it is to be distinctly understood that the foregoing descriptive matter to be implemented merely as illustrative of the invention and not as limitation.

[0192] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications are intended to be comprehended within the meaning and range of equivalents of the disclosed embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the scope of the embodiments as described herein.

Claims

1.A method (300) for generating a personalized audio stream for a user, the method comprising:transmitting (304) one or more audio tones to one or more audio devices corresponding to each ear of the user;receiving (306), from the user, a response based on the transmitted one or more audio tones;generating (308) an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user;determining (310) one or more audio masking parameters based on the audiogram;receiving (312) an input audio stream; andgenerating (314) the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream.2.The method of claim 1, wherein the method (300) further comprises generating the one or more audio tones at a plurality of loudness levels for a frequency among a pre-defined set of frequencies.3.The method of any one of claims 1 or 2, wherein the response is one of an audible response and an inaudible response.4.The method of any one of claims 1 to 3, wherein the threshold loudness level indicates a minimum loudness level from a set of loudness levels at which the one or more audio tones are audible at the each ear of the user.5.The method of any one of claims 1 to 4, wherein generating (308) the audiogram comprises normalizing the audiogram based on at least one of a set of physiological parameters of the user and a set of environmental parameters related to the user.6.The method of any one of claims 1 to 5, wherein generating (314) the personalized audio stream for the user comprises:receiving the input audio stream at the one or more audio devices,determining an amplification factor for the one or more audio devices based on the audiogram, wherein the amplification factor is related to the received input audio stream, andapplying the one or more audio masking parameters based on at least one of a set of real-time physiological parameters of the user, a set of real-time environmental parameters related to the user, the determined amplification factor or a threshold loudness level corresponding to one or more frequencies, wherein the one or more audio masking parameters comprises a set of parameters related to a current auditory profile information of the user.7.The method of any one of claims 1 to 6, wherein the amplification factor is determined based on a difference between a loudness level of the input audio stream received at a particular frequency and a threshold loudness level corresponding to the particular frequency.8.The method of any one of claims 1 to 7, wherein the input audio stream comprises one or more sub-streams, and the personalized audio stream is generated based on an implementation of the one or more audio masking parameters on the one or more sub-streams.9.The method of any one of claims 1 to 8, wherein the method (300) further comprises:identifying at least one of one or more prominent streams and one or more non-prominent streams from the one or more sub-streams, andapplying the one or more audio masking parameters to at least the one or more prominent streams.10.The method of any one of claims 1 to 9, wherein the one or more sub-streams are identified as at least one of the one or more prominent streams, or the one or more non-prominent streams based on one or more parameters of the received input audio stream.11.The method of any one of claims 1 to 10, wherein the personalized audio stream is generated based on at least one of an audiogram shifting technique, a frequency shifting technique, a bitwise modification technique, or a spatial audio interaural level difference (ILD) modification technique to provide a three-dimensional (3D) audio effect.12.The method of any one of claims 1 to 11, wherein the method comprises generating a compressed audio file for the personalized audio stream based on a bitwise encoding technique.13.The method of any one of claims 1 to 12, wherein the one or more audio tones are transmitted to one or more audio devices at one or more sets of loudness levels and one or more frequencies, and wherein each of the one or more sets of loudness levels corresponds to a frequency among the one or more frequencies.14.The method of any one of claims 1 to 13, wherein the threshold loudness level corresponding to the each ear of the user is determined at each frequency among the one or more frequencies based on the response.15.A system for generating a personalized audio stream for a user, the system comprising:a memory unit, anda processing unit connected to the memory unit, wherein the processing unit is configured to:transmit one or more audio tones to one or more audio devices corresponding to each ear of the user,receive, from the user, a response based on the transmitted one or more audio tones,generate an audiogram for the user based on a threshold loudness level corresponding to the each ear of the user,determine one or more audio masking parameters based at least on the audiogram,receive an input audio stream, andgenerate the personalized audio stream for the user by applying the one or more audio masking parameters to the received input audio stream.

Citation Information

Patent Citations

  • Audio playing method, electronic equipment and computer readable storage medium

    CN117093182A

  • Individual audio receiver programmer

    US20100020988A1

  • Binaural audio calibration

    US20180249271A1

  • Method and electronic device for personalized audio enhancement

    US20230260526A1

  • Customized selective attenuation of game audio

    US20240064487A1