Method and electronic apparatus for controlling membrane excursion
The system addresses micro-speaker distortions by using a trained distortion classifier and AI models to generate distortion-free audio, improving audio quality and speaker durability.
Patent Information
- Application Number
- PCT/KR2025/002884
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-12
- Filing Date
- 2025-03-04
- Publication Date
- 2025-09-18
AI Technical Summary
Micro-speakers in mobile devices face challenges with low power efficiency, non-linear behavior leading to audible distortions and mechanical damage due to irregular non-linear modes, which existing methods struggle to effectively address.
A system and method using a distortion classifier trained with a synthetic audio dataset and AI models to identify and correct audio frames causing distortions, combined with AI speaker models to generate distortion-free audio frames.
Significantly reduces audible distortions, enhances audio clarity and richness, and minimizes mechanical stress on the speaker diaphragm, while allowing user customization of sound output.
Smart Images

Figure KR2025002884_18092025_PF_FP_ABST
Abstract
Description
METHOD AND ELECTRONIC APPARATUS FOR CONTROLLING MEMBRANE EXCURSION
[0001] The present disclosure generally relates to the field of audio signal processing in electronic devices. More particularly, the present disclosure relates to a method and a system for membrane excursion control of a speaker device.
[0002] The increasing use of mobile devices for multimedia applications such as music playback, video calls, gaming, and podcasts has elevated the importance of micro-speakers in delivering immersive audio experiences. These compact electroacoustic transducers or the micro-speakers, typically measuring around 15 mm x 11 mm, are essential for providing sound output within the limited physical and power constraints of mobile devices. However, the micro-speakers face significant challenges due to their low power efficiency and the need to produce loud audio output with rich bass. To achieve these requirements, the micro-speaker diaphragm frequently operates beyond its linear range, leading to non-linear behaviour. This non-linearity can introduce two major issues: audible distortions that degrade audio quality and mechanical damage to the diaphragm and surrounding components over time.
[0003] The operation of micro-speakers can be classified into linear, regular non-linear, and irregular non-linear modes. In linear mode, the diaphragm operates within safe displacement limits, maintaining sound quality with minimal distortion. Regular non-linear mode, which occurs at up to about 10% of the maximum rated displacement, involves deterministic behaviours such as changes in the spring constant and force factor, which are theoretically controllable. However, irregular non-linear mode, caused by larger displacements and / or manufacturing defects, results in semi-deterministic or stochastic behaviours like rub and buzz effects and airflow noise, which are harder to predict and control.
[0004] Existing methods for mitigating these issues typically rely on current-voltage (I-V) sensing and physical modelling of the speaker's characteristics, but they face limitations. These include the need for complex hardware, sensitivity to model assumptions, and a focus primarily on speaker protection, often at the expense of audio quality in terms of timbre, loudness, and bass response. Despite advancements in adaptive control strategies, irregular non-linearities such as rub and buzz remain particularly challenging to address, highlighting the limitations of current approaches.
[0005] Some prior arts in this domain include methods such as predicting the diaphragm amplitude of a micro-speaker based on an excitation signal and determining whether this amplitude exceeds a preset threshold. If the amplitude exceeds the present threshold, these methods apply a target adjustment gain via an equalizer algorithm to reduce the diaphragm amplitude below the threshold. This process aims to avoid airflow noise, prevent loudness fluctuations, and improve the efficiency of audio signal adjustment.
[0006] Additionally, other approaches utilize neural networks by inputting data samples of current and prior displacements of moving loudspeaker components into successive layers. These methods generate data vectors via nonlinear activation functions, transform these vectors through affine operations, and output predictions for the voltage required to control loudspeaker displacement. While these methods provide advanced techniques for managing speaker operation, they still face challenges in addressing irregular non-linear distortions and maintaining overall audio quality.
[0007] Therefore, in view of the above-mentioned problems, it is desirable to provide a system and a method that may eliminate, or at least, mitigate one or more of the above-mentioned problems associated with the existing solutions.
[0008] This summary is provided to introduce a selection of concepts, in a simplified format, that are further described in the detailed description of the disclosure. This summary is neither intended to identify key or essential inventive concepts of the invention nor is it intended to determine the scope of the disclosure.
[0009] According to one embodiment of the present disclosure, a method for membrane excursion control of an electronic apparatus is disclosed. The method includes receiving an audio signal input from one or more audio sources. The method includes extracting one or more audio parameters derived from the audio signal input. The one or more audio parameters include one or more speaker characteristics and signal features. The method includes inputting the one or more audio parameters into a distortion classifier to identify a first set of audio frames in the audio signal input causing audible distortion. The distortion classifier is trained using a synthetic audio dataset generated using a diffusion model with a custom loss function. The method further includes generating a first output, comprising a second set of audio frames, by inputting the identified first set of audio frames causing audio distortion into a plurality of speaker models. The second set of audio frames is distortion less and the plurality of speaker models are artificial intelligence (AI) models trained using the synthetic audio dataset.
[0010] According to another embodiment of the present disclosure, an electronic apparatus for controlling membrane excursion is disclosed. The electronic apparatus includes a memory configured to store executable instructions and at least one processor coupled to the memory and configured to receive an audio signal input from one or more audio sources. The at least one processor is configured to extract one or more audio parameters derived from the audio signal input. The one or more audio parameters include one or more speaker characteristics and signal features. The at least one processor is configured to input the one or more audio parameters into a distortion classifier to identify a first set of audio frames in the audio signal input causing audible distortion. The distortion classifier is trained using a synthetic audio dataset generated using a diffusion model with a custom loss function. The at least one processor is configured to generate a first output, comprising a second set of audio frames, by inputting the identified first set of audio frames causing audio distortion into a plurality of speaker models. The second set of audio frames is distortion less and the plurality of speaker models are artificial intelligence (AI) models trained using the synthetic audio dataset.
[0011] To further clarify the advantages and features of the present disclosure, a more particular description of the disclosure will be rendered by reference to specific embodiments thereof, which are illustrated in the appended drawing. It is appreciated that these drawings depict only typical embodiments of the disclosure and are therefore not to be considered limiting its scope. The disclosure will be described and explained with additional specificity and detail with the accompanying drawings.
[0012] These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:
[0013] FIG. 1 illustrates an architecture of a system for membrane excursion control of a speaker device, according to an embodiment of the present disclosure;
[0014] FIG. 2 illustrates a process-flow for generating a target audio using a diffusion model , according to an embodiment of the present disclosure;
[0015] FIG. 3a-3b illustrate a process-flow for training and inference of the distortion classifier, according to an embodiment of the present disclosure;
[0016] FIG.s 4a illustrate a process-flow for training the first speaker model or the second speaker model, according to an embodiment of the present disclosure;
[0017] FIG. 4b illustrates a process-flow for inference of the first speaker model and the second speaker model, according to an embodiment of the present disclosure;
[0018] FIG. 5 illustrates a system for membrane excursion control of the speaker device according to an embodiment of the present disclosure; and
[0019] FIG. 6 illustrates a flowchart depicting a method for membrane excursion control of the speaker device, according to an embodiment of the present disclosure.
[0020] Further, skilled artisans will appreciate that elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent steps involved to help to improve understanding of aspects of the present disclosure. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding the embodiments of the present disclosure so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
[0021] For the purpose of promoting an understanding of the principles of the disclosure, reference will now be made to the various embodiments, and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the disclosure as illustrated therein being contemplated as would normally occur to one skilled in the art to which the disclosure relates.
[0022] The term "some" or "one or more" as used herein is defined as "one", "more than one", or "all." Accordingly, the terms "more than one," "one or more" or "all" would all fall under the definition of "some" or "one or more". The term "an embodiment", "another embodiment", "some embodiments", or "in one or more embodiments" may refer to one embodiment or several embodiments, or all embodiments. Accordingly, the term "some embodiments" is defined as meaning "one embodiment, or more than one embodiment, or all embodiments."
[0023] The terminology and structure employed herein are for describing, teaching, and illuminating some embodiments and their specific features and elements and do not limit, restrict, or reduce the spirit and scope of the claims or their equivalents. The phrase "exemplary" may refer to an example.
[0024] More specifically, any terms used herein such as but not limited to "includes," "comprises," "has," "consists," "have" and grammatical variants thereof do not specify an exact limitation or restriction and certainly do not exclude the possible addition of one or more features or elements, unless otherwise stated, and must not be taken to exclude the possible removal of one or more of the listed features and elements, unless otherwise stated with the limiting language "must comprise" or "needs to include".
[0025] Whether or not a certain feature or element was limited to being used only once, either way, it may still be referred to as "one or more features", "one or more elements", "at least one feature", or "at least one element." Furthermore, the use of the terms "one or more" or "at least one" feature or element does not preclude there being none of that feature or element unless otherwise specified by limiting language such as "there needs to be one or more " or "one or more element is required."
[0026] Unless otherwise defined, all terms, and especially any technical and / or scientific terms, used herein may be taken to have the same meaning as commonly understood by one having ordinary skill in the art.
[0027] FIG. 1 illustrates an architecture 100 of a system for membrane excursion control of a speaker device, according to an embodiment of the present disclosure.
[0028] In an embodiment, the architecture 100 of the system for membrane excursion control of the speaker device includes, a windowing module 104 receiving an input PCM audio waveform 102, a distortion classifier 106, a first speaker model 108, and a second speaker model 110 producing a modified PCM audio waveform 112. In an embodiment, the first speaker model may be ideal speaker model and the second speaker model may be an inverse speaker model.
[0029] In one embodiment, an input pulse code modulation (PCM) audio waveform 102 is provided as an input to the windowing module 104. In one example the input PCM audio waveform 102 of an audio track played during a video call or music playback is provided as input to the windowing module 104. The windowing module 104 may be configured to extract one or more audio parameters derived from the audio signal input. The one or more audio parameters include one or more speaker characteristics and signal features. The speaker characteristics indicates a physical and acoustic characteristics of the speaker and the signal features include amplitude, frequency, a root mean square energy, and the like.
[0030] Upon extracting one or more audio parameters derived from the input PCM audio signal, the distortion classifier 106 may be configured to receive the extracted one or more audio parameters to identify a first set of audio frames in the audio signal input causing audible distortion. The first set of audio frames in the audio signal input causing audible distortion indicates a specific portions or segments (a one or more audio frames) of the overall input PCM audio signal that are responsible for producing noticeable, undesired sound distortions when played through the speaker.
[0031] Further, the distortion classifier 106 may be configured to evaluate the one or more audio parameters received from the windowing module 104 and consequently determine whether the one or more audio frames are likely to cause audible distortions based on measured speaker characteristics and signal features. The one or more audio parameters may include the one or more speaker characteristics such as a physical and acoustic characteristics of the speaker and the signal features such as amplitude, frequency, and the like. Furthermore, the distortion classifier 106 may be trained using a synthetic audio dataset generated using a diffusion model with a custom loss function. The synthetic audio dataset is a type of audio dataset trained using synthetic data or duplicate data.
[0032] In an example, when a user plays a high-bass song on their smartphone at maximum volume, the distortion classifier 106 may be configured to analyze the incoming audio waveform to extract key parameters such as frequency content, amplitude, and transients. For instance, the distortion classifier may detect that certain frames of the audio signal, such as the bass drops at 150 Hz, have high amplitude low-frequency components that could exceed the speaker's excursion limits. The distortion classifier 106 evaluates the one or more audio parameters and identify the bass-heavy frames as likely to cause audible distortion.
[0033] Upon identifying the first set of audio frames in the audio signal input causing audible distortion by the distortion classifier 106, the first speaker model 108 may be configured to generate a second output by receiving the identified first set of audio frames. Further, the second speaker model 110 may be configured to receive the second output from the first speaker model 108 to generate the first output i.e., modified PCM audio waveform 112. The modified PCM audio waveform 112 may comprise of the a second set of audio frames which are distortion-less or distortion free audio frames.
[0034] Additionally, to generate a second output, the first speaker model 108 may be configured to determine an acoustic pressure of the identified first set of audio frames. The first speaker model 108 may then be configured to generate the second output based on the first set of audio frames. The second output indicates the determined acoustic pressure of the identified first set of audio frames.
[0035] The determined acoustic pressure of the first set of audio frames refers to the variation in pressure caused by the sound waves generated by the audio signal, specifically for the frames identified as distortion-prone. When an audio signal drives the speaker, it produces mechanical vibrations in the speaker membrane, consequently creating sound waves in the surrounding air. These sound waves result in localized pressure changes, measured as acoustic pressure, which corresponds to the perceived loudness and intensity of the sound. For the first set of audio frames, which are flagged for causing audible distortion, the acoustic pressure is typically higher and often exhibits irregular behavior due to excessive displacement of the speaker membrane. This irregularity can lead to nonlinear responses, such as distorted or harsh sounds.
[0036] In an example, when a user streams a high-quality movie on a smartphone, the distortion classifier 106 may be configured to analyze the input audio signal to identify the first set of audio frames that may cause audible distortion, such as loud explosion effects with low-frequency components. The identified first set of audio frames is then processed by the first speaker model 108, which simulates the speaker's ideal acoustic response and generates the second output, representing the corrected version of the distortion-prone frames. The second output is subsequently fed into the second speaker model 110, which applies further refinement to match the actual speaker's physical characteristics and generates the distortion less modified PCM audio waveform 112.
[0037] FIG. 2 illustrates a process-flow 200 for generating a target audio 210 using a diffusion model 202, according to an embodiment of the present disclosure.
[0038] In an embodiment, the input pulse code modulation (PCM) audio waveform 102 is provided as an input to the diffusion model 202. The input PCM audio waveform 102 may generate various imperfections, such as distortions or artifacts, that affect the sound quality when played back over micro-speakers.
[0039] The diffusion model 202 may be trained specifically to generate high-quality audio waveforms which do not produce audible distortions over micro-speakers (the target audio 210) by iteratively refining the input PCM audio waveform 102. Further, custom loss functions are introduced, which are designed to measure audio quality metrics such as partial loudness, harmonic signal ratio errors, and distortion. The custom loss functions may enable the diffusion model 202 to minimize distortions while maximizing sound fidelity. The diffusion model is a type of generative machine learning model designed to generate new data by learning the underlying distribution of a given dataset. The diffusion model is particularly effective in creating high-quality images, audio, or other types of structured data.
[0040] The selection module 204 may be configured to select a second set of audio tracks based on audio quality metrics derived from the custom loss functions such that the identified second set of audio tracks is free from distortion in terms of sound quality. The audio quality metrics refer to measurable parameters or criteria used to evaluate the quality of the audio signals. The audio quality metrics may be useful to assess the audio characteristics such as clarity, fidelity, and absence of distortions. The examples of the audio quality metrics may include signal-to-noise ratio (SNR), total harmonic distortion (THD), spectral flatness, perceptual evaluation of audio quality (PEAQ), loudness, timbre, and the like.
[0041] Further, the method 200 includes inputting the selected second set of audio tracks to an equalization filter 206, wherein the equalization filter 206 is configured to emulate a perceptual filter bank-inspired filter, with filter parameters adjustable by a user listening to the audio track. In other words, the selected audio tracks (by the selection module 204) are processed through one or more equalization filters 206 (EQ filters). The one or more equalization filters are configured to mimic perceptual filter banks, such as 1 / 3rd octave or equal ERB filters, and may be adjusted for parameters such as gain, bandwidth, and quality factor. The set of audio tracks are fine-tuned by the equalization filter 206 to meet the audio quality metrics and the listener preferences.
[0042] Once the one or more equalization filters 206 have been applied, the processed set of audio tracks undergo subjective listening tests. The listening tests may involve human evaluators who rate the audio tracks based on widely accepted paradigms, such as MUSHRA (Multiple Stimuli with Hidden Reference and Anchor) or MOS (Mean Opinion Score). The evaluators assess parameters like loudness, timbre, and audible distortions.
[0043] Further, a rank determiner 208 may be configured to determine ranks of the second set of audio tracks based on audio quality metrics. The audio quality metrics of the second set of audio tracks are determined based on testing of the second set of audio tracks. The subjective listening test results are analyzed using statistical methods, such as ANOVA (Analysis of Variance), to rank the audio tracks by a ranking module 208.
[0044] The tests based on the audio quality metrics and the listener preferences may help to identify the tracks with the least distortion and the highest audio quality. Only the best-performing tracks are selected for the final output. The final stage of the process-flow 200 generates a modified PCM audio waveform i.e., the target audio 210 which is based on selecting at least one of the second set of audio tracks such that the selected at least one of the second set of audio tracks is free from distortion. The target audio 210 has significantly reduced rub and buzz artifacts and improved sound quality. The target audio 210 is distortion-free, with enhanced loudness and timbre, making it suitable for high-quality playback in speaker systems.
[0045] FIG. 3a and 3b illustrate a process-flow 300 for training and inference of the distortion classifier 106, according to an embodiment of the present disclosure.
[0046] Referring to FIG. 3a, the process-flow 300 begins with the input PCM audio waveform 102 to be provided as an input to the windowing module 104. The windowing module 104 may be configured to segment the input PCM audio waveform 102 into smaller, manageable chunks or frames using a windowing process. The smaller frames or the segmented frames may be configured to allow the distortion classifier 106 to analyze the signal at a finer resolution, focusing on individual time segments for distortion analysis.
[0047] Further, a feature extractor 302 may be configured to extract the signal features and the signal-based speaker characteristics. In an embodiment, the signal features may include a Short-Time Fourier Transform (STFT) and peak RMS, which may be extracted from the audio frames to capture spectral and temporal properties. Speaker features may include lumped parameter-based responses (e.g., stiffness, inductance, and force factor) and non-linear state-space responses, which are extracted from the signal segments together with features pre-computed for speakers under test. The speaker features and the signal features are critical in determining how the audio interacts with the speaker system, particularly under stress conditions.
[0048] In an embodiment, the distortion classifier 106 is a neural network model, such as a fully connected neural network, trained using the extracted features. The distortion classifier 106 is a binary classifier that is configured to determine whether each frame is likely or not likely to cause audible distortions.
[0049] The distortion classifier 106 may be trained using synthetic data generated by simulating speaker responses under various conditions, as described in FIG. 2. The synthetic data includes distorted data and high-quality data, ensuring robust training. The target audio is the audio which is generated based on synthetic data for training the distortion classifier 106.
[0050] Referring now to FIG. 3b, similar to the training, the inference process begins with the input PCM audio waveform 102 being inputted to the windowing module 104. The windowing module 104 may be configured to segment the input PCM waveform 102 into smaller frames for detailed analysis during inference.
[0051] The signal features and the speaker features are extracted from the audio frames during inference by the feature extractor which is the same as the feature extraction performed during training. The trained distortion classifier 106 may be used to analyze each frame. The trained distortion classifier may be configured to predict whether a frame is likely to cause audible distortion or not by using the extracted features. The distortion classifier 106 may be configured to produce two outputs namely audible frames not leading to audible distortion 304 which are distortion-free frames that can be directly processed or used without correction and audible frames leading leading to audible distortion 306 which are distortion-prone frames identified by the distortion classifier 106. These frames may require further processing to reduce distortions.
[0052] FIG. 4a and 4b illustrate a process-flow 400 for training the first speaker model 108 and the second speaker model 110, according to an embodiment of the present disclosure.
[0053] Referring to FIG. 4a, the method 400 begins with the input PCM audio waveform 102 to be provided as an input to the windowing module 104. The windowing module 104 may be configured to segment the input PCM audio waveform 102 into smaller, manageable chunks or frames using the windowing process.
[0054] Further, the segmented audio frames may be analyzed to extract speaker features and acoustic pressure using the feature extractor 302. The features may include linear lumped parameters, such as stiffness, inductance, and force factor.
[0055] The first speaker model 108 and the second speaker model 110 may be trained using the extracted features and the target audio 212. In an embodiment, the first speaker model 108 may be trained to predict the acoustic pressure and displacement for the given input PCM audio frames 102. The first speaker model serves as a benchmark for optimal audio playback, generating acoustic pressure for audio signals free from distortion when played back over target micro-speakers. In an example, the first speaker model may be an ideal speaker model. The ideal speaker model trained to estimate the acoustic pressure of the identified first set of audio frames.
[0056] In an embodiment, the second speaker model 110 may be trained to generate modified PCM waveforms 112 by compensating for the distortions introduced by the speaker's characteristics. The second speaker model 110 essentially "inverts" the speaker's response to reduce distortion and produce the output PCM waveforms. In an example, the second speaker model 110 is an inverse speaker model. The inverse speaker model is trained to generate the first output using the estimated acoustic pressure of the identified first set of audio frames.
[0057] FIG. 4b illustrates a process-flow 400 for inferenceing the first speaker model 108 and the second speaker model 110, according to an embodiment of the present disclosure.
[0058] Referring now to FIG. 4b, similar to the training, the inference process begins with the input PCM audio waveform 102 being inputted to the windowing module 104. The windowing module 104 may be configured to segment the input PCM waveform 102 into smaller frames for detailed analysis during inference.
[0059] The first speaker model 108 may be configured to receive the audio frames as input and may be configured to estimate the displacement and acoustic pressure. This step of receiving the audio frames as the input provides a reference for the speaker behaviour in an ideal, distortion-free scenario. The second speaker model or the inverse speaker model may be configured to modify the input PCM waveform 102 by applying adjustments based on the estimated displacement and acoustic pressure which may include timbre scaling and loudness scaling.
[0060] The final output (i.e., the first output) is the modified PCM audio waveform 112 with reduced distortions and improved audio quality. The modified PCM waveform 112 is suitable for playback through real-world speaker systems, ensuring an enhanced listener experience.
[0061] FIG. 5 illustrates a system 500 for membrane excursion control of the speaker device according to an embodiment of the present disclosure.
[0062] The system 500 includes a processor 502, a memory 504, a storage component 506, an input component 508, an output component 510, a communication interface 512, and a bus 514.
[0063] The processor 502, as used herein, means any type of computational circuit that may comprise hardware elements and software elements. The processor 502 may be embodied as a multi-core processor, a single core processor, or a combination of one or more multi-core processors and / or one or more single core processors, a distributed processing system, or the like. The processor 502 may be a Central Processing Unit (CPU) a graphics processing unit (GPU), an accelerated processing unit (APU), an application-specific integrated circuit (ASIC), or another type of processing component.
[0064] The memory 504 includes a non-transitory computer readable medium. The memory 504 includes a random-access memory (RAM), a read only memory (ROM), and / or another type of dynamic or static storage device (e.g., a flash memory, a magnetic memory, and / or an optical memory) that stores information and / or instructions for use by processor 502. The memory 504 comprises machine-readable instructions which are executable by the processor 502. These machine-readable instructions when executed by the processor 502 cause the processor 502 to perform one or more method steps of an embodiment described above.
[0065] The storage component 506 stores information and / or software related to the operation and use of the device. For example, the storage component 506 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid-state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium, along with a corresponding drive.
[0066] The input component 508 may be configured to receive information, such as user input. For example, the input component 508 may include, but not be limited to, a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, and / or a microphone. The output component 510 is configured to convey information from the system to the user or other systems, utilizing a variety of devices and technologies tailored to specific application needs. The output component 510 may include visual output devices such as display screens (LCD, LED, OLED), projectors, and heads-up displays (HUDs) for presenting graphical or textual information. Additionally, auditory output through speakers and headphones provides audio feedback and alerts, while haptic output devices, like vibration motors in smartphones or game controllers, offer tactile feedback. Functionally, the output component serves multiple roles, including displaying graphical user interface (GUI) elements for user interaction, delivering notifications and alerts through sound, visual indicators, or vibrations, and rendering complex data visualizations like charts and graphs for easier comprehension.
[0067] In an embodiment, the output component 510 may be configured to receive processed data from the processor 502, which determines the information to be communicated, and the output component 510 may access memory 504 and storage component 506 to retrieve and display stored information such as documents, media files, or application states. Furthermore, the output component 510 may be configured to meet the specific requirements of different applications, such as high-resolution visual output and immersive audio for gaming systems or clear and precise data visualization and alert mechanisms for industrial control systems. Through these varied output methods, the output component 510 ensures effective communication of information, enhancing both system functionality and user experience.
[0068] FIG. 6 illustrates a flowchart depicting a method 600 for membrane excursion control of the speaker device, according to an embodiment of the present disclosure.
[0069] At step 602, the method 600 includes receiving an audio signal input from one or more audio sources.
[0070] At step 604, the method 600 includes extracting one or more audio parameters derived from the audio signal input. The one or more audio parameters include one or more speaker characteristics and signal features.
[0071] At step 606, the method 600 includes inputting the one or more audio parameters into a distortion classifier to identify a first set of audio frames in the audio signal input causing audible distortion. The distortion classifier is trained using a synthetic audio dataset generated using a diffusion model with custom loss function.
[0072] At step 608, the method includes generating a first output, comprising a second set of audio frames, by inputting the identified first set of audio frames causing audio distortion into a plurality of speaker models. The second set of audio frames are distortion less
[0073] and the plurality of speaker models are artificial intelligent (AI) models trained using the synthetic audio dataset.
[0074] The present invention provides various advantages:
[0075] The present invention significantly reduces audible distortions, improving clarity and richness in audio playback.
[0076] The present invention minimizes mechanical stress on the speaker diaphragm, enhancing its lifespan by controlling membrane excursion.
[0077] The integration of equalization filters allows for adjustable parameters, enabling users to customize the sound output according to their preferences.
[0078] The use of AI-based speaker models (e.g., ideal speaker model and inverse speaker model) enables intelligent compensation for distortions in real time.
[0079] The invention incorporates a the diffusion model which generates a high-quality synthetic audio dataset, reducing the reliance on large-scale real-world data collection.
[0080] While specific language has been used to describe the disclosure, any limitations arising on account of the same are not intended. As would be apparent to a person in the art, various working modifications may be made to the method in order to implement the inventive concept as taught herein.
[0081] The drawings and the foregoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one embodiment may be added to another embodiment. For example, orders of processes described herein may be changed and are not limited to the manner described herein.
[0082] Moreover, the actions of any flow diagram need not be implemented in the order shown; nor do all of the acts necessarily need to be performed. Also, those acts that are not dependent on other acts may be performed in parallel with the other acts. The scope of embodiments is by no means limited by these specific examples. Numerous variations, whether explicitly given in the specification or not, such as differences in structure, dimension, and use of material, are possible. The scope of embodiments is at least as broad as given by the following claims.
[0083] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any component(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or component of any or all the claims.
Claims
1.A method for membrane excursion control of an electronic apparatus, the method comprising:receiving an audio signal input from one or more audio sources;extracting one or more audio parameters derived from the audio signal input, wherein the one or more audio parameters include one or more speaker characteristics and signal features;inputting the one or more audio parameters into a distortion classifier to identify a first set of audio frames in the audio signal input causing audible distortion, wherein the distortion classifier is trained using a synthetic audio dataset generated using a diffusion model with custom loss function; andgenerating a first output, comprising a second set of audio frames, by inputting the identified first set of audio frames causing audio distortion into a plurality of speaker models, wherein the second set of audio frames are distortion less, and wherein the plurality of speaker models are artificial intelligent (AI) models trained using the synthetic audio dataset.2.The method as claimed in claim 1, wherein the plurality of speaker models include at least one of a first speaker model and a second speaker model.3.The method as claimed in claim 1, wherein generating the first output, the method comprising:generating a second output by inputting the identified first set of audio frames into the first speaker model; andgenerating the first output by inputting the second output to the second speaker model.4.The method as claimed in claim 3, wherein generating the second output, the method comprising:determining an acoustic pressure of the identified first set of audio frames; andgenerating the second output based on the first set of audio frames, wherein the second output indicates the determined acoustic pressure of the identified first set of audio frames.5.The method as claimed in claim 3, wherein generating the first output, the method comprising:providing the acoustic pressure of the identified first set of audio frames to the second speaker model; andgenerating the first output based on the acoustic pressure of the identified first set of audio frames, wherein the first output indicates the second set of audio frames.6.The method as claimed in claim 1, generating the synthetic audio dataset, the method comprising:training the diffusion model with the custom loss function to generate an audio waveform with reduced distortion, wherein the custom loss function is based on differentiable functions associated with the one or more audio quality parameters; andgenerating the synthetic audio dataset using the trained diffusion model.7.The method as claimed in claim 1, wherein the training of the diffusion model with custom loss function comprising:receiving a first set of audio tracks by the diffusion model;selecting a second set of audio tracks based on audio quality metrics derived from the custom loss functions such that the identified second set of audio tracks is free from distortion in terms of sound quality;inputting the selected second set of audio tracks to an equalization filter, wherein the equalization filter is designed to emulate a perceptual filter bank-inspired filter, with filter parameters adjustable by listeners;determining ranks of the second set of audio tracks based on audio quality metrics, wherein the audio quality metrics of the second set of audio tracks are determined based on testing of the second set of audio tracks; andselecting at least one of the second set of audio tracks such that the selected at least one of the second set of audio tracks is free from distortion.8.The method as claimed in claim 1, wherein the first speaker model is an ideal speaker model trained to estimate the acoustic pressure of the identified first set of audio frames, and wherein the second speaker model is an inverse speaker model trained to generate the first output using the estimated acoustic pressure of the identified first set of audio frames.9.An electronic apparatus for controlling membrane excursion, comprises:a memory configured to store executable instructions;at least one processor coupled to the memory and configured to execute the instructions to:receive an audio signal input from one or more audio sources;extract one or more audio parameters derived from the audio signal input, wherein the one or more audio parameters include one or more speaker characteristics and signal features;input the one or more audio parameters into a distortion classifier to identify a first set of audio frames in the audio signal input causing audible distortion, wherein the distortion classifier is trained using a synthetic audio dataset generated using a diffusion model with custom loss function; andgenerate a first output, comprising a second set of audio frames, by inputting the identified first set of audio frames causing audio distortion into a plurality of speaker models, wherein the second set of audio frames are distortion less, and wherein the plurality of speaker models are artificial intelligent (AI) models trained using the synthetic audio dataset.10.The electronic apparatus of claim 9, wherein the plurality of speaker models include at least one of a first speaker model and a second speaker model.11.The electronic apparatus of claim 9, wherein the at least one processor is further configured to:generate a second output by inputting the identified first set of audio frames into the first speaker model; andgenerate the first output by inputting the second output to the second speaker model.12.The electronic apparatus of claim 11, wherein the at least one processor is further configured to:determine an acoustic pressure of the identified first set of audio frames; andgenerate the second output based on the first set of audio frames, wherein the second output indicates the determined acoustic pressure of the identified first set of audio frames.13.The electronic apparatus of claim 11, whereinthe at least one processor is further configured to:provide the acoustic pressure of the identified first set of audio frames to the second speaker model; andgenerate the first output based on the acoustic pressure of the identified first set of audio frames, wherein the first output indicates the second set of audio frames.14.The electronic apparatus of claim 9, wherein the at least one processor is further configured to:train the diffusion model with the custom loss function to generate an audio waveform with reduced distortion, wherein the custom loss function is based on differentiable functions associated with the one or more audio quality parameters; andgenerate the synthetic audio dataset using the trained diffusion model.15.The electronic apparatus of claim 9, wherein the at least one processor is further configured to:receive a first set of audio tracks by the diffusion model;select a second set of audio tracks based on audio quality metrics derived from the custom loss functions such that the identified second set of audio tracks is free from distortion in terms of sound quality;input the selected second set of audio tracks to an equalization filter, wherein the equalization filter is designed to emulate a perceptual filter bank-inspired filter, with filter parameters adjustable by listeners;determine ranks of the second set of audio tracks based on audio quality metrics, wherein the audio quality metrics of the second set of audio tracks are determined based on testing of the second set of audio tracks; andselect at least one of the second set of audio tracks such that the selected at least one of the second set of audio tracks is free from distortion.
Citation Information
Patent Citations
Method and audio device for volume control in speaker
KR100835955B1
Method and apparatus for compensating for nonlinear distortion of speaker system
US20050047606A1
Auto-equalization, in-room low-frequency sound power optimization
US20180262175A1
Energy limiter for loudspeaker protection
US20190281385A1
Reluctance force compensation for loudspeaker control
US20200344548A1