A method of audio processing based on a remote hearing aid
By processing audio signals using a hydrodynamic field theory model in a remote hearing aid system, the problems of signal transmission delay and signal distortion are solved, thereby improving audio-visual synchronization and auditory comfort.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAMEN WENATONE MEDICAL TECH CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-28
AI Technical Summary
Remote hearing aid systems suffer from auditory lag due to signal transmission delays, as well as signal distortion and artifacts introduced by traditional noise reduction methods.
Audio signals are acquired by the wearable audio acquisition and playback unit, short-time Fourier transform is performed and converted into binaural complex spectrum data, physical property mapping is performed by the remote fluid computing core unit to generate a dynamic vector field, differential geometric operations and adaptive viscous sound field rectification are performed, the acoustic spectrum of the next moment is predicted and a frequency domain gain mask is generated, and the data is sent to the wearable unit for processing in real time.
It achieves audio-visual synchronization, eliminates signal distortion, improves auditory comfort and naturalness, and reduces the misjudgment rate of speech recognition.
Smart Images

Figure CN121462960B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal processing technology, specifically to an audio processing method based on a remote hearing aid. Background Technology
[0002] With the rapid development of hearing aid technology, remote audio processing architectures based on cloud or mobile terminal computing power are becoming increasingly popular, aiming to utilize powerful computing resources to cope with complex acoustic environments. Although this architecture improves processing capabilities, it faces severe challenges in real-time signal transmission and processing in practical applications.
[0003] Currently, conventional remote hearing aid systems transmit audio data via wireless links. The physical transmission link and complex signal processing algorithms inevitably introduce time delays, causing the auditory signal perceived by the wearer to lag behind the visual environment, thus disrupting the sense of audio-visual synchronization. In addition, traditional noise reduction methods mostly rely on spectrum truncation techniques based on hard thresholds. This discontinuous processing method is prone to causing spectral structure damage, producing artifacts such as musical noise, and it is difficult to accurately remove high-energy non-steady-state noise while preserving the naturalness of speech. Therefore, how to overcome the auditory lag caused by physical delays in remote hearing aid scenarios and eliminate signal distortion caused by hard spectrum truncation has become an urgent problem to be solved in this field. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides an audio processing method based on a remote hearing aid. Specifically, the technical solution of this invention includes:
[0005] The wearable audio acquisition and playback unit acquires the original audio signal, performs a short-time Fourier transform on the original audio signal to obtain binaural complex spectrum data, and transmits the binaural complex spectrum data to the remote fluid computing core unit.
[0006] The remote fluid computing core unit performs physical property mapping operations on the received binaural complex spectrum data, converting the binaural complex spectrum data into a two-dimensional compressible fluid field, and generating a dynamic vector field containing energy distribution and motion trends.
[0007] Perform differential geometric operations on the dynamic vector field to calculate the vorticity value in the local time-frequency region. Based on the vorticity value, identify the distribution areas of speech signals and noise signals, and generate a flow pattern distribution map.
[0008] A dynamic coupling mechanism for the viscosity of the medium driven by the flow pattern distribution is established. Adaptive viscous acoustic field rectification is performed on the dynamic vector field according to the flow pattern distribution diagram. The energy of the noise signal distribution area is eliminated by the physical diffusion process, and the cleanroom flow field data is output.
[0009] The momentum and inertia of the speech signal in the cleanroom flow field data are extracted, and fluid advection prediction calculation is performed based on the momentum and inertia to predict the dynamic vector field at the next moment.
[0010] The predicted dynamic vector field is subjected to acoustic inverse mapping to reconstruct the acoustic spectrum at the prediction time, and the frequency domain gain mask is calculated based on the acoustic spectrum at the prediction time.
[0011] The frequency domain gain mask is sent back to the wearable audio acquisition and playback unit in real time, and the wearable audio acquisition and playback unit processes and plays the real-time acquired audio by applying the frequency domain gain mask.
[0012] Preferably, the received binaural complex spectrum data is subjected to a physical property mapping operation by a remote fluid computing core unit, converting the binaural complex spectrum data into a two-dimensional compressible fluid field, generating a dynamic vector field containing energy distribution and motion trends, including:
[0013] The energy amplitude of the binaural complex spectrum data is extracted and defined as the density property of the fluid.
[0014] Extract the rate of change of binaural complex spectrum data on the time and frequency axes, and define the rate of change as the velocity vector property of the fluid;
[0015] Based on density and velocity vector properties, a dynamic vector field is constructed, which serves as the computational foundation for subsequent fluid dynamics operators.
[0016] Preferably, differential geometric operations are performed on the dynamic vector field to calculate the vorticity value in the local time-frequency region. Based on the vorticity value, the distribution regions of the speech signal and the noise signal are identified, and a flow regime distribution map is generated, including:
[0017] Calculate the vorticity value of each local time-frequency region in the dynamic vector field, and use the vorticity value as the basis for feature discrimination;
[0018] Feature discrimination of local time-frequency regions based on vorticity values;
[0019] When a local time-frequency region exhibits low vorticity and strong momentum characteristics, it is determined that the local time-frequency region is a laminar flow region corresponding to the speech signal.
[0020] When a local time-frequency region exhibits high vorticity and weak momentum characteristics, it is determined that the local time-frequency region is a turbulent region corresponding to the noise signal.
[0021] Based on the results of feature discrimination, a flow state distribution map is generated, which serves as a mask representation of noise and speech signals.
[0022] Preferably, a dynamic coupling mechanism is established to drive the viscosity of the medium based on the flow pattern distribution. Adaptive viscous acoustic field rectification is performed on the dynamic vector field according to the flow pattern distribution diagram. The energy in the noise signal distribution area is eliminated using a physical diffusion process, and cleanroom flow field data is output, including:
[0023] For coordinate points identified as turbulent regions in the flow distribution diagram, the virtual dynamic viscosity coefficient of the coordinate points in the turbulent regions is increased, forcing the energy of the coordinate points in the turbulent regions to be smoothly dissipated to the surroundings through the diffusion equation.
[0024] For coordinate points in the flow distribution map that are identified as laminar flow regions, the virtual dynamic viscosity coefficient of the coordinate points in the laminar flow regions is reduced, allowing the signals of the coordinate points in the laminar flow regions to pass through without loss due to their own momentum;
[0025] The diffusion process ensures a smooth transition of spectral energy in both time and frequency dimensions, generating clean laminar flow field data with noise energy removed.
[0026] Preferably, the momentum and inertia of the speech signal are extracted from the cleanroom flow field data, and fluid advection prediction calculation is performed based on the momentum and inertia to predict the dynamic vector field at the next moment, including:
[0027] Momentum of speech signals is extracted from cleanroom flow field data; momentum is the trend vector of fluid density evolution over time.
[0028] Based on the current velocity vector properties and momentum, perform fluid advection prediction calculations to obtain the dynamic vector field for the next frame.
[0029] By leveraging the continuity of fluids, fluid advection prediction calculations can offset the physical time delay caused by remote transmission and processing.
[0030] Preferably, an acoustic inverse mapping is performed on the predicted dynamic vector field to reconstruct the acoustic spectrum at the prediction time, and a frequency domain gain mask is calculated based on the acoustic spectrum at the prediction time, including:
[0031] Perform the inverse operation of physical property mapping on the predicted dynamic vector field for the next frame.
[0032] By performing inverse operations, the density values in the dynamic vector field at the predicted next frame time are restored to the energy amplitude of the spectrum;
[0033] By performing inverse operations, the velocity vector properties in the predicted dynamic vector field of the next frame are restored to the phase change rate of the spectrum;
[0034] Based on the energy amplitude and phase change rate obtained from the reconstruction, the acoustic spectrum at the predicted moment is reconstructed;
[0035] A frequency domain gain mask is generated based on the reconstructed acoustic spectrum at the predicted time.
[0036] Preferably, the frequency domain gain mask is sent back to the wearable audio acquisition and playback unit in real time, and the wearable audio acquisition and playback unit processes and plays the real-time acquired audio by applying the frequency domain gain mask, including:
[0037] The frequency domain gain mask is sent to the wearable audio acquisition and playback unit in advance;
[0038] The wearable audio acquisition and playback unit applies a frequency domain gain mask to the actual sound signal being acquired.
[0039] By applying frequency domain gain masks, zero-latency synchronization in terms of auditory perception is achieved.
[0040] Compared with the prior art, the present invention has the following beneficial effects:
[0041] 1. This method extracts the momentum and inertia of the speech signal from the cleanroom flow field data and uses fluid advection prediction to calculate the dynamic vector field at the next moment, thus constructing a negative delay signal processing system. This mechanism utilizes the predictability provided by the law of conservation of fluid momentum to perform preemptive calculation on the time axis, sending the frequency domain gain mask to the wearable audio acquisition and playback unit in advance before the passage of physical time. This completely cancels the physical time delay caused by remote transmission and processing at the user's auditory perception level, achieving audio-visual synchronization and effectively solving the lag problem in distributed hearing aid systems.
[0042] 2. This method establishes an audio processing architecture based on a fluid dynamics field theory model, transforming the discrete signal processing process into a physical evolution process of a continuous medium. It replaces the hard threshold cutoff of traditional filtering with a physical diffusion process. Through the dynamic coupling mechanism of fluid distribution driving medium viscosity, the virtual diffusion coefficient is increased for noise distribution areas, and the dissipation term of the diffusion equation forces energy to attenuate smoothly in all directions. This processing method ensures a smooth transition of spectral energy in the time and frequency dimensions, eliminates music noise artifacts caused by spectral discontinuity, and makes the residual background noise sound naturally attenuated, significantly improving the listening comfort and naturalness.
[0043] 3. This method introduces differential geometric operators to analyze the dynamic vector field, uses vorticity values to quantify the phase inconsistency and disorder of the signal in the time-frequency plane, and generates a flow distribution map to identify speech and noise. Compared with traditional detection methods based on energy thresholds, vorticity-based fluid topology classification has extremely high sensitivity to non-steady-state noise. Even in high-energy noise scenarios, it can accurately eliminate noise interference by identifying high vorticity features, thereby significantly reducing the false recognition rate of speech recognition and effectively improving the signal-to-noise ratio in complex acoustic environments.
[0044] 4. This method performs physical property mapping operations, assigning virtual physical properties to dimensionless spectral data, mapping energy amplitude to fluid density, and mapping the rate of change of time frequency to flow velocity vector, thus constructing a two-dimensional compressible fluid field. This deterministic physical mapping allows the signal processing process to inherit the mathematical constraints of continuity and conservation in fluid mechanics, ensuring the smoothness of the signal envelope in subsequent processing, avoiding phase abrupt changes caused by traditional discrete processing, and providing a solid computational foundation for high-precision adaptive viscous sound field rectification and momentum prediction in virtual physical space. Attached Figure Description
[0045] The present invention will be further explained below with reference to the accompanying drawings and embodiments:
[0046] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.
[0048] Example 1:
[0049] Please see Figure 1 An audio processing method based on a remote hearing aid includes:
[0050] The wearable audio acquisition and playback unit acquires the original audio signal, performs a short-time Fourier transform on the original audio signal to obtain binaural complex spectrum data, and transmits the binaural complex spectrum data to the remote fluid computing core unit.
[0051] The remote fluid computing core unit performs physical property mapping operations on the received binaural complex spectrum data, converting the binaural complex spectrum data into a two-dimensional compressible fluid field, and generating a dynamic vector field containing energy distribution and motion trends.
[0052] Perform differential geometric operations on the dynamic vector field to calculate the vorticity value in the local time-frequency region. Based on the vorticity value, identify the distribution areas of speech signals and noise signals, and generate a flow pattern distribution map.
[0053] A dynamic coupling mechanism for the viscosity of the medium driven by the flow pattern distribution is established. Adaptive viscous acoustic field rectification is performed on the dynamic vector field according to the flow pattern distribution diagram. The energy of the noise signal distribution area is eliminated by the physical diffusion process, and the cleanroom flow field data is output.
[0054] The momentum and inertia of the speech signal in the cleanroom flow field data are extracted, and fluid advection prediction calculation is performed based on the momentum and inertia to predict the dynamic vector field at the next moment.
[0055] The predicted dynamic vector field is subjected to acoustic inverse mapping to reconstruct the acoustic spectrum at the prediction time, and the frequency domain gain mask is calculated based on the acoustic spectrum at the prediction time.
[0056] The frequency domain gain mask is sent back to the wearable audio acquisition and playback unit in real time, and the wearable audio acquisition and playback unit processes and plays the real-time acquired audio by applying the frequency domain gain mask.
[0057] A distributed audio processing architecture based on a fluid dynamics field theory model is constructed. The core logic of this architecture is to transform the discrete signal processing process into the physical evolution process of a continuous medium in order to solve the technical contradiction of signal distortion and transmission delay coexisting in remote hearing aid scenarios.
[0058] The wearable audio acquisition and playback unit is responsible for capturing ambient acoustic signals in real time. The unit is equipped with a digital signal processor and is programmed to perform a short-time Fourier transform on the acquired raw audio signal. During this process, the system sets a specific frame length and overlap rate to convert the time-domain waveform into binaural complex spectrum data containing amplitude and phase information. To reduce the transmission bandwidth usage, the unit transmits the above frequency domain data to the remote fluid computing core unit via Bluetooth or a proprietary wireless protocol.
[0059] After receiving the data, the remote fluid computing core unit initiates a physical property mapping operation. This operation is not only a conversion of the data format, but also a migration of the computing domain. The system uses a specific mapping algorithm to assign virtual physical properties to the dimensionless spectral data, constructing a two-dimensional compressible fluid field. In this field, the audio signal is no longer regarded as an independent pixel, but as a fluid medium that follows the continuity equation, thereby generating a dynamic vector field containing energy density distribution and velocity vector field.
[0060] The system introduces differential geometric operators to analyze the dynamic vector field; for each local time-frequency unit in the field, the system calculates its vorticity value; the vorticity value is defined as the curl of the flow velocity vector, which quantifies the degree of phase inconsistency and disorder of the signal in the time-frequency plane; based on this quantification index, the system performs binary classification logic: regions exhibiting low curl characteristics are marked as speech signal distribution regions, regions exhibiting high curl characteristics are marked as noise signal distribution regions, and a flow distribution map describing the signal properties of the entire field is generated accordingly;
[0061] Based on the above distribution map, the system activates a dynamic coupling mechanism that drives the viscosity of the medium by adjusting the physical parameters of the virtual medium. This mechanism performs signal processing by adjusting the physical parameters of the virtual medium. For regions marked as noise, the system adjusts the diffusion coefficient at that location to a high level and forces energy attenuation using the dissipation term of the physical diffusion equation. For speech regions, the diffusion coefficient is suppressed to a low level, and the convection term is used to maintain signal integrity. This process is called adaptive viscous sound field rectification, and its output is cleanroom flow field data with turbulence interference removed.
[0062] To compensate for the physical delay caused by the remote link, the system enters the prediction phase; from the cleanroom flow field data, the system analyzes the momentum inertia of the voice signal, that is, the gradient trend of the fluid density evolution over time; based on this physical inertia, the system performs fluid advection prediction calculation to calculate the fluid field state at future moments.
[0063] The system performs acoustic inverse mapping on the predicted dynamic vector field, restores the fluid parameters to the acoustic spectrum, and calculates the frequency domain gain mask for gain control. This mask is sent back to the front end in real time. When the wearable audio acquisition and playback unit receives and applies the mask, its actual effective time point coincides precisely with the time point when the ambient sound reaches the eardrum on the time axis, thus achieving zero-delay synchronization at the auditory perception level.
[0064] This solution introduces fluid dynamics equations and utilizes the isotropic characteristics of physical diffusion to replace the hard threshold cutoff of traditional filtering, eliminating music noise caused by spectral discontinuities. By leveraging the predictability provided by the law of conservation of fluid momentum, it achieves negative delay compensation based on physical laws, effectively solving the problem of audio-visual synchronization in distributed hearing aid systems.
[0065] Example 2:
[0066] The remote fluid computing core unit performs physical property mapping operations on the received binaural complex spectrum data, converting the binaural complex spectrum data into a two-dimensional compressible fluid field, generating a dynamic vector field containing energy distribution and motion trends, including:
[0067] The energy amplitude of the binaural complex spectrum data is extracted and defined as the density property of the fluid.
[0068] Extract the rate of change of binaural complex spectrum data on the time and frequency axes, and define the rate of change as the velocity vector property of the fluid;
[0069] Based on density and velocity vector properties, a dynamic vector field is constructed, which serves as the computational foundation for subsequent fluid dynamics operators.
[0070] This paper details the implementation logic and parameter definition methods of the physical property mapping operation, aiming to construct a virtual physical computing environment that accurately corresponds to acoustic characteristics. It should be noted that the system independently executes the following physical property mapping and subsequent fluid calculation processes for the spectral data of the left and right ears, generating their respective corresponding fluid field data. The following description uses single-channel processing as an example:
[0071] The remote fluid computing core unit processes the input binaural complex spectrum data. Perform modulus calculation to extract the energy amplitude that represents the signal strength. The system performs a differential operation on the spectrum data, calculating its time dimension separately. and frequency dimension Partial derivatives on the time axis; to eliminate the difference in physical dimensions between the time axis and the frequency axis, the system pre-performs coordinate normalization; a time normalization constant is set. For frame shift time, the frequency normalization constant is... Frequency interval; define dimensionless time coordinates. and dimensionless frequency coordinates All subsequent physics calculations are performed in It is performed in a dimensionless space; thereby obtaining the rate of change of the spectral data on the time and frequency axes.
[0072] The system establishes a strict mapping relationship: the energy amplitude is directly assigned as the density attribute of the virtual fluid. This establishes an equivalent relationship between acoustic energy and fluid mass; the vector formed by the rate of change... Assign the value to the velocity vector attribute of the virtual fluid. This characterizes the direction and rate of signal evolution in the time and frequency domain;
[0073] To construct a dynamic vector field capable of characterizing the evolution of acoustic textures and possessing non-zero curl, the system defines a scalar field. Unlike the gradient definition, which can only generate irrotational fields, this system uses the optical flow method to calculate the velocity vector. It is assumed that acoustic energy follows the brightness conservation assumption between consecutive spectral frames, i.e., the total derivative. Then we have the optical flow equation:
[0074]
[0075] in, and These represent the spatial gradients of the spectral energy along the normalized time axis and frequency axis, respectively. This represents the difference between the current frame and the previous frame; the system locally... Within the neighborhood, the Lucas-Kanade algorithm is used to establish and solve an overdetermined system of equations. Specifically, the system in the local neighborhood... Internally, by minimizing the weighted residuals Seeking a solution :
[0076]
[0077] in, Weights defined for the Gaussian kernel function; by... beg By taking the partial derivatives of the equations and setting them to zero, we obtain the system of linear equations. :
[0078]
[0079] Solving this system of equations will yield the result. ;
[0080] Obtain the flow velocity vector This vector represents the true direction and speed of motion of the spectral energy texture on the time-frequency plane, and it can naturally generate non-zero vorticity components.
[0081] Based on the definitions of the two core physical quantities mentioned above, the system constructs a continuous dynamic vector field through a gridded interpolation algorithm; in this field, any coordinate point All of them have clearly defined density scalars and velocity vectors; this dynamic vector field is not only a collection of data, but also the direct computational foundation for all subsequent fluid dynamics operators;
[0082] Through this deterministic physical mapping, massless and inertial acoustic signals are transformed into fluid media with virtual mass and virtual momentum. This transformation allows the signal processing process to inherit the mathematical constraints of continuity and conservation in fluid mechanics, ensuring the smoothness of the signal envelope in subsequent processing and avoiding phase abrupt changes caused by traditional discrete processing.
[0083] Example 3:
[0084] Perform differential geometric operations on the dynamic vector field to calculate the vorticity value in the local time-frequency region. Based on the vorticity value, identify the distribution regions of the speech signal and noise signal, and generate a flow regime distribution map, including:
[0085] Calculate the vorticity value of each local time-frequency region in the dynamic vector field, and use the vorticity value as the basis for feature discrimination;
[0086] Feature discrimination of local time-frequency regions based on vorticity values;
[0087] When a local time-frequency region exhibits low vorticity and strong momentum characteristics, it is determined that the local time-frequency region is a laminar flow region corresponding to the speech signal.
[0088] When a local time-frequency region exhibits high vorticity and weak momentum characteristics, it is determined that the local time-frequency region is a turbulent region corresponding to the noise signal.
[0089] Based on the results of feature discrimination, a flow state distribution map is generated, which serves as a mask representation of noise and speech signals.
[0090] The feature discrimination logic adopts a nonlinear classification method based on fluid topology, the core of which is to use vorticity as a physical parameter to characterize the coherence of the signal;
[0091] The system performs curl calculations on the dynamic vector field; for each local time-frequency grid cell, it calculates its velocity vector. curl The resulting modulus value is the vorticity value of that region. This parameter quantifies the intensity of the fluid element's rotation around its own axis, reflecting the degree of randomness of the local phase at the signal level.
[0092] The classification threshold is not a fixed value, but is adaptively calibrated based on the current ambient noise floor; the system extracts the preceding information. The frame is used as the initial window, and the mean value of the vorticity across the entire field within that window is calculated. and standard deviation Set vorticity threshold ,in, In this embodiment, the confidence coefficient is used. The preferred value range is... to To adapt to sensitivity requirements under different signal-to-noise ratio environments; at the same time, define the local momentum modulus. Before calculation Momentum magnitude values of all grid points in the frame The average value is denoted as Momentum threshold The design is based on statistical analysis of the momentum characteristics of clean speech signals. Experimental results show that speech signals typically possess higher-than-average momentum. Based on momentum and inertia, this setting can effectively distinguish speech signals with clear motion trends from disordered background noise; the system is based on a set vortex threshold. With momentum threshold Perform logical checks across the entire field:
[0093] When the calculation results for a certain local region satisfy and When the flow is orderly and the energy has a clear directionality, it is consistent with the characteristics of a dense harmonic structure and continuous phase of a speech signal; the system determines it to be a laminar flow region.
[0094] When a certain local area satisfies and When the signal is strong, it indicates that there are strong local vortices in the region and a lack of macroscopic flow trends, which is consistent with the characteristics of chaotic and disordered phase of environmental noise; the system identifies it as a background turbulent region.
[0095] Based on the overall judgment results, the system generates a binary or continuous gradient flow distribution map; this map, as a physical mask, precisely defines the signal region that needs to be retained and the noise region that needs to be suppressed in subsequent processing.
[0096] Compared to traditional energy threshold-based speech activity detection, the vorticity feature introduced in this scheme has extremely high sensitivity to non-steady-state noise. Even in high-energy noise scenarios, it can still be accurately identified and eliminated due to its high vorticity feature, thus significantly reducing the false positive rate of speech recognition and improving the signal-to-noise ratio in complex acoustic environments. To achieve smoother gradual control of medium parameters, this embodiment further refines and expands upon the aforementioned binary classification:
[0097] In addition to the two typical states mentioned above, for those that satisfy and The system classifies the region as a transition region; for regions that meet the requirements... and The system classifies the region as a region of strong turbulence; the system establishes a system from flow regime classification to virtual viscosity coefficient. Complete mapping table: Laminar flow region correspondence Turbulent and strongly turbulent regions correspond to Linear interpolation is used in the transition region. This is to ensure the continuity of the flow field medium parameters.
[0098] Example 4:
[0099] A dynamic coupling mechanism is established to drive the viscosity of the medium based on the flow pattern distribution. Adaptive viscous acoustic field rectification is performed on the dynamic vector field according to the flow pattern distribution diagram. The energy in the noise signal distribution area is eliminated using a physical diffusion process, and cleanroom flow field data is output, including:
[0100] For coordinate points identified as turbulent regions in the flow distribution diagram, the virtual dynamic viscosity coefficient of the coordinate points in the turbulent regions is increased, forcing the energy of the coordinate points in the turbulent regions to be smoothly dissipated to the surroundings through the diffusion equation.
[0101] For coordinate points in the flow distribution map that are identified as laminar flow regions, the virtual dynamic viscosity coefficient of the coordinate points in the laminar flow regions is reduced, allowing the signals of the coordinate points in the laminar flow regions to pass through without loss due to their own momentum;
[0102] The diffusion process ensures a smooth transition of spectral energy in both time and frequency dimensions, generating clean laminar flow field data with noise energy removed.
[0103] The adaptive viscous sound field rectification module achieves precise rectification of signal energy by adjusting the physical parameters of the virtual medium; the working mechanism of this module relies on the tight coupling of data flow and control flow.
[0104] The system defines a spatially varying variable—the virtual dynamic viscosity coefficient. This variable simultaneously serves as both the dynamic viscosity coefficient in fluid mechanics and the diffusion coefficient in the physical diffusion equation, used to uniformly control the viscous dissipation and energy diffusion processes of the fluid.
[0105] The value of this coefficient establishes a functional mapping relationship with the classification results in the flow regime distribution map:
[0106] For coordinate points in the flow regime distribution map that are identified as turbulent regions, the system will Set to an extremely high value This high dynamic viscosity coefficient significantly enhances the diffusion term. The weight of the turbulent energy density in the region causes a strong damping effect, which physically forces the turbulent energy density in the region to dissipate rapidly and smoothly into the surrounding low-energy regions and eventually annihilate, in accordance with the second law of thermodynamics.
[0107] For coordinate points identified as being within the laminar flow region, the system will Set to an extremely low value close to zero Under these conditions, the fluid approaches an ideal fluid state, the dynamic viscosity is negligible, and the voice signal can pass through the flow field without loss by relying on its own momentum, maintaining its original spectral structure and energy intensity.
[0108] The system on the virtual timeline Upper fluid density field implement This iteration updates the process to achieve the physical diffusion process; the number of iterations... Settings and system physical latency Relevant, and should satisfy the iteration stopping condition:
[0109] Based on delay time: The lower limit should guarantee the virtual time step. ;
[0110] Based on stable convergence: the system sets a very small convergence threshold. When the density field changes between two adjacent iterations The iteration terminates at that time.
[0111] Number of iterations Usually taken arrive This allows for a sufficiently smooth diffusion rectification effect while ensuring numerical stability; the specific discretized evolution equation is expressed using a five-point difference scheme as follows:
[0112]
[0113] in, and All are normalized unit grid step sizes; This is the anisotropic balance coefficient, used to compensate for the difference between the time axis and the frequency axis on the physical scale, and its value is [value missing]. ; For the virtual evolution time step, the numerical stability condition must be met. ;
[0114] Through the above-described rectification process based on physical diffusion, the system eliminates noise while utilizing the continuity characteristics of the diffusion equation to form a natural transition gradient at the boundary between noise and speech; this processing logic ultimately outputs cleanroom flow field data.
[0115] This mechanism uses physical diffusion instead of traditional hard threshold shearing. Since the diffusion process is mathematically continuous and differentiable, the processed spectrum has no breaks or holes, eliminating common music noise artifacts in digital signal processing. This makes the residual background noise sound like it has decayed naturally, greatly improving the comfort and naturalness of the listening experience.
[0116] Example 5:
[0117] The momentum-inertia of the speech signal in the cleanroom flow field data is extracted. Based on the momentum-inertia, fluid advection prediction calculations are performed to predict the dynamic vector field at the next moment, including:
[0118] Momentum of speech signals is extracted from cleanroom flow field data; momentum is the trend vector of fluid density evolution over time.
[0119] Based on the current velocity vector properties and momentum, perform fluid advection prediction calculations to obtain the dynamic vector field for the next frame.
[0120] By leveraging the continuity of fluids, fluid advection prediction calculations can offset the physical time delay caused by remote transmission and processing.
[0121] The predicted dynamic vector field is subjected to acoustic inverse mapping to reconstruct the acoustic spectrum at the prediction time, and a frequency domain gain mask is calculated based on the acoustic spectrum at the prediction time, including:
[0122] Perform the inverse operation of physical property mapping on the predicted dynamic vector field for the next frame.
[0123] By performing inverse operations, the density values in the dynamic vector field at the predicted next frame time are restored to the energy amplitude of the spectrum;
[0124] By performing inverse operations, the velocity vector properties in the predicted dynamic vector field of the next frame are restored to the phase change rate of the spectrum;
[0125] Based on the energy amplitude and phase change rate obtained from the reconstruction, the acoustic spectrum at the predicted moment is reconstructed;
[0126] A frequency domain gain mask is generated based on the reconstructed acoustic spectrum at the predicted time.
[0127] The frequency domain gain mask is sent back to the wearable audio acquisition and playback unit in real time. The wearable audio acquisition and playback unit then applies the frequency domain gain mask to process and play the real-time acquired audio, including:
[0128] The frequency domain gain mask is sent to the wearable audio acquisition and playback unit in advance;
[0129] The wearable audio acquisition and playback unit applies a frequency domain gain mask to the actual sound signal being acquired.
[0130] By applying frequency domain gain masks, zero-latency synchronization in terms of auditory perception is achieved.
[0131] The core of this mechanism lies in using the inertial characteristics of the fluid model to make short-term predictions in order to offset the inherent physical delay of the system.
[0132] Semi-Lagrange inverse tracing is performed based on the normalized velocity vector; the total system delay is known to be... Convert it to a normalized time span Based on the semi-Lagrange inverse tracking principle, the time is predicted. Spectral energy density The calculation formula is:
[0133]
[0134] in, The frequency-axis normalized velocity components are obtained through optical flow calculation. This represents the distance the spectral texture moves along the frequency axis during the delay time; this formula ensures that both the subtrahend and minuend are dimensionless coordinate values, thus correctly canceling transmission delay at the physical level; when the coordinate points are traced in reverse... When the coordinates are not on integer grid coordinates, the system uses bilinear interpolation to calculate. The method uses a distance-weighted average of the four nearest grid points to ensure the density field's numerical value; Spatial continuity allows for the accurate acquisition of the required fluid density values for prediction;
[0135] The system performs acoustic inverse mapping; it converts the predicted density values in the fluid field into... The energy amplitude is restored to the acoustic spectrum, and the flow velocity vector is converted. The phase change rate of the spectrum is restored; through this inverse transformation, the system mathematically reconstructs the acoustic spectrum at the predicted time; based on the predicted spectrum, the system calculates the corresponding frequency domain gain mask, which precisely indicates the gain adjustment amount at each frequency point in the future time.
[0136] Based on the predicted clean spectrum energy And the noise energy estimate at the current moment obtained from the statistics of the turbulent regions identified in the flow distribution map of step (3). Calculate the frequency domain gain mask The mask calculation uses a soft threshold function to avoid abrupt changes in auditory perception. The specific formula is as follows:
[0137]
[0138] in, This is an overestimation factor for noise. The spectral attenuation index; this is the gain mask. The range of values is strictly limited to Within the range, the gain adjustment ratio for each frequency point at future times is precisely indicated;
[0139] The remote core unit has completed the passage of physical time. Previously, the frequency domain gain mask was sent in advance to the wearable audio acquisition and playback unit via a wireless link; when the wearable unit received the instruction, the actual physical time had just elapsed. At this point, the predicted time corresponding to the received mask completely coincides with the wearer's current actual time; the wearing unit then directly applies the mask to the currently acquired sound signal.
[0140] This solution constructs a negative delay signal processing system. By preemptively calculating on the time axis, the system enables the processing parameters to arrive at the terminal in advance to wait for the audio signal, thereby completely offsetting the lag caused by the remote processing link in the user's subjective hearing, achieving audio-visual synchronization, and making remote hearing aid solutions based on the powerful computing power of the cloud or mobile device truly practical.
[0141] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An audio processing method based on a remote hearing aid, characterized in that, include: The wearable audio acquisition and playback unit acquires the original audio signal, performs a short-time Fourier transform on the original audio signal to obtain binaural complex spectrum data, and transmits the binaural complex spectrum data to the remote fluid computing core unit. The remote fluid computing core unit performs physical property mapping operations on the received binaural complex spectrum data, converting the binaural complex spectrum data into a two-dimensional compressible fluid field, and generating a dynamic vector field containing energy distribution and motion trends. Perform differential geometric operations on the dynamic vector field to calculate the vorticity value in the local time-frequency region. Based on the vorticity value, identify the distribution areas of speech signals and noise signals, and generate a flow pattern distribution map. A dynamic coupling mechanism for the viscosity of the medium driven by the flow pattern distribution is established. Adaptive viscous acoustic field rectification is performed on the dynamic vector field according to the flow pattern distribution diagram. The energy of the noise signal distribution area is eliminated by the physical diffusion process, and the cleanroom flow field data is output. The momentum and inertia of the speech signal in the cleanroom flow field data are extracted, and fluid advection prediction calculation is performed based on the momentum and inertia to predict the dynamic vector field at the next moment. The predicted dynamic vector field is subjected to acoustic inverse mapping to reconstruct the acoustic spectrum at the prediction time, and the frequency domain gain mask is calculated based on the acoustic spectrum at the prediction time. The frequency domain gain mask is sent back to the wearable audio acquisition and playback unit in real time, and the wearable audio acquisition and playback unit processes and plays the real-time acquired audio by applying the frequency domain gain mask.
2. The audio processing method based on a remote hearing aid according to claim 1, characterized in that, The remote fluid computing core unit performs physical property mapping operations on the received binaural complex spectrum data, converting the binaural complex spectrum data into a two-dimensional compressible fluid field, generating a dynamic vector field containing energy distribution and motion trends, including: The energy amplitude of the binaural complex spectrum data is extracted and defined as the density property of the fluid. Extract the rate of change of binaural complex spectrum data on the time and frequency axes, and define the rate of change as the velocity vector property of the fluid; A dynamic vector field is constructed based on density and velocity vector properties.
3. The audio processing method based on a remote hearing aid according to claim 1, characterized in that, Perform differential geometric operations on the dynamic vector field to calculate the vorticity value in the local time-frequency region. Based on the vorticity value, identify the distribution regions of the speech signal and noise signal, and generate a flow regime distribution map, including: Calculate the vorticity value of each local time-frequency region in the dynamic vector field, and use the vorticity value as the basis for feature discrimination; Feature discrimination of local time-frequency regions based on vorticity values; When a local time-frequency region exhibits low vorticity and strong momentum characteristics, it is determined that the local time-frequency region is a laminar flow region corresponding to the speech signal. When a local time-frequency region exhibits high vorticity and weak momentum characteristics, it is determined that the local time-frequency region is a turbulent region corresponding to the noise signal. Based on the results of feature discrimination, a flow state distribution map is generated, which serves as a mask representation of noise and speech signals.
4. The audio processing method based on a remote hearing aid according to claim 3, characterized in that, A dynamic coupling mechanism is established to drive the viscosity of the medium based on the flow pattern distribution. Adaptive viscous acoustic field rectification is performed on the dynamic vector field according to the flow pattern distribution diagram. The energy in the noise signal distribution area is eliminated using a physical diffusion process, and cleanroom flow field data is output, including: For coordinate points identified as turbulent regions in the flow distribution diagram, the virtual dynamic viscosity coefficient of the coordinate points in the turbulent regions is increased, forcing the energy of the coordinate points in the turbulent regions to be smoothly dissipated to the surroundings through the diffusion equation. For coordinate points in the flow distribution map that are identified as laminar flow regions, the virtual dynamic viscosity coefficient of the coordinate points in the laminar flow regions is reduced, allowing the signals of the coordinate points in the laminar flow regions to pass through without loss due to their own momentum; The diffusion process ensures a smooth transition of spectral energy in both time and frequency dimensions, generating clean laminar flow field data with noise energy removed.
5. The audio processing method based on a remote hearing aid according to claim 1, characterized in that, The momentum-inertia of the speech signal in the cleanroom flow field data is extracted. Based on the momentum-inertia, fluid advection prediction calculations are performed to predict the dynamic vector field at the next moment, including: Momentum of speech signals is extracted from cleanroom flow field data; momentum is the trend vector of fluid density evolution over time. Based on the current flow velocity vector properties and momentum, fluid advection prediction calculation is performed to obtain the dynamic vector field at the next frame.
6. The audio processing method based on a remote hearing aid according to claim 5, characterized in that, The predicted dynamic vector field is subjected to acoustic inverse mapping to reconstruct the acoustic spectrum at the prediction time, and a frequency domain gain mask is calculated based on the acoustic spectrum at the prediction time, including: Perform the inverse operation of physical property mapping on the predicted dynamic vector field for the next frame. By performing inverse operations, the density values in the dynamic vector field at the predicted next frame time are restored to the energy amplitude of the spectrum; By performing inverse operations, the velocity vector properties in the predicted dynamic vector field of the next frame are restored to the phase change rate of the spectrum; Based on the energy amplitude and phase change rate obtained from the reconstruction, the acoustic spectrum at the predicted moment is reconstructed; A frequency domain gain mask is generated based on the reconstructed acoustic spectrum at the predicted time.
7. The audio processing method based on a remote hearing aid according to claim 1, characterized in that, The frequency domain gain mask is sent back to the wearable audio acquisition and playback unit in real time. The wearable audio acquisition and playback unit then applies the frequency domain gain mask to process and play the real-time acquired audio, including: The frequency domain gain mask is sent to the wearable audio acquisition and playback unit in advance; The wearable audio acquisition and playback unit applies a frequency domain gain mask to the actual sound signal being acquired.
Citation Information
Patent Citations
Hearing aid device based on smart watch
CN118921613A
Method, apparatus and system for low latency audio enhancement
US12231851B1