Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Masking threshold" patented technology

The masking threshold is the sound pressure level of a sound needed to make the sound audible in the presence of another noise called a "masker". This threshold depends upon the frequency, the type of masker, and the kind of sound being masked. The effect is strongest between two sounds close in frequency.

Partial discharge acoustic frequency dynamic mapping method and system based on multi-modal intelligent perception

PendingCN122469118AAccurately grasp the distribution of noise intensityPreserve time domain feature informationNoiseCarrier signal
This application discloses a method and system for dynamic mapping of partial discharge audio signals based on multimodal intelligent sensing, relating to the field of signal processing. The method includes: acquiring ambient audio signals and partial discharge signals; performing frequency domain conversion on the ambient audio signals to obtain a frequency domain energy distribution; calculating the masking sound pressure level based on the frequency domain energy distribution and constructing a global masking threshold curve; traversing the global masking threshold curve to extract the center frequency of the frequency band, and determining the center frequency of the frequency band with the lowest value as the dynamic carrier frequency; extracting the discharge pulse amplitude and pulse rise time characteristics from the partial discharge signal; mapping the inverse ratio of the pulse rise time characteristics to the modulation frequency, and mapping the direct ratio of the discharge pulse amplitude to the frequency modulation index; generating a synthesized audio stream; and outputting the synthesized audio stream to external headphones. This application can effectively sense and identify the converted partial discharge audio signal in complex noise backgrounds.
Owner:WUHAN MOEN INTELLIGENT ELECTRIC CO LTD

Audio processing method and system based on auditory model optimization

The invention relates to the technical field of audio signal processing, and discloses an audio processing method and system based on auditory model optimization. The method comprises the following steps: performing multi-scale frequency band decomposition processing on an input audio signal to generate a sub-band signal set covering different frequency ranges; the sub-band signal set is input into an auditory masking effect calculation model, and the model calculates masking threshold distribution of each sub-band signal according to frequency sensitivity characteristics of a human auditory system. And performing dynamic gain adjustment processing on the sub-band signal set based on the calculated masking threshold distribution to generate a gain-optimized sub-band signal set. And performing nonlinear distortion suppression processing on the sub-band signal set after gain optimization to generate a sub-band signal set after distortion suppression. And inputting the sub-band signal set after distortion suppression into a frequency band synthesis model to reconstruct a complete time domain audio signal. According to the invention, audio optimization processing conforming to human ear perception characteristics is realized.
Owner:中央广播电视总台 +1

Audio coding method, device, equipment, medium and product

The invention relates to the technical field of audio coding, and particularly discloses an audio coding method and device, equipment, a medium and a product, and the method comprises the steps: carrying out the spectrum analysis of an input audio frame, and judging whether a single-frequency signal meeting a preset condition exists or not; if yes, determining that a scale factor band where the single-frequency signal is located and N scale factor bands adjacent to the scale factor band form a target scale factor band combination; setting the masking threshold value of each scale factor band in the target combination to be smaller than the initial masking threshold value, and setting the masking threshold value of each scale factor band in the non-target combination to be larger than the initial masking threshold value; and selecting a scaling factor based on the masking threshold of each scaling factor band, and quantifying each scaling factor band according to the scaling factor to obtain a coded bit stream. Therefore, under the condition that the bit budget is limited, the coding fidelity of the single-frequency signal is improved, and the objective distortion index (for example, the objective distortion index of THD + N is improved) is improved.
Owner:ZGMICRO HEFEI LTD

Information processing apparatus and method

This disclosure relates to information processing apparatuses and methods that can suppress both the reduction in subjective quality of reproduced sound and the reduction in coding efficiency. In this disclosure, a direction-defined bitstream is generated by encoding sound data using a bit allocation based on a direction-defined spatial masking threshold, where the direction-defined spatial masking threshold is a spatial masking threshold that defines the directional range of a sound source. The direction-defined bitstream generated by encoding sound data using a bit allocation based on the direction-defined spatial masking threshold is acquired, and the acquired direction-defined bitstream is decoded, where the direction-defined spatial masking threshold is a spatial masking threshold that defines the directional range of a sound source. This disclosure can be applied, for example, to information processing apparatuses or information processing methods.
Owner:SONY GROUP CORP

A sound dynamic regulation and optimization method and system based on a convolutional neural network

This invention relates to the field of audio signal processing technology, and discloses a method and system for dynamic audio control optimization based on convolutional neural networks. The method includes: separating the mixed audio signals using a source separation convolutional neural network to obtain independent estimated spectra for each source category; extracting energy demand prediction vectors and transient feature descriptors for each source category using a spectral residual convolutional neural network; calculating the driving characteristic matching score between each source category and each speaker unit; performing propagation calculations based on the room transfer function and constructing an arrival power prediction matrix; calculating the masking threshold matrix and perceived loudness estimate for each listener position; solving for the optimal power allocation coefficient through a joint optimization objective function; extracting position-specific perception sensitivity coefficients using an auditory masking perception convolutional neural network; generating adaptive dynamic compression gain values ​​and finally outputting the driving signals for each speaker unit.
Owner:DONGGUAN JINWEIJU TECH CO LTD

Masking threshhold determinator, quantization step size determinator, audio encoder, methods and computer program applying a comodulation dependant post-masking modeling

PCT designated stageWO2026114512A1Speech analysisEngineeringMasking threshold
A masking threshold determinator for providing a masking threshold information on the basis of an input audio signal is configured to obtain envelope information describing envelopes of different frequency ranges of the input audio signal. The masking threshold determinator is configured to determine a comodulation strength information describing a comodulation between different frequency ranges of the input audio signal. The masking threshold determinator is configured to apply a post-masking modeling to the envelope information, or to a pre-processed version thereof, to obtain a post-masking-processed envelope information, which describes a masking threshold. The masking threshold determinator is configured to adapt a post-masking decay time of the post-masking modeling in dependence on the comodulation strength information. A quantization step size determinator and an encoder, corresponding methods and a computer program are also disclosed.
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV +1

Audio signal compression processing method and related device

The invention relates to an audio signal compression processing method and a related device thereof. The method comprises the following steps: acquiring feature information of an audio signal to be compressed; wherein the feature information at least comprises time domain information, frequency domain information and classification information; inputting the feature information of the audio signal into a preset psychological acoustic model to obtain a perception masking threshold value of the audio signal, and determining a compression parameter of the audio signal based on the feature information and the perception masking threshold value; and based on the compression parameter, compressing the audio signal to be compressed, and generating a compressed target audio signal. According to the scheme provided by the invention, the compression parameter can be dynamically adjusted according to the audio signal, redundant information is reduced, and the compression efficiency is improved while key audio information perceived by human ears is reserved.
Owner:XIAMEN LEYUNRUI TECHNOLOGY CO LTD

Hybrid masking threshold-based perceptual slack for audio watermarking

PendingUS20260188332A1Audio watermarkMasking threshold
Techniques are described for hybrid masking threshold-based perceptual slacks for audio watermarking. In some embodiments, the techniques include identifying an audio signal, determining perceptual slacks for the audio signal, generating a watermarked audio signal that includes an audio watermark based on the perceptual slacks, and outputting the watermarked audio signal using one or more speakers, for localization of the one or more speakers.
Owner:HARMAN INT IND INC

Psychoacoustic analysis method, apparatus, device, and storage medium

The embodiment of the disclosure discloses a psychoacoustic analysis method, device, equipment and storage medium, which can be applied to a communication system. The method comprises the following steps: determining a plurality of masking sources of an audio signal; and analyzing a masking threshold of the audio signal according to part of the masking sources. By implementing the method of the disclosure, the calculation amount of psychoacoustic analysis can be effectively reduced, and the calculation complexity can be reduced by selecting part of the masking sources from all the masking sources of the audio signal to participate in the analysis and calculation of the masking threshold.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

An artificial intelligence-based blended teaching evaluation system

PendingCN122311937AAlgorithmEngineering
This invention relates to the field of teaching evaluation technology, specifically to a hybrid teaching evaluation system based on artificial intelligence. The system includes an audiovisual coupling tagging module, a frequency domain fingerprint embedding module, a gaze vector gating module, an interaction intent verification module, and a teaching effectiveness rating module. In this invention, by extracting the pixel displacement modulus of the blackboard and cross-correlation with the audio energy envelope, the system accurately segments substantive teaching segments using physical signal synchronicity, automatically eliminates invalid silences to focus on high-value intervals, embeds markers at high frequencies in the audio using auditory masking thresholds, binds sound waves to the real space to prevent the absence of online users, constructs a three-dimensional normal vector angle model between the face and the screen to determine orthogonal focusing of gaze, combines screen displacement logic verification to eliminate false data related to mechanical gaze, calculates the ratio of effective cognitive injection time to teaching segments, quantifies net teaching effectiveness, and accurately maps the actual conversion efficiency of knowledge transfer.
Owner:JINAN PRESCHOOL TEACHERS COLLEGE

Masking threshold determinator, method and computer program for determining a masking threshold information

ActiveEP4552121C0Masking thresholdThresholding
Owner:FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV

Method for processing audio signal using psychoacoustic model and electronic device for performing same

PCT designated stageWO2026063616A1Speech analysisBiological modelsNerve networkMasking threshold
Disclosed are an audio signal processing method using a psychoacoustic model and an electronic device for performing same. The audio signal processing method may comprise: an operation of receiving an input audio signal; an operation of converting the input audio signal into an audio signal in the frequency domain; an operation of dividing the audio signal in the frequency domain into audio signals of a plurality of subbands; an operation of determining a masking threshold corresponding to each of the subbands by using a psychoacoustic model that takes the audio signals of the subbands as inputs; an operation of acquiring audio vector data corresponding to each of the subbands by using neural network–based encoders; and an operation of generating a compressed audio signal by quantizing the audio vector data corresponding to each of the subbands on the basis of the masking threshold corresponding to each of the subbands.
Owner:SAMSUNG ELECTRONICS CO LTD

Audio watermark embedding and extraction method and system

PCT designated stageWO2026076926A1Speech analysisPattern recognitionAudio watermark
Disclosed in the present invention are an audio watermark embedding and extraction method and system. The method comprises: acquiring audio data, calculating a masking threshold for each frequency point of the audio data, and preprocessing the audio data; performing encoding on the basis of watermark information to generate an audio watermark, and embedding the audio watermark into the audio data on the basis of the masking threshold of the audio data; and locating the position of the audio watermark in the audio data into which the audio watermark is embedded, and extracting the audio watermark on the basis of the masking threshold of the original audio data. In the present invention, an audio watermark is embedded on the basis of human auditory masking theory; the masking threshold of audio data is calculated, and the audio watermark is added to audio in a frequency domain on the basis of the masking threshold, thereby allowing the audio watermark to better resist these attacks, ensuring the robustness of the audio watermark, and ensuring that the detectability can be maintained even under harsh conditions. The present invention solves the problem whereby existing audio watermarks have insufficient robustness and are prone to affect the quality of audio itself.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

A psychoacoustic speech masking method for active defense against voice cloning

PendingCN122337245AMasking thresholdTime frequency decomposition
This application discloses a psychoacoustic speech masking method for proactive defense against speech cloning, relating to the field of speech information security. The method includes: performing time-frequency decomposition on the original reference speech using time-frequency analysis to obtain speech information in the speech representation space; adding a look-ahead perturbation to the speech information to obtain look-ahead speech; determining a speaker confusion loss using a speaker coding network based on the original reference speech and the look-ahead speech; determining a total masking threshold using a psychoacoustic masking model based on the speech information in the speech representation space; determining a speech quality constraint loss based on the perturbation power spectrum according to the total masking threshold; determining a total loss based on the speaker confusion loss and the speech quality constraint loss; generating a final perturbation using an iterative optimization algorithm based on the total loss; and generating a final protected reference speech based on the final perturbation and the original reference speech. This application weakens the ability to extract the identity of the real speaker, thereby reducing the risk of speech cloning and voiceprint impersonation.
Owner:QINGHAI UNIV FOR NATITIES

Audio compression model processing method and system, equipment and program product

PendingCN121545533ASpeech analysisFrequency spectrumMasking threshold
The invention discloses an audio compression model processing method and system, equipment and a program product, and the method comprises the steps: compressing a sample audio through an initial audio compression model, and obtaining a predicted audio; respectively dividing the sample audio and the predicted audio into a sensitive frequency band and a non-sensitive frequency band; dividing sensitive frequency bands of the sample audio and the predicted audio into four frequency bands based on a first auditory masking threshold value, and determining a weight corresponding to a frequency spectrum square difference value of the frequency bands in the same range of the sample audio and the predicted audio, dividing the non-sensitive evaluation of the sample audio and the predicted audio into four frequency bands based on a second auditory masking threshold, and determining a weight corresponding to a frequency spectrum square difference value of the frequency bands of the same range of the sample audio and the predicted audio; obtaining the frequency spectrum Mel sub-band loss by using each frequency spectrum square difference value and the corresponding weight; and optimizing the initial audio compression model based on the spectrum Mel sub-band loss to obtain a target audio compression model.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

An audio feature extraction method and processing terminal based on perception-semantic dual domains

PendingCN122337221AFrequency spectrumMasking threshold
This invention discloses an audio feature extraction method based on a perceptual-semantic dual-domain approach, comprising the following steps: Step 1: Obtaining an audio signal A and its critical frequency band feature B; Step 2: Calculating the simultaneous masking effect of the audio signal A, characterized by a joint masking threshold. Spectral components below the joint masking value will not be perceived by the human ear. After calculating the joint masking threshold, the perceptual importance of the audio signal A is determined by comparing the joint masking threshold with the actual energy of the audio signal A; Step 3: Inputting the critical frequency band feature B into an audio semantic encoder, and projecting the audio features of the audio signal A, which incorporates type labels and emotional attributes, into a 256-dimensional compact semantic space through a semantic embedding generation network to generate a semantic embedding vector. This invention effectively obtains both the perceptual importance and semantic embedding vector features of audio.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Low-noise design method for an acoustic fan based on fan noise

PendingCN122328402ANoiseNoise power spectrum
This invention relates to the field of fan, speaker, and light design and control technology, and discloses a low-noise design method for a speaker fan based on fan noise. The method includes receiving the original audio signal, generating an audio power spectrum, and extracting low-frequency transient features; calculating the instantaneous masking threshold using a psychoacoustic model, and then inversely shaping the allowable noise power spectrum of the fan motor to generate a target fan spectrum; generating a fan drive signal based on the target spectrum using a random spread spectrum algorithm, and constructing an airflow velocity vector by combining the angular velocity fed back from the motor and the duct parameters; calculating the phase compensation amount based on the airflow velocity vector to generate a phase correction signal, and then combining the low-frequency transient features with an anti-phase aerodynamic impedance compensation component to synthesize a speaker signal; finally, driving the fan motor and speaker unit to operate synchronously. This invention utilizes the auditory masking effect to achieve invisible fan operation and cancels the interference of airflow on sound waves through flow field mapping and phase pre-distortion, achieving synergistic optimization of the sound field and airflow field.
Owner:HEFEI FENGZHILUODONG INTELLIGENT TECHNOLOGY CO LTD