Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

25 results about "Auditory masking" patented technology

Auditory masking occurs when the perception of one sound is affected by the presence of another sound. Auditory masking in the frequency domain is known as simultaneous masking, frequency masking or spectral masking. Auditory masking in the time domain is known as temporal masking or non-simultaneous masking.

An aac security steganography algorithm and system based on masking effect

ActiveCN115620733BSpeech analysisVocal intensityAlgorithm
The application discloses an AAC security steganography algorithm and system based on a masking effect. The human ear hearing masking effect can cause some audio signals with low intensity to be masked by signals with high intensity, and there is space for hiding secret information. Therefore, the application analyzes a quantization process of MDCT coefficients in AAC coding, records the masked audio signals as modifiable positions, and combines an STC adaptive steganography framework to realize embedding of secret information. Experimental results show that the algorithm can achieve a maximum embedding capacity of 13.61 kbps, can guarantee auditory concealment of speech, and has good security.
Owner:WUHAN UNIV

Multi-channel audio signal hybrid coding method and device, and storage medium

PendingCN121306149ASpeech analysisData streamDynamic range compression
The invention relates to the field of audio signal processing, in particular to a multi-channel audio signal hybrid coding method and device and a storage medium, and the method comprises the steps: collecting digital audio signals through a plurality of audio input channels, carrying out the anti-aliasing filtering of the collected digital audio signals, so as to obtain audio signals in a 48kHz / 16bit format, and outputting the audio signals in the 48kHz / 16bit format. An audio signal in a 48kHz / 16bit format is processed through a loudness weighted average method, the audio signal after loudness weighted average is processed through a frequency domain mixing technology to obtain a mixed signal, and dynamic range compression is performed on the mixed signal based on calculation of an auditory masking curve to obtain a compressed signal. And carrying out hybrid coding on the compressed signal to obtain a compressed data stream, and packaging and outputting the compressed data stream. Through dynamic mixing and compression processing, the problem of signal distortion caused by a traditional linear superposition method is effectively solved, reasonable superposition of all paths of signals is ensured through dynamic mixing, and level overload caused by simple addition is avoided.
Owner:HUBEI DONGWEIYUNSHI TECHNOLOGY CO LTD

Intelligent sensing-based adaptive real-time reverberation processing method and system

The application provides an adaptive real-time reverberation processing method and system based on intelligent sensing. The method comprises the following steps: dividing input audio into n frequency band regions according to sensing importance, and dynamically adjusting the frequency division point position; performing differential processing on the signals of each frequency band, and dynamically adjusting the operation parameters in the differential processing in combination with the real-time load of the system; performing frequency domain block convolution processing on the n frequency band signals based on a pre-constructed multi-level cache architecture; mixing the convolution results of the n frequency band signals according to the human ear hearing masking effect, and outputting the final reverberation audio through a dynamic range control mechanism while ensuring the phase continuity among the frequency bands. The application successfully realizes the dynamic balance between professional sound quality and high real-time performance in the limited computing resource scene in the field of real-time reverberation processing, significantly improves the real-time response speed of audio reverberation processing and the multi-device heterogeneous adaptation capability on the basis of ensuring the reverberation sound quality, and maximally reduces the invalid consumption of computing resources.
Owner:WANSHENG MUSIC TECH (SHENZHEN) CO LTD

In-vehicle multi-working-condition active sound control method and system based on masking effect

The invention discloses an in-vehicle multi-working-condition active sound control method and system based on a masking effect. The method comprises the steps that the current driving working condition of a vehicle is recognized according to vehicle speed changes; dynamically controlling the duration of the active sound in the vehicle based on the working condition, so that the time characteristic of the active sound is continuously adjusted along with the change of the working condition of the vehicle; meanwhile, background noise information in the vehicle is collected, and the amplitude of the active sound is adaptively adjusted based on an auditory masking effect, so that the active sound keeps good perceptibility and auditory comfort under different working conditions and different noise levels; and in the working condition switching process, smooth transition control is performed on the duration time and the amplitude of the active sound so as to realize continuous output of the active sound. By means of the mode, the problems of sound sudden change, masking unbalance or auditory interference occurring when the working condition changes in an existing active sound production scheme can be effectively avoided, and the stability and comfort of the in-vehicle sound environment under different operation conditions are improved.
Owner:JILIN UNIVERSITY

Noise suppression and frequency response equalization dynamic optimization method of sound system in complex sound field environment

PendingCN122340424AFrequency spectrumNoise
This invention provides a dynamic optimization method for noise suppression and frequency response equalization of an audio system in complex sound field environments. The method includes acquiring spatial acoustic signals, auxiliary sensing signals, and system reference signals; constructing a feature fusion dataset based on a physical-acoustic coupling mapping mechanism; performing spatiotemporal modeling of the current sound field environment; and generating a sound field state vector and a sound field complexity index. Based on the multimodal feature fusion dataset and the sound field state vector, a frequency domain energy residual surrogate model is introduced as a physical embedded constraint to perform feature-level bidirectional joint solution for noise suppression and dynamic frequency response equalization, generating an initial filter coefficient set. This invention, through soft decision-making and convex set projection, ensures that the final output filter coefficients simultaneously satisfy auditory masking protection and physical frequency response compensation, effectively avoiding the conflict between equalization amplification of background noise and noise reduction creating spectral holes.
Owner:GUANGDONG HUIDU ELECTRONICS CO LTD

A forgery speech detection algorithm and system based on principal component filtering

The existing counterfeit speech detection method has weak robustness in the re-encoding and noise mismatch scene, in order to improve the robustness of the existing method, the counterfeit speech detection research work puts forward the strategy of data augmentation to the training data set. However, the data augmentation strategy can increase the training data quantity, reduce the model training efficiency, and can only be used for known encoding algorithm and noise difference scene. The present application relates to the field of counterfeit speech detection, especially to the field of counterfeit speech detection for re-encoding and noise interference scene, specifically relates to a counterfeit speech detection algorithm and system based on subject filtering, mainly designs a subject signal filtering module based on the relationship between the human ear hearing masking effect and the signal-to-noise energy ratio, which can eliminate the part causing the distribution difference in the spectrogram feature, at the same time, without increasing the training data quantity, can improve the robustness of the model in the unknown encoding algorithm and noise difference scene, and has good universality.
Owner:WUHAN UNIV

Noise howling component perception saliency evaluation method, device, equipment and medium

The invention relates to the technical field of vehicle vibration noise analysis, in particular to a noise howling component perception saliency evaluation method, device and equipment and a medium, and the method comprises the steps: obtaining an original signal with a howling component, and calculating the feature loudness of the original signal; removing a howling component in the original signal to obtain a reference signal, and calculating the characteristic loudness of the reference signal; calculating a feature loudness contribution amount caused by the howling component according to the two feature loudnesses; and calculating the reference loudness of the reference signal in the howling influence interval so as to calculate the howling component perception saliency according to the feature loudness contribution amount caused by the howling component and the reference loudness. Therefore, the problems that in a related howling problem evaluation index method starting from a howling perception mechanism, the sound pressure level cannot accurately represent human ear perception, and the influence of the masking effect is not considered, that is, the difference between the frequency band energy where the howling component is located and the lower frequency band energy has decisive influence on the prominence perception of the howling component are solved.
Owner:CHINA FAW CO LTD

Sparse auditory pulse coding method and device fusing masking effect and dynamic threshold

PendingCN121905196ASpeech analysisNeural architecturesNeuromorphic hardwarePsychoacoustics
The invention is suitable for the technical field of audio signal processing, and provides a sparse auditory pulse coding method and device fusing a masking effect and a dynamic threshold. According to the auditory rarefaction method provided by the embodiment of the invention, a human auditory perception mechanism can be combined, and the threshold group characteristics are utilized to realize time sequence pulse coding, so that the audio signal can directly adapt to the pulse neural network model, the coding efficiency and fidelity are improved, and meanwhile, the data redundancy is reduced. While the perception quality is ensured, the number of pulse events can be remarkably reduced, the energy efficiency and accuracy of a recognition task are improved, and the method is naturally adaptive to brain-like hardware and a pulse neural network. According to the embodiment of the invention, the masking effect of psychoacoustics is fused into the pulse coding process for the first time, and is highly consistent with the feeling and processing mode of organisms on sound signals. The biological reasonability is relatively high.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL

Frequency-based compensation filter for in-ear monitors or headphones

PendingUS20260122437A1Deaf amplification systemsFrequency response correctionAudio power amplifierTransducer
An in-ear audio transducer such as a single earphone or a pair of earphones produces audio which accounts for frequency detected hearing impairment and auditory masking conditions spanning frequencies expected to be encountered by the user, thereby improving auditory perception without having to increase overall volume as much as in the prior art. The audio transducer(s) communicates with an audio base unit configured to output modified multichannel digital audio data. The in-ear audio transducer element has a processing unit, a plurality of transducer amplifiers and a plurality of acoustic elements.
Owner:SOUND DEVICES LLC

A sound dynamic regulation and optimization method and system based on a convolutional neural network

This invention relates to the field of audio signal processing technology, and discloses a method and system for dynamic audio control optimization based on convolutional neural networks. The method includes: separating the mixed audio signals using a source separation convolutional neural network to obtain independent estimated spectra for each source category; extracting energy demand prediction vectors and transient feature descriptors for each source category using a spectral residual convolutional neural network; calculating the driving characteristic matching score between each source category and each speaker unit; performing propagation calculations based on the room transfer function and constructing an arrival power prediction matrix; calculating the masking threshold matrix and perceived loudness estimate for each listener position; solving for the optimal power allocation coefficient through a joint optimization objective function; extracting position-specific perception sensitivity coefficients using an auditory masking perception convolutional neural network; generating adaptive dynamic compression gain values ​​and finally outputting the driving signals for each speaker unit.
Owner:DONGGUAN JINWEIJU TECH CO LTD

Audio processing method and device, electronic equipment and storage medium

The present disclosure relates to an audio processing method and device, electronic equipment and storage medium. The method comprises: determining a first audio signal from an initial audio based on a signal value of an audio signal in the initial audio; reducing a signal value of a second audio signal adjacent to the first audio signal to obtain an output audio signal; wherein the first audio signal and the second audio signal produce a masking effect; and obtaining a to-be-played audio based on the output audio signal. By reducing the second audio signal adjacent to the first audio signal to the output audio signal, the influence of the masking effect can be reduced, hearing protection is achieved by using the masking effect in psychoacoustics, the objective output energy of the sound is effectively reduced, the sound radiation of the music playing device to the human ear is reduced, and the perception of the human ear to the sound volume reduction is not caused, that is, the subjective listening experience is unchanged; and no frequency band processing is required, and the computing power requirement is reduced.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Multi-source Bluetooth audio stream arbitration system based on vehicle state perception and anti-interference method

The invention provides a multi-source Bluetooth audio stream arbitration system based on vehicle state perception and an anti-interference method, and relates to the technical field of vehicle-mounted communication. The system comprises a vehicle state and environment sensing module, a multi-source Bluetooth communication management module and a central arbitration processing unit. The method comprises the following steps: collecting vehicle kinetic parameters, monitoring electrical characteristics of an execution part by using an LC oscillation detection unit, and constructing a dynamic electromagnetic interference potential energy model to predict a future interference trend; when the predicted value exceeds the standard, preemptive link reconstruction is triggered before interference occurs, the type of an underlying data packet is forcibly adjusted to be a high-anti-interference single-slot packet, and an anti-interference frequency hopping table is locked; meanwhile, the audio stream weight is calculated in combination with the vehicle speed and the safety index, and dynamic sorting and output arbitration are carried out on multi-source audios. According to the invention, spectral bandwidth replacement is carried out by using a wind noise masking effect, and joint optimization of active anti-interference of a physical layer and intelligent scheduling of an application layer is realized by cooperating with user interaction feedback.
Owner:SHENZHEN HENGCHANGTONG ELECTRONICS CO LTD

An artificial intelligence-based blended teaching evaluation system

PendingCN122311937AAlgorithmEngineering
This invention relates to the field of teaching evaluation technology, specifically to a hybrid teaching evaluation system based on artificial intelligence. The system includes an audiovisual coupling tagging module, a frequency domain fingerprint embedding module, a gaze vector gating module, an interaction intent verification module, and a teaching effectiveness rating module. In this invention, by extracting the pixel displacement modulus of the blackboard and cross-correlation with the audio energy envelope, the system accurately segments substantive teaching segments using physical signal synchronicity, automatically eliminates invalid silences to focus on high-value intervals, embeds markers at high frequencies in the audio using auditory masking thresholds, binds sound waves to the real space to prevent the absence of online users, constructs a three-dimensional normal vector angle model between the face and the screen to determine orthogonal focusing of gaze, combines screen displacement logic verification to eliminate false data related to mechanical gaze, calculates the ratio of effective cognitive injection time to teaching segments, quantifies net teaching effectiveness, and accurately maps the actual conversion efficiency of knowledge transfer.
Owner:JINAN PRESCHOOL TEACHERS COLLEGE

Audio watermark embedding and extraction method and system

PCT designated stageWO2026076926A1Speech analysisPattern recognitionAudio watermark
Disclosed in the present invention are an audio watermark embedding and extraction method and system. The method comprises: acquiring audio data, calculating a masking threshold for each frequency point of the audio data, and preprocessing the audio data; performing encoding on the basis of watermark information to generate an audio watermark, and embedding the audio watermark into the audio data on the basis of the masking threshold of the audio data; and locating the position of the audio watermark in the audio data into which the audio watermark is embedded, and extracting the audio watermark on the basis of the masking threshold of the original audio data. In the present invention, an audio watermark is embedded on the basis of human auditory masking theory; the masking threshold of audio data is calculated, and the audio watermark is added to audio in a frequency domain on the basis of the masking threshold, thereby allowing the audio watermark to better resist these attacks, ensuring the robustness of the audio watermark, and ensuring that the detectability can be maintained even under harsh conditions. The present invention solves the problem whereby existing audio watermarks have insufficient robustness and are prone to affect the quality of audio itself.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Audio compression model processing method and system, equipment and program product

PendingCN121545533ASpeech analysisFrequency spectrumMasking threshold
The invention discloses an audio compression model processing method and system, equipment and a program product, and the method comprises the steps: compressing a sample audio through an initial audio compression model, and obtaining a predicted audio; respectively dividing the sample audio and the predicted audio into a sensitive frequency band and a non-sensitive frequency band; dividing sensitive frequency bands of the sample audio and the predicted audio into four frequency bands based on a first auditory masking threshold value, and determining a weight corresponding to a frequency spectrum square difference value of the frequency bands in the same range of the sample audio and the predicted audio, dividing the non-sensitive evaluation of the sample audio and the predicted audio into four frequency bands based on a second auditory masking threshold, and determining a weight corresponding to a frequency spectrum square difference value of the frequency bands of the same range of the sample audio and the predicted audio; obtaining the frequency spectrum Mel sub-band loss by using each frequency spectrum square difference value and the corresponding weight; and optimizing the initial audio compression model based on the spectrum Mel sub-band loss to obtain a target audio compression model.
Owner:TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD

An audio feature extraction method and processing terminal based on perception-semantic dual domains

PendingCN122337221AFrequency spectrumMasking threshold
This invention discloses an audio feature extraction method based on a perceptual-semantic dual-domain approach, comprising the following steps: Step 1: Obtaining an audio signal A and its critical frequency band feature B; Step 2: Calculating the simultaneous masking effect of the audio signal A, characterized by a joint masking threshold. Spectral components below the joint masking value will not be perceived by the human ear. After calculating the joint masking threshold, the perceptual importance of the audio signal A is determined by comparing the joint masking threshold with the actual energy of the audio signal A; Step 3: Inputting the critical frequency band feature B into an audio semantic encoder, and projecting the audio features of the audio signal A, which incorporates type labels and emotional attributes, into a 256-dimensional compact semantic space through a semantic embedding generation network to generate a semantic embedding vector. This invention effectively obtains both the perceptual importance and semantic embedding vector features of audio.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Earphone

The utility model relates to a karaoke earphone, which comprises a left sounding unit, a right sounding unit, a microphone assembly and an earphone control box, the left sounding unit is connected with the earphone control box through a left audio cable, the right sounding unit is connected with the earphone control box through a right audio cable, the microphone assembly is connected with the earphone control box, and the earphone control box is connected with intelligent equipment. The earphone control box is used for carrying out noise reduction, gain adjustment, analog-to-digital conversion and sound effect processing on signals input by the microphone, the human voice frequency band can be enhanced, the environmental sound and the human voice frequency in accompaniment can be inhibited, the masking effect is avoided, meanwhile, digital audio signals input by the intelligent equipment are decoded into analog signals with independent left and right sound channels, and the intelligent equipment is more intelligent. The audio signal can be accurately restored, the sound effect of a stereo field is presented, the ear return sound can be accurately and rapidly transmitted, and the singing experience is improved.
Owner:SHENZHEN WEIDONG ACOUSTIC TECH CO LTD

Low-noise design method for an acoustic fan based on fan noise

PendingCN122328402ANoiseNoise power spectrum
This invention relates to the field of fan, speaker, and light design and control technology, and discloses a low-noise design method for a speaker fan based on fan noise. The method includes receiving the original audio signal, generating an audio power spectrum, and extracting low-frequency transient features; calculating the instantaneous masking threshold using a psychoacoustic model, and then inversely shaping the allowable noise power spectrum of the fan motor to generate a target fan spectrum; generating a fan drive signal based on the target spectrum using a random spread spectrum algorithm, and constructing an airflow velocity vector by combining the angular velocity fed back from the motor and the duct parameters; calculating the phase compensation amount based on the airflow velocity vector to generate a phase correction signal, and then combining the low-frequency transient features with an anti-phase aerodynamic impedance compensation component to synthesize a speaker signal; finally, driving the fan motor and speaker unit to operate synchronously. This invention utilizes the auditory masking effect to achieve invisible fan operation and cancels the interference of airflow on sound waves through flow field mapping and phase pre-distortion, achieving synergistic optimization of the sound field and airflow field.
Owner:HEFEI FENGZHILUODONG INTELLIGENT TECHNOLOGY CO LTD

Self-adaptive real-time reverberation processing method and system based on intelligent perception

The invention provides a self-adaptive real-time reverberation processing method and system based on intelligent perception, and the method comprises the steps: dividing an input audio into n frequency band regions according to the perception importance, and dynamically adjusting the position of a frequency division point; differential processing is carried out on the signals of all the frequency bands, and operation parameters in differential processing are dynamically adjusted in combination with the real-time load of the system; performing frequency domain block convolution processing on the n frequency band signals based on a pre-constructed multi-level cache architecture; the convolution results of the n frequency band signals are mixed according to the auditory masking effect of human ears, and the final reverberation audio is output through a dynamic range control mechanism while the phase continuity between the frequency bands is guaranteed. According to the invention, dynamic balance between professional tone quality and high real-time performance in a computing resource limited scene is successfully realized in the field of real-time reverberation processing, the real-time response speed and multi-device heterogeneous adaptation capability of audio reverberation processing are remarkably improved on the basis of guaranteeing the reverberation tone quality, and invalid consumption of computing resources is reduced to the greatest extent.
Owner:WANSHENG MUSIC TECH (SHENZHEN) CO LTD

A method and system for active sound control in a vehicle cabin under multiple working conditions based on masking effect

The application discloses a kind of based on masking effect's in-vehicle multi-working condition active sound control method and system.The method includes: according to vehicle speed variation, the current driving condition of vehicle is identified;And based on the condition, the duration of active sound in vehicle is dynamically controlled, so that the time characteristics of active sound are continuously adjusted with the change of vehicle condition;While collecting the background noise information in vehicle, based on the amplitude of the adaptive adjustment of active sound of auditory masking effect, so that active sound keeps good perceptibility and auditory comfort under different conditions and different noise levels;During the switching process of condition, the duration and amplitude of active sound are smoothly transitioned to control, to realize the continuous output of active sound.Through the above mode, the application can effectively avoid the sound mutation, masking imbalance or auditory interference problem of existing active sound generation scheme when condition changes, improve the stability and comfort of in-vehicle sound environment under different operating conditions.
Owner:JILIN UNIVERSITY

Frequency-based compensation filter for in-ear monitors or headphones

PCT designated stageWO2026096446A1Deaf amplification systemsFrequency response correctionAudio power amplifierTransducer
An in-ear audio transducer such as a single earphone or a pair of earphones produces audio which accounts for frequency detected hearing impairment and auditory masking conditions spanning frequencies expected to be encountered by the user, thereby improving auditory perception without having to increase overall volume as much as in the prior art. The audio transducer(s) communicates with an audio base unit configured to output modified multichannel digital audio data. The in-ear audio transducer element has a processing unit, a plurality of transducer amplifiers and a plurality of acoustic elements.
Owner:SOUND DEVICES LLC

In-vehicle rough sound evaluation method and device, electronic equipment and storage medium

PendingCN121720739AVehicle testingSustainable transportationHuman auditory systemNoise
The invention discloses an in-vehicle rough sound evaluation method and device, electronic equipment and a storage medium, and relates to the technical field of noise evaluation.The method comprises the steps that in-vehicle noise signals of a vehicle are collected, filtering processing is conducted on the in-vehicle noise signals, and noise sub-signals corresponding to a plurality of continuous critical frequency bands are obtained; performing signal processing on the noise sub-signals based on a preset masking effect model to obtain a roughness contribution value of each critical frequency band; and determining an in-vehicle coarse noise evaluation result according to the roughness contribution value of each critical frequency band. Therefore, according to the method, the masking effect model in psychoacoustics is introduced, so that the actual working principle of a human auditory system during complex sound processing is simulated, roughness analysis is carried out on the collected in-vehicle noise signals, the accuracy and credibility of an in-vehicle noise evaluation result are improved, meanwhile, a large amount of time and labor cost are saved, and the accuracy and reliability of in-vehicle noise evaluation are improved. And the test efficiency is improved.
Owner:CHINA FAW CO LTD

Speech enhancement method and system based on spectral decomposition

PendingCN122067548ASpeech analysisFrequency spectrumGabor atom
The invention discloses a speech enhancement method and system based on spectral decomposition. The method comprises the following steps: acquiring an amplitude spectrum and a phase spectrum of noisy speech through short-time Fourier transform; constructing a harmonic extraction model combining basis tracking and an auditory masking effect, and adaptively separating structured harmonic components through a non-convex optimization problem; performing sparse decomposition on the residual frequency spectrum by using an over-complete dictionary formed by Gabor atoms and Dirac atoms, and extracting transient detail components; synthesizing an enhanced amplitude spectrum by adopting a sensing weighted fusion strategy based on signal-to-noise ratio dynamic adjustment; and reconstructing a phase spectrum through a U-Net neural network and outputting a time domain signal. The method provided by the invention solves the problems of significant music noise, phase distortion and insufficient transient feature retention in the traditional method, and significantly improves the voice signal-to-noise ratio and perception quality in a complex noise environment.
Owner:FOURTH MILITARY MEDICAL UNIVERSITY +1

Vehicle-mounted multi-sensory collaborative interaction method and system based on multi-dimensional audio features

The invention discloses a vehicle-mounted multi-sensory collaborative interaction method and system based on multi-dimensional audio features, and the method comprises the steps: obtaining a vehicle-mounted original audio source signal, employing a double-flow parallel analysis architecture, extracting the semantic features of lyrics through employing an NLP model, and extracting the acoustic physical features through employing a CNN model; weighted arbitration is carried out on the double-flow features based on the vehicle driving state, and a basic sensory control vector is generated; processing an in-vehicle environment audio signal by using an adaptive noise cancellation algorithm, calculating an environment masking coefficient, and performing dynamic gain correction on a control vector; and finally, in combination with reverse time sequence scheduling and PID closed-loop control, executing units such as fragrance, an atmosphere lamp and a seat vibration motor are driven to work cooperatively. According to the method, the problem of emotion deficiency of single physical feature control is solved through double-flow fusion, the road noise masking effect is eliminated through environmental perception compensation, and accurate immersive interactive experience is achieved through hardware time sequence synchronization.
Owner:ZHEJIANG HEQIAN ELECTRONIC TECH CO LTD