Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

12 results about "Time frequency masking" patented technology

Time-frequency masking (TFM) method [1] separates sound sources by masking unwanted sounds in the time-frequency domain. The method primarily relies on clustering of the mixed signals with respect to their amplitudes and time delays.

A bearing composite fault diagnosis method and system

The application discloses a bearing composite fault diagnosis method and system, first, a bearing composite fault vibration observation signal is acquired; then, the observation signal is subjected to short-time Fourier transform and cepstrum threshold processing, two-dimensional time-frequency mask blind source separation, inverse short-time Fourier transform is performed on the obtained time-frequency domain separation signal and band-pass filtering is performed, independent time domain estimation signals are obtained; then, the estimation signals are subjected to Hilbert envelope demodulation and Fourier transform, envelope spectra of the filtered signals are obtained; finally, each order characteristic frequency band of a target fault is searched, a harmonic energy impact index is constructed, and an impact pulse value is calculated; the calculated impact pulse value is compared with a threshold value, and whether the bearing has a target type fault is diagnosed.
Owner:HUNAN UNIV +1

Voice separation method and device, electronic equipment and storage medium

The invention provides a voice separation method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the first-stage sound region division of multiple groups of voice signals, obtaining the first-stage masks of a plurality of first-stage sound regions, enabling the multiple groups of voice signals to be obtained through the synchronous pickup of a plurality of microphone arrays, and enabling the first-stage sound regions to be in one-to-one correspondence with the microphone arrays; performing secondary sound area division on each group of voice signals to obtain secondary masks of a plurality of secondary sound areas under each primary sound area; based on the primary masks of the plurality of primary sound regions, the secondary masks of the plurality of secondary sound regions and the mapping relationship between the secondary sound regions and the recombination sound region, determining a time frequency mask of the recombination sound region; the recombination sound area is a dynamic sound area adapted to personnel distribution; and performing voice separation on the multiple groups of voice signals based on the time-frequency mask of the recombined voice area. According to the method, the device, the electronic equipment and the storage medium provided by the invention, dynamic self-adaptive adjustment of the sound area in the space is realized, so that the accuracy of voice interaction and the user experience are ensured.
Owner:IFLYTEK CO LTD

Sound source separation method, device and equipment, storage medium and vehicle

The invention discloses a sound source separation method, device and equipment, a storage medium and a vehicle. The method comprises the following steps: acquiring a time-frequency graph corresponding to a mixed audio; frequency band division is carried out on the time-frequency diagram through a Mel filter bank, K sub-band time-frequency diagrams are obtained, an overlapping part exists between every two adjacent sub-band time-frequency diagrams, and K is a positive integer larger than or equal to 1; performing feature extraction on the K sub-band time-frequency graphs to obtain a target feature tensor; performing mask estimation on the target feature tensor by using a target sound source separation model to obtain a target time frequency mask corresponding to a target audio, the target sound source separation model being used for separating the target audio from the mixed audio; and determining the audio corresponding to the target time frequency mask in the mixed audio as the target audio. According to the sound source separation method provided by the embodiment of the invention, the accuracy of sound source separation can be ensured, and the sound quality of the target audio obtained by sound source separation is improved.
Owner:BEIJING CO WHEELS TECH CO LTD

Apparatus, system and / or method for device localization and optimization utilizing a predetermined audible signal

In at least one embodiment, an audio system including a first loudspeaker and a second loudspeaker and at least one controller is provided. The first loudspeaker transmits a first audio signal including a first signature tone into a listening environment. The second loudspeaker transmits a second audio signal including a second signature tone into the listening environment and receive the first audio signal including and the first signature tone. The second loudspeaker receives the second audio signal including the second signature tone after transmitting the second signature tone into the listening environment and determines an estimated distance between the first loudspeaker and the second loudspeaker based at least on the first signature tone and the second signature tone. The second loudspeaker performs a time frequency masking operation to extract a least one of the first signature tone and the second signature tone from the noisy and reverberant mixture.
Owner:HARMAN INT IND INC

Method for adjusting human voice in song

PCT designated stageWO2026077161A1Speech analysisTime domainFeature extraction
A method and apparatus for adjusting human voice in a song, and a terminal device, a medium and a product. The method comprises: converting a first time-domain signal in a song into a first time-frequency-domain signal (102); performing feature extraction on the first time-frequency-domain signal, so as to obtain a time-frequency-domain audio feature (104); inputting the time-frequency-domain audio feature into a human voice extraction model, so as to obtain a time-frequency mask outputted by the human voice extraction model, wherein the time-frequency mask is used for representing the strength of a human voice signal at corresponding time and frequency points in the first time-frequency-domain signal (106); on the basis of the time-frequency mask, adjusting the first time-frequency-domain signal, so as to obtain a second time-frequency-domain signal (108); and converting the second time-frequency-domain signal into a second time-domain signal, so as to obtain an adjusted song (110).
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Active noise reduction method, device, equipment, medium, program product and vehicle

The invention discloses an active noise reduction method and device, equipment, a medium, a program product and a vehicle. The method comprises the following steps: acquiring a reference signal transmitted from a noise source to a first position; performing short-time Fourier transform on the reference signal to obtain a reference signal matrix; inputting the reference signal matrix into a time-frequency masking model to obtain a time-frequency masking value output by the time-frequency masking model; processing the reference signal matrix based on the time-frequency masking value to obtain a noise reduction signal; and superposing the noise reduction signal and an original noise signal transmitted to the second position by the noise source to realize noise reduction at the second position. According to the embodiment of the invention, active noise reduction can be carried out on equipment such as a vehicle, the inhibition capability of an ANC system on complex noise is improved, noise reduction is carried out by adopting a mode of generating the time-frequency masking value by the time-frequency masking model, and compared with a mode of carrying out noise reduction by directly outputting a signal for driving a secondary loudspeaker through the model, the complexity of the model is reduced, and the noise reduction efficiency is improved. And a possible solution is provided for hardware landing.
Owner:BEIJING CO WHEELS TECH CO LTD

Opportunity signal self-positioning method based on time-frequency multi-scale mask pre-training large model

The invention discloses an opportunity signal self-positioning method based on a time-frequency multi-scale mask pre-training large model, belongs to the technical field of opportunity signal self-positioning, and solves the problem of insufficient positioning precision caused by limited label data of an existing positioning task. The method comprises the following steps: acquiring time-frequency representation of wavelet packet transform domain sequences of a plurality of signal segments corresponding to each opportunity signal sample; constructing a pre-training set and a labeled data set; training a pre-training neural network based on the random-time frequency mask by using the pre-training set; cutting the pre-trained neural network, adding an average pooling layer and a multi-layer perceptron, and constructing a positioning model; training and verifying the positioning model by using a labeled data set to obtain a successfully verified positioning model; opportunity signals of a point to be positioned are collected in real time, time-frequency representation of the wavelet packet transform domain sequences of the corresponding multiple signal segments is obtained, and coordinates of the point to be positioned are obtained through prediction after processing of the verified positioning model.
Owner:36TH RES INST OF CETC

Local sound amplification method

PCT designated stageWO2026077160A1Public address systemsTime domainFeature extraction
A local sound amplification method and apparatus, a terminal device, a medium, and a product. The method may comprise: performing feature extraction on a first time-frequency domain signal to be played back, and obtaining a first time-frequency domain feature; performing feature extraction on a second time-frequency domain signal to be played back after the first time-frequency domain signal, and obtaining a second time-frequency domain feature; inputting the first time-frequency domain feature and the second time-frequency domain feature into a human voice extraction model, and obtaining a time-frequency mask outputted by the human voice extraction model, the time-frequency mask being set to represent the intensity of a human voice time-frequency domain signal at corresponding time and frequency points in the first time-frequency domain signal; on the basis of the time-frequency mask, suppressing acoustic factors in the first time-frequency domain signal, and obtaining a human voice time-frequency domain signal corresponding to the first time-frequency domain signal; and converting the human voice time-frequency domain signal corresponding to the first time-frequency domain signal into a human voice time-domain signal, and performing local sound amplification on the human voice time-domain signal.
Owner:ZHEJIANG GEELY HLDG GRP CO LTD +1

Voice processing method and device, electronic equipment and storage medium

The embodiment of the invention discloses a voice processing method and device, electronic equipment and a storage medium, and the method comprises the steps: sampling a reference position in a target voice region, determining a target phase difference when a voice signal from the reference position is collected, carrying out the signal alignment of the collected to-be-processed voice signal based on the target phase difference, obtaining a sensor signal pair, and carrying out the recognition of the to-be-processed voice signal. Determining an observation signal difference according to the difference of the sensor signal pair in the time-frequency domain, determining an observation phase difference of the to-be-processed voice signal on the acoustic sensor pair, and performing time-frequency masking prediction based on the to-be-processed voice signal, the observation signal difference and the observation phase difference to obtain sound area masking information, and the target voice signal of the target voice region is extracted from the to-be-processed voice signal based on the voice region masking information, so that the target voice signal is extracted by introducing the observation signal difference on the basis of the observation phase difference, the amount of information carried by the spatial features is enriched, and the accuracy of voice processing is improved. The accuracy can be effectively improved when the target voice signal corresponding to the target voice area is extracted.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Voice signal denoising method and device, electronic equipment and program product

The invention discloses a voice signal denoising method and device, electronic equipment and a program product, and relates to the technical field of artificial intelligence, and the denoising method comprises the steps: collecting voice data of a target user based on a preset collection strategy, and carrying out the denoising of the voice data through employing a preset back diffusion model, and obtaining initial voice data; inputting the initial voice data into a preset dual-path recursive network, and outputting a voice feature sequence; generating a time frequency mask for the speech feature sequence, and determining a target speech spectrum based on the time frequency mask; and based on the target voice spectrum, performing voice waveform reconstruction by adopting inverse short-time Fourier transform to obtain a target voice signal. According to the invention, the technical problem of low speech recognition accuracy under multi-noise and multi-speaker interference in the prior art is solved.
Owner:TRAVELSKY TECHNOLOGY LIMITED

Speech enhancement method of amplitude-phase mixed feature cross

The application discloses a deep learning speech enhancement method based on amplitude-phase mixed feature intersection; according to the collected noisy speech signal, an enhanced mixed intersection feature is obtained; according to the collected clean speech signal and the corresponding noisy speech signal, a label intersection compressed complex mask used for training of an amplitude-phase noise reduction network APNSN is calculated; the enhanced mixed intersection feature is input into the trained APNSN network to obtain an estimated intersection compressed complex mask; according to the estimated intersection compressed complex mask and the spectrum of the noisy speech signal, a time domain reconstructed signal is obtained; compared with a single feature method, such as amplitude spectrum mapping and time-frequency masking based on amplitude spectrum features, the method can further improve speech quality and intelligibility under the condition of the same model size; under a relatively small model, the method can obtain speech quality and intelligibility comparable to the single feature method.
Owner:XIHUA UNIV