Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

26 results about "Audio restoration" patented technology

Audio restoration is a generalized term for the process of removing imperfections (such as hiss, impulse noise, crackle, wow and flutter, background noise, and mains hum) from sound recordings. Audio restoration can be performed directly on the recording medium (for example, washing a gramophone record with a cleansing solution), or on a digital representation of the recording using a computer (such as an AIFF or WAV file). Record restoration is a particular form of audio restoration that seeks to repair the sound of damaged gramophone records.

Audio restoration method and apparatus, and program, medium and device

PCT designated stageWO2025194883A1Speech analysisBiological modelsAlgorithmAudio restoration
Provided in the present disclosure are an audio restoration method, an audio restoration apparatus, a computer program product, a computer-readable storage medium and an electronic device. The method comprises: inputting damaged audio into a first-stage model to obtain a complex spectrum; and inputting the complex spectrum into a second-stage model to obtain restored audio, wherein the first-stage model comprises an encoder used for performing down-sampling on the complex spectrum, and a decoder used for performing up-sampling on the complex spectrum, and the second-stage model comprises a full-band module used for modeling the complex spectrum on all bands, and a sub-band module used for modeling the complex spectrum on a plurality of sub-bands. By means of the present disclosure, complex and diverse audio distortion problems can be addressed.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Audio restoration method and apparatus, and electronic device

The invention discloses an audio restoration method and device and electronic equipment, and belongs to the technical field of artificial intelligence. The method comprises the following steps: performing feature extraction on a to-be-repaired first audio clip, a previous-text audio clip and a next-text audio clip of the first audio clip to obtain a first acoustic feature vector of the first audio clip, a second acoustic feature vector of the previous-text audio clip and a third acoustic feature vector of the next-text audio clip; and inputting the first acoustic feature vector, the second acoustic feature vector and the third acoustic feature vector into an audio large model, performing autoregression prediction through the audio large model to obtain a predicted discrete coding vector, and decoding the predicted discrete coding vector to obtain a repaired second audio clip.
Owner:VIVO MOBILE COMM HANGZHOU CO LTD

Low-power communication method and apparatus for wireless audio device, soc chip, and storage medium

PCT designated stageWO2026175268A1Source encodingData stream
A low-power communication method and apparatus for a wireless audio device, an SoC chip, and a storage medium. The low-power communication method for a wireless audio device comprises: sampling acoustic features of heterogeneous audio data from multiple target audio devices, calculating speech spectra corresponding to device characteristics, extracting multi-device acoustic features, and performing multi-frequency sub-band speech enhancement, heterogeneous segmentation and grouping of the multiple target audio devices, check value calculation, and differentiated encoding on each extracted acoustic feature sequence to obtain a multi-source encoded speech segment; dividing the multi-source encoded speech segment into transmission units, and performing cooperative signal modulation, cooperative resource allocation, and data stream transmission on each transmission unit to obtain a transmitted speech data stream; and acquiring the transmitted speech data stream transmitted with low power, and performing audio recovery on the transmitted speech data stream to obtain a target audio data stream transmitted by the multiple target audio devices, thereby reducing the communication power consumption of the multiple target audio devices and ensuring the audio transmission quality.
Owner:SHENZHEN HESHENGCHENG TECHNOLOGY CO LTD

Lightweight lossless audio encoding and decoding method and system

The invention discloses a lightweight lossless audio encoding and decoding method and system, and belongs to the technical field of audio signal processing. The method comprises the following steps: converting an original audio waveform into a compact coding expression through an encoder formed by connecting a TConv unit, a convolution unit, a down-sampling unit and a Local-Transform unit in series; a single finite scalar quantizer is adopted to discretize the coded representation, and a mixed quantization strategy is introduced to generate a single and consistent discrete token stream; and finally, a high-fidelity audio waveform is reconstructed through a decoder formed by connecting a Local-Transform unit, a TEConv unit, an up-sampling unit and an audio recovery unit in series. Through unit structure optimization and joint loss function training, indexes such as SDR, PESQ, STOI, WER and the like under multiple different code rates are all superior to those of a baseline model, model parameters are only 10.31-11.29 M, collaboration of lightweight deployment and high-fidelity reconstruction is achieved, resource-constrained equipment is adapted to downstream scenes such as voice communication, ASR and the like, and good engineering practicability is achieved.
Owner:XI AN JIAOTONG UNIV

Audio restoration method and device

The embodiment of the invention provides an audio restoration method and device, and relates to the technical field of audio restoration. The method comprises the following steps: carrying out sonic boom detection on a to-be-repaired audio to obtain a sonic boom proportion of the to-be-repaired audio; if the detonation sound proportion is greater than a first threshold value, performing detonation sound restoration on the to-be-restored audio to obtain a first audio; performing voice detection on the first audio to obtain a voice ratio of the first audio; if the voice ratio is greater than a second threshold value, converting the first audio frequency into a first time-frequency domain signal, segmenting the first time-frequency domain signal into a first number of sub-band signals with non-overlapped frequency bands according to the resolution of the first audio frequency, and respectively acquiring spectrum features of the first number of sub-band signals, and performing voice separation on the first audio according to the spectrum features of the sub-band signals to obtain a second audio, and performing tone quality restoration on the second audio to obtain a restoration result of the to-be-restored audio. The embodiment of the invention is used for audio restoration.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Foundation ai model for improving audio quality and enhancing audio characteristics

A system and method for enhancing or restoring audio data utilizing an artificial intelligence module, and more particularly utilizing deep neural networks and generative adversarial networks. The system and method are both able to train the artificial intelligence module to provide for different format and other characteristic-specific transforms for determining how to restore audio to source quality and even beyond. The present invention includes the steps of acquiring source data, pre-processing the source data, implementing the artificial intelligence module, indexing the data, applying transforms, and optimizing the data for a particular audio modality.
Owner:ZON GLOBAL IP INC

End-to-end speech separation algorithm based on speech language model

PendingCN121583282ASpeech analysisAudio restorationVoice source
The invention discloses an end-to-end speech separation algorithm based on a speech language model, and the algorithm comprises the steps: discretizing a continuous audio into a 32-order discrete codebook sequence through a residual vector quantization coder-decoder, and introducing a transcription start symbol lt through an SOT strategy; sOSgt, SOSgt; a special separator is lt; sCgt; and a termination symbol lt; eOSgt, EOSgt; splicing a multi-person voice sequence; extracting audio depth features by using a pre-trained WavLM model, and guiding an autoregression decoder to output a separated zero-order codebook sequence in combination with a cross attention mechanism; predicting a high-order codebook sequence step by step through a non-autoregression model, configuring an independent embedding layer to fuse low-order information, and introducing a task embedding mechanism to optimize modeling; based on a special separator lt; sCgt; and slicing the multi-order discrete codebook sequence, and outputting an independent voice source through an Encodec decoder. According to the method, the intelligibility of voice separation and the audio restoration quality can be effectively improved, the decoding speed is high, the subjective hearing experiment result and the downstream task performance are excellent, the scene that the number of speakers is unknown is supported, and the industrialization application prospect is wide.
Owner:SHANGHAI JIAOTONG UNIV

Audio encoding and decoding method, electronic equipment and program product

The invention provides an audio encoding and decoding method, electronic equipment and a program product, and relates to the technical field of audio processing, and the method comprises the steps: obtaining to-be-processed audio data, and converting the to-be-processed audio data into a spectrogram; inputting the spectrogram into an encoder of a pre-trained encoding and decoding neural network model, and sequentially performing local acoustic feature extraction, intra-frame dependency modeling and inter-frame dependency modeling on the spectrogram to obtain an encoding vector output by the encoder; inputting the coding vector into a residual vector quantizer of a coding and decoding neural network model, and performing multi-layer quantization processing on the coding vector to obtain a codebook index output by the residual vector quantizer; and inputting the codebook index into a decoder of the coding and decoding neural network model, and decoding the codebook index to obtain a reconstructed spectrogram output by the decoder. According to the invention, high-quality audio restoration can be realized in an ultra-low bit rate communication environment.
Owner:IFLYTEK CO LTD

Spatial audio recovery apparatus, spatial audio recovering method, and program

PendingUS20260129390A1Speech analysisCharacter and pattern recognitionMonauralAudio restoration
A spatial audio restoration device of an embodiment includes a video feature amount calculation unit that calculates a video feature amount on the basis of video information, an audio feature amount calculation unit that calculates an audio feature amount on the basis of audio information that is a monaural sound corresponding to the video information, and a coefficient calculation unit that calculates a high-order Ambisonics coefficient on the basis of the video feature amount and the audio feature amount.
Owner:NT T INC

Audio processing method, audio processing device, electronic equipment and storage medium

PendingCN120656473ASpeech analysisSound source separationAudio restoration
The invention relates to an audio processing method, an audio processing device, electronic equipment and a storage medium. The audio processing method comprises the following steps: acquiring an audio to be processed, and acquiring audio features of the audio to be processed; the audio features are subjected to audio separation through a separator, target audio features are obtained, the separator comprises a plurality of separation layers, and each separation layer sequentially executes a processing mode based on expansion convolution, a processing mode based on multi-scale fusion and a processing mode based on a multi-channel attention mechanism; the target audio feature comprises an audio feature corresponding to each sound source in the plurality of sound sources; and performing audio reduction on the target audio feature to obtain a target audio, the target audio comprising an audio corresponding to each sound source in the plurality of sound sources. According to the sound source separation method and device, when sound source separation is executed, a good separation effect is guaranteed, meanwhile, the requirement for equipment performance is lowered, and a sound source separation algorithm can stably run in portable mobile equipment.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Audio coding method, electronic device, and program product

The application provides an audio coding method, an electronic device and a program product, and relates to the technical field of audio processing. The method comprises the following steps: obtaining audio data to be processed, and converting the audio data to be processed into a spectrum graph; inputting the spectrum graph into an encoder of a pre-trained coding and decoding neural network model, sequentially performing local acoustic feature extraction, intra-frame dependence modeling and inter-frame dependence modeling on the spectrum graph, and obtaining an encoding vector output by the encoder; inputting the encoding vector into a residual vector quantizer of the coding and decoding neural network model, performing multi-layer quantization processing on the encoding vector, and obtaining a codebook index output by the residual vector quantizer; and inputting the codebook index into a decoder of the coding and decoding neural network model, performing decoding processing on the codebook index, and obtaining a reconstructed spectrum graph output by the decoder. The application can realize high-quality audio restoration in a communication environment with an ultra-low code rate.
Owner:IFLYTEK CO LTD

Audio restoration method and apparatus

PendingUS20260141909A1Speech analysisFrequency spectrumAudio restoration
Embodiments of this application provide an audio restoration method and apparatus, and relates to the technical field of audio restoration. The method includes: performing pop detection on audio to be restored to obtain a pop proportion of the audio to be restored; performing pop restoration on the audio to be restored to obtain first audio in a case where the pop proportion is greater than a first threshold; performing speech detection on the first audio to obtain a speech proportion of the first audio; converting, in a case where the speech proportion is greater than a second threshold, the first audio into a first time-frequency domain signal, segmenting the first time-frequency domain signal into a first number of sub-band signals with non-overlapping frequency bands according to a resolution of the first audio, respectively obtaining spectrum features of the first number of sub-band signals.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

A valve audio pre-stage processing system

The application relates to the technical field of audio processing, and particularly discloses an electronic tube audio pre-stage processing system, which comprises an audio source input module, an MCU intelligent distribution module, a DSP audio processing module, a feedback suppression module, an echo and reverberation processing module, an electronic tube audio restoration module and an audio output module; in the application, the electronic tube audio restoration module utilizes the characteristics of even harmonics to give the audio a warm and soft listening feeling, obviously improves the sound performance of a traditional digital audio processor which lacks humanization due to excessive precision, a multi-frequency band dynamic compression algorithm guarantees the detail performance of the audio, and effectively prevents distortion, and the application is particularly suitable for processing of high dynamic range audio sources, such as live music recording and high-fidelity sound playing; through real-time spectrum analysis of FFT combined with an adaptive gain control algorithm, the system can dynamically eliminate audio howling and feedback noise, and the performance is particularly outstanding in a complex sound environment, and the system is provided with reverberation denoising and equalization optimization.
Owner:ENPING SHENGLAI ELECTRONIC TECH CO LTD

Audio restoration model processing method, audio restoration method, device, equipment, storage medium and program product

PendingCN121725797ASpeech analysisFrequency spectrumAudio restoration
The invention relates to an audio restoration model processing method, an audio restoration method and device, equipment, a storage medium and a program product, and relates to the technical field of audio and artificial intelligence. Processing the first audio sample of the first sampling frequency according to a second sampling frequency to obtain a second audio sample of the second sampling frequency, and performing distortion simulation processing on the second audio sample to obtain a third audio sample of the second sampling frequency, inputting the third audio sample into a to-be-trained frequency spectrum information restoration network in the audio restoration model to obtain first restoration frequency spectrum information, and processing the first restoration frequency spectrum information and the first frequency spectrum information of the second audio sample according to the first sampling frequency to obtain second restoration frequency spectrum information and second frequency spectrum information, and training the to-be-trained spectrum information repair network based on the second repair spectrum information and the second spectrum information. According to the invention, the audio restoration accuracy of the audio restoration model can be improved.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Audio Restoration Method Based on Nonlinear Interleaving Mapping and Video Microphone System

The present invention provides an audio recovery method based on non - linear interleaving mapping, including: S1. Obtain a video to be processed, preprocess the video to be processed to generate a continuous sequence of grayscale images, where the video to be processed includes vibration images generated by an object affected by surrounding sound sources; S2. Perform non - linear interleaving mapping of pixel coordinates and image brightness for each image in the grayscale image sequence with a two - dimensional trigonometric function interleaving auxiliary function, so as to obtain a set of two - dimensional 0 - 1 matrices corresponding to the current image, and perform dimensionality reduction processing on each two - dimensional 0 - 1 matrix to generate one - dimensional data; S3. After filtering and denoising each of the one - dimensional data, the recovered audio is obtained. The present invention has better scale scalability from a global perspective, can improve the overall detail resolution, and helps to improve the recovery quality of sound details.
Owner:DALIAN UNIV

Low-power communication method, device, SoC chip and storage medium for wireless audio devices

The present invention relates to the field of wireless communication technologies, and discloses a low-power communication method, apparatus, SoC chip and storage medium for wireless audio devices. The method includes: sampling the acoustic features of heterogeneous audio data of multiple target audio devices, calculating the voice spectra corresponding to the device characteristics, and extracting the acoustic features of multiple devices, and performing voice enhancement in multiple frequency subbands, heterogeneous segmentation and grouping of multiple target audio devices, check value calculation and differential coding on the extracted acoustic feature sequences to obtain multi-source coded voice segments; dividing the multi-source coded voice segments into transmission units and performing cooperative signal modulation, cooperative resource allocation and data stream transmission for each transmission unit to obtain a transmitted voice data stream; obtaining the transmitted voice data stream for low-power transmission and performing audio restoration on the obtained data to obtain a target audio data stream transmitted by multiple target audio devices. This application reduces the communication power consumption of multiple target audio devices and ensures the audio transmission quality.
Owner:SHENZHEN HESHENGCHENG TECH CO LTD

Method for repairing sound quality, electronic equipment, storage medium and program product

The invention provides a sound quality repairing method, electronic equipment, a storage medium and a program product, and relates to the technical field of audio. In the method, firstly, acoustic features of first audio data are extracted, and the first audio data can be audio data obtained by processing original audio data; performing local feature extraction and tone quality restoration processing on the acoustic features of the first audio data through a diffusion model to obtain first frequency spectrum data, and sending the first frequency spectrum data to a generation model; performing audio recovery processing on the first spectrum data through the generation model to obtain second audio data; and performing voice filtering processing on the second audio data to obtain repaired third audio data. Through the method provided by the invention, the quality of repairing the audio data is improved, so that the retention rate of voice data processing can be greatly improved.
Owner:IFLYTEK CO LTD

A binaural audio generation method and system that fuses position and audio general representations

ActiveCN117789692BStereophonic systemsSpeech synthesisSound sourcesAudio restoration
The application discloses a kind of binaural audio generation method and system of fusion position and audio general representation, it is characterized in that, including, S1, making video frame data set and audio data set;S2, short-time Fourier transform and calculation are carried out to audio data set, obtain corresponding complex spectrogram, amplitude spectrogram and phase spectrogram;S3, video frame data set, audio data set and its corresponding spectrogram are input into binaural audio restoration model containing relative position information extractor, audio general representation extractor, mask generation module and are trained and optimized;S4, based on the binaural audio restoration model of well-trained, carries out binaural audio restoration.The network model proposed in the present application can effectively extract the relative position information of sound source in video frame, obtain more effective audio general representation, for guiding the generation of binaural audio, to improve system performance.
Owner:XIAMEN UNIV

Foundation ai model for audio recovery and enhancement employing novel segmented recurrent network indexing with a temporal GAN ensemble ai

A system and method for enhancing or restoring audio data utilizing an artificial intelligence module, and more particularly utilizing deep neural networks and generative adversarial networks. The system and method are both able to train the artificial intelligence module to provide for different format and other characteristic-specific transforms for determining how to restore audio to source quality and even beyond. The present invention includes the steps of acquiring source data, pre-processing the source data, implementing the artificial intelligence module, indexing the data, applying transforms, and optimizing the data for a particular audio modality.
Owner:ZON GLOBAL IP INC

Music Audio Restoration Method and System Based on DCT-DDPM

The present invention discloses a music audio repair method and system based on DCT-DDPM, belonging to the fields of speech processing and deep learning, and solving the problem that the prior art can only perform unconditional modification and cannot restore the original segment information. The present invention includes: 1) obtaining music audio data; 2) transforming the audio data into the frequency domain; 3) processing to obtain a Mel spectrogram with a Mask; 4) training DCT-DDPM; 5) repairing the audio based on the trained DCT-DDPM; 6) transforming the repaired Mel spectrogram into the time domain. The present invention is used for music audio repair.
Owner:SICHUAN UNIV

A model training method, an audio processing method, and related devices

PendingCN122337222AAudio restorationSound quality
This invention discloses a model training method, an audio processing method, and related apparatus. First, an audio processing model is constructed, including a pre-trained audio compression and reconstruction module and an audio generation module to be adjusted. Then, the audio compression and reconstruction module processes the original audio features of the fine-tuned sample audio to obtain predicted audio features and their pitch features. Next, the predicted audio features and their pitch features are input into the audio generation module, which outputs predicted audio. Finally, the audio generation module is adjusted based on the difference between the predicted audio and the fine-tuned sample audio to obtain the final audio processing model. Therefore, this invention, by introducing pitch features and predicted audio features for joint training, optimizes the module training logic and parameter adjustment criteria, significantly improving the sound quality and completeness of audio restoration and effectively optimizing the overall audio processing experience.
Owner:CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1

Model training and audio repairing method and device, electronic equipment and storage medium

The invention relates to a model training and audio repairing method and device, electronic equipment and a storage medium, which are used for detecting and repairing an audio abnormal problem in real time. The model training method comprises the following steps: selecting a first sample which comprises a positive sample audio and a corresponding negative sample audio; extracting frequency spectrum image features and time domain features of the negative sample audio, and inputting the features into a compression network of a to-be-trained audio restoration model for dimension reduction compression to obtain first compression features; on the basis of a reconstruction network in an audio restoration model, dimension raising reconstruction is carried out on the first compression feature, and a reconstructed audio is obtained; based on the difference between the reconstructed audio and the positive sample audio, parameter adjustment is carried out on the audio restoration model; inputting the frequency domain feature of the sample audio in the second sample and the second compression feature of the sample audio extracted by the compression network into a to-be-trained audio detection model to obtain an anomaly detection result; and performing parameter adjustment on the audio detection model based on the difference between the abnormal detection result and the corresponding sample tag.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Audio restoration method and system based on multi-dimensional collaboration

The invention discloses an audio restoration method and system based on multi-dimensional collaboration, and belongs to the technical field of audio restoration, and the method comprises the steps: carrying out the feature extraction and damage detection of an input to-be-restored audio, and obtaining a damage distribution map of the to-be-restored audio; performing noise removal, distortion correction and audio repair on the to-be-repaired audio according to the damage type based on the damage distribution map to obtain an audio repair result; and performing quality evaluation on the audio repair result through a preset repair evaluation rule, adjusting repair parameters of the audio repair result according to an evaluation result, and outputting a final repair result. Therefore, by implementing the method and the device, the problems that in the prior art, a large amount of training data is needed to realize audio restoration, and only single damage can be restored can be solved.
Owner:GUANGZHOU BAOLUN ELECTRONICS CO LTD

Audio transmission quality determination method, electronic equipment, storage medium and product

The invention provides an audio transmission quality determination method, electronic equipment, a computer readable storage medium and a computer program product, and relates to the technical field of Internet. The audio transmission quality determination method comprises the steps that transmission information of a target audio between a source end and a target end and playing information of the target audio on the target end are acquired, one of the source end and the target end is a cloud desktop server, and the other one is a cloud desktop client; according to the transmission information and / or the playing information, transmission quality parameters of the target audio are determined, and the transmission quality parameters comprise at least one of an audio transmission delay parameter, an audio reduction degree and an audio lagging rate; the audio transmission delay parameter is used for representing the delay of the target audio from the time when the target audio is sent from the source end to the time when the target audio is played by the target end, and the audio reduction degree is used for representing the difference between the target audio played by the target end and the target audio sent by the source end; the audio lagging rate is used for representing the playing lagging of the target audio by the target end.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Audio data processing method and apparatus, electronic device and storage medium

Provided in the embodiments of the present disclosure are an audio data processing method and apparatus, an electronic device and a storage medium. The method comprises: acquiring Mel spectrum data corresponding to initial audio data; using a band extension module to process the Mel spectrum data, so as to generate enhanced Mel spectrum data, the band extension module comprising a residual denoising diffusion model, and being used to predict a corresponding high-frequency feature on the basis of a low-frequency feature of the Mel spectrum data so as to generate the enhanced Mel spectrum data having the low-frequency feature and the high-frequency feature; and, on the basis of an audio restoration module, processing the enhanced Mel spectrum data to obtain optimized audio data. After the initial audio data is converted into the Mel spectrum data, the band extension module comprising the residual denoising diffusion model is used to extend the high-frequency feature of the Mel spectrum data, so as to form enhanced Mel spectrum data having the high-frequency feature and then, restoration is performed on the basis of the enhanced Mel spectrum data.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Foundation AI model for improving audio quality and enhancing audio characteristics

A system and method for enhancing or restoring audio data utilizing an artificial intelligence module, and more particularly utilizing deep neural networks and generative adversarial networks. The system and method are both able to train the artificial intelligence module to provide for different format and other characteristic-specific transforms for determining how to restore audio to source quality and even beyond. The present invention includes the steps of acquiring source data, pre-processing the source data, implementing the artificial intelligence module, indexing the data, applying transforms, and optimizing the data for a particular audio modality.
Owner:ZON GLOBAL IP INC