Audio data processing method and device, electronic equipment and storage medium

By acquiring the Mel spectrum of audio data and using a residual denoising diffusion model to predict high-frequency features, the problem of audio quality degradation in existing technologies is solved, achieving higher audio resolution and better listening experience in audio data processing.

CN121963753APending Publication Date: 2026-05-01BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-10-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies, in scenarios such as voice calls, video conferencing, and live video streaming, compress audio data to reduce the overhead of device and network resources, resulting in a decline in audio quality and affecting the user's listening experience.

Method used

By acquiring the Mel spectrum data of the initial audio data, the high-frequency features are predicted using the residual denoising diffusion model in the band extension module to generate enhanced Mel spectrum data. This enhanced Mel spectrum data is then processed by the audio restoration module to generate optimized audio data with higher audio resolution.

Benefits of technology

It improves the audio resolution of audio data, enhances sound details and texture, and improves the user's listening experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963753A_ABST
    Figure CN121963753A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an audio data processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining Mel spectrum data corresponding to initial audio data, processing the Mel spectrum data through a frequency band extension module, and generating enhanced Mel spectrum data, the high-frequency feature prediction module is used for predicting corresponding high-frequency features according to low-frequency features of the Mel-spectrum data so as to generate enhanced Mel-spectrum data with the low-frequency features and the high-frequency features; and processing the enhanced Mel spectrum data based on an audio reduction module to obtain optimized audio data, converting initial audio data into Mel spectrum data, and expanding the high-frequency characteristics of the Mel spectrum data by using a frequency band expansion module comprising a residual denoising diffusion model to form enhanced Mel spectrum data comprising the high-frequency characteristics, and based on the enhanced Mel spectrum data, reduction is carried out to generate optimized audio data with higher audio resolution, so that the listening feeling of a user is improved.
Need to check novelty before this filing date? Find Prior Art