模型训练方法和基于模型的音频处理方法及对应装置

By performing temporal addition and phase information determination on the left and right channel background audio samples of stereo audio, training data is constructed to train the model, solving the problem of blurred supervision signal in mono-to-stereo rendering and achieving better stereo effect generation.

CN121096362BActive Publication Date: 2026-07-17BEIJING ZITIAO NETWORK TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2025-09-16
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, the training process of mono-to-stereo rendering models is hampered by the combination of the mid-side signal and the prediction side signal, which leads to fuzzy supervision signals and makes it difficult to clearly deconstruct the stereo effect structure, thus affecting the spatial perception restoration effect of stereo audio.

Method used

By extracting the left and right channel background audio samples of stereo audio and adding them in the time domain, the phase information is determined, and a mono input signal and a stereo effect target signal are constructed to form audio training data. Based on this data, a preset model is trained, and the phase information is preserved to enhance the stereo effect.

Benefits of technology

It improves the spatial perception reproduction of stereo audio, enhances the model's ability to learn the differences between the left and right channels, reduces amplitude loss, and improves the generation quality of stereo effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096362B_ABST
    Figure CN121096362B_ABST
Patent Text Reader

Abstract

一种模型训练方法和基于模型的音频处理方法及对应装置,该方法包括:提取立体声音频样本中的第一背景音音频,对左声道背景音音频和右声道背景音音频进行时域相加,得到声道混合音频,确定声道混合音频的第一相位信息和第一背景音音频的第二相位信息;基于第一相位信息和第一背景音音频确定单声道输入信号,基于第二相位信息和第一背景音音频确定立体声效目标信号;基于音频训练数据对初始的预设模型进行模型训练,训练完成的预设模型作为基于单声道音频预测立体声效音频的音频处理模型。通过构造立体声效目标信号,提取出背景音音频在每个声道独特的空间特性,增强立体声效果,并作为监督信号使得预设模型能够更好地学习左右声道的差异特征。
Need to check novelty before this filing date? Find Prior Art