AI-based panoramic audio generation method, system, and storage medium

By extracting multimodal features and modeling spherical harmonic functions, combined with adaptive rendering algorithms, high-precision panoramic sound audio is automatically generated, solving the problems of low efficiency and inaccurate positioning in traditional panoramic sound generation, and achieving efficient adaptation and immersive experience in complex acoustic environments.

CN121547723BActive Publication Date: 2026-05-26SHANGHAI RUIHEFENG ELECTRONIC TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI RUIHEFENG ELECTRONIC TECHNOLOGY CO LTD
Filing Date
2025-12-17
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Traditional panoramic sound audio generation relies on manual intervention, resulting in long production cycles, high costs, and unstable spatial positioning accuracy. Existing AI technology lacks accurate modeling of three-dimensional spatial characteristics and device adaptability, making it unable to effectively reproduce the sound field in complex acoustic environments.

Method used

By using multimodal feature extraction, spherical harmonic function modeling, and adaptive rendering algorithms, the system achieves automated fusion and conversion of the original audio signal and scene parameters, constructs a high-precision 3D spatial sound field model, and adaptively renders multi-channel panoramic audio based on the target device parameters.

Benefits of technology

It achieves efficient and high-precision adaptive generation of panoramic audio, solving the problems of low efficiency, high cost and insufficient spatial positioning accuracy in traditional methods, and providing an immersive auditory experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547723B_ABST
    Figure CN121547723B_ABST
Patent Text Reader

Abstract

This invention discloses an artificial intelligence-based method, system, and storage medium for generating panoramic sound audio. The method includes the following steps: S1. Inputting the original audio signal and scene parameters, obtaining the time-domain-frequency domain features and spatial scene features of the original audio signal through a multimodal feature extraction algorithm, and weighted fusing the time-domain-frequency domain features and spatial scene features to obtain fused features; S2. Constructing the mathematical expression basis of the 3D spatial sound field model based on the spherical harmonic function, mapping the fused features to three-dimensional spatial coordinates through a sound field modeling algorithm to obtain the 3D spatial sound field model; S3. Using an adaptive rendering algorithm, converting the 3D spatial sound field model into multi-channel panoramic sound audio output according to the parameters of the target playback device. This invention achieves automated panoramic sound audio generation by inputting the original audio signal and scene parameters for feature extraction, fusion, and sound field modeling, and converting the output based on adaptive rendering. It has the advantages of automation, high efficiency, and high precision.
Need to check novelty before this filing date? Find Prior Art