Audio generation method, apparatus, and electronic device
By performing phoneme conversion, prosody recognition, and emotion recognition on text in the in-vehicle system, synthetic audio is generated, which solves the problem of insufficient anthropomorphic features and improves the naturalness and realism of the audio.
CN122116867APending Publication Date: 2026-05-29BEIJING CO WHEELS TECH CO LTD
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CO WHEELS TECH CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Technical Problem
In vehicle systems with limited computing power, existing TTS technology suffers from insufficient human-like features and poor realism in synthesized speech.
Method used
By performing phoneme conversion, prosody recognition, and emotion recognition on the text, phoneme data, prosody data, and emotion label data are generated. These data are then input into an acoustic model, and the duration of the initial pronunciation interval between phonemes is adjusted to generate synthesized audio.
Benefits of technology
It enhances the anthropomorphism of synthesized audio in terms of tone, pauses, and emotions, improves the naturalness of synthesized audio expression, and enhances the user listening experience.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN122116867A_ABST
Abstract
The application provides an audio generation method, device and electronic equipment, and belongs to the technical field of audio processing. The audio generation method comprises: performing phoneme conversion processing on a text to be processed to obtain phoneme data of the text; performing prosody recognition processing on the text to obtain prosody data of the text; performing emotion recognition processing on the text to obtain emotion label data of the text; inputting the phoneme data, the prosody data and the emotion label data into an acoustic model to obtain acoustic feature data of the text and length data of each phoneme in the phoneme data; inputting the acoustic feature data into a vocoder for speech conversion processing to generate synthesized audio of the text; and adjusting a starting pronunciation interval duration between phonemes in the synthesized audio according to the length data of each phoneme to obtain adjusted synthesized audio. The synthesized audio can improve the degree of personification in terms of tone, pause and emotion, improve the naturalness of expression of the synthesized audio, and improve the listening experience of a user.
Need to check novelty before this filing date? Find Prior Art