Voice data processing method and computer program product
By generating and fusing target multi-turn dialogue data and human voice alignment data, the problem of high-cost and inefficient acquisition of speech training data is solved, and high-quality and diverse speech training data is achieved, providing stronger recognition support for speech agents in noisy environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHENGDU BOSS INNOVATION TECH CO LTD
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies suffer from high costs, low efficiency, lack of real-world scenario information, and low fidelity when acquiring high-quality speech training data, especially in noisy environments where the recognition accuracy of speech agents is insufficient.
By acquiring target multi-turn dialogue data and human voice alignment data, initial clean speech data is generated and fused with real environmental noise and background human voice data to simulate a complex acoustic environment and generate high-fidelity, strongly aligned speech training data.
It significantly reduces data acquisition costs, improves the realism and diversity of data, enhances the robustness and recognition accuracy of voice training data, and supports stronger voice agent recognition capabilities in noisy environments.
Smart Images

Figure CN122417044A_ABST