Desktop robot interaction control method and device based on deep learning
By combining sound source information and visual information through deep learning, the problems of complex hardware structure and inaccurate positioning of desktop robots have been solved, achieving efficient sound source localization and hardware simplification, and improving the user interaction experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-13
AI Technical Summary
Existing desktop robot hardware structures are complex and require the deployment of a large number of microphones to achieve accurate sound source localization, resulting in high hardware deployment difficulty and waste of resources. Furthermore, voice wake-up is prone to recognition errors in multi-person scenarios, affecting user experience.
By employing a deep learning-based approach, combining sound source information and visual information to generate lateral and longitudinal deflection information, and using three audio acquisition units and at least one image acquisition unit, the current orientation of the interactive unit is determined, and initial and secondary deflections are performed, thereby improving positioning accuracy and hardware deployment efficiency.
It improves the accuracy of sound source localization and the efficiency of hardware deployment, reduces hardware costs and algorithm design difficulty, and enhances the realism and immersion of user interaction.
Smart Images

Figure CN121649984A_ABST