Desktop robot interaction control method and device based on deep learning

By combining sound source information and visual information through deep learning, the problems of complex hardware structure and inaccurate positioning of desktop robots have been solved, achieving efficient sound source localization and hardware simplification, and improving the user interaction experience.

CN121649984APending Publication Date: 2026-03-13SHENZHEN TIANJING YUHONG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-04
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing desktop robot hardware structures are complex and require the deployment of a large number of microphones to achieve accurate sound source localization, resulting in high hardware deployment difficulty and waste of resources. Furthermore, voice wake-up is prone to recognition errors in multi-person scenarios, affecting user experience.

Method used

By employing a deep learning-based approach, combining sound source information and visual information to generate lateral and longitudinal deflection information, and using three audio acquisition units and at least one image acquisition unit, the current orientation of the interactive unit is determined, and initial and secondary deflections are performed, thereby improving positioning accuracy and hardware deployment efficiency.

Benefits of technology

It improves the accuracy of sound source localization and the efficiency of hardware deployment, reduces hardware costs and algorithm design difficulty, and enhances the realism and immersion of user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121649984A_ABST
    Figure CN121649984A_ABST
Patent Text Reader

Abstract

The invention provides a desktop robot interaction control method and device based on deep learning, and the method comprises the steps: determining the current orientation of an interaction unit when a wake-up audio is received; wherein the current orientation comprises a current horizontal orientation and a current vertical orientation; determining first deflection information of the interaction unit and sound source horizontal position information of the wake-up audio according to the current horizontal orientation and the reception time of the wake-up audio received by each audio acquisition unit, and driving the interaction unit to deflect for the first time according to the first deflection information; and acquiring image information of a sound source target, generating second deflection information of the interaction unit according to the image information and the current vertical orientation, and driving the interaction unit to perform secondary deflection according to the second deflection information. The accuracy of the deflection angle is improved, and meanwhile the deployment difficulty of internal hardware of the desktop robot is lowered.
Need to check novelty before this filing date? Find Prior Art