A multimodal speech and vision fusion ultrasonic robot control method based on human demonstration
By combining multimodal data acquisition and end-to-end neural networks with human demonstration learning, intelligent closed-loop control of the ultrasound robot is achieved, solving the problems of consistency and intelligence in ultrasound examinations, reducing the workload of doctors, and adapting to individual differences and environmental changes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UESTC (SHENZHEN) ADVANCED RES INST
- Filing Date
- 2026-04-24
- Publication Date
- 2026-07-17
AI Technical Summary
Current ultrasound examinations rely on doctors' experience, resulting in poor operational consistency and high workload for doctors. Existing ultrasound robots lack flexibility and multimodal perception fusion capabilities, making them unable to adapt to individual patient differences and dynamic environments, and their level of intelligence is insufficient.
By collecting multimodal data from ultrasound physicians and using end-to-end deep neural networks for time alignment and spatial calibration, a multimodal fusion ultrasound robot control method is constructed. This method enables precise synchronous acquisition and calibration of voice and visual information, and, combined with human demonstration and learning mechanisms, forms an intelligent closed-loop control.
It improves the standardization and consistency of ultrasound examinations, reduces the risk of missed or misdiagnosed cases, alleviates the workload of doctors, adapts to individual patient differences and dynamic environmental changes, and enables robots to automatically generate precise control commands.
Smart Images

Figure CN122401461A_ABST