A multimodal speech and vision fusion ultrasonic robot control method based on human demonstration

By combining multimodal data acquisition and end-to-end neural networks with human demonstration learning, intelligent closed-loop control of the ultrasound robot is achieved, solving the problems of consistency and intelligence in ultrasound examinations, reducing the workload of doctors, and adapting to individual differences and environmental changes.

CN122401461APending Publication Date: 2026-07-17UESTC (SHENZHEN) ADVANCED RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UESTC (SHENZHEN) ADVANCED RES INST
Filing Date
2026-04-24
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Current ultrasound examinations rely on doctors' experience, resulting in poor operational consistency and high workload for doctors. Existing ultrasound robots lack flexibility and multimodal perception fusion capabilities, making them unable to adapt to individual patient differences and dynamic environments, and their level of intelligence is insufficient.

Method used

By collecting multimodal data from ultrasound physicians and using end-to-end deep neural networks for time alignment and spatial calibration, a multimodal fusion ultrasound robot control method is constructed. This method enables precise synchronous acquisition and calibration of voice and visual information, and, combined with human demonstration and learning mechanisms, forms an intelligent closed-loop control.

Benefits of technology

It improves the standardization and consistency of ultrasound examinations, reduces the risk of missed or misdiagnosed cases, alleviates the workload of doctors, adapts to individual patient differences and dynamic environmental changes, and enables robots to automatically generate precise control commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122401461A_ABST
    Figure CN122401461A_ABST
Patent Text Reader

Abstract

本发明公开了一种基于人类示范的多模态语音与视觉融合超声机器人控制方法。该方法由超声医生完成超声扫查形成示范操作过程,同步采集多模态数据并经对齐标定形成样本对;通过视觉与语言编码网络提取对应特征,对空间位姿轨迹处理得到动作监督标签;构建含多模态融合模块的端到端深度神经网络,以语义特征为调制条件融合视觉特征,经损失函数优化网络参数;在线控制时,实时采集环境图像与语音指令,编码后融合可选预存示范特征生成控制指令,驱动机械臂更新探头位姿,循环执行闭环流程直至完成扫描。本发明使超声机器人学习医生操作经验、响应语义指令并适配复杂环境,降低医生劳动强度,提升超声检查的一致性与精准性。
Need to check novelty before this filing date? Find Prior Art