This invention discloses a method, device, and
robot for non-dialogue intent rejection in
robot interaction based on multimodal
processing. The method is executed by a
robot system, which includes a multimodal
perception module, an analysis module, a
memory module, and a
response control module. This invention constructs a three-dimensional human-
machine spatial relationship through
millisecond-level synchronization of multi-source heterogeneous data and cross-perspective spatial modeling, accurately filtering
background noise and non-target commands, significantly improving anti-interference capabilities. Simultaneously, it integrates multi-dimensional features such as
semantics, emotion, and body language, and combines spatiotemporal trajectory memory maps for historical context analysis, deeply analyzing the user's implicit needs. Furthermore, it optimizes response timing based on dialogue
rhythm and dynamically adjusts strategies by monitoring
negative feedback in real time through a
feedback loop mechanism, achieving natural, smooth, highly adaptive, and self-correcting intelligent human-
machine interaction.