A method, device, and robot for dialogue-free intent rejection in robot interaction based on multimodal approaches.
By combining multimodal perception, modeling, and memory modules, the robot can accurately identify non-dialogue intentions in complex environments, solving the problem of false responses in existing technologies and achieving higher rejection accuracy and natural interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BENZIJUZHI (SHANGHAI) TECHNOLOGY CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies struggle to effectively identify non-dialogue intentions in robot interactions within complex and dynamic scenarios, leading to false responses and insufficient intelligence, as well as a lack of comprehensive judgment capabilities regarding the scenario, the object, and the topic of the dialogue.
A multimodal perception module is used to collect multidimensional data. Through cross-view spatial modeling and multidimensional feature analysis, combined with a memory module to record interaction history, the effectiveness of the interaction intent is evaluated by a response control module, and rejection or response is executed.
It improves the anti-interference capability and rejection accuracy in complex environments, realizes deep intent understanding and personalized services, and enhances the natural fluency of interaction and user comfort.
Smart Images

Figure CN122125673A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method for rejecting robot dialogue intent, and more particularly to a method, apparatus, and robot for rejecting robot dialogue intent based on multimodal interaction. Background Technology
[0002] With the development of artificial intelligence technology, human-like natural voice interaction has become an important function of intelligent robots. To achieve human-computer interaction, existing technologies typically employ three main approaches to process voice commands and attempt to filter invalid input: first, voice rejection based on sound source localization, which uses a microphone array to calculate the angle and distance of the sound source to determine its location; second, intent judgment based on contextual association, which combines historical dialogues and application scenarios to analyze semantic relevance; and third, rejection based on multimodal information, which attempts to integrate sound features, voiceprints, text semantics, and visual information such as facial orientation to assist in judgment. These methods can, to some extent, distinguish valid commands from environmental noise, improving the accuracy of interaction in fixed scenarios.
[0003] However, existing technologies still have significant shortcomings in anthropomorphic interactions in complex and dynamic scenarios. First, the interaction logic of robots generally exhibits a passive "hear and respond" mode, lacking the ability to comprehensively judge the scene, object, and topic of conversation. This makes them prone to misinterpreting casual conversation among family members or background chatter as commands, resulting in ineffective responses. Second, existing multimodal methods are mostly limited to the binary distinction between commands and noise. The application of visual information is limited to facial recognition or simple direction judgment, lacking a deep understanding of body language, facial expressions, tone of voice, and the rhythm of conversation. This makes it difficult to handle commands with hesitant tones or flexible timing of interjections. Most importantly, existing technologies do not effectively utilize "spatial intelligence" to assist decision-making. They cannot combine changes in spatial positional relationships across perspectives and multimodal historical memory to accurately determine non-conversational intentions, resulting in a rigid and insufficiently intelligent interactive experience when robots face non-target users or complex environments. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a method, device and robot for rejecting non-dialogue intentions in robot interaction based on multimodal interaction, which solves the problem that robots cannot effectively reject non-dialogue intentions when facing complex environments.
[0005] The technical problem to be solved by this invention is achieved by the following technical solution: This invention provides a method for rejecting non-dialogue intent in robot interaction based on multimodal interaction, which is based on the execution of the robot system; The robot system includes a multimodal perception module, an analysis module, a memory module, and a response control module; The method includes the following steps: The multimodal perception module collects multidimensional data about the environment and the user, and performs time synchronization processing. The analysis module performs cross-perspective spatial modeling and multi-dimensional feature analysis based on the synchronized data, and then integrates and generates the probability of interactive intent. The memory module records the spatiotemporal trajectory and semantic information during the interaction process, and provides historical association information as a reference based on the current scene characteristics; The response control module combines the probability of the interaction intent with the historical association information to evaluate the effectiveness and priority of the interaction intent. If it is determined that there is no intention to engage in dialogue, then rejection will be enforced. If the intent is deemed valid, a corresponding interactive response is generated and executed.
[0006] As a preferred embodiment of the present invention, the multimodal sensing module includes a first acquisition module, a second acquisition module, a sound pickup module, a motion state acquisition module, and a data synchronization module; The process of acquiring multidimensional data about the environment and the user through the multimodal perception module and performing time synchronization processing specifically includes: The first acquisition module is used to acquire global environmental image data from multiple indoor perspectives; The second acquisition module is used to acquire depth and distance data from the robot's perspective. The sound pickup module is used to acquire sound source localization and speech signal data; The motion state acquisition module is used to acquire the robot's own posture and acceleration data; The data synchronization module aligns and synchronizes the data streams collected by the different sensors based on a unified timestamp.
[0007] As a preferred embodiment of the present invention, the analysis module includes a spatial relationship modeling module, a semantic understanding module, a recognition module, an action analysis module, and a fusion strategy module; The process involves using the analysis module to perform cross-perspective spatial modeling and multi-dimensional feature analysis based on synchronized data, and then fusing these to generate interaction intent probabilities. Specifically, this includes: The spatial relationship modeling module maps two-dimensional images into three-dimensional spatial relationships and calculates the relative positions of humans and machines. The semantic understanding module parses the explicit meaning and implicit requirements of voice commands. The recognition module extracts facial features to analyze the user's emotional state. The motion analysis module tracks key points of the limbs to identify user behavior patterns; the fusion strategy module integrates the above spatial, semantic, emotional and motion features to calculate the probability distribution of the user's current interaction intent.
[0008] As a preferred technical solution of the present invention, the memory module includes a trajectory recording module, a relationship archiving module, a semantic content indexing module, a timeline archive module, and a multimodal association storage module; The recording of the spatiotemporal trajectory and semantic information during the interaction process through the memory module specifically includes: The trajectory recording module stores the dynamic changes in the sound source location, user distance, and gaze direction. The relationship archiving module stores the sequence of changes in the human-machine spatial position relationship. The semantic content indexing module associates and stores the dialogue content and its background information. The timeline archive module records the fluctuation trend of user emotions over time. The multimodal association storage module packages the above information into memory units for retrieval and retrieval during subsequent interactions.
[0009] As a preferred technical solution of the present invention, the steps performed by the response control module 400 include intent priority evaluation S401, response timing optimization S402, response content generation S403, action planning and execution S404, and feedback closed loop S405. The specific process for generating and executing the corresponding interactive response is as follows: First, the intent priority assessment (S401) and response timing optimization (S402) are performed. Then, the response content is generated (S403), and the action planning and execution are performed based on the generated content (S404). After the action planning and execution step S404, the feedback loop step S405 is entered, where the user's feedback characteristics are monitored and judged in real time: If negative feedback from users is detected, return to step S403 to generate response content, adjust the strategy and regenerate the response; If no negative feedback is detected, proceed to step S406 to complete the response and end the interaction.
[0010] The present invention also provides an apparatus for performing the aforementioned method for rejecting non-dialogue intent in robot interaction based on multimodal methods, the apparatus being installable on the robot system or externally.
[0011] The present invention also provides a robot for performing the aforementioned method for non-dialogue intent rejection in multimodal robot interaction, having or being connected to the aforementioned device.
[0012] The beneficial effects of this invention are: 1. Significantly Enhances Anti-interference Capability and Rejection Accuracy in Complex Environments. This invention achieves millisecond-level alignment of multi-source heterogeneous data, including the global indoor view, the robot's local view, auditory signals, and its own motion state, through a data synchronization mechanism. Unlike traditional rejection methods that rely solely on sound source localization or single vision, this solution utilizes cross-view spatial modeling technology to construct a three-dimensional human-machine spatial relationship. This means that the robot can not only "hear" sounds but also accurately determine whether the sound source is within the effective interaction area and whether the speaker exhibits accompanying interactive behaviors such as "gazing" or "pointing." This allows for precise filtering of background noise, idle chatter, or untargeted self-talk in noisy home or office environments, effectively solving the problem of accidental triggering of "interrupting conversations upon hearing sounds." 2. Achieving Deep Intent Understanding with "Spatial Intelligence": This invention breaks through the limitations of traditional voice interaction, which only focuses on textual semantics. Through an analysis module, it integrates multi-dimensional features such as semantics, emotion, body language, and spatial relationships. By recognizing subtle changes in the user's facial expressions (such as hesitation or aversion) and body language (such as waving in refusal or leaning forward), the system can parse the implicit needs behind voice commands. This deep integration of non-verbal cues enables the robot to comprehensively judge the validity of interactive intentions like a human, avoiding misjudgments caused by the lack of information from a single modality. 3. This invention innovatively designs a multimodal memory module capable of recording the spatiotemporal trajectory (such as movement paths and gaze patterns) and emotional history of interactive objects. This "memory map" enables the robot to make decisions not by viewing a single instruction in isolation, but by combining historical context for comprehensive judgment. For example, the system can use the user's consistent habits or recent spatial changes to help determine the true meaning of a vague instruction, thereby providing more personalized and coherent intelligent services. 4. The robot possesses a natural and fluid interaction rhythm and adaptive capabilities. Through the intent priority evaluation and response timing optimization mechanism in the response control module, the robot can intelligently choose when to interrupt or respond based on the conversation rhythm, avoiding interrupting the user and making human-computer dialogue closer to natural human-to-human interaction. More importantly, this invention introduces a feedback closed-loop mechanism, which can capture negative user feedback (such as frowning or dissatisfaction in tone) in real time and dynamically adjust the response strategy accordingly. This self-correcting ability not only improves the error tolerance rate but also significantly enhances user comfort and satisfaction during prolonged interactions. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the structure and process of the present invention; Figure 2This is a schematic diagram of the architecture of the present invention; In the diagram: 100, Multimodal Perception Module; 101, First Acquisition Module; 102, Second Acquisition Module; 103, Sound Pickup Module; 104, Motion State Acquisition Module; 105, Data Synchronization Module; 200, Analysis Module; 201, Spatial Relationship Modeling Module; 202, Semantic Understanding Module; 203, Recognition Module; 204, Action Analysis Module; 205, Fusion Strategy Module; 300, Memory Module; 301, Trajectory Recording Module; 302, Relationship Archiving Module; 303, Semantic Content Indexing Module; 304, Timeline Archive Module; 305, Multimodal Association Storage Module; 400, Response Control Module; S401, Intent Priority Evaluation; S402, Response Timing Optimization; S403, Response Content Generation; S404, Action Planning and Execution; S405, Feedback Loop; S4051, Feedback Processing; S406, Response Completion. Detailed Implementation
[0014] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0015] Example 1 like Figure 1 As shown, this embodiment provides a method for rejecting non-dialogue intent in robot interaction based on multimodal perception. This method is executed by a robot system. The robot system includes a multimodal perception module 100, an analysis module 200, a memory module 300, and a response control module 400. The method includes the following steps: the multimodal perception module 100 collects multidimensional data of the environment and the user and performs time synchronization processing; the analysis module 200 performs cross-view spatial modeling and multidimensional feature analysis based on the synchronized data, and fuses and generates an interaction intent probability; the memory module 300 records the spatiotemporal trajectory and semantic information during the interaction process, and provides historical association information as a reference based on the current scene characteristics; the response control module 400 combines the interaction intent probability and historical association information to evaluate the validity and priority of the interaction intent; if it is determined to be a non-dialogue intent, rejection is performed; if it is determined to be a valid intent, a corresponding interaction response is generated and executed.
[0016] The multimodal perception module 100 includes a first acquisition module 101, a second acquisition module 102, a sound pickup module 103, a motion state acquisition module 104, and a data synchronization module 105. The multimodal perception module 100 acquires multidimensional data of the environment and the user, and performs time synchronization processing. Specifically, it uses the first acquisition module 101 to acquire global environmental image data from multiple indoor perspectives; uses the second acquisition module 102 to acquire depth and distance data from the robot's perspective; uses the sound pickup module 103 to acquire sound source localization and voice signal data; uses the motion state acquisition module 104 to acquire the robot's own posture and acceleration data; and uses the data synchronization module 105 to align and synchronize the data streams acquired by the different sensors based on a unified timestamp.
[0017] The analysis module 200 includes a spatial relationship modeling module 201, a semantic understanding module 202, a recognition module 203, an action analysis module 204, and a fusion strategy module 205. Based on synchronized data, the analysis module 200 performs cross-perspective spatial modeling and multi-dimensional feature analysis, and fuses these to generate an interaction intent probability. Specifically, this includes: mapping two-dimensional images to three-dimensional spatial relationships using the spatial relationship modeling module 201 to calculate the relative position of the human and machine; parsing the explicit meaning and implicit needs of voice commands using the semantic understanding module 202; extracting facial features to analyze the user's emotional state using the recognition module 203; tracking key points of the limbs using the action analysis module 204 to identify user behavior patterns; and integrating the above spatial, semantic, emotional, and action features using the fusion strategy module 205 to calculate the probability distribution of the user's current interaction intent.
[0018] Regarding cross-view spatial relationship modeling, this embodiment uses a multi-view camera system to reconstruct the coordinates of a person in 3D space from 2D image points of multiple viewpoints using known camera intrinsic and extrinsic parameters. The algorithm first involves camera calibration, employing methods such as Zhang Zhengyou's calibration to obtain the intrinsic and extrinsic parameters of each camera. In each frame processing, the system uses pose estimation algorithms such as OpenPose, MediaPipe, or YOLO-Pose to detect human keypoints on each camera image and obtain their homogeneous coordinates. Subsequently, a 2D-to-3D mapping method is used to solve for the 3D coordinates of these keypoints in a unified world coordinate system. Specifically, the system constructs a projection matrix for each camera, projecting 3D points in the world coordinate system onto the camera image, and establishes a system of direct linear transformation (DLT) linear equations accordingly. By stacking the constraints of two or more observation cameras, a corresponding coefficient matrix is formed. To obtain the optimal solution, the system performs singular value decomposition (SVD) on this matrix, selecting the column vector corresponding to the minimum singular value as the 3D point in homogeneous coordinates, and finally converting it to Euclidean coordinates. By repeating this process for all key points of the human body, a complete skeletal model can be reconstructed in global 3D space. This skeleton can be directly used to calculate all spatial relationships, including the pelvic center position, shoulder orientation vector, interpersonal distance, and relative pose between the human and the robot, providing accurate spatial data support for subsequent intent analysis, as detailed below: Consider using a multi-view camera system, assuming there are N cameras, each with known intrinsic and extrinsic parameters. The goal is to reconstruct the coordinates of a person in 3D space using 2D image points from multiple viewpoints. The algorithm flow is as follows: Camera calibration For each camera, the intrinsic and extrinsic parameters are obtained using, for example, the Zhang Zhengyou calibration method. Its intrinsic parameter matrix is External reference is
[0019] 2. 2D Human Key Point Detection In each frame, a set of human keypoints is detected on the image from each camera using a reliable 2D human pose estimation algorithm (such as OpenPose, MediaPipe, YOLO-Pose, etc.). For example, the coordinates of keypoint k (such as "nose") on the image from camera i are... (homogeneous coordinates).
[0020] 3. 3D Keypoint Reconstruction The two-dimensional to three-dimensional mapping method involves solving for the three-dimensional coordinates of the key point k in a unified world coordinate system W.
[0021] Construct the projection matrix for each camera
[0022]
[0023] 3D points in the world coordinate system The image is projected onto camera ii.
[0024] (2) Establish the DLT linear equation system
[0025] For each observation camera, by projection relation Two linear constraints can be derived.
[0026]
[0027] in , , yes of three lines.
[0028] Stacking the constraints of all cameras with N≥2, we get
[0029] (3) Solving for 3D points
[0030] Perform singular value decomposition on A:
[0031] Take the last column of V (corresponding to the smallest singular value), which is the 3D point in homogeneous coordinates:
[0032] Convert to Euclidean coordinates
[0033] This shows the 3D positions of the key points of the human body in a unified world coordinate system. Repeating the above process for K key points of a human body will yield its complete skeleton in global 3D space.
[0034] This skeleton can be directly used to calculate: human position (such as pelvic center), orientation (shoulder → head vector), human-to-human distance, human-robot relative pose, and all other spatial relationships.
[0035] Regarding the multimodal fusion strategy, this embodiment designs a multimodal fusion framework based on graph neural networks (GNNs). By mapping heterogeneous sensor data such as visual, auditory, and spatial data to a graph structure, it overcomes the limitations of single-modality fusion in complex environments and improves the robustness of intent recognition. The framework defines four core node types: visual nodes contain body movements and facial expression features; auditory nodes cover speech content, tone, and sound source localization information; spatial nodes record the human-machine three-dimensional distance, global coordinates, and robot posture angles; and semantic nodes reflect the user's emotional state and intent probability distribution. Nodes are associated through spatiotemporal edges, modal interaction edges, and semantic edges, and weights are calculated using time intervals, spatial distances, and multi-head attention mechanisms to reflect the interaction strength and spatiotemporal continuity between different information. The graph construction process sequentially involves node initialization, edge generation, weight calculation, and structural optimization using the HGSampling algorithm. At the feature processing level, the system performs targeted dimensionality reduction and fusion on the high-dimensional original features: limb movement features are reduced to 30 dimensions through principal component analysis (PCA), spatial features are compressed to 5 dimensions through multilayer perceptron (MLP), and auditory features are composed of directly invoked speech content vectors and dynamically weighted speech localization features. Finally, the features from each modality are concatenated to form a 91-dimensional final feature vector, thereby achieving deep integration of multimodal information and accurate intent analysis, as detailed below: 1. Multimodal fusion framework based on graph neural networks In intelligent robot interaction systems, accurately understanding user intent is crucial for achieving natural interaction. Traditional methods typically employ single-modal data (such as visual or auditory) for intent analysis, but single-modal approaches have limitations; for example, visual data can be affected by occlusion, and auditory data can be interfered with by environmental noise. Multimodal fusion can combine the advantages of different sensors, improving the robustness and accuracy of intent recognition. However, multisensor data exhibits heterogeneity, temporal sequence, and spatial correlation, making it difficult for traditional fusion methods to effectively capture these complex relationships. Graph Neural Networks (GNNs) can model the relationships between nodes through graph structures, making them particularly suitable for handling interactions between multimodal data. This invention designs a multimodal fusion framework based on graph neural networks for spatial intelligent multimodal perception systems based on multisensor fusion. This framework maps heterogeneous sensor data such as visual, auditory, and spatial data onto a graph structure, and by defining appropriate node types, edge types, and feature vectors, achieves effective fusion and intent analysis of multimodal information.
[0036] 2. Definition of Graph Structure 2.1 Node Types Node types are the foundation of graph structures. Based on system requirements, the following four node types are defined:
[0037] Visual nodes: Composed of the 3D coordinates (x, y, z) and confidence scores (c ∈ [0, 1]) of 25 human keypoints detected by OpenPose, these are reconstructed and standardized to form limb motion nodes (75-dimensional). Simultaneously, the expression analysis module extracts the Valence-Arousal dimension (2-dimensional) of facial expressions.
[0038] Auditory nodes: Speech content is generated into 768-dimensional sentence vectors through the BERT model. Speech intonation features are extracted as MFCC (20-dimensional), fundamental frequency (F0) (1-dimensional), and energy features (1-dimensional). Sound source localization includes direction angle (θ, φ) (2-dimensional), distance (d) (1-dimensional), and signal-to-noise ratio (SNR) (1-dimensional).
[0039] Spatial nodes: user-robot 3D distance (dx, dy, dz) measured by ToF camera (3D), global coordinates (X, Y, Z) generated by SLAM algorithm (3D), and robot attitude angles (pitch, roll, yaw) output by IMU sensor (3D).
[0040] Semantic nodes: The emotional state (such as relaxed, tense) output by the sentiment analysis module is extracted by the Valence-Arousal dimension (2-dimensional) through the CNN-RNN model, and the intention probability distribution (num_intents dimension) output by the intention classification module.
[0041] 2.2 Edge Types Edge types define the relationships between nodes. Based on system requirements, the following three edge types are defined:
[0042] Spatiotemporal edges: connect the limb action nodes and spatial location nodes of the same user at different times. The weight is calculated by the weighted reciprocal of the time interval (Δt) and spatial distance (Δs), reflecting the continuity in time and space.
[0043] Modal interaction edges: These connect nodes across different modalities, such as physical action nodes and voice command nodes, facial expression nodes and sentiment analysis nodes, etc. Weights are calculated using a multi-head attention mechanism to reflect the intensity of interaction between modalities.
[0044] Semantic edges: connect semantically related nodes, such as sentiment analysis nodes and spatial orientation nodes. The weights are calculated using relative time encoding (RTE) and spatial distance to reflect the degree of semantic association.
[0045] 2.3 Graph Construction Process The graph construction process is as follows: Node initialization: Extract feature vectors from each sensor and construct nodes of the corresponding type.
[0046] Edge generation: Spatiotemporal edges are generated based on node type and timestamp; modal interaction edges and semantic edges are generated based on semantic association.
[0047] Edge weight calculation: Calculate the weight for each edge to reflect the strength of the association between nodes.
[0048] Graph structure optimization: The graph structure is optimized using the HGSampling algorithm, while retaining key nodes and edges.
[0049] 3. Input Feature Vector Design 3.1 Feature Extraction and Processing Visual features: The 3D coordinates (x, y, z) and confidence scores (c ∈ [0, 1]) of 25 body keypoints detected by OpenPose were reconstructed (with the neck as the origin) and standardized to form limb motion node features (75-dimensional). The expression analysis module extracted the Valence-Arousal dimension (2-dimensional) of facial expressions.
[0050] Auditory features: The speech content is generated into a 768-dimensional sentence vector through the BERT model. The speech intonation features are extracted as MFCC (20-dimensional), fundamental frequency (F0) (1-dimensional), and energy features (1-dimensional). The sound source localization includes direction angle (θ, φ) (2-dimensional), distance (d) (1-dimensional), and signal-to-noise ratio (SNR) (1-dimensional).
[0051] Spatial features: user-robot 3D distance (dx, dy, dz) measured by ToF camera (3D), global coordinates (X, Y, Z) generated by SLAM algorithm (3D), and robot attitude angles (pitch, roll, yaw) output by IMU sensor (3D).
[0052] Semantic features: The emotional state (such as relaxed, tense) output by the sentiment analysis module is extracted by the Valence-Arousal dimension (2-dimensional) through the CNN-RNN model, and the intention probability distribution (num_intents dimension) output by the intention classification module.
[0053] 3.2 Feature Dimensionality Reduction and Fusion Since the original features have high dimensionality, dimensionality reduction is required. Visual feature dimensionality reduction: PCA was used to reduce the limb action node features (75 dimensions) to 30 dimensions, while retaining 95% of the variance.
[0054] Spatial feature dimensionality reduction: through MLP (such as nn.Linear(9,5) The spatial node features (9 dimensions) are compressed to 5 dimensions.
[0055] Auditory feature processing: Speech content node features (768 dimensions) are used directly without additional dimensionality reduction; speech intonation features (22 dimensions) and sound source localization features (4 dimensions) are concatenated to form 26 dimensions, which are dynamically weighted through an attention mechanism.
[0056] Final feature vector: The reduced visual features (30 dimensions), the processed auditory features (26 dimensions), and the spatial features (5 dimensions) are concatenated to form a final feature vector of 91 dimensions.
[0057] The memory module 300 includes a trajectory recording module 301, a relationship archiving module 302, a semantic content indexing module 303, a timeline archive module 304, and a multimodal association storage module 305. The memory module 300 records the spatiotemporal trajectory and semantic information during the interaction process. Specifically, the trajectory recording module 301 stores the dynamic changes in sound source location, user distance, and eye direction; the relationship archiving module 302 saves the sequence of changes in human-computer spatial positional relationships; the semantic content indexing module 303 associates and stores dialogue content and its background information; the timeline archive module 304 records the fluctuation trend of user emotions over time; and the multimodal association storage module 305 packages the above information into memory units for retrieval and recall during subsequent interactions.
[0058] Regarding the processing of memory units, the core of this embodiment lies in encapsulating heterogeneous data such as audio, video, text, and spatial relationships within the same time period into "memory units," and utilizing an efficient association mechanism to achieve unified storage and rapid retrieval of cross-modal information. This design simulates the characteristics of human memory networks, providing robots with detailed historical references and solving the problems of traditional systems in heterogeneous data integration and query efficiency. A memory unit represents a complete interaction state within a specific time period. It achieves precise alignment of multimodal data through timestamps and entity identifiers, and leverages metadata and association relationships to achieve efficient management of cross-modal information. In terms of data structure, the memory unit adopts a layered design: the top layer records metadata such as unique identifiers, timestamps, interaction subjects, and environment, including semantic embedding vectors and contextual associations; the middle layer records the features and embedding vectors of each modality, such as audio sampling rate, video facial embedding, text sentiment scoring, and specific parameters such as 3D position; the bottom layer stores compressed original data paths or references. This architecture improves the system's flexibility in acquiring information and its query efficiency. In terms of packaging strategy, the system adopts a dual-anchoring association method of "timestamp + entity ID" to ensure precise spatiotemporal alignment. It establishes an entity relationship graph by assigning unique identities and voice IDs to users, achieving entity-centric memory storage. Furthermore, the system introduces an event-driven mechanism, using specific events such as user silence, action completion, or command execution as boundaries of memory units, achieving comprehensive and orderly recording of interaction history. Specific settings are as follows: 1. Meaning of memory units The core of multimodal associative storage lies in packaging heterogeneous data such as audio, video, text, and spatial relationships within the same time period into "memory units," and achieving unified storage and rapid retrieval of cross-modal information through an efficient data association mechanism. This design not only simulates the characteristics of human memory networks but also provides robots with rich historical references. Traditional systems typically store this data separately in different databases, leading to difficulties in data integration and low query efficiency when cross-modal analysis is required. This invention designs a "memory unit" system that can align multiple modal data temporally and semantically, enabling comprehensive recording and efficient retrieval of interaction history.
[0059] 2. Memory Unit Architecture Design (1) Definition of memory unit A memory unit represents the complete interactive state at a specific point in time or within a time period. Each memory unit contains the following core elements:
[0060] The core design of the memory unit achieves precise alignment of multimodal data through timestamps and entity identifiers, and efficient management of cross-modal information through metadata and relationships. This architecture enables robots to recall past interaction scenarios based on temporal cues and spatial relationships, just like humans.
[0061] (2) Memory unit data structure The memory unit adopts a hierarchical data structure design, which includes the following layers: # Top-level metadata of memory units: Records the basic attributes and associated information of memory units. { "memory_id": "UUID", # Unique identifier "timestamp": "timestamp", # Interaction time "duration": 3000, # Interaction duration (milliseconds) "user_id": "User's unique identifier", # User ID "robot_id": "Unique identifier for the robot", # Robot ID "environment": "living room", # Environment description "modality_types": ["audio", "video", "text", "spatial"], # List of modality types "next_memory_id": "UUID", # Next memory unit ID "previous_memory_id": "UUID", # Previous memory cell ID "related_memory_ids": ["UUID1", "UUID2"], # List of related memory unit IDs "embedding": "vector representation" # semantic embedding vector } # Mid-level relational data: Records the identifiers and characteristics of different modalities of data { "audio": { "audio_id": "UUID", # Audio data ID "duration": 5000,# Audio duration (milliseconds) "sample_rate": 16000, # Sampling rate (Hz) "embedding": "vector representation" # Audio embedding vector }, "video": { "video_id": "UUID", # Video data ID "frame_count": 150, # Frame count "fps": 30,# Frame rate (frames per second) "spatial_embeddings": ["vector1", "vector2"], # List of spatially embedded vectors "facial_embeddings": ["vector1", "vector2"]# List of facial embedding vectors }, "text": { "text_id": "UUID", # Text data ID "content": "Turn on the living room lights", # Text content Sentiment: 0.75, # Sentiment score (0-1) "embedding": "vector representation" # Text embedding vector }, "spatial": { "spatial_id": "UUID", # Spatial data ID "position": {"x": 2.5, "y": 1.2, "z": 0.0}, # 3D position coordinates "orientation": {"pitch": 0.0, "roll": 0.0, "yaw": 35.0}, # Orientation angle (radians) "distance": 1.8,# Distance from the main body (meters) "environment_map": "Environment map reference" # Environment map path } } # Underlying raw data: Stores compressed raw data { "audio_data": "Audio file path or reference", # Compress audio data "video_data": "video file path or reference", # Compress video data "environment_data": "Environment map path or reference"# Environment data } This hierarchical structure allows the system to flexibly obtain information at different levels according to needs, improving query efficiency and data management flexibility.
[0062] 3. Multimodal data packaging The core of multimodal data packaging is a dual-anchoring strategy of timestamp + entity ID, ensuring that data of different modalities are precisely aligned in time and space: Time anchoring: using a unified timestamp as the basis for data association. Entity anchoring: using entity IDs (users, robots, etc.) as another basis for data association. Each user is assigned a unique identity ID (face_id) and a voice ID (voice_id) to achieve entity-centric memory storage. Establish an entity relationship graph to record the interaction relationships between different entities. Event-driven: Using specific events as the boundaries of memory units, defining event triggering conditions, such as user silence for more than 3 seconds, completion of an action, completion of instruction execution, etc.
[0063] The response control module 400 executes the following steps: intent priority assessment S401, response timing optimization S402, response content generation S403, action planning and execution S404, and feedback loop S405. The specific process for generating and executing the corresponding interactive response is as follows: First, intent priority assessment S401 and response timing optimization S402 are performed, followed by response content generation S403, and action planning and execution S404 are performed based on the generated content. After action planning and execution S404, the feedback loop S405 step is entered, where the user's feedback characteristics are monitored in real time and a judgment is made: if negative feedback from the user is detected, the process returns to the response content generation S403 step to adjust the strategy and regenerate the response; if no negative feedback is detected, the process enters the response completion S406 step to end the interaction.
[0064] Regarding intent priority assessment and response timing optimization, this embodiment achieves intelligent decision-making through a multi-dimensional weighted model and dialogue rhythm analysis. The intent priority assessment module integrates the outputs of the perception, understanding, and memory modules, focusing on three core factors: confidence (recognition accuracy), urgency (time sensitivity), and user preference (historical habits). The system sets different confidence thresholds based on task type. If the confidence level is below the threshold, it proactively asks for confirmation and calculates a comprehensive priority score using a weighted scoring model. Simultaneously, it dynamically adjusts the weights based on special situations such as user emotions and multi-task triggering to ensure that the decision aligns with the urgency and importance of the user's needs. The response timing optimization module, by real-time monitoring of speech rate, pause patterns, and tone changes, combined with multimodal information such as visual attention, predicts the end point of the user's speech and determines the optimal time to interrupt. The system categorizes speech speed into three levels—fast, medium, and slow—and matches them with corresponding response thresholds. Combined with rules such as long pauses and eye contact, it typically responds within 100 to 1400 milliseconds after the user finishes speaking, choosing the most natural moment to respond. This effectively avoids mechanical interruptions or delays in the interaction, significantly improving the naturalness and intelligence of human-computer dialogue, as detailed below: I. Intent Priority Assessment Module The intent priority assessment module is responsible for calculating the confidence score of each possible intent based on the combined outputs of the spatial intelligent multimodal perception module, the spatial intelligent understanding and intent analysis module, and the spatial intelligent multimodal memory module, and determining the response priority according to preset rules. This module achieves intelligent sorting of intents through a multi-dimensional weighted model, ensuring that the robot can make reasonable decisions based on the urgency of the user's needs, confidence level, and historical preferences.
[0065] Intent prioritization assessment employs a multi-dimensional model based on weighted scoring, comprehensively considering the following three core factors: Confidence level: The accuracy with which the intent analysis module identifies the current intent. Urgency level: the time sensitivity and importance level of the intention User preferences: User habits and preference weights based on historical interaction data Through dynamic weight allocation, the system can intelligently adjust the weight of each factor according to the current context and user characteristics, so as to achieve more accurate intent priority assessment. The specific judgment criteria and thresholds are as follows.
[0066] 1. Confidence assessment Confidence assessment is based on the output of the intent analysis module and uses the following threshold criteria:
[0067] When the confidence level of an intent is lower than the corresponding threshold, the system will proactively ask the user to confirm the intent, for example: "Do you mean to turn on the living room lights?" 2. Urgency assessment Urgency assessment is based on task classification and execution time, using the following weighting criteria:
[0068] The urgency level weight is dynamically adjusted based on the task execution time; the shorter the task execution time, the higher the weight. For example, turning on the TV (1 minute) has a higher priority than opening the curtains (2 minutes).
[0069] 3. User Preference Assessment User preference assessment is based on historical interaction data in the spatial intelligent multimodal memory module, and is calculated using the following formula: Preference score = (Frequency of historical requests × Emotional satisfaction) / Total number of interactions Wherein: - Historical request frequency: the number of times the user requested this task in the past 10 interactions - Emotional satisfaction: the user's emotional rating (0-1, based on emotional timeline profile) after completing this task The user preference weight is 0.2, which is used to adjust the execution order of tasks within the same priority category.
[0070] 4. Intent Priority Comprehensive Calculation Model Intent priority is calculated using the following weighted formula: Priority score = Confidence level × ω 1 + Urgency level × ω 2 + Preference level × ω 3 The weighting parameters are: - ω1 (confidence weight) = 0.3 - ω2 (urgency weight) = 0.3 - ω3 (preference weight) = 0.2 The priority score ranges from 0 to 1, with a higher score indicating that the intent requires a higher priority response.
[0071] 5. Dynamic weight adjustment mechanism In special circumstances, the system will dynamically adjust the weight parameters:
[0072] The dynamic weight adjustment mechanism ensures that the system can flexibly adjust its decision-making strategy according to changes in user characteristics and context.
[0073] II. Response Timing Optimization Module The response timing optimization module analyzes the rhythmic characteristics of user conversations to determine the optimal response time, avoiding interruptions or delayed responses. This module analyzes features such as speech rate, pause patterns, and intonation variations, combined with multimodal information, to achieve more natural conversational interaction. Response timing optimization is based on natural dialogue analysis and employs the following core features: Speech rate analysis: Determines the speed and rhythm of the user's speech. Pause Pattern Recognition: Analyzing Pause Patterns in User Conversations intonation change detection: Recognizing emotional changes in user speech Multimodal collaboration: Combining visual, auditory, and other multimodal information for comprehensive judgment. By analyzing these characteristics, the system can predict when a user will end their speech and select the best time to respond. The specific judgment criteria and thresholds are as follows.
[0074] 1. Speech rate threshold The speech rate threshold is based on statistical analysis of family conversation scenarios, using the following criteria:
[0075] When a user speaks at a rate exceeding 20 words per second, the system will shorten the response threshold to 600ms to accommodate the fast pace of conversation.
[0076] 2. Pause Mode Standard The pause pattern standard is based on natural dialogue analysis and adopts the following classification:
[0077] When a long pause (>1 second) is detected and the user's gaze is fixed on the robot, the system will determine it as the optimal time to respond.
[0078] 3. Threshold for intonation changes The intonation change threshold is based on research in speech emotion recognition and adopts the following criteria:
[0079] When a sudden change in intonation is detected, the system will reassess the timing of the response, which may be earlier or later.
[0080] 4. Optimize decision-making process based on response timing The following decision-making process is used to optimize response timing: Real-time monitoring: Continuously analyzes changes in user speech rate, pauses, and intonation. Feature extraction: Identifying key features (such as sudden changes in speech rate and long pauses). Contextual Judgment: Determining User Status by Combining Multimodal Information Timing decision: Determine the optimal response time based on comprehensive analysis. The system will respond within 100-1400ms after the user finishes speaking, depending on the speaking speed and pause mode.
[0081] 5. Multimodal cooperative response rules The multimodal collaborative response rule combines visual and auditory information and employs the following strategy:
[0082] Multimodal collaborative response rules ensure that the system can more accurately understand user intent and provide a more natural interactive experience.
[0083] Specifically, when this invention is applied to complex home interaction scenarios (e.g., a user issuing commands to a robot in the living room), the system first performs comprehensive environmental perception through the multimodal perception module 100: the first acquisition module 101 captures real-time images and gestures of the user from multiple indoor perspectives; the second acquisition module 102 obtains the precise three-dimensional distance between the user and the robot using ToF technology; the sound pickup module 103 captures voice commands (such as "Turn on the TV for me, turn the volume down") and performs sound source localization through a microphone array; and the motion state acquisition module 104 monitors the robot's own six-axis posture data. The above heterogeneous data is then integrated into the data synchronization module 105 and aligned at the millisecond level based on a unified timestamp to eliminate spatiotemporal misalignment caused by different sensor sampling rates.
[0084] In this embodiment, the time synchronization scheme uses the clock frequency of the robot's main computing box as the unified benchmark for the entire system, achieving efficient alignment of multi-source heterogeneous data through a strategy combining local and network methods. For local sensors such as depth cameras, microphone arrays, gyroscopes, and IMUs, data acquisition is completed directly within the computing box, and precise timestamps are marked in real time using their local clocks, thus achieving natural synchronization and completely eliminating clock deviations caused by network transmission. For indoor multi-camera systems requiring network synchronization, a dedicated NTP server is deployed within the local area network to provide a high-precision clock source, and each camera node is periodically calibrated to maintain internal clock consistency. Simultaneously, the main computing box serves as the benchmark for event understanding, periodically synchronizing with the server by running an NTP client to ensure high alignment between the main computing box's local time, the NTP server's time, and the multi-camera system's time, successfully establishing a cross-device time communication channel. During the data processing phase, the main computing box aggregates remote image data with synchronized timestamps and local sensor data, achieving correlation of information across various dimensions through timestamp comparison. This synchronization scheme primarily serves scene event understanding, strictly controlling errors to the millisecond level through the NTP mechanism. In practical applications, when the overall synchronization error is kept within 50 milliseconds, it can accurately identify people walking, objects moving, and simple interactive actions, ensuring the system's keen perception and timely response to environmental changes. Since this accuracy is sufficient for judgment requirements, the system no longer needs to perform sensor frequency alignment and cumbersome difference processing, effectively improving overall information processing efficiency, as detailed below: Within the main computing power unit, information collected by the indoor multi-camera system, robot depth camera, microphone array, gyroscope, and IMU sensors needs to be synchronized. This synchronization primarily serves scene event understanding, therefore frequency alignment and interpolation processing between different sensors are unnecessary. The system clock is based on the clock frequency of the robot's main computing power box. Since the robot's depth camera, microphone array, gyroscope, and IMU sensors all collect data locally within the robot's computing power box, there are no network synchronization issues. The only network synchronization involved is the synchronization between the indoor multi-camera system and the sensor data collected on the computing power box. This is achieved by deploying a dedicated NTP server on a local area network to provide a unified clock for the relevant devices, controlling the synchronization error to the millisecond level. In practical scene judgment applications, a synchronization error range of no more than 50 milliseconds is sufficient to accurately understand and judge common scene events (such as people walking, objects moving, and simple interactive actions), ensuring the system can promptly and accurately perceive and respond to environmental changes. Although the main computing box serves as the time reference for the local sensors, it must also periodically synchronize its time with a dedicated NTP server within the same local area network to achieve time alignment with the remote multi-camera system. The main computing box runs an NTP client to align its system clock with the NTP server. This step ensures that: the main computing box's local time ≈ the NTP server's time ≈ the multi-camera system's time, thus establishing a time channel between the local and remote sensors. The main computing device information synchronization process is described in detail below: 1. Establishment of system clock reference: The clock frequency of the robot's main computing box is used as the unified clock reference for the entire system.
[0085] 2. Local Sensor Data Acquisition and Time Stamping: The robot's depth camera, microphone array, gyroscope, and IMU sensors all acquire data locally on the robot's computing power box. When acquiring data, these devices directly use the main computing power box's local clock to precisely timestamp each piece of data, thus achieving natural synchronization between local sensor data and the main computing power box's clock, eliminating clock deviation issues caused by network transmission.
[0086] 3. Indoor multi-camera system clock synchronization deployment: a. Setting up a local area network (LAN) NTP server: Deploy a dedicated NTP (Network Time Protocol) server within a local area network (LAN) environment.
[0087] b. Unified clock source: Configure the NTP server to obtain a high-precision, stable external time source (or, when no external source is available, to serve as the authoritative time source within the local area network).
[0088] c. Multi-camera system clock calibration: Each camera node in the indoor multi-camera system periodically synchronizes its time with the dedicated NTP server via the local area network to ensure that the internal clock of all camera nodes is consistent with the clock of the NTP server.
[0089] 4. Main computing box and NTP server clock synchronization: The main computing box also synchronizes its time with the aforementioned dedicated NTP server through the local area network, so that the system clock of the main computing box (which serves as the reference clock for data fusion and event understanding of the entire system) is consistent with the clock of the NTP server, thereby indirectly achieving clock synchronization with the indoor multi-camera system.
[0090] 5. Multi-source data timestamp alignment and event understanding: a. Data aggregation: Image data acquired by the indoor multi-camera system (already containing timestamps synchronized with the NTP server) is transmitted to the main computing box via the network. Data acquired by the robot's depth camera, microphone array, gyroscope, and IMU sensors (already containing local timestamps synchronized with the main computing box) is directly available within the main computing box.
[0091] b. Timestamp comparison: When the main computing box is understanding scene events, it will correlate and align all the data received from different devices (indoor multiple cameras, various sensors of the main body) in time by using the timestamps they carry.
[0092] c. Synchronization Error Control and Application Judgment: Through the aforementioned NTP synchronization mechanism, the synchronization error between the indoor multi-camera system and the main computing box is controlled within milliseconds. In practical scene event judgment applications, when the timestamp difference (i.e., synchronization error range) among all sensor data involved in event understanding does not exceed 50 milliseconds, it can meet the requirements for accurate understanding and judgment of common scene events (such as people walking, objects moving, simple interactive actions, etc.), ensuring that the system can perceive and respond to environmental changes in a timely and accurate manner. During this process, since the synchronization accuracy already meets the requirements for event understanding, there is no need for additional frequency alignment and interpolation processing of different sensor data. The multi-source data time synchronization flowchart is as follows: Figure 2 As shown; The synchronized data enters the analysis module 200. The spatial relationship modeling module 201 uses geometric projection to map the two-dimensional image into a three-dimensional spatial relationship, confirming that the user is facing the robot and within the effective interaction distance. The semantic understanding module 202 combines the context to parse the explicit instruction of "turn on the TV" and the implicit need of "turn down the volume". The recognition module 203 and the action analysis module 204 extract the user's facial expressions (such as whether they are tired) and limb key points (such as whether they make pointing gestures). The fusion strategy module 205 calculates the probability distribution of the current interaction intention by combining the above information. At the same time, the memory module 300 is called. The trajectory recording module 301 and the relationship archiving module 302 trace back the user's sound source movement trajectory and dwell position at past moments. The semantic content indexing module 303 retrieves historical dialogue habits. The multimodal association storage module 305 provides historical preferences in similar scenarios (such as the user's habitual volume value) to assist in accurate decision-making. Subsequently, the response control module 400 takes over the process: First, in step S401, it assesses the priority by combining intent probability and historical memory to confirm that the instruction is a high-priority valid intent rather than background chatter; in step S402, it determines the best time to interrupt by detecting pauses in the user's speech; in step S403, it generates response content and instruction code that match the user's current mood (such as a caring tone); then, in step S404, it plans the robot's actions (such as turning towards the user) and executes the operation of turning on the TV; after execution, it immediately enters the feedback closed loop step S405, which monitors the user's reaction in real time after execution. If negative feedback such as the user frowning or waving again is detected, the system determines that the response has not met expectations, and immediately returns to step S403 to adjust the strategy (such as further reducing the volume or asking whether to continue from the playback record), regenerates and executes the response; if no negative feedback is detected in step S405, the interaction is determined to be successful, and the system enters step S406 to complete the response, thereby achieving accurate, natural, and self-correcting intelligent interaction.
[0093] In this embodiment, regarding the negative feedback process, the negative feedback detection and adjustment strategy module constructs a closed-loop optimization process by monitoring user reactions in real time and integrating multimodal information to achieve accurate emotion recognition and dynamic strategy adjustment. In terms of detection parameters, the system combines facial key points and voice signals for multidimensional analysis: when the eyebrow is lowered at an angle greater than 15 degrees for more than 2 seconds, or the eyes are closed for more than 2 seconds, the system determines it as negative feedback in the facial expression dimension; if the voice fundamental frequency changes abruptly by more than 30 Hz per second within three consecutive frames, or the speech rate suddenly drops by more than 20 words per second, it is determined as negative feedback in the voice dimension. For the identified signals, the system presets classification-triggered adjustment rules, such as lowering the volume and switching to a more tactful tone when a frown is detected, and pausing the action when the eyes are closed; in multimodal collaborative mode, if both a frown and a fundamental frequency change occur simultaneously, the system will immediately pause the task and actively ask the user. The entire execution process follows a standardized path from feedback detection, emotion judgment, rule matching to strategy execution and secondary monitoring, ensuring a rapid and appropriate response under high-confidence feedback. In addition, the system introduces an effectiveness evaluation mechanism. Quantitative indicators such as negative feedback elimination rate (target value > 80%), task completion rate (target value > 90%), user satisfaction, and adjustment response time (target value < 1.5 seconds) are used to evaluate the effectiveness of the strategy in real time and drive continuous rule optimization, thereby continuously improving the quality and intelligence level of the robot's interaction. Specifically: The negative feedback detection and adjustment strategy module monitors user reactions in real time. When negative feedback (such as frowning or a dissatisfied tone) is detected, the strategy is immediately adjusted and the response content is regenerated, forming a closed-loop optimization process. This module achieves more accurate negative feedback identification and more effective strategy adjustment through multimodal information fusion.
[0094] 1. Negative Feedback Detection Parameters (1) Facial feature detection parameters Facial expression feature detection is based on facial key point analysis and uses the following parameters:
[0095] When the system detects a frown (eyebrows downturned by more than 15 degrees and lasting for more than 2 seconds) or closed eyes for more than 2 seconds, it will be judged as negative feedback.
[0096] (2) Speech feature detection parameters Speech feature detection is based on speech signal analysis and uses the following parameters:
[0097] When a fundamental frequency change greater than 30 Hz / second (for 3 consecutive frames) or a speech rate drop greater than 20 words / second is detected, the system will determine it as negative feedback.
[0098] 2. Negative Feedback Adjustment Strategy Rules Table (1) Facial expression trigger adjustment rules
[0099] (2) Voice trigger adjustment rules
[0100] (3) Multimodal cooperative adjustment rules The multimodal collaborative adjustment rules combine facial expression and voice information, and adopt the following strategies:
[0101] Multimodal collaborative adjustment rules ensure that the system can more comprehensively understand user emotions and provide more effective coping strategies.
[0102] 3. Adjust the strategy execution process The adjustment strategy is implemented using the following process: Negative feedback detection: Real-time analysis of facial expressions and voice features Emotion type identification: Identify negative emotion types (such as dissatisfaction, confusion). Adjustment rule matching: Match corresponding adjustment strategies based on emotion type. Strategy execution: Adjusting response content or robot behavior Secondary feedback monitoring: Continuously monitor user feedback to verify the effectiveness of adjustments. When a high-confidence negative feedback is detected, the system will immediately pause the current task and adjust the strategy.
[0103] 4. Evaluation of the effectiveness of the adjusted strategy The effectiveness of the adjusted strategy was evaluated based on subsequent user feedback, using the following metrics:
[0104] The system will dynamically optimize the adjustment rules based on the evaluation results of the adjustment strategy, thereby improving the quality of interaction.
[0105] The following is a typical implementation example. When this invention is applied to a complex interactive scenario in a family living room, the system first activates the spatial intelligent multimodal perception module. Utilizing three RGB-D cameras distributed indoors in collaboration with the robot's own sensors, the synchronization error of the multi-source heterogeneous data is precisely controlled to within 32 milliseconds (far better than the 50-millisecond tolerance standard) via an NTP server, ensuring the consistency of the spatiotemporal reference for scene understanding. At this time, the user, located 2.5 meters from the robot on the sofa, issues the voice command "Open the news for me," accompanied by a hand gesture. The spatial intelligent understanding module quickly intervenes, using cross-view 3D modeling technology to reconstruct the two-dimensional image into a three-dimensional skeleton, confirming that the user's facial orientation yaw angle is +5 degrees (i.e., facing the robot). Visual and auditory features are then input into a graph neural network (GNN) for fusion, calculating an intent confidence score of 0.94 for "Start movie viewing mode." Next, the memory module retrieves the user's "memory unit," recalling historical data showing a preference for watching financial channels at 8 PM with a habitual volume of 25%. Combined with the current scenario, the intent priority score is calculated to be 0.87 (exceeding the response threshold of 0.7). The response control module then took over, detecting that the user's speech rate was approximately 12 words per second with a long pause of 0.8 seconds, which was determined to be the optimal time to interrupt. Within 450 milliseconds, it generated a command and controlled the robot to turn towards the TV to perform the activation operation. Immediately after the action was executed, the feedback loop mechanism detected that the user's eyebrows were lowered at an angle of 18 degrees (exceeding the negative feedback threshold of 15 degrees) for 2.1 seconds. The system determined that the response did not fully meet expectations (e.g., the volume was too loud) and immediately triggered an adjustment strategy, automatically lowering the volume by 10% within 1.2 seconds and switching to an alternative channel. Subsequently, it detected that the user's facial expression had relaxed, ultimately determining that the interaction was successfully completed. The entire process demonstrated a closed-loop intelligent effect, from accurate perception and deep understanding to adaptive adjustment.
[0106] Here is another typical implementation example: when this invention is applied to a bedroom scenario and detects an elderly user falling accidentally, the system immediately initiates an emergency response process. First, the indoor multi-camera system captures drastic changes in the user's posture. The spatial relationship modeling module quickly intervenes, constructing projection relationships using the intrinsic and extrinsic parameter matrices of each camera, targeting key points on the human body. (e.g., "head") Establish a system of DLT linear equations And by analyzing the matrix Singular Value Decomposition (SVD) is performed, selecting the right singular vector corresponding to the minimum singular value. The 3D Euclidean coordinates of the keypoint in the world coordinate system are calculated, showing a sudden drop in height from Z=1.65m to Z=0.15m, thus determining the user's posture as "lying flat" accompanied by a violent fall. Simultaneously, the audio pickup module captures the user's groans, and the sentiment analysis module identifies negative emotional characteristics of "pain / panic," with a fundamental frequency mutation rate reaching 45Hz / second. The decision control module then initiates intent priority assessment. Based on the identified "cry for help / fall" intent, the following core data is extracted and substituted into the formula: confidence P=0.96 (high multimodal consistency), urgency U=1.0 (highest level of emergency), and user preference H=0.5 (default value). Given the detected strong negative emotions, the system automatically triggers a dynamic weight adjustment mechanism, adjusting the urgency weight... The value was increased from the default 0.3 to 0.5. The system adjusts this according to the formula. Calculations are performed to obtain priority scores. This score far exceeded the response threshold for routine tasks, and the system classified it as a highest priority event. The response control module immediately skipped the usual interruption timing detection and directly generated an emergency response strategy within 200 milliseconds: on the one hand, it controlled the robot to quickly move to the user's side and adjust the camera angle to maintain continuous monitoring; on the other hand, it automatically initiated a remote call and played a reassuring voice message, "We have detected that you may have fallen and are contacting your family." The entire process, from the fall to the triggering of assistance, took only 1.5 seconds, fully verifying the accuracy and timeliness of this invention based on spatial modeling formulas and dynamic decision-making algorithms in extreme scenarios.
[0107] Example 2 This embodiment provides an apparatus for performing a multimodal robot interaction non-dialogue intent rejection method as described in Embodiment 1. The apparatus can be installed on a robot system or externally.
[0108] Example 3 This embodiment provides a robot for performing a multimodal robot interaction non-dialogue intent rejection method based on embodiment 1, having or being connected to one of the devices in embodiment 2.
[0109] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments, and various changes and modifications can be made without departing from the spirit and scope of the invention, all of which fall within the scope of protection claimed by the present invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for rejecting dialogue-free intent in robot interaction based on multimodal communication, characterized in that, This method is based on the execution of a robotic system; The robot system includes a multimodal perception module (100), an analysis module (200), a memory module (300), and a response control module (400); The method includes the following steps: The multi-modal perception module (100) collects multi-dimensional data of the environment and the user, and performs time synchronization processing. The analysis module (200) performs cross-perspective spatial modeling and multi-dimensional feature analysis based on the synchronized data, and integrates and generates the probability of interactive intent. The memory module (300) records the spatiotemporal trajectory and semantic information during the interaction process, and provides historical association information as a reference based on the current scene characteristics; The response control module (400) evaluates the effectiveness and priority of the interaction intent by combining the probability of the interaction intent with the historical association information. If it is determined that there is no intention to engage in dialogue, then rejection will be enforced. If the intent is deemed valid, a corresponding interactive response is generated and executed.
2. The method for non-dialogue intent rejection in robot interaction based on multimodal communication according to claim 1, characterized in that, The multimodal sensing module (100) includes a first acquisition module (101), a second acquisition module (102), a sound pickup module (103), a motion state acquisition module (104), and a data synchronization module (105); The process of collecting multidimensional data of the environment and the user through the multimodal perception module (100) and performing time synchronization processing specifically includes: The first acquisition module (101) is used to acquire global environmental image data from multiple indoor perspectives; The second acquisition module (102) is used to acquire depth and distance data from the robot's perspective; The sound pickup module (103) is used to acquire sound source localization and speech signal data; The motion state acquisition module (104) is used to acquire the robot's own posture and acceleration data; The data synchronization module (105) aligns and synchronizes the data streams collected by the different sensors based on a unified timestamp.
3. The method for non-dialogue intent rejection in robot interaction based on multimodal communication according to claim 1, characterized in that, The analysis module (200) includes a spatial relationship modeling module (201), a semantic understanding module (202), a recognition module (203), an action analysis module (204), and a fusion strategy module (205); The step of using the analysis module (200) to perform cross-perspective spatial modeling and multi-dimensional feature analysis based on synchronized data, and then fusing the data to generate interaction intent probabilities, specifically includes: The spatial relationship modeling module (201) maps the two-dimensional image into a three-dimensional spatial relationship and calculates the relative position of the human and the machine. The semantic understanding module (202) parses the explicit meaning and implicit requirements of the voice commands; Facial features are extracted through the recognition module (203) to analyze the user's emotional state; The motion analysis module (204) tracks key points of the limbs to identify user behavior patterns; the fusion strategy module (205) integrates the above spatial, semantic, emotional and motion features to calculate the probability distribution of the user's current interaction intent.
4. The method for non-dialogue intent rejection in robot interaction based on multimodal communication according to claim 1, characterized in that, The memory module (300) includes a trajectory recording module (301), a relationship archiving module (302), a semantic content indexing module (303), a timeline archive module (304), and a multimodal association storage module (305); The recording of the spatiotemporal trajectory and semantic information during the interaction process through the memory module (300) specifically includes: The trajectory recording module (301) stores the dynamic changes in the sound source location, user distance, and eye direction. The relationship archiving module (302) stores the sequence of changes in the human-machine spatial position relationship; The semantic content indexing module (303) associates and stores the dialogue content and its background information. The timeline archive module (304) records the fluctuation trend of user emotions over time; The multimodal association storage module (305) packages the above information into a memory unit for retrieval and calling during subsequent interactions.
5. The method for non-dialogue intent rejection in robot interaction based on multimodal communication according to claim 1, characterized in that, The steps performed by the response control module (400) include intent priority assessment (S401), response timing optimization (S402), response content generation (S403), action planning and execution (S404), and feedback loop (S405). The specific process for generating and executing the corresponding interactive response is as follows: First, the intent priority assessment (S401) and response timing optimization (S402) are performed, followed by response content generation (S403), and action planning and execution are carried out based on the generated content (S404). Following the action planning and execution (S404), the feedback loop (S405) step is entered, where the user's feedback characteristics are monitored and judged in real time: If negative feedback from users is detected, return to the response content generation (S403) step to adjust the strategy and regenerate the response; If no negative feedback is detected, proceed to the Complete Response (S406) step to end the interaction.
6. An apparatus for performing a multimodal robot interaction non-dialogue intent rejection method as described in any one of claims 1 to 5, characterized in that, The device can be installed on the robot system or externally.
7. A robot for performing a multimodal robot interaction method for rejecting dialogue-free intentions as described in any one of claims 1 to 5, characterized in that, It has a device as described in claim 6 or is connected to such device.