Body interaction robot control method and device, computer equipment and storage medium
By combining voice, vision, and motion information, the embodied interactive robot control method improves the accuracy and naturalness of understanding, solves the ambiguity problem of single voice interaction, and achieves more accurate and natural interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PENG CHENG LAB
- Filing Date
- 2025-12-19
- Publication Date
- 2026-05-08
AI Technical Summary
Existing embodied interactive robots rely on a single voice interaction method in complex operation scenarios, resulting in low accuracy of understanding and low naturalness of interaction, which can easily lead to ambiguity.
By acquiring audio of the target object's voice interaction, image frame sequences of the object's facial expressions, image frame sequences of the object's posture, and image frame sequences of the scene, and combining visual and voice information, the target object is identified, pitch and semantic vectors are acquired and encoded, and emotional and intentional features are fused to generate an interaction sequence to control the robot's actions and feedback.
It improves the accuracy and naturalness of embodied interactive robots in the interaction process. Through multimodal information collaborative understanding, it enhances the collaborative inference of users' emotional state and true intentions, and generates responses that are more in line with human interaction habits.
Smart Images

Figure CN121997245A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, and in particular to a control method, device, computer equipment and storage medium for an embodied interactive robot. Background Technology
[0002] Embossed interactive robots integrate mechanical bodies and intelligent algorithms, enabling them to perform tasks in real-world environments through their physical entities (such as robotic arms and mobile chassis) and to identify user intentions by using interaction information with users.
[0003] In related technologies, the interaction between users and embodied interactive robots mainly relies on speech recognition and semantic understanding technologies. Specifically, users issue commands to the embodied interactive robot via voice, the robot converts the speech into text through a speech recognition module, then uses natural language processing technology to interpret the user's intent, and finally provides feedback through voice or actions.
[0004] However, since the interaction between users and embodied interactive robots mainly relies on a single voice interaction method to understand the user's instructions, this method is prone to ambiguity in complex operation scenarios, resulting in low accuracy of understanding and naturalness of interaction for embodied interactive robots. Summary of the Invention
[0005] This application proposes a control method, device, computer equipment, and storage medium for an embodied interactive robot, which can improve the understanding accuracy and naturalness of the interactive robot during the interaction process.
[0006] To achieve the above objectives, a first aspect of this application proposes a method for controlling an embodied interactive robot, the method comprising: Acquire the target object's voice interaction audio, the object's facial expression image frame sequence, the object's pose image frame sequence, and the scene image frame sequence; When the voice interaction audio contains referential information, the center coordinates of each candidate object contained in the scene image frame sequence are determined, the hand movement direction of the target object is determined according to the object posture image frame sequence, and the target object to which the target object is pointing is determined according to the angle difference between the hand movement direction and the center coordinates of each object. The pitch vector and semantic vector corresponding to the voice interaction audio are obtained, and the pitch vector and the object's facial expression image frame sequence are encoded to obtain the corresponding emotional features. The semantic vector and the object's pose image frame sequence are encoded to obtain the corresponding intention features. Based on the fusion of the emotional features and the intention features, the corresponding target fusion features are obtained; The target fusion features are classified to obtain emotion classification results and intent classification results. Based on the location coordinates of the target object, the emotion classification results and the intent classification results, a corresponding target interaction sequence is generated. Based on the target interaction sequence, the embodied interactive robot is controlled to perform action interaction and voice feedback.
[0007] Accordingly, a second aspect of the embodiments of this application proposes an embodied interactive robot control device, the device comprising: The acquisition module is used to acquire the target object's voice interaction audio, the object's facial expression image frame sequence, the object's pose image frame sequence, and the scene image frame sequence; The determination module is used to determine the object center coordinates of each candidate object contained in the scene image frame sequence when the voice interaction audio contains referential information, determine the hand movement direction of the target object according to the object posture image frame sequence, and determine the target object to which the target object is pointing according to the angle difference between the hand movement direction and the center coordinates of each object. The encoding module is used to obtain the pitch vector and semantic vector corresponding to the voice interaction audio, and to encode the pitch vector and the object expression image frame sequence to obtain the corresponding emotional features, and to encode the semantic vector and the object posture image frame sequence to obtain the corresponding intention features; The fusion module is used to fuse the emotional features and the intention features to obtain the corresponding target fusion features; The classification module is used to classify the target based on the target fusion features to obtain the emotion classification result and the intent classification result, and to generate the corresponding target interaction sequence based on the position coordinates of the target object, the emotion classification result and the intent classification result; The control module is used to control the embodied interactive robot to perform action interactions and provide voice feedback based on the target interaction sequence.
[0008] In some embodiments, the hand includes the base of the fingers and the fingertips, and the determining module is further configured to: The first coordinates of the root of the finger and the second coordinates of the fingertip of the target object are obtained in the object pose image frame sequence, and the hand movement direction vector is determined based on the difference between the first coordinates and the second coordinates; Calculate the object pointing vector based on the difference between the first coordinate and the center coordinate of each object; Based on the hand movement direction vector and the object pointing vector, calculate the angular difference between the hand movement direction of the target object and the center coordinates of each object; Based on the magnitude relationship between multiple angle differences corresponding to multiple candidate objects, the candidate object with the smallest angle difference is determined from the multiple candidate objects as the target object to which the target object points.
[0009] In some embodiments, the encoding module is further configured to: The pitch vector and the sequence of image frames of the object's facial expression are encoded by the self-attention calculation module of the preset target model to obtain the corresponding emotional features; The semantic vector and the object pose image frame sequence are encoded through the spatiotemporal recurrent network module of the target model to obtain the corresponding intent features.
[0010] In some implementations, the fusion module is further configured to: The attention fusion module determines the emotional auxiliary features associated with the intent features, and defines the intent features as a query vector, the emotional features as a key vector, and the emotional auxiliary features as a value vector. Cross-modal attention is calculated based on the query vector, the key vector, and the value vector to obtain the corresponding initial fusion vector; The corresponding target fusion feature is obtained by weighted fusion of the emotional features and the initial fusion vector.
[0011] In some implementations, the fusion module is further configured to: The gating fusion unit outputs a first gating weight based on the emotional features and the initial fusion vector. Obtain a reference value, and calculate a second gating weight based on the difference between the reference value and the first gating weight; The initial fusion vector is adjusted using the first gating weight to obtain the first target vector; The emotional features are adjusted using the second gating weight to obtain the second target vector; The target fusion feature is obtained by fusing the first target vector and the second target vector.
[0012] In some implementations, the classification module is further configured to: The target fusion features are classified using the pre-defined intent classification head of the target model to obtain the intent classification result; The target fusion features are classified using the sentiment classification head of the target model to obtain the sentiment classification result.
[0013] In some embodiments, the embodied interactive robot control device further includes a trigger module for: The target confidence level corresponding to the target fusion feature is output through the preset confidence estimation head of the target model. Obtain a preset confidence threshold; when the target confidence level is less than the confidence threshold, trigger a clarification command. Based on the clarification instruction, the historical interaction context associated with the voice interaction audio is obtained, and based on the historical interaction context, the location coordinates of the target object, the emotion classification result, and the intent classification result, a corresponding target interaction sequence is generated; Based on the target interaction sequence, the embodied interactive robot is controlled to perform action interaction and voice feedback.
[0014] Accordingly, a third aspect of the embodiments of this application proposes a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the embodied interactive robot control method of any one of the embodiments of the first aspect of this application.
[0015] Accordingly, a fourth aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the embodied interactive robot control method of any one of the embodiments of the first aspect of this application.
[0016] This application embodiment acquires the target object's voice interaction audio, object facial expression image frame sequence, object posture image frame sequence, and scene image frame sequence. When the voice interaction audio contains referential information, the object center coordinates of each candidate object in the scene image frame sequence are determined. The direction of the target object's hand movement is determined based on the object posture image frame sequence, and the target object pointed to by the target object is determined based on the angular difference between the hand movement direction and the center coordinates of each object. The pitch vector and semantic vector corresponding to the voice interaction audio are acquired, and the pitch vector and the object facial expression image frame sequence are encoded to obtain the corresponding emotional features. The semantic vector and the object posture image frame sequence are encoded to obtain the corresponding intention features. The emotional features and intention features are fused to obtain the corresponding target fusion features. The target fusion features are classified to obtain emotional classification results and intention classification results. Based on the target object's position coordinates, emotional classification results, and intention classification results, the corresponding target interaction sequence is generated. The embodied interactive robot is controlled to perform action interaction and voice feedback based on the target interaction sequence. In this way, through the synchronous acquisition and fusion understanding of multimodal information, the transformation from single voice interaction to collaborative understanding of voice, vision, action, and scene can be realized. Specifically, by combining gesture pointing with the coordinate analysis of scene objects to interpret referential information, the visual information collected by the embodied robot can be fully utilized. This allows visual information to be used not only for object manipulation but also for improving the accurate understanding of ambiguous or omitted instructions. Simultaneously, by encoding and fusing multimodal features such as tone of voice, facial expressions, and posture, the collaborative inference of the user's emotional state and true intentions is enhanced, thereby generating an interaction sequence that takes into account target location, emotion, and intention. Finally, based on the fused understanding results, a coordinated and consistent multimodal response is generated, making the embodied robot's feedback more in line with human interaction habits. In summary, this application can improve the accuracy of understanding and the naturalness of interaction in the interactive robot process. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the architecture of the embodied interactive robot control system provided in an embodiment of this application; Figure 2 This is a flowchart of the embodied interactive robot control method provided in the embodiments of this application; Figure 3 This is a flowchart of the embodied interactive robot control method provided in the embodiments of this application; Figure 4 This is a schematic diagram of the functional modules of the embodied interactive robot control device provided in the embodiments of this application; Figure 5 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0019] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0021] Embossed interactive robots integrate mechanical bodies and intelligent algorithms, enabling them to perform tasks in real-world environments through their physical entities (such as robotic arms and mobile chassis) and to identify user intentions by using interaction information with users.
[0022] In related technologies, the interaction between users and embodied interactive robots mainly relies on speech recognition and semantic understanding technologies. Specifically, users issue commands to the embodied interactive robot via voice, the robot converts the speech into text through a speech recognition module, then uses natural language processing technology to interpret the user's intent, and finally provides feedback through voice or actions.
[0023] However, since the interaction between users and embodied interactive robots mainly relies on a single voice interaction method to understand the user's instructions, this method is prone to ambiguity in complex operation scenarios, resulting in low accuracy of understanding and naturalness of interaction for embodied interactive robots.
[0024] Based on this, embodiments of this application provide a method, apparatus, computer device, and storage medium for controlling an embodied interactive robot, which can improve the understanding accuracy and naturalness of the interactive robot during the interaction process.
[0025] The embodied interactive robot control method, device, computer equipment, and storage medium provided in this application are specifically described through the following embodiments. First, the embodied interactive robot control system in this application embodiment is described.
[0026] Please refer to Figure 1In some embodiments, this application provides an embodied interactive robot control system, including a terminal 11 and a server 12.
[0027] In some implementations, terminal 11 can be used to collect multimodal interaction data of users in real time and perform local preprocessing and response actions. For example, it can be an embodied robot body or an embedded device integrating a microphone array, RGB-D camera, computing unit and motion mechanism. Terminal 11 can synchronously collect users' voice signals, facial expression videos, body movements and environmental scene images through sensors, and perform preliminary feature extraction and time alignment to upload the multimodal feature vector to server 12.
[0028] In some implementations, server 12 can be used to achieve deep fusion of multimodal information, contextual understanding and intelligent decision-making. For example, it can be a cloud server, a high-performance computing platform or a dedicated AI server equipped with a GPU cluster. Server 12 can receive feature data uploaded by terminal 11 and run referential resolution, intent-emotion joint reasoning, scene graph construction and response generation models to output fusion understanding results and multimodal response instructions.
[0029] Furthermore, the terminal 11 and the server 12 communicate and synchronize data via wired or wireless networks. The terminal 11 is responsible for real-time perception and action execution, while the server 12 undertakes high-load fusion computing and decision generation. Together, they form a distributed embodied robot control system to achieve a low-latency, highly intelligent humanoid robot interaction experience.
[0030] The embodied interactive robot control method in this application can be illustrated through the following embodiments.
[0031] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user will be obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent will the necessary user-related data for the normal operation of the embodiments of this application be obtained.
[0032] In this application embodiment, the description will focus on the embodied interactive robot control device, which can be integrated into a computer device. See also Figure 2 , Figure 2This is a flowchart illustrating the steps of the embodied interactive robot control method provided in this application embodiment. Taking the embodied interactive robot control device specifically integrated into a terminal or server as an example, the specific process when the processor on the terminal or server executes the program instructions corresponding to the embodied interactive robot control method is as follows: Step 101: Obtain the target object's voice interaction audio, object facial expression image frame sequence, object pose image frame sequence, and scene image frame sequence.
[0033] In some implementations, a comprehensive perceptual input that includes auditory, visual, and environmental context can be constructed. This can be achieved by synchronously collecting and acquiring user voice interaction audio, object facial expression image frame sequence, object posture image frame sequence, and dynamic images of the surrounding environment. This provides a precise and synchronous multimodal data foundation for subsequent context-aware interactive control.
[0034] The target object can be the user who interacts with the embodied interactive robot, such as the operator who issues operating instructions or the user who expresses emotions. The target object is the source of multimodal information.
[0035] The audio for voice interaction can be a continuous sound signal containing semantic and paralinguistic information (such as intonation and rhythm) emitted by the target object during interaction with the robot. It can be the user's speech content collected through a microphone array mounted on the robot.
[0036] The sequence of facial expression images can be a collection of images of the target object's facial region captured sequentially over time. For example, it can be a video stream captured and output by an RGB-D camera, reflecting the user's facial muscle movements and expression changes, which can be used to extract and analyze the user's emotional state features.
[0037] Among them, the object pose image frame sequence can be a collection of images of the target object's body (especially hands and arms) pose and movements captured sequentially in time. For example, it can be a video stream that reflects the user's gestures, pointing, body orientation and other limb movements captured and output by the RGB-D camera of the embodied robot.
[0038] The scene image frame sequence can be a collection of images of the embodied robot's surrounding environment captured sequentially in time. For example, it can be a video stream captured and output by an RGB-D camera mounted on the robot, containing objects, layouts, and spatial relationships in the environment, which can be used to provide the environmental context in which interactions occur and information about operable objects.
[0039] For example, a robot can acquire voice interaction audio from a target object using an onboard microphone array. The microphone array (e.g., a circular 6-microphone array) has beamforming and noise suppression capabilities, enabling it to directionally amplify voice signals from the target object's location while suppressing ambient background noise. After preprocessing the raw audio signal acquired by the robot, such as pre-emphasis, framing, and windowing, the system can obtain a digital audio stream suitable for subsequent speech recognition and pitch analysis, i.e., voice interaction audio.
[0040] In some implementations, the robot can simultaneously capture visual data streams using one or more RGB-D cameras (such as the Intel RealSense D435i), which can simultaneously output high-resolution RGB color image frame sequences and corresponding depth information. The system can then extract object expression image frame sequences, object pose image frame sequences, and scene image frame sequences from the continuous RGB color image frame sequences in real time.
[0041] Furthermore, by performing face detection on the RGB color image frame sequence and cropping out the image region containing only the user's face, a continuous object expression image frame sequence can be obtained; by performing human body detection on the RGB color image frame sequence of the same time period, an image region containing the user's whole body or upper body can be obtained, resulting in a continuous object pose image frame sequence; by preserving the complete RGB environment image of the RGB color image frame sequence of the time period, the RGB color image frame sequence can be used as a scene image frame sequence.
[0042] In some implementations, while ensuring timestamp alignment, the audio of voice interaction, the sequence of image frames of object facial expressions, the sequence of image frames of object poses, and the sequence of image frames of scene can be acquired by different cameras. This application does not limit the specific method of acquisition.
[0043] In some implementations, to ensure the spatiotemporal consistency of cross-modal information, hardware-triggered synchronization or software timestamp alignment mechanisms can be employed. For example, all sensors (microphones, cameras) can be connected to the same hardware trigger signal to ensure they begin acquiring data at the same physical moment. Alternatively, each sensor data packet can be timestamped based on the system's high-precision clock. During backend processing, interpolation algorithms can be used to align data from different modalities to a unified timeline, achieving millisecond-level synchronization accuracy. This provides a precisely aligned data foundation for subsequent cross-modal fusion, avoiding misunderstandings caused by time delays.
[0044] In some implementations, in addition to acquiring visual information through the aforementioned RGB-D camera, a tactile sensor array can be integrated into the end effector (such as a robotic arm) of the embodied robot. When the robot physically interacts with a target object or environmental objects, the tactile sensors can acquire information such as contact force, texture, and slippage in real time. This tactile modal information can be synchronized with visual and auditory information and input as an additional perceptual dimension into subsequent fusion modules. For example, when a user says "lighter," combined with the larger grasping force detected by the tactile sensors, the robot can more accurately understand the specific force threshold of the intention "lighter," thereby making more precise force control adjustments. In this way, the dimensions of interaction can be further enriched, achieving deep fusion interaction of vision, hearing, and touch, and improving the safety and intelligence of the robot in physical collaborative tasks.
[0045] In some implementations, the synchronization mechanism can also incorporate visual sampling based on sound events. Specifically, energy changes in the audio of voice interaction can be detected in real time. When the user begins to speak (sound event), the visual sensor is immediately triggered to perform a high-frame-rate intensive sampling. This ensures that the most matching image frame sequence is obtained the instant the user issues a key command, which is particularly suitable for scenarios where interactive actions are fleeting, thereby enhancing semantic relevance at the data source and improving the accuracy of intent capture.
[0046] By using the above methods, a comprehensive and dynamic perception of user status and interaction environment can be established simultaneously, overcoming the shortcomings of insufficient single voice input information. This can lay an effective data foundation for subsequent interactive control.
[0047] Step 102: When there is referential information in the voice interaction audio, determine the center coordinates of each candidate object contained in the scene image frame sequence, determine the direction of the hand movement of the target object according to the object posture image frame sequence, and determine the target object that the target object is pointing to according to the angle difference between the direction of the hand movement and the center coordinates of each object.
[0048] In some implementations, in order to accurately resolve ambiguous spatial reference information in user voice commands, when the reference information is detected in the voice interaction audio, the three-dimensional center coordinates of each candidate object in the environment can be calculated simultaneously using visual information and the direction of the user's hand pointing can be identified. Then, the specific target can be locked by calculating the angle difference between the direction and the object, so as to accurately associate the abstract voice reference with the specific visual spatial entity.
[0049] The referential information can be words or phrases contained in the audio of voice interaction that cannot be independently identified as specific referents, such as pronouns "this" or "that," or instruction fragments that omit the target name. These can be used to trigger the system to start a vision-based spatial referential resolution process.
[0050] Among them, candidate objects can be a set of physical objects that may be referred to by the user in the scene image frame sequence identified by the object detection and recognition module. For example, objects such as tables, cups, and books segmented and identified from the scene image frame sequence by object detection algorithms such as YOLO can be used to form a target candidate set for spatial reference resolution.
[0051] Among them, the object center coordinates can be data used to characterize the position information of the candidate object in three-dimensional space. For example, the three-dimensional coordinates (X,Y,Z) of the geometric center point of the candidate object are obtained by back projection calculation by combining the two-dimensional detection box and depth information of the RGB-D camera, which is also the world coordinates of the geometric center point of the candidate object.
[0052] The direction of hand movements can be a vector representing the pointing tendency of the target object's hand (especially the index finger) in three-dimensional space. For example, it can be calculated by obtaining the key point coordinates of the base and tip of the index finger and using the vector obtained from these key point coordinates. The direction of hand movements can be used as a visual indication of the user's spatial intention.
[0053] The target object can be the specific object that the user intends to point to, which is ultimately determined based on the angle difference between the direction of the hand movement and the center coordinates of each candidate object. For example, the cup with the smallest included angle in the angle difference calculation is the task object that the robot will subsequently perform specific operations such as grasping and moving.
[0054] In some implementations, this can be accomplished collaboratively by an Automatic Speech Recognition (ASR) module and a Natural Language Processing (NLP) module. The ASR module converts the audio stream into a text command (e.g., "Please give me that"). Subsequently, the NLP module performs grammatical and part-of-speech analysis on the text, detecting whether it contains predefined referential words, such as pronouns or locative phrases like "this," "that," "it," or "this way." Once such referential information is detected, a visuospatial resolution process is immediately triggered to determine the target object to which the target object is pointing.
[0055] In some implementations, a deep learning-based 2D object detection model (such as YOLO, Faster R-CNN, etc.) can be used to process the scene image frame sequence to identify the bounding boxes of all interactive objects in each frame and obtain the category label (such as "cup", "book", "remote control") and the pixel coordinates (u,v) of the center point of the bounding box for each object. Subsequently, combined with the synchronously acquired depth map information (from the RGB-D camera) and the camera's intrinsic parameter matrix, the 2D pixel coordinates are back-projected into the 3D space under the camera coordinate system. Thus, the object center coordinates of each candidate object in the world coordinate system can be obtained.
[0056] Furthermore, a 3D hand pose estimation model (such as MediaPipe Hands) can be used to process the image region containing the user's hand. This model can output the 3D coordinates of 21 key points of the hand in the image pixel coordinate system and depth information. In order to obtain a stable and accurate hand movement direction, this application selects the root joint point (Index MCP, key point 5) and the fingertip joint point (Index Tip, key point 8) of the index finger for calculation.
[0057] Furthermore, the vector v_dir corresponding to the direction of the hand movement can be calculated using the following formula: v_dir=P_tip-P_mcp=(x_t-x_m,y_t-y_m,z_t-z_m); Here, P_tip=(x_t,y_t,z_t) represents the 3D coordinates of the index fingertip, and P_mcp=(x_m,y_m,z_m) represents the 3D coordinates of the base of the index finger (MCP joint). The vector v_dir points from the base to the fingertip, representing the pointing axis of the finger. In this way, the user's subtle pointing gestures can be quantified into a specific spatial vector.
[0058] Furthermore, for each candidate object in the candidate object list, a vector v_obj = P_obj - P_mcp can be calculated from the starting point P_obj (base of the index finger) to the object, using its 3D center coordinates P_obj and the starting point P_mcp (base of the index finger).
[0059] Therefore, the angle difference between the hand movement direction and the center coordinates of each object can be calculated using v_dir and v_obj, and the candidate object with the smallest angle difference can be determined as the target object.
[0060] In some implementations, to handle ambiguity in direction (such as multiple objects having very close angles) or situations where the pointed area has no significant object, this application introduces a scene semantic attention mechanism. This mechanism not only calculates the angular difference between the hand movement direction and the center coordinates of each object, but also considers the semantic information of the object and the correlation between the audio and the voice interaction. Specifically, a lightweight semantic matching network can be used to receive the object category label of each candidate object and the embedding vector of the instruction text corresponding to the voice interaction audio, and output a semantic relevance score. Then, based on the semantic relevance score of each candidate object, the weights of different candidate objects are determined. For example, the relevance scores of multiple candidate objects can be normalized to obtain corresponding weights. In this way, the corresponding angular differences can be weighted according to the weights to obtain the final score for each candidate object, and based on the final score of each candidate object, the target object to which the target object points can be determined.
[0061] For example, when a user says "get that to drink," even if the angle between the finger and the book and the water glass is similar, the water glass has a higher weight because its semantic meaning is closer to "get that to drink." Therefore, the system will assign a higher attention weight to the water glass (due to its higher semantic match with "drink"), ultimately selecting the water glass as the target object. This can further improve the system's robustness and accuracy in understanding complex and crowded scenarios.
[0062] In some implementations, if there is no referential information in the voice interaction audio, the embodied robot can directly determine the target object to be operated on. For example, if the semantic information of the voice interaction audio is "take the water cup", the robot can directly locate and take the water cup.
[0063] By using the above methods, vague language references from users (such as "that") can be transformed into precise localizations of specific objects in the environment. This significantly improves the robot's accuracy in understanding natural language commands and the reliability of its operations, thereby providing a clear and unique operational target for the subsequent generation and execution of embodied interactive responses that conform to the user's true intentions.
[0064] In some implementations, to accurately map a user's gesture to a specific object in the environment for spatial reference, a pointing vector can be calculated by extracting the coordinates of key finger points from the gesture sequence, and the angle between the vectors can be calculated by combining the object's center coordinates. This allows for the unique identification of the target object pointed to by the user using the minimum angle criterion. For example, the hand may include the base and tip of the fingers, and step 102, "determining the target object pointed to by the target object based on the angular difference between the hand movement direction and the center coordinates of each object," may include: (102.1) Obtain the first coordinate of the root of the finger and the second coordinate of the fingertip in the object pose image frame sequence, and determine the hand movement direction vector based on the difference between the first coordinate and the second coordinate; (102.2) Calculate the object pointing vector based on the difference between the first coordinate and the center coordinate of each object; (102.3) Based on the hand movement direction vector and the object pointing vector, calculate the angle difference between the hand movement direction of the target object and the center coordinates of each object; (102.4) Based on the magnitude relationship between the multiple angle differences corresponding to multiple candidate objects, the candidate object with the smallest angle difference is determined from the multiple candidate objects as the target object to which the target object points.
[0065] Among them, the base of the finger can be a key point representing the position of the base of the finger in the hand posture, such as the spatial position of the metacarpophalangeal joint of the index finger identified by a three-dimensional human posture estimation model.
[0066] Among them, the fingertip can be a key point representing the position of the fingertip in the hand posture, such as the spatial position of the index fingertip identified by a three-dimensional human posture estimation model.
[0067] The first coordinate can be the position data of the base of the finger in three-dimensional space, such as the three-dimensional coordinates (x_w, y_w, z_w) of the key point of the base of the index finger obtained from the RGB-D depth camera coordinate system, which can be used to characterize the spatial position of the starting point of the hand pointing.
[0068] The second coordinate can be the position data of the fingertip in three-dimensional space, such as the three-dimensional coordinates (x_t, y_t, z_t) of the key point of the index fingertip obtained from the RGB-D depth camera coordinate system, which can be used to characterize the spatial position of the end point of the hand pointing.
[0069] Among them, the hand movement direction vector can be a three-dimensional spatial vector pointing from the coordinates of the base of the finger to the coordinates of the fingertip. For example, the vector v_dir=(x_t-x_m,y_t-y_m,z_t-z_m) obtained by subtracting the first coordinate from the second coordinate can be used to represent the precise pointing direction of the user's hand.
[0070] The object pointing vector can be a three-dimensional spatial vector pointing from the coordinates of the base of the finger to the coordinates of the center of the candidate object. For example, for each candidate object, the vector v_obj obtained by calculating the difference between its center point coordinates and the first coordinate can be used; v_obj is the object pointing vector.
[0071] The angle difference can be the angle between the hand movement direction vector and the object pointing vector, such as the angle θ calculated using the vector dot product formula, and can be expressed in degrees. The angle difference can be used to quantify the alignment between the hand pointing direction and the directions of various candidate objects.
[0072] In some implementations, a pre-trained 3D hand keypoint detection model (e.g., the MediaPipeHands model) can be used to analyze the hand region detected in the object pose image frame. This model can output the three-dimensional coordinates of 21 hand keypoints in the image coordinate system and associated depth information. To achieve stable and accurate pointing direction estimation, this application selects the metacarpophalangeal joint (MCP, typically the 5th keypoint) and the tip of the index finger (typically the 8th keypoint) as the calculation reference points. Optionally, the wrist joint and the tip of the index finger can be selected as the calculation reference points, or other finger bases and fingertips can be selected, etc.
[0073] Furthermore, the first coordinate P_mcp represents the three-dimensional coordinates (x_m, y_m, z_m) of the base of the index finger, and the second coordinate P_tip represents the three-dimensional coordinates (x_t, y_t, z_t) of the tip of the index finger. The hand movement direction vector v_dir can be obtained by calculating the vector difference from the base of the finger to the tip: v_dir=P_tip-P_mcp=(x_t-x_m,y_t-y_m,z_t-z_m); v_dir physically represents the direction the index finger points along the axis.
[0074] Furthermore, to facilitate subsequent geometric calculations, v_dir can be normalized to a unit vector: v_normalized = v_dir / ||v_dir||. This allows the user's intuitive gestures to be transformed into mathematical vectors that can be used for precise spatial calculations.
[0075] In some implementations, in addition to using single-frame static coordinates to calculate the direction vector, multi-frame smoothing and intent start-point detection mechanisms can be introduced. Specifically, the system can maintain a sequence of P_mcp and P_tip coordinates within a short-term window (e.g., the most recent 5 frames). When a pronoun is detected in a voice command, the system backtracks and analyzes the hand movement trajectory to identify the starting frame where the finger transitions from a relaxed state to a stable pointing state. The final P_mcp and P_tip values can be estimated using a weighted average of the coordinates from several frames after the starting frame or after Kalman filtering, thereby effectively eliminating noise caused by minor hand tremors and making the estimation of the hand movement direction vector and object pointing vector more stable and reliable.
[0076] Furthermore, the object center coordinates P_obj=(X_o,Y_o,Z_o) of the kth candidate object identified and calculated from the scene image frame sequence are used to calculate the corresponding object pointing vector v_obj as follows (the coordinates of the finger root used here should be the same as the first coordinate used to calculate the hand movement direction vector): v_obj=P_obj-P_mcp=(X_o-x_m,Y_o-y_m,Z_o-z_m); The vector v_obj_k represents the direction and distance from the base of the user's finger directly to the geometric center of the object. The system iterates through all N candidate objects and calculates a set of object pointing vectors {v_obj_1, v_obj_2, ..., v_obj_N}, providing basic data for the next step of angle comparison.
[0077] In some implementations, the angular difference θ for each candidate object k is calculated based on the geometric principle of vector dot product. Specifically, this is achieved using the following formula: cosθ=(v_normalized·v_obj) / (||v_normalized|| ||v_obj||); cosθ=(v_normalized·v_obj) / ||v_obj||; The angular difference θ = arccos(cosθ), with units of radians or degrees.
[0078] Where · denotes the dot product operation of vectors, and ||v_obj|| represents the magnitude (i.e., length) of the object pointing to the vector v_obj. ||v_normalized|| is the normalized hand movement direction vector with a magnitude of 1. The value of θ ranges from [0,π], and the unit can be radians. The system performs this calculation in parallel or iteratively on all candidate objects, obtaining a set of angle difference values {θ_1,θ_2,...,θ_N}.
[0079] In some implementations, the calculation of the angle difference can incorporate an adaptive threshold or consider distance decay weights. Specifically, for objects very far from the user, even a small angle difference may not represent the user's intended target. Therefore, when calculating the final score, the angle difference angle can be combined with the object distance ||v_obj||, for example, using the formula score=θ. Log(1+||v_obj||) means that the greater the distance, the more the angle difference is amplified. Alternatively, the system can dynamically set a reasonable angle threshold (such as 30 degrees) and only include objects with θ less than this threshold in the final candidate set, thereby eliminating objects that are obviously not within the pointing range in advance, improving computational efficiency and anti-interference ability.
[0080] In some implementations, all calculated angle difference values {θ_1, θ_2, ..., θ_N} can be compared to find the minimum value. That is, the candidate object with the smallest index k that minimizes θ is selected and identified as the target object the user intends to point to. For example, assuming the scene contains a cup (θ=10°), a book (θ=25°), and a mobile phone (θ=8°), the system will select the mobile phone as the target object because its angle difference with the finger direction is the smallest. This allows for a precise conversion from vague reference to a unique identifier.
[0081] By using the above methods, the user's gesture can be quantified into a spatial angular relationship with various objects in the environment. This allows for an objective and quantitative comparison and determination of the object that best matches the pointing direction, solving the problem of ambiguous reference caused by relying on a single voice in traditional methods. This provides a reliable spatial target input for the robot to perform precise operations in the future.
[0082] Step 103: Obtain the pitch vector and semantic vector corresponding to the voice interaction audio, and encode the pitch vector and the object's facial expression image frame sequence to obtain the corresponding emotional features, and encode the semantic vector and the object's pose image frame sequence to obtain the corresponding intent features.
[0083] In some implementations, in order to separate and deeply extract the user's emotional state and operational intent features from multimodal interaction information, pitch vectors and semantic vectors can be obtained by parsing the audio of the voice interaction, and then encoded and fused with the object's facial expression image frame sequence and the object's pose image frame sequence, respectively, to generate high-dimensional features that independently represent emotions and intentions, providing structured input for subsequent cross-modal joint reasoning.
[0084] Among them, pitch vector can be a feature vector that represents the acoustic characteristics of speech, extracted from the audio of speech interaction through speech signal processing technology. It is a set of acoustic parameters such as fundamental frequency, energy, and intonation profile, and can be used to reflect the speaker's emotional state and tone intensity.
[0085] Among them, semantic vectors can be feature vectors that represent the semantic information of speech text content, which are converted and extracted from speech interaction audio through automatic speech recognition and natural language processing technologies. For example, word embedding or sentence embedding vectors can be used to understand the specific meaning and operation purpose of user instructions.
[0086] Among them, the emotional features can be a deep feature representation that comprehensively represents the emotional state of the target object, obtained by jointly encoding the tone vector and the image frame sequence of the object's facial expression through a convolutional neural network and a long short-term memory network.
[0087] Among them, the intent feature can be a deep feature representation that comprehensively represents the operational intent of the target object, obtained by jointly encoding the semantic vector and the object pose image frame sequence through the Transformer encoder.
[0088] Specifically, pitch vectors can be used to represent paralinguistic information in speech. By analyzing preprocessed speech interaction audio (e.g., after framing and windowing), acoustic features including, but not limited to, the following can be extracted: fundamental frequency (F0, reflecting pitch), energy (loudness), Mel-frequency cepstral coefficients (MFCC), spectral centroid, zero-crossing rate, etc. These low-dimensional features are concatenated in temporal order to form a dynamic feature sequence. Subsequently, a temporal model (such as a one-dimensional convolutional network or a recurrent neural network) can be used to encode and pool this feature sequence, outputting a fixed-dimensional pitch vector V_Prosody to encapsulate the emotional nuances (such as urgency, calmness) and emphasis information in the speech interaction audio.
[0089] In some implementations, semantic vectors can be used to represent the textual content information of speech, i.e., "what is said". First, an automatic speech recognition engine can convert the audio stream into a text string. Then, a pre-trained language model, such as the encoder part of BERT or RoBERTa, or a dedicated sentence embedding model, such as Sentence-BERT, is used to convert the text string into a high-dimensional, context-aware vector representation, i.e., a semantic vector V_Semantic, to capture the lexical, grammatical, and primary semantic information of the instruction. In this way, continuous speech signals can be decoupled into independent emotional carriers (pitch) and semantic carriers (text vectors).
[0090] In some implementations, the sequence of facial expression image frames is first processed by a lightweight convolutional neural network (CNN, such as MobileNetV2) to extract features from each frame, resulting in a series of facial expression feature maps. Considering the dynamic nature of emotion, these temporal feature maps and a synchronized tone vector sequence need to be jointly modeled. Therefore, this application employs a dual-stream coding network. Specifically, the visual stream uses a CNN-LSTM structure: the CNN extracts the spatial features of each facial expression image frame, while the LSTM captures the evolution pattern of the expression over time, outputting a visual emotion feature H_Visual_Emo. The speech stream (tone vector sequence) is also input into an LSTM network, outputting an auditory emotion feature H_Audio_Emo. Finally, H_Visual_Emo and H_Audio_Emo are fused through a cross-modal attention module or a simple feature concatenation followed by a fully connected layer to generate a unified emotion feature h_emotion. In this way, the complementarity and enhancement of visual and auditory emotional cues are achieved, resulting in a more robust emotion representation than a single modality.
[0091] In some implementations, 2D or 3D keypoint coordinates of the user in each frame of an object pose image frame sequence can be extracted using human pose estimation models (such as OpenPose and HRNet). These keypoint coordinate sequences constitute temporal data of limb movements. Simultaneously, the semantic vector V_Semantic provides the context of language instructions. Thus, the keypoint coordinate sequence can be treated as a spatiotemporal graph data, processed using a spatiotemporal graph convolutional network (ST-GCN) or a Transformer encoder to capture dynamic patterns of posture (such as waving, pointing, bending), and output action features H_Action. The semantic vector can be used directly as input or adapted through a fully connected layer. Subsequently, a modal interaction module (e.g., an attention mechanism using semantic features as queries and action features as keys and values) guides attention to relevant action patterns through semantic information and refines semantic understanding using action information. The output of the modal interaction module is the intent feature vector h_intent. This enables deep alignment and joint understanding of language instructions and accompanying limb movements.
[0092] By using the above methods, the pitch and semantic information in speech and the facial expressions and posture information in vision can be specifically encoded to form independent feature representations of emotion and intention, respectively. In this way, the emotional tendencies and intention goals in user interactions can be captured more accurately and in a more structured manner, overcoming the shortcomings of single-modality or simple splicing of features. This provides a high-quality and differentiated feature foundation for achieving deep collaborative integration and reasoning of emotion and intention in subsequent steps.
[0093] In some implementations, to achieve differentiated feature encoding using deep neural networks for different modal data characteristics, and to fully extract the semantic information contained therein, a pre-defined integrated target model can be used. Its internal self-attention computing module can be used to specifically process temporal tone and facial expression data to extract emotional features, and a spatiotemporal recurrent network module can be used to specifically process posture data with both semantic and spatiotemporal structure to extract intention features. This achieves more accurate and efficient multimodal feature extraction and representation. For example, step 103, "encoding the tone vector and the object's facial expression image frame sequence to obtain the corresponding emotional features, and encoding the semantic vector and the object's posture image frame sequence to obtain the corresponding intention features," may include: (103.1) The pitch vector and the image frame sequence of the object's facial expression are encoded by the self-attention calculation module of the preset target model to obtain the corresponding emotional features; (103.2) The semantic vector and the object pose image frame sequence are encoded by the spatiotemporal recurrent network module of the target model to obtain the corresponding intention features.
[0094] The target model can be a pre-built and trained integrated neural network model for multimodal feature extraction and fusion, such as a deep learning model containing multiple dedicated encoder branches and fusion layers, which can be used to process multi-source heterogeneous input data in a unified and efficient manner.
[0095] The self-attention computation module can be a network component in the target model that is specifically designed to process sequential data and capture long-distance dependencies, such as the multi-head self-attention mechanism layer in the Transformer architecture.
[0096] Among them, the spatiotemporal recurrent network module can be a network component in the target model that is specifically designed to process spatial structure and time series data simultaneously, such as a hybrid model that combines convolutional neural networks and long short-term memory networks (CNN-LSTM network).
[0097] In some implementations, emotion feature encoding is performed by the emotion encoder branch in the target model. For example, pitch vectors and sequences of image frames of the subject's facial expressions can be mapped to a unified feature dimension through linear projection layers, and positional encoding can be added to preserve temporal information before concatenation to obtain a multimodal sequence. This multimodal sequence is then input into a self-attention computation module. The self-attention computation module consists of multiple identical stacked layers, each containing a self-attention sublayer and a feedforward neural network sublayer. In the self-attention computation module, each position (corresponding to a pitch or facial expression feature at a given moment) aggregates global information by calculating attention weights with all positions in the multimodal sequence. In this way, the self-attention computation module can automatically learn the intrinsic relationship between pitch changes and facial expression changes.
[0098] Furthermore, after multi-layer Transformer encoding within the self-attention computation module, the output vector corresponding to the special classification flag ([CLS]) is taken, or the entire sequence output is pooled (e.g., average pooling), to obtain the emotion feature vector h_emotion that integrates tone and facial expression information. For example, when a user's tone rises and their face displays a surprised expression, the self-attention mechanism strengthens the interaction between these two features, making the output emotion feature more biased towards the surprise category. In this way, deep modeling and fusion of nonverbal emotional cues are achieved.
[0099] In some implementations, a human pose estimation model can be used to extract the coordinates of 2D human key points in each frame of an object pose image frame sequence. These coordinates are standardized and organized into a tensor of dimension (T, J, C), where T is the number of time steps (frames), J is the number of key points, and C is the number of coordinate channels (e.g., x, y). This tensor constitutes a spatiotemporal sequence of limb movements.
[0100] Furthermore, the J keypoints of each frame can be processed spatially using a two-dimensional convolutional neural network (2D-CNN) within a spatiotemporal recurrent neural network module (CNN-LSTM). The (J,C) data of a single frame is treated as an image, where the J keypoints are analogous to pixel locations, and the C coordinate channels are analogous to color channels. The CNN learns spatial association patterns between joints (e.g., the coordinated movement of the elbow and wrist) through its convolutional kernels and outputs a sequence of spatial feature vectors.
[0101] Subsequently, this spatial feature vector sequence is input into a Long Short-Term Memory (LSTM) network. The LSTM unit processes the sequence in chronological order, and its internal gate control mechanism (input gate, forget gate, output gate) enables it to capture the long-term temporal dependencies and dynamic evolution patterns of actions. The hidden state at the final time step of the LSTM or the aggregation of its outputs at all time steps (such as average pooling) constitutes the action feature vector H_Action that encodes spatiotemporal information.
[0102] Furthermore, H_Action can be fused with the semantic vector V_Semantic. Fusion can be achieved by concatenating features and then adding a fully connected layer, or by using an attention mechanism (with V_Semantic as the query and H_Action as the key), ultimately generating the intent feature vector h_intent.
[0103] By using the above methods, the most suitable deep network architecture can be matched for feature learning of emotion-related modalities (tone, expression) and intention-related modalities (semantics, posture). This allows for more complete extraction and preservation of key information and complex patterns within each modality, improving the representational ability and discriminative power of the generated emotion and intention features. Consequently, it provides an accurate and efficient feature foundation for high-quality cross-modal feature fusion and joint inference in subsequent steps.
[0104] Step 104: Based on the fusion of emotional features and intention features, the corresponding target fusion features are obtained.
[0105] In some implementations, in order to dynamically weigh and integrate emotional and intentional information, emotional features and intentional features can be fused to obtain a target fusion feature that can more comprehensively and accurately characterize the user's overall interaction state.
[0106] Among them, the target fusion feature can be a joint feature vector that uniformly represents the user's emotional tendency and operational intention, obtained by cross-modal fusion of emotional features and intention features.
[0107] For example, the intent feature h_intent can be used as the query, and the emotion feature h_emotion, along with emotion-related action features (such as waving frequency and body tension) extracted from the original pose sequence, can be used as the key and value. By calculating the attention weights Attention(Query, Key, Value), an initial fusion vector h_intent_enhanced with enhanced emotion can be output. This allows intent understanding to possess emotional sensitivity.
[0108] Furthermore, the initial fusion vector h_intent_enhanced, which enhances emotion, can be weighted and fused with the original emotion feature h_emotion. Specifically, different weights can be assigned to the initial fusion vector and the emotion feature. These weights can be set according to the actual situation; for example, the weight corresponding to the initial fusion vector could be 0.6, and the weight corresponding to the emotion feature could be 0.4, and so on.
[0109] By using the above methods, emotions and intentions can be deeply integrated at the feature level to generate a more complete joint representation. This allows subsequent classification and decision-making modules to simultaneously consider the user's emotional state (such as eagerness or calmness) and operational intention (such as requesting items or asking to stop), thereby providing accurate and comprehensive state input for generating human-like, context-adaptive robot responses.
[0110] In some implementations, to dynamically enhance and correct the understanding of intent using emotional information, an attention fusion module can be used. This module uses intent features as queries and emotional features and their derived auxiliary features as keys and values to generate an initial fusion vector that has been weighted and adjusted by emotional information. This initial fusion vector is then weighted and fused with the original emotional features to obtain a target fusion feature that more accurately reflects the intent after emotional modification. For example, step 104 may include: (104.1) Through the attention fusion module, the sentiment auxiliary features associated with the intent features are determined, and the intent features are determined as query vectors, the sentiment features are determined as key vectors, and the sentiment auxiliary features are determined as value vectors; (104.2) Perform cross-modal attention calculation based on query vector, key vector and value vector to obtain the corresponding initial fusion vector; (104.3) Based on the emotional features and the initial fusion vector, a weighted fusion is performed to obtain the corresponding target fusion features.
[0111] The attention fusion module can be a neural network component in the target model used to perform cross-modal attention computation.
[0112] The emotional auxiliary features can be sub-feature vectors extracted or derived from the emotional features to assist in representing the emotional state during attention computation. They can be generated by the attention fusion module or obtained by mapping from the emotional features through a fully connected layer.
[0113] The query vector can be a feature vector used in attention computation to retrieve and match other information. For example, using the intent feature as a query vector can be used to find the most relevant emotional context in the emotional information space.
[0114] Here, the key vector can be a feature vector used in attention calculation to perform similarity matching with the query vector. For example, using sentiment features as key vectors can be used to measure the relevance of various dimensions of sentiment state to the current intent.
[0115] Here, the value vector can be a feature vector that is weighted and aggregated based on the matching results during attention calculation. For example, using sentiment auxiliary features as value vectors can integrate relevant sentiment information into the intent features.
[0116] The initial fusion vector can be an intermediate feature vector obtained through cross-modal attention calculation and preliminarily weighted by sentiment information. For example, it can be a feature representation obtained by calculating attention weights from the query vector and key vector, and then weighting and summing the value vectors.
[0117] In some implementations, emotion-enhancing features can be dynamic motion features strongly correlated with emotional expression, further extracted from the object's pose image frame sequence. Emotion-enhancing features may include, but are not limited to, the frequency and amplitude of hand gestures, the forward / backward tilt angle of the body posture, and the urgency of the movement (calculated through optical flow or acceleration features). Emotion-enhancing features can be extracted from the pose sequence using a lightweight temporal analysis module (such as differentiating the pose keypoint sequence and calculating statistical features) or a dedicated emotion-motion encoding sub-network, serving as a supplement to h_emotion to provide specific manifestations of emotion in body movements.
[0118] Specifically, the query vector can be obtained by passing the intent feature h_intent through a linear projection layer; the key vector can be obtained by passing the sentiment feature h_emotion through another linear projection layer; and the value vector can be obtained by passing the sentiment auxiliary feature through a third linear projection layer. Each of the three linear projection layers corresponds to a different weight matrix. In this way, an attention-based retrieval framework can be constructed, using intent as the question, sentiment state as the index, and sentiment action as the answer, laying the foundation for deep cross-modal interaction.
[0119] Furthermore, the attention fusion module can calculate the corresponding attention weights, as follows:
[0120] in, , , These are the query, key, and value vector matrices obtained in the previous step; Represents the transpose of the key matrix; It is the dimension of the key vector. The softmax function is used to scale the dot product result to prevent the gradient from becoming too small; it normalizes the similarity score to a weight distribution. The resulting vector is the initial fusion vector h_intent_enhanced. For example, when h_intent points to "request to stop," and h_emotion and the initial fusion vector indicate "urgent" and "rapid waving," the attention weights will cause the initial fusion vector to strongly incorporate the action feature of "rapid waving," thus obtaining an enhanced vector with the semantics of "urgent stop request." In this way, the contextual enhancement and correction of intention understanding by sentiment information is achieved.
[0121] Furthermore, the initial fusion vector h_intent_enhanced, which enhances emotion, can be weighted and fused with the original emotion feature h_emotion. Specifically, different weights can be assigned to the initial fusion vector and the emotion feature. These weights can be set according to the actual situation; for example, the weight corresponding to the initial fusion vector could be 0.6, and the weight corresponding to the emotion feature could be 0.4, and so on.
[0122] By using the above methods, attention mechanisms can be used to allow the model to autonomously select the most critical parts from emotional information for understanding the current intention, and adjust the representation of the intention accordingly. This can effectively solve the problem of intention comprehension deviation caused by ambiguous voice commands or strong emotional coloring, and thus provide a higher quality and more complete fusion feature foundation for generating emotionally adapted and intentionally accurate robot responses.
[0123] In some implementations, to adaptively adjust the contribution of sentiment information and intent information to the final fused features, a gated fusion unit can be used to perform weighted fusion of the initial fusion vector and the original sentiment features, thereby achieving a flexible, balanced, and dynamically adjustable sentiment-intent feature fusion based on the interaction context. For example, (104.3) may include: (104.3.1) The first gating weight is output through the gating fusion unit based on the emotional features and the initial fusion vector; (104.3.2) Obtain the reference value and calculate the second gating weight based on the difference between the reference value and the first gating weight; (104.3.3) The initial fusion vector is adjusted by the first gating weight to obtain the first target vector; (104.3.4) The sentiment features are adjusted by the second gating weight to obtain the second target vector; (104.3.5) Based on the first target vector and the second target vector, the target fusion feature is obtained by fusing them.
[0124] The gated fusion unit can be a neural network component used to dynamically generate weight coefficients based on input features, such as a combination of a fully connected layer and an activation function similar to a long short-term memory network gated structure.
[0125] The first gating weight can be a weight coefficient calculated by the gating fusion unit to adjust the initial fusion vector. For example, it can be a scalar value between 0 and 1, which can be used to control the intensity of the emotion-weighted intent information (i.e., the initial fusion vector) in the final fusion feature.
[0126] The reference value can be a preset constant used to calculate the weight that is complementary to the first gating weight, such as the value 1.
[0127] The second gating weight can be used to control the intensity of sentiment features in the final fused features.
[0128] The first target vector can be an adjusted feature vector obtained by multiplying the initial fusion vector by the first gating weight.
[0129] The second target vector can be an adjusted feature vector obtained by multiplying the sentiment feature by the second gating weight.
[0130] In some implementations, the emotion feature h_emotion and the initial fusion vector h_intent_enhanced are first concatenated to form a joint feature vector [h_intent_enhanced; h_emotion](link to h_emotion). This joint vector is then non-linearly transformed through one or more fully connected layers (also known as dense layers), and finally passed through a sigmoid activation function to output the first gating weight g. The sigmoid function maps the output value to the (0,1) interval, allowing g to be considered a scaling factor.
[0131] The calculation process of the first gating weight g can be expressed as: g=σ(W_g [h_intent_enhanced;h_emotion]); Here, σ represents the Sigmoid activation function, and W_g is the learnable weight matrix of the corresponding layer in the gated fusion unit. g can be a scalar (global gating) or a vector with the same dimension as the feature vector (element-by-element gating). The magnitude of g reflects the relative importance of the sentiment-modulated intent information (i.e., h_intent_enhanced) to the original sentiment information in the current context. For example, when h_intent_enhanced has already well fused the sentiment information, g may be large, indicating that subsequent actions will rely more on this vector.
[0132] In some implementations, the reference value is a constant 1. The second gating weight can be obtained by calculating the difference between the reference value and the first gating weight g, and can be represented as 1-g.
[0133] As g increases, 1-g automatically decreases, and vice versa, ensuring that the fused features do not suffer from scaling issues due to unconstrained weights. This achieves an automatic balance of the contributions of the two information sources.
[0134] Furthermore, the first target vector can be obtained by element-wise multiplying the first gating weight g with the initial fusion vector h_intent_enhanced; and the second target vector can be obtained by element-wise multiplying the second gating weight 1-g with the sentiment feature h_emotion. Finally, the first and second target vectors are added together to synthesize the independent and important parts of the initial fusion vector and the sentiment feature, resulting in the target fusion feature h_joint, which dynamically reflects the joint state of sentiment and intent. The specific process is as follows: h_joint=g⊙h_intent_enhanced+(1-g)⊙h_emotion; By using the above methods, it is possible to decide whether to lean towards the emotionally modified understanding of intent or to retain the purer emotional state, depending on the specific situation. This avoids information suppression or distortion during the fusion process, generating a target fusion feature that more accurately and evenly reflects the user's overall state. In turn, this provides the optimal feature basis for the robot to make response decisions that are both in line with the operational goals and rich in emotional empathy.
[0135] Step 105: Classify according to the target fusion features to obtain sentiment classification results and intent classification results, and generate the corresponding target interaction sequence based on the target object's position coordinates, sentiment classification results and intent classification results.
[0136] In some implementations, in order to transform the results of multimodal fusion understanding into specific and executable robot response planning, clear emotion and intention labels can be obtained by parallel classification of target fusion features, and combined with the precise spatial coordinates of the target object, a sequence of instructions that coordinates multimodal outputs (such as voice, action, and visual feedback) can be generated to achieve closed-loop control from intelligent understanding to human-like execution.
[0137] Among them, the emotion classification result can be a discrete label representing the emotional state of the target object after classifying the target fusion features, such as the categories such as eagerness, calmness, and confusion output by the classification model, which can be used to guide the robot to generate a response with corresponding emotional adaptability.
[0138] The intent classification result can be a discrete label representing the operational intent of the target object, obtained by classifying the target fusion features. For example, the categories output by the classification model, such as requesting an item, requesting to move, or expressing dissatisfaction, can be used to determine the core task type that the robot needs to perform.
[0139] Among them, position coordinates can be data used to characterize the specific position of the target object in the robot's operating space, such as the three-dimensional center coordinates (X_o, Y_o, Z_o) of the target object, which can be used to plan the robot's movement trajectory, robotic arm grasping pose, and other physical actions.
[0140] The target interaction sequence can be a sequence of instructions generated based on emotion classification results, intent classification results, and target object position coordinates. It is used to control the robot to perform multimodal interactions. For example, it can include a set of ordered instructions such as turning to the target object, playing a confirmation voice, and displaying a focused expression on the screen. It can be used to coordinate the robot's voice, actions, and interface feedback to complete a complete interaction loop.
[0141] In some implementations, the target model's classification head can consist of independent, pre-trained sentiment and intent classification heads. Each classification head can consist of one or more fully connected layers and a softmax activation function at the end.
[0142] For example, for sentiment classification, a probability distribution vector can be output across all predefined sentiment categories (such as calm, anxious, confused, and satisfied), and the category with the highest probability is taken as the final sentiment classification result. Similarly, for intent classification, the parameters of the intent classification head are used to calculate the probability distribution across predefined intent categories (such as picking up an object, placing it, moving it, asking a question, and stopping), and the category with the highest probability is selected as the intent classification result. In this way, a precise and joint determination of the user's overall state (sentiment + intent) is achieved.
[0143] Furthermore, after obtaining the emotion classification results and intent classification results, the position coordinates of the target object can be input into the embodied robot's response generator. The response generator can then retrieve or select the most matching response strategy based on the emotion and intent classification results. Subsequently, the response strategy is instantiated with specific contextual parameters (the position coordinates of the target object) to generate a structured target interaction sequence.
[0144] For example, the target interaction sequence can be a set containing multiple parallel or serial instructions, and may include: Voice feedback: Generate natural language statements based on intent and emotion (e.g., when calm: "Okay, I'll get you the cup."; when urgent: "Bring the cup right away!").
[0145] Physical motion planning: Based on the position coordinates of the target object and the intention classification results, calculate the motion trajectory of the embodied robot (such as the grasping path of the robotic arm, the moving path of the chassis, and the turning of the head gimbal to gaze at the target object).
[0146] Visual feedback: Based on the emotion, display corresponding expression icons or status prompts on the robot's facial screen.
[0147] Transitional responses: If the task is complex or takes a long time to execute, insert transitional actions (such as nodding) or voice (such as "planning the route, please wait") to maintain a smooth interaction.
[0148] By using the above methods, the abstract, multi-dimensional understanding (emotion, intention, space) in the upstream steps can be visualized into a structured action plan that can directly drive the robot's various execution units. This ensures that the robot's final response is semantically consistent with the user's intention, emotionally matches the user's state, and spatially precisely manipulates the target object, thereby achieving a seamless connection and a highly unified human-like interactive experience from perception and understanding to physical interaction.
[0149] In some implementations, to efficiently and accurately decode the user's explicit emotional state and operational intent labels, parallel intent classification heads and emotion classification heads in a preset target model can be used to independently classify the same target fusion feature, thereby simultaneously obtaining structured and interpretable emotion classification results and intent classification results, thus achieving a precise mapping from fusion representation to specific semantic labels. For example, step 105, "classifying according to the target fusion feature to obtain emotion classification results and intent classification results," may include: (105.1) The target fusion features are classified using the pre-defined sentiment classification head of the target model to obtain the sentiment classification result; (105.2) The target fusion features are classified by the intent classification head of the target model to obtain the intent classification result.
[0150] The intent classification head can be a network component in the target model specifically designed to predict the category of the operation intent from the target fusion features. For example, it can be an output layer consisting of a fully connected layer and a Softmax activation function, which can be used to map the fusion features to specific intent category labels such as requesting an item or inquiring about a status.
[0151] The sentiment classification head can be a network component in the target model specifically designed to predict the sentiment state category from the target fused features. For example, it can be an output layer consisting of a fully connected layer and a Softmax activation function, which can be used to map the fused features to specific sentiment category labels such as eagerness, calmness, and confusion.
[0152] For example, the intent classification head can be a lightweight classification network, which can consist of one or more fully connected layers and a final Softmax activation function. Specifically, the target fusion feature h_joint is first non-linearly transformed and dimensionality reduced through fully connected layers, mapping it to a dimension equal to the number of predefined intent categories. Then, the probability distribution of each category is output through the Softmax function, and the category with the highest probability is taken as the final intent classification result.
[0153] Understandably, the structure of the sentiment classification head is similar to that of the intent classification head, but its parameters are independent, and the output dimension corresponds to the number of sentiment categories. It receives the same target fusion feature h_joint as the intent classification head, calculates the parameter matrix independently, and outputs a probability distribution vector, such as [P(calm) = 0.70, P(urgency) = 0.25, P(confusion) = 0.05]. The sentiment category with the highest probability is taken as the final sentiment classification result (e.g., calm).
[0154] By using the above method, two parallel and functionally specialized classification heads can be used to simultaneously decode key user state semantics from fusion features rich in multimodal information. This enables efficient and accurate fine-grained analysis of complex interaction states (integrating emotion and intent), avoiding information confusion or accuracy loss that may occur with a single classification head. This provides direct and reliable decision input for generating semantically clear and emotion-adaptive robot interaction sequences in subsequent steps.
[0155] Step 106: Control the embodied interactive robot to perform action interaction and voice feedback based on the target interaction sequence.
[0156] In some implementations, in order to ultimately translate intelligent understanding and decision-making into human-like, multimodal responses of the robot in the physical world, the target interaction sequence can be parsed and executed, the motion mechanism of the embodied interactive robot can be coordinated and controlled to perform corresponding physical actions, and its voice module can be driven to output voice with content and tone matching, thereby realizing a complete closed loop from information understanding to physical interaction, significantly improving the naturalness and completeness of the interaction.
[0157] Among them, embodied interactive robots can be physical robots that integrate a mechanical body and a central execution controller. For example, physical robot platforms equipped with a mobile chassis, robotic arm, microphone, speaker and display screen can be used to interact with users in real physical environments through multimodal methods such as action and voice.
[0158] Among them, motion interaction can be the movement executed by the embodied interactive robot according to the target interaction sequence, which generates physical interaction with the environment and the user, such as controlling the robotic arm to grasp the target object, controlling the gimbal to rotate to look at the user or target, and controlling the robot chassis to move to a designated position.
[0159] Among them, voice feedback can be voice information containing semantic content and emotional tone generated by the embodied interactive robot based on the target interaction sequence and output through a speaker. For example, playing confirmatory or explanatory statements such as "Okay, I will get you the red cup" can be used to convey clear semantic information and emotional state to the user.
[0160] In some implementations, after the response generator module generates the corresponding target interaction sequence, the central execution controller of the embodied robot can receive and parse the target interaction sequence. Specifically, it can parse out individual instruction elements and distribute them to the corresponding underlying execution modules according to their type (motion control, speech synthesis, graphics display): motion planning module, speech synthesis module, human-machine interface module, etc.
[0161] Furthermore, to ensure that voice, motion, and interface feedback form a natural and harmonious whole, the controller needs to coordinate the timing. For time-dependent instructions, the central execution controller calculates and issues precise trigger timestamps or synchronization signals. For example, it can plan for the robot to start playing voice as soon as (or slightly after) its robotic arm begins moving towards the target, and when the voice announces a keyword (such as "cup"), control the robot's head gimbal to precisely focus on the target object, with a focus icon simultaneously displayed on the screen. In this way, it can effectively avoid disconnection or conflict between different modal feedbacks, creating a human-like experience that is consistent with speech and action.
[0162] Furthermore, after receiving motion commands (such as grasp coordinates and waypoints), the motion planning module can utilize the robot's kinematic model and path planning algorithms, such as Rapidly-exploring Random Tree (RRT) and trajectory optimization, to calculate the smooth trajectory of each joint of the robot and send it to the motor driver for execution. During motion execution, the central execution controller can perform closed-loop detection through the robot's joint encoders, visual servos, or force sensors. If an anomaly is detected (such as path occlusion or grasp slippage), the current motion will be interrupted, and an execution failure event will be reported to the decision layer, which may trigger subsequent actions such as retrying or requesting human assistance.
[0163] Furthermore, the speech synthesis module can convert text into fluent speech with a specified intonation and emotion, and play it through a speaker. The human-machine interface module controls the screen to display corresponding facial animations or status information.
[0164] This application embodiment acquires the target object's voice interaction audio, object facial expression image frame sequence, object posture image frame sequence, and scene image frame sequence. When the voice interaction audio contains referential information, the object center coordinates of each candidate object in the scene image frame sequence are determined. The direction of the target object's hand movement is determined based on the object posture image frame sequence, and the target object pointed to by the target object is determined based on the angular difference between the hand movement direction and the center coordinates of each object. The pitch vector and semantic vector corresponding to the voice interaction audio are acquired, and the pitch vector and the object facial expression image frame sequence are encoded to obtain the corresponding emotional features. The semantic vector and the object posture image frame sequence are encoded to obtain the corresponding intention features. The emotional features and intention features are fused to obtain the corresponding target fusion features. The target fusion features are classified to obtain emotional classification results and intention classification results. Based on the target object's position coordinates, emotional classification results, and intention classification results, the corresponding target interaction sequence is generated. The embodied interactive robot is controlled to perform action interaction and voice feedback based on the target interaction sequence. In this way, through the synchronous acquisition and fusion understanding of multimodal information, the transformation from single voice interaction to collaborative understanding of voice, vision, action, and scene can be realized. Specifically, by combining gesture pointing with the coordinate analysis of scene objects to interpret referential information, the visual information collected by the embodied robot can be fully utilized. This allows visual information to be used not only for object manipulation but also for improving the accurate understanding of ambiguous or omitted instructions. Simultaneously, by encoding and fusing multimodal features such as tone of voice, facial expressions, and posture, the collaborative inference of the user's emotional state and true intentions is enhanced, thereby generating an interaction sequence that takes into account target location, emotion, and intention. Finally, based on the fused understanding results, a coordinated and consistent multimodal response is generated, making the embodied robot's feedback more in line with human interaction habits. In summary, this application can improve the accuracy of understanding and the naturalness of interaction in the interactive robot process.
[0165] In some implementations, to improve the decision-making security and interaction robustness of the system in complex and uncertain interaction scenarios, a confidence estimation mechanism can be introduced to quantitatively assess the reliability of the fusion understanding results. When the confidence level is insufficient, a clarification process is proactively triggered, and historical interaction context is integrated for re-decision-making. This avoids erroneous operations due to misunderstandings or insufficient information, thereby constructing a reliable human-computer interaction closed loop with risk perception and proactive clarification capabilities. For example, the embodied interactive robot control method may further include: (A.1) Output the target confidence corresponding to the target fusion feature through the preset target model confidence estimation head; (A.2) Obtain the preset confidence threshold. When the target confidence is less than the confidence threshold, trigger the clarification command. (A.3) Based on the clarification instruction, obtain the historical interaction context associated with the voice interaction audio, and generate the corresponding target interaction sequence based on the historical interaction context, the location coordinates of the target object, the emotion classification result and the intent classification result; (A.4) Control the embodied interactive robot to perform action interaction and voice feedback based on the target interaction sequence.
[0166] The confidence estimation head can be a network component in the pre-defined target model specifically used to evaluate the reliability of the classification or inference results corresponding to the target fusion features. For example, it can be a regression layer parallel to the sentiment classification head and the intent classification head, which can be used to output a score representing the overall understanding confidence.
[0167] The target confidence score can be a numerical value calculated using a confidence estimation head to quantify the reliability of the current multimodal fusion understanding result. For example, it can be a scalar score between 0 and 1, used to determine whether the system should adopt the current understanding result and perform subsequent operations.
[0168] The confidence threshold can be a pre-set numerical threshold used to compare with the target confidence level to decide whether to trigger the clarification process. For example, 0.7 can be used as a standard to judge whether the understanding result is reliable enough and executable.
[0169] Among them, the clarification instruction can be a special control instruction generated internally by the system when the target confidence level is lower than the confidence level threshold. It can be used to stop the current execution process and initiate active interaction to clarify the user's intention, such as triggering the generation of the rhetorical question "Are you referring to the cup on the left?" or displaying the selection menu.
[0170] Among them, the historical interaction context can be a record of dialogue and scene information that occurred in time before the current voice interaction audio, associated with the current voice interaction audio. For example, instructions issued by the user in the previous few minutes, robot responses, and sequences of state changes of environmental objects can be used to provide additional background reference for understanding the current ambiguous instructions when clarifying or re-deciding.
[0171] In some implementations, the confidence estimation head can output a target confidence score calculated based on probability distribution entropy or a specific branch. When the target confidence score is less than a confidence threshold, it indicates that the current target fused features are fuzzy or at the class boundary, and the classification result may be unreliable. In this case, the system can refrain from generating subsequent actions and instead trigger an active learning or clarification instruction (e.g., "I'm not sure, do you want me to pick it up or just take a look?"), recording this interaction data for subsequent model optimization to improve security and user experience in actual deployments.
[0172] Furthermore, when a clarification command is triggered, the system can retrieve recent dialogue turns and interaction results associated with the current voice interaction audio from its short-term memory module. For example, it can retrieve objects mentioned by the user in the past few minutes, tasks performed, and preferences expressed. Then, a target interaction sequence is generated. The strategies for generating the target interaction sequence include, but are not limited to: Clarification of referential meaning: Based on uncertain target object location coordinates and intent classification results (such as "retrieving an object"), but with low target confidence, the system generates a selection query containing multiple candidate objects. For example, if there are red and blue cups in the scene, and the direction is ambiguous, the system generates the voice: "Are you referring to the red cup or the blue cup?", and controls the robot's head to look at the two candidate objects in turn.
[0173] Intent Confirmation and Clarification: Combining intent classification results and sentiment classification results, generate intent confirmation questions. For example: "Do you want me to take it, or just look at it?"
[0174] Personalized clarification based on history: Use historical interaction context to narrow down the scope of clarification. For example, if the history shows that the user has frequently interacted with the remote control in the past 5 minutes, then even if the current direction is ambiguous, the clarification question can be designed as: "Do you need the remote control?".
[0175] Understandably, when the target confidence level is less than the confidence level threshold, the main purpose of generating the target interaction sequence is to obtain user confirmation rather than to directly execute physical operations.
[0176] Furthermore, the robot will precisely execute the generated target interaction sequence. For example, it will verbally recite the aforementioned question while simultaneously pointing its head or the end effector of its robotic arm sequentially to several of the most likely candidate objects (based on the target object), and displaying a puzzled or expectant emoticon on the screen. This process returns the initiative of the interaction to the user, awaiting secondary input from the user (such as verbal confirmation of "the red one," a nod, or a more specific pointing gesture). After receiving confirmation feedback from the user, the system will update its understanding, regenerate the updated target interaction sequence with high confidence, and execute the updated target interaction sequence.
[0177] By employing the above methods, the system can be endowed with the ability to quantitatively perceive the uncertainty of its own understanding, and proactively intervene at high-risk (low-confidence) decision points. By introducing historical context for clarification or reanalysis, it can effectively prevent the forced execution of potentially erroneous operations when information is insufficient or ambiguous. This significantly improves the system's security and interactive reliability in real and complex environments, thereby building a more robust, intelligent, and trustworthy embodied interactive experience.
[0178] Please refer to Figure 3 The following is combined with Figure 3 The overall process of the embodiments of this application will be introduced.
[0179] For example, the system can synchronously schedule and collect user input through various sensor resources (such as microphones and cameras), specifically including voice interaction audio, object facial expression image frame sequences, object pose image frame sequences, and scene image frame sequences. These raw multimodal data constitute the basic resources for subsequent processing.
[0180] Furthermore, computing resources can be allocated to perform multimodal feature extraction. For speech, pitch vectors and semantic vectors are extracted in parallel; for visual data, facial expression feature extraction is performed simultaneously, and the angular differences between the hand movement direction and the center coordinates of each object are determined based on the pose sequence. Simultaneously, scene semantic segmentation and scene graph construction are performed to understand the environment. The extracted features are organized into a unified multimodal feature vector.
[0181] Subsequently, the system enters the core fusion and decision-making phase, allocating computational resources such as attention mechanisms for deep information integration. Through the attention fusion module, the system uses intent features as queries and sentiment features and related auxiliary information as keys and values to calculate cross-modal attention, initially determining the target object pointed to by the user and fusing sentiment and intent information. The preliminary fusion result is further adaptively weighted by a gating fusion unit, ultimately generating high-quality target fusion features. These target fusion features embody the fused user state and environmental understanding.
[0182] Furthermore, based on target fusion features, the system can allocate decision-making resources for decision-making and response. This includes a dialogue management and task planning module, which combines emotion, intent, and target object information from the fusion features to plan interaction logic and task steps. Then, a multimodal response generator is invoked, which, based on the decision results, collaboratively generates a response sequence containing natural language responses and specific action instructions, and may trigger visual feedback (such as screen expression updates). Finally, the generated specific instruction sequence is sent to the robot body, scheduling its physical execution resources to achieve voice feedback and action interactions including head turning, gestures, grasping, and movement, thereby completing a complete and natural multimodal embodied interaction closed loop. The entire process embodies a systematic and hierarchical resource scheduling and execution logic, from raw perceptual data acquisition, computational resource collaborative feature extraction and fusion, to intelligent decision-making and response generation, and finally, scheduling physical action resources.
[0183] Please see Figure 4 This application also provides an embodied interactive robot control device, which can implement the above-described embodied interactive robot control method. The embodied interactive robot control device includes: The acquisition module 41 is used to acquire the target object's voice interaction audio, the object's facial expression image frame sequence, the object's pose image frame sequence, and the scene image frame sequence; The determination module 42 is used to determine the center coordinates of each candidate object in the scene image frame sequence when there is referential information in the voice interaction audio, determine the direction of the hand movement of the target object according to the object posture image frame sequence, and determine the target object to which the target object is pointing according to the angle difference between the direction of the hand movement and the center coordinates of each object. The encoding module 43 is used to obtain the pitch vector and semantic vector corresponding to the voice interaction audio, and to encode the pitch vector and the object expression image frame sequence to obtain the corresponding emotional features, and to encode the semantic vector and the object posture image frame sequence to obtain the corresponding intention features. Fusion module 44 is used to fuse emotion features and intention features to obtain the corresponding target fusion features; The classification module 45 is used to classify based on the target fusion features to obtain the emotion classification result and the intent classification result, and to generate the corresponding target interaction sequence based on the target object's position coordinates, the emotion classification result and the intent classification result; The control module 46 is used to control the embodied interactive robot to perform action interaction and voice feedback based on the target interaction sequence.
[0184] The specific implementation of the embodied interactive robot control device is basically the same as the specific embodiment of the embodied interactive robot control method described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this application, the embodied interactive robot control device may also be equipped with other functional modules to implement the embodied interactive robot control method in the above embodiments.
[0185] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described embodied interactive robot control method. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0186] Please see Figure 5 , Figure 5 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes: The processor 51 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 52 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 52 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 52 and is called and executed by the processor 51 using the embodied interactive robot control method of the embodiments of this application. Input / output interface 53 is used to implement information input and output; The communication interface 54 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 55 transmits information between various components of the device (e.g., processor 51, memory 52, input / output interface 53, and communication interface 54); The processor 51, memory 52, input / output interface 53, and communication interface 54 are connected to each other within the device via bus 55.
[0187] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described embodied interactive robot control method.
[0188] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0189] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0190] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0191] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0192] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0193] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0194] It should be understood that in this application, "at least one" and "several" refer to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0195] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0196] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0197] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0198] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0199] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A method for controlling an embodied interactive robot, characterized in that, The method includes: Acquire the target object's voice interaction audio, object facial expression image frame sequence, object pose image frame sequence, and scene image frame sequence; When the voice interaction audio contains referential information, the center coordinates of each candidate object contained in the scene image frame sequence are determined, the hand movement direction of the target object is determined according to the object posture image frame sequence, and the target object to which the target object is pointing is determined according to the angle difference between the hand movement direction and the center coordinates of each object. The pitch vector and semantic vector corresponding to the voice interaction audio are obtained, and the pitch vector and the object's facial expression image frame sequence are encoded to obtain the corresponding emotional features. The semantic vector and the object's pose image frame sequence are encoded to obtain the corresponding intention features. Based on the fusion of the emotional features and the intention features, the corresponding target fusion features are obtained; The target fusion features are classified to obtain emotion classification results and intent classification results. Based on the location coordinates of the target object, the emotion classification results and the intent classification results, a corresponding target interaction sequence is generated. Based on the target interaction sequence, the embodied interactive robot is controlled to perform action interaction and voice feedback.
2. The embodied interactive robot control method according to claim 1, characterized in that, The hand includes the base of the fingers and the fingertips. Determining the target object pointed to by the target object based on the angular difference between the direction of the hand movement and the center coordinates of each object includes: The first coordinates of the root of the finger and the second coordinates of the fingertip of the target object are obtained in the object pose image frame sequence, and the hand movement direction vector is determined based on the difference between the first coordinates and the second coordinates; Calculate the object pointing vector based on the difference between the first coordinate and the center coordinate of each object; Based on the hand movement direction vector and the object pointing vector, calculate the angular difference between the hand movement direction of the target object and the center coordinates of each object; Based on the magnitude relationship between multiple angle differences corresponding to multiple candidate objects, the candidate object with the smallest angle difference is determined from the multiple candidate objects as the target object to which the target object points.
3. The embodied interactive robot control method according to claim 1, characterized in that, The process of encoding the pitch vector and the object's facial expression image frame sequence to obtain corresponding emotion features, and encoding the semantic vector and the object's pose image frame sequence to obtain corresponding intent features, includes: The pitch vector and the sequence of image frames of the object's facial expression are encoded by the self-attention calculation module of the preset target model to obtain the corresponding emotional features; The semantic vector and the object pose image frame sequence are encoded through the spatiotemporal recurrent network module of the target model to obtain the corresponding intent features.
4. The embodied interactive robot control method according to claim 1, characterized in that, The process of fusing the emotional features and the intention features to obtain the corresponding target fusion features includes: The attention fusion module determines the emotional auxiliary features associated with the intent features, and defines the intent features as a query vector, the emotional features as a key vector, and the emotional auxiliary features as a value vector. Cross-modal attention is calculated based on the query vector, the key vector, and the value vector to obtain the corresponding initial fusion vector; The corresponding target fusion feature is obtained by weighted fusion of the emotional features and the initial fusion vector.
5. The embodied interactive robot control method according to claim 4, characterized in that, The step of weighted fusion based on the emotional features and the initial fusion vector to obtain the corresponding target fusion features includes: The gating fusion unit outputs a first gating weight based on the emotional features and the initial fusion vector. Obtain a reference value, and calculate a second gating weight based on the difference between the reference value and the first gating weight; The initial fusion vector is adjusted using the first gating weight to obtain the first target vector; The emotional features are adjusted using the second gating weight to obtain the second target vector; The target fusion feature is obtained by fusing the first target vector and the second target vector.
6. The embodied interactive robot control method according to claim 1, characterized in that, The classification based on the target fusion features to obtain emotion classification results and intent classification results includes: The target fusion features are classified using the pre-defined intent classification head of the target model to obtain the intent classification result; The target fusion features are classified using the sentiment classification head of the target model to obtain the sentiment classification result.
7. The embodied interactive robot control method according to claim 6, characterized in that, The method further includes: The target confidence level corresponding to the target fusion feature is output through the preset confidence estimation head of the target model. Obtain a preset confidence threshold; when the target confidence level is less than the confidence threshold, trigger a clarification command. Based on the clarification instruction, the historical interaction context associated with the voice interaction audio is obtained, and based on the historical interaction context, the location coordinates of the target object, the emotion classification result, and the intent classification result, a corresponding target interaction sequence is generated; Based on the target interaction sequence, the embodied interactive robot is controlled to perform action interaction and voice feedback.
8. A control device for an embodied interactive robot, characterized in that, The device includes: The acquisition module is used to acquire the target object's voice interaction audio, the object's facial expression image frame sequence, the object's pose image frame sequence, and the scene image frame sequence; The determination module is used to determine the object center coordinates of each candidate object contained in the scene image frame sequence when the voice interaction audio contains referential information, determine the hand movement direction of the target object according to the object posture image frame sequence, and determine the target object to which the target object is pointing according to the angle difference between the hand movement direction and the center coordinates of each object. The encoding module is used to obtain the pitch vector and semantic vector corresponding to the voice interaction audio, and to encode the pitch vector and the object expression image frame sequence to obtain the corresponding emotional features, and to encode the semantic vector and the object posture image frame sequence to obtain the corresponding intention features; The fusion module is used to fuse the emotional features and the intention features to obtain the corresponding target fusion features; The classification module is used to classify the target based on the target fusion features to obtain the emotion classification result and the intent classification result, and to generate the corresponding target interaction sequence based on the position coordinates of the target object, the emotion classification result and the intent classification result; The control module is used to control the embodied interactive robot to perform action interactions and provide voice feedback based on the target interaction sequence.
9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the embodied interactive robot control method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the embodied interactive robot control method according to any one of claims 1 to 7.