Multi-modal data-based intelligent man-machine interaction evaluation method, edge computing device and medium
By using an embodied intelligent human-computer interaction evaluation method based on multimodal data, interactive event data and multimodal signals are acquired, and interactive evaluation indicators are identified in stages. This solves the problem of low evaluation accuracy in existing technologies and realizes a multi-dimensional quantitative evaluation of the interaction quality of intelligent robots.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- KINGFAR INTERNATIONAL INC
- Filing Date
- 2025-12-26
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies in the evaluation of intelligent robots cannot cover the complex human-computer interaction characteristics by using single-indicator evaluation methods based on interface interaction, button behavior, or voice input and output. This results in low accuracy and incomplete evaluation of interaction quality.
An embodied intelligent human-computer interaction evaluation method based on multimodal data is adopted to acquire interaction event data and multimodal signals during the interaction between the user and the robot. By identifying the target interaction stage in stages, interaction evaluation indicators are determined, and the indicator values are determined based on the interaction event data and multimodal signals, and finally the evaluation results of human-computer interaction are determined.
It enables phased and multi-dimensional quantitative evaluation of the interaction quality of intelligent robots, improving the accuracy and comprehensiveness of the evaluation and reflecting the interaction quality of embodied intelligent systems more comprehensively.
Smart Images

Figure CN121958044A_ABST
Abstract
Description
Embodying intelligent human-computer interaction evaluation method based on multimodal data, edge computing devices and media Technical Field
[0001] This disclosure relates to the fields of artificial intelligence, robotics and human-computer interaction, and in particular to an embodied intelligent human-computer interaction evaluation method, edge computing device and medium based on multimodal data. Background Technology
[0002] With the rapid development of artificial intelligence technology, intelligent robots are gradually being applied to multiple industrial fields. Therefore, performance evaluation of intelligent robots is of paramount importance.
[0003] Existing UX (User Experience Evaluation), HMI (Human-Computer Interface Evaluation), or traditional HRI (Robotics Technology and Human-Computer Interaction Evaluation) methods are mostly based on interface interaction, button behavior, or voice input and output to evaluate intelligent robots. They also use single indicators such as task success rate, reaction time, and voice recognition rate for manual data analysis, which has problems such as single test dimensions, low evaluation accuracy, and low efficiency. Summary of the Invention
[0004] One of the technical problems that this disclosure aims to solve is that existing technologies evaluate intelligent robots based on simple data and single indicators such as interface interaction and button behavior, which cannot cover the complex human-computer interaction characteristics, resulting in low accuracy and incomplete evaluation of the interaction quality of intelligent robots.
[0005] To address the aforementioned technical problems, this disclosure provides a method for evaluating embodied intelligent human-computer interaction based on multimodal data, comprising: acquiring interaction event data and multimodal signals collected from the user during the interaction process between the user and the robot; determining a target interaction stage from a preset multi-stage interaction based on the interaction event data, wherein the target interaction stage includes at least one of the following: user intent issuance stage, robot perception and intent understanding stage, robot execution and human-computer collaboration stage, user feedback stage, and robot strategy adjustment stage; determining the interaction evaluation index corresponding to the target interaction stage; determining the index value corresponding to the interaction evaluation index based on the interaction event data and / or multimodal signals; and determining the evaluation result of the human-computer interaction based on the index value.
[0006] In some embodiments, determining the interaction evaluation indicators corresponding to the target interaction stage includes: for at least one of the user intent issuance stage and the robot perception and intent understanding stage, determining the corresponding interaction evaluation indicators including the robot perception and understanding quality indicators; for the robot execution and human-machine collaboration stage, determining the corresponding interaction evaluation indicators including at least one of the behavior planning and design rationality indicators and the user physiological load and collaboration comfort indicators; for at least one of the user intent issuance stage, the user feedback stage, and the robot strategy adjustment stage, determining the corresponding interaction evaluation indicators including at least one of the intent clarification strategy quality indicators and the feedback-driven self-learning performance indicators.
[0007] In some embodiments, for robot perception and understanding quality indicators, determining the corresponding indicator values for interaction evaluation indicators based on interaction event data includes: determining the robot intent recognition result based on interaction event data, and comparing the consistency of the intent recognition result with the prediction results of multiple models to determine the multi-model intent consistency indicator value; determining the time difference between the user intent input time and the robot intent parsing completion time based on interaction event data, and determining the intent parsing delay indicator value based on the time difference; determining the number of times the robot triggers clarification and the number of times it successfully clarifies the user's ambiguous intent based on interaction event data, and determining the ambiguous intent re-determination capability indicator value based on the ratio between the number of successful clarifications and the number of triggers; and weighting the multi-model intent consistency indicator value, the intent parsing delay indicator value, and the ambiguous intent re-determination capability indicator value to determine the indicator value of the robot perception and understanding quality indicators.
[0008] In some embodiments, for the behavior planning and design rationality index, the multimodal signals include electroencephalogram (EEG) signals and eye-tracking signals. Determining the index value corresponding to the interaction evaluation index based on interaction event data and multimodal signals includes: determining robot behavior based on interaction event data, determining the user's expected behavior based on EEG signals, and determining the user's neural mismatch response index value based on the behavioral deviation between the robot behavior and the user's expected behavior; determining expected visual data based on interaction event data, determining the user's visual data based on eye-tracking signals, and determining the visual attention shift index value based on the visual deviation between the expected visual data and the user's visual data; and determining the index value of the behavior planning and design rationality index based on the user's neural mismatch response index value and the visual attention shift index value.
[0009] In some embodiments, for user physiological load and collaborative comfort indicators, multimodal signals include electroencephalogram (EEG) signals, heart rate variability (HRV) signals, electrodermal transfer (EDT) signals, electromyography (EMG) signals, speech information, and facial expression information. Determining the corresponding indicator values for the interaction evaluation indicators based on the multimodal signals includes: identifying user load status based on EEG signals, HRV signals, and EDT signals to determine load status indicator values; identifying user emotions based on EEG signals, HRV signals, EDT signals, speech information, and facial expression information to determine emotion indicator values; determining the comfort level of user-robot collaboration based on EMG signals to determine collaborative comfort indicator values; and determining the indicator values for user physiological load and collaborative comfort based on load status indicator values, emotion indicator values, and collaborative comfort indicator values.
[0010] In some embodiments, for the quality index of intent clarification strategy, the index value corresponding to the interaction evaluation index is determined based on interaction event data, including: determining the actual number of times the robot clarified the user's ambiguous intent and the theoretical number of times the robot clarified the user's ambiguous intent based on the interaction event data, and determining the clarification rationality index value based on the ratio between the actual number of times the robot clarified the user's ambiguous intent; determining the efficiency of the robot in clarifying the user's ambiguous intent based on the interaction event data, and determining the clarification efficiency index value; determining the number of times the robot triggered clarification and the number of times the robot successfully clarified the user's ambiguous intent based on the interaction event data, and determining the clarification success index value based on the ratio between the number of times the robot successfully clarified and the number of times the robot triggered clarification; and determining the index value of the quality index of intent clarification strategy based on the clarification rationality index value, the clarification efficiency index value, and the clarification success index value.
[0011] In some embodiments, the interaction event data includes interaction event data before the robot update and interaction event data after the robot update. For feedback-driven self-learning performance indicators, the indicator values corresponding to the interaction evaluation indicators are determined based on the interaction event data, including: determining the indicator values for the improvement in intent recognition error rate, the change in intent clarification frequency and efficiency, and the indicator values for error correction memory retention capability before and after the robot update based on the interaction event data before and after the robot update; and determining the indicator values for the feedback-driven self-learning performance indicators based on the indicator values for the improvement in intent recognition error rate, the change in intent clarification frequency and efficiency, and the error correction memory retention capability.
[0012] In some embodiments, acquiring interaction event data includes: acquiring robot motion event data and robot voice data; acquiring user motion event data and user voice data; acquiring process event data of the robot performing interactive tasks; and / or, after determining the target interaction stage based on the interaction event data, the method further includes: aligning multimodal signals to the target interaction stage.
[0013] Another embodiment of this disclosure provides an edge computing device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the method described in any of the above embodiments.
[0014] Another embodiment of this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method of any of the above embodiments.
[0015] Through the above technical solution, the embodied intelligent human-computer interaction evaluation method based on multimodal data provided in this disclosure can combine the interaction event data of human-computer interaction and the user's multimodal signals to carry out phased and multi-dimensional quantitative evaluation of human-computer interaction, which effectively improves the accuracy and comprehensiveness of the evaluation of the interaction quality of intelligent robots. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 is a flowchart of the embodied intelligent human-computer interaction evaluation method based on multimodal data disclosed in this embodiment; Figure 2 is a block diagram of the intelligent robot evaluation device based on multimodal signals disclosed in this embodiment; Figure 3 is a block diagram of the edge computing device disclosed in this embodiment. Detailed Implementation
[0018] The embodiments of this disclosure will be further described in detail below with reference to the accompanying drawings and examples. The detailed description of the embodiments and the accompanying drawings are used to illustrate the principles of this disclosure by way of example, but should not be used to limit the scope of this disclosure. This disclosure can be implemented in many different forms and is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
[0019] These embodiments are provided to make the disclosure thorough and complete, and to fully express the scope of the disclosure to those skilled in the art. It should be noted that, unless otherwise specifically stated, the relative arrangement of components and steps, material composition, numerical expressions, and values set forth in these embodiments should be interpreted as exemplary only and not as limiting.
[0020] All terms used in this disclosure have the same meaning as understood by one of ordinary skill in the art to which this disclosure pertains, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and not as idealized or highly formalized, unless expressly defined herein.
[0021] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.
[0022] The physiological data involved in this disclosure are various measurable data signals in the human body, including but not limited to electrocardiogram (ECG) signals, skin temperature (SKT) signals, photoplethysmogram (PPG) signals, electrodermal activity (EDA) signals, heart rate (HR) signals, electromyogram (EMG) signals, electroencephalogram (EEG) signals, and peripheral capillary oxygen saturation (SPO2) signals.
[0023] With the rapid development of artificial intelligence technology, intelligent robots are gradually being applied to multiple industrial fields. Therefore, performance evaluation of intelligent robots is of paramount importance.
[0024] Existing UX, HMI, or traditional HRI evaluation methods are mostly based on interface interaction, button behavior, or voice input and output to evaluate intelligent robots. They also use single indicators such as task success rate, reaction time, and voice recognition rate for manual data analysis, which has problems such as single test dimensions, low evaluation accuracy, and low efficiency.
[0025] Specifically, existing methods for evaluating the interaction of intelligent robots have the following shortcomings: They lack evaluation schemes specific to embodied intelligent systems. Most existing UX, HMI, or traditional HRI evaluation methods target interface interaction, button behavior, or voice input / output, and are not suitable for embodied intelligent robots, which possess autonomous decision-making, human-machine collaboration, and dynamic behavior planning capabilities. These methods cannot evaluate the agent's comprehension ability, decision-making rationality, collaborative fluency, and feedback predictability, thus failing to accurately reflect the interaction quality of embodied intelligent systems.
[0026] The evaluation dimensions are too narrow and the indicators lack integration: Most existing technologies use single indicators such as task success rate, reaction time, and speech recognition rate for evaluation, which cannot cover the multi-dimensional capabilities involved in intelligent interaction, such as intent understanding, reasoning and decision-making, emotional load, collaborative rhythm, and feedback mechanisms. At the same time, there is a lack of a unified indicator system to combine behavioral data with user status, making it difficult to form a comprehensive and structured assessment of interaction quality.
[0027] The data synchronization and alignment capabilities of multimodal assessment methods are insufficient: Although some systems have attempted to use physiological data such as EEG, EDA, and eye tracking to assist in analysis, the lack of a unified time synchronization mechanism and event labeling method between different modalities makes it impossible to align the data to the specific stage of intelligent interaction. Clock offsets, sampling differences, and structural inconsistencies exist between multimodal signals, making it difficult to accurately correlate behavior and physiological responses, thus affecting the accuracy of the assessment.
[0028] The lack of automated indicator extraction and evaluation report generation mechanisms is a significant issue: existing technologies generally rely on manual processing of physiological characteristics, manual analysis of behavioral logs, and manual writing of evaluation reports, resulting in low efficiency and a lack of scalability. Furthermore, the absence of automated data processing, indicator fusion, scoring calculation, and report generation systems for intelligent interaction makes it difficult to meet the needs of rapid iterative verification of intelligent robot interaction systems.
[0029] In view of this, this disclosure proposes an embodied intelligent human-computer interaction evaluation method, edge computing device and storage medium based on multimodal data, which can combine human-computer interaction behavior data and user status to conduct phased and multi-dimensional quantitative evaluation of human-computer interaction, effectively improving the accuracy and comprehensiveness of the evaluation of intelligent robot interaction quality.
[0030] In the technical solution of this disclosure embodiment, firstly, interaction event data and multimodal signals collected from the user during the interaction process between the user and the robot are acquired. Then, based on the interaction event data, a target interaction stage is determined in a preset multi-stage interaction process. Next, the interaction evaluation index corresponding to the target interaction stage is determined. Then, based on the interaction event data and / or multimodal signals, the index value corresponding to the interaction evaluation index is determined. Finally, the evaluation result of human-computer interaction is determined based on the index value. This allows for a phased and multi-dimensional quantitative evaluation of human-computer interaction by combining behavioral data and user status, effectively improving the accuracy and comprehensiveness of the evaluation of the interaction quality of intelligent robots.
[0031] Figure 1 is a flowchart of an embodied intelligent human-computer interaction evaluation method based on multimodal data provided in an embodiment of this disclosure.
[0032] As shown in Figure 1, the embodied intelligent human-computer interaction evaluation method 100 based on multimodal data provided in this embodiment includes steps S110 to S150, where multimodal data includes interaction event data and multimodal signals collected from users.
[0033] Step S110: Acquire interaction event data and multimodal signals collected from the user during the interaction process between the user and the robot.
[0034] Specifically, the interaction event data can include various behavioral events generated during the interaction between the user and the robot. For example, it can include robot behavior events such as various robot actions, robot voice output data, and robot voice recognition results of user input during human-computer interaction, as well as user behavior events such as user actions and user voice input to the robot, and interaction process events of the robot performing human-computer interaction tasks. The user's multimodal signals can include electrophysiological signals, as well as user overt actions and voice multimodal signals during the interaction between the user and the robot. The above-mentioned robot and user-related interaction event data and user multimodal signals during the human-computer interaction process can be collected and acquired by relevant devices. The robot can be, for example, an embodied intelligent robot, that is, an artificial intelligence system with a physical entity. The robot involved in the embodiments of this disclosure is an embodied intelligent robot as an example, and will not be described in detail below.
[0035] Step S120: Based on the interaction event data, determine the target interaction stage from the preset multi-stage interaction.
[0036] Specifically, the target interaction stage includes the interaction process between the user and the robot. By integrating the robot behavior events, user behavior events and interaction flow events of the robot performing the interaction task, the corresponding structured division of the interaction stage is determined. The target interaction stage may include one of the following: user intent issuance stage, robot perception and intent understanding stage, robot execution and human-robot collaboration stage, user feedback stage, and robot strategy adjustment stage. By dividing the interaction stage, the start and end boundaries of key stages in the interaction process can be automatically identified for targeted evaluation.
[0037] Step S130: Determine the interaction evaluation indicators corresponding to the target interaction stage.
[0038] Specifically, interaction evaluation metrics can select the most suitable evaluation modality based on the processing characteristics of each interaction stage as defined above, and construct a multi-source indicator system for behavior, physiology, and cognition in the human-computer interaction process to achieve multi-dimensional quantification of interaction quality. The interaction evaluation metrics corresponding to the target interaction stage are used to reflect the human-computer interaction capabilities of the target stage. For example, when the target stage is the user intention issuance stage and the robot perception and intention understanding stage, the corresponding interaction evaluation metrics can be relevant indicators used to evaluate the embodied intelligent robot's ability to perceive, analyze, and clarify input signals after the user's intention is issued, such as the Perception-Understanding Quality Index (PUQI).
[0039] Step S140: Determine the index value corresponding to the interaction evaluation index based on the interaction event data and / or multimodal signals.
[0040] Specifically, based on the defined target interaction stage and test metrics, and using the acquired corresponding interaction event data or multimodal signals, metric values can be determined. These metric values represent a score for the interaction capability corresponding to the target stage. For example, when the target stage is the user intent issuance stage and the robot perception and intent understanding stage, the test metric is the robot perception and understanding quality metric. This metric includes multiple sub-metrics, and the metric value can include a weighted combination of multiple sub-metrics calculated based on the corresponding interaction event data or multimodal signals.
[0041] Step S150: Determine the evaluation results of human-computer interaction based on the indicator values.
[0042] Specifically, the interaction evaluation results can be results related to the interaction capabilities represented by the corresponding indicator values. For example, when the test indicator is the robot's perception and understanding quality indicator, the results related to the robot's perception, analysis, and clarification capabilities after the user issues an intention can be obtained based on the indicator value. For example, whether the robot system understands accurately, whether the robot system reacts in a timely manner, and whether the robot system can successfully recover the true intention in ambiguous situations.
[0043] In the technical solution of this disclosure embodiment, interaction event data and user multimodal signals during the interaction between the user and the robot are first acquired. Then, based on the interaction event data, the target interaction stage is determined. Next, the interaction evaluation index corresponding to the target interaction stage is determined. Then, based on the interaction event data or multimodal signals, the index value corresponding to the interaction evaluation index is determined. Finally, the interaction evaluation result is determined based on the index value. This enables a phased and multi-dimensional quantitative evaluation of human-computer interaction by combining human-computer interaction behavior data and user status, effectively improving the accuracy and comprehensiveness of the evaluation of the interaction quality of intelligent robots.
[0044] In one example, the embodied intelligent human-computer interaction evaluation method based on multimodal data of this disclosure is executed based on an intelligent robot interaction design evaluation system with multimodal electrophysiological signals. The system includes multiple modules and aims to achieve multi-dimensional quantitative evaluation of the embodied intelligent system in each link of the interaction chain, such as understanding, decision-making, collaboration, and feedback. The function of each module is described below.
[0045] The intelligent robot interaction design and evaluation system based on multimodal electrophysiological signals includes six core modules: embodied interaction scenario and event management module, multimodal electrophysiological and behavioral data acquisition and synchronization module, interaction stage identification module, phased comprehensive indicator system construction module, intelligent interaction capability assessment and design comparison module, and automated evaluation report generation and visualization module.
[0046] I. Embodied Interaction Scenarios and Event Management Module: This module is used to uniformly record robot action commands, voice prompts, feedback behaviors, and user response behaviors, constructing a complete human-computer interaction event flow. This module uniformly records and structurally manages various behavioral events generated during the interaction process of the embodied intelligent robot. Its core significance lies in providing a complete timeline foundation for subsequent multimodal data alignment, interaction stage identification, and the construction of stage-based indicators. Unlike traditional HMI systems that only record user operations or interface events, this module integrates robot-side action commands, voice prompts, and feedback behaviors, as well as multi-source behavioral information from the user side, such as actions, voice, and gaze, and combines this with task logic to generate a replayable and analyzable sequence of interaction events. By constructing a complete interaction event flow, this module enables the system to accurately identify the stage boundaries of the interaction process and achieve refined and interpretable stage-based evaluation.
[0047] II. Multimodal Electrophysiological and Behavioral Data Acquisition and Synchronization Module: This module synchronously acquires multi-source physiological and behavioral signals, including EEG, fNIRS, EMG, EDA, and eye tracking, and achieves high-precision synchronization of cross-modal data through a unified timestamp. During embodied interaction, this module synchronously acquires EEG, ECG, EMG, EDA, eye tracking, and user-manifested movements and speech signals. It uses an NTP (Network Time Protocol) reference clock and high-precision crystal oscillator calibration technology to achieve unified time alignment of multimodal signals, establishing a consistent time reference across devices and modalities. Through this mechanism, the system can accurately map the user's neural responses, emotional fluctuations, physical load, attention allocation, and behavioral performance to specific interaction stages, providing a high-precision multimodal data foundation for subsequent stage recognition and the construction of stage-specific indicators.
[0048] III. The Interaction Stage Identification Module automatically identifies key stages in the interaction process based on robot behavior, user behavior, and task logic. These stages include intent presentation, intent understanding, information addressing, decision execution, collaborative actions, feedback confirmation, and error recovery, thereby achieving phased alignment of multimodal data. This module is based on a five-stage interaction structured partitioning method specifically designed for embodied intelligent robots. By integrating robot behavior events, user behavior logs, and task logic states, it automatically identifies the start and end boundaries of key stages in the interaction process. Unlike traditional UX / HMI evaluations based on interface events, the stage model constructed by this module fully covers the intelligent interaction chain of "intent expression—perception understanding—embodied execution—human-robot collaboration—feedback adjustment," reflecting the real operating mechanism of the embodied intelligent system in dynamic scenarios. This stage model provides a unified structural foundation for subsequent multimodal data alignment, staged indicator construction, and interaction capability scoring, and is a crucial technical link in achieving interpretable evaluation using the method disclosed in this paper.
[0049] Fourth, the phased comprehensive indicator system construction module is used to select the most suitable evaluation modality based on the processing characteristics of each interaction stage, construct a multi-source indicator system including behavior, physiology and cognition, and realize multi-dimensional quantification of interaction quality.
[0050] V. Intelligent Interaction Capability Assessment and Design Comparison Module: This module is used to score core interaction capabilities such as intent understanding ability, human-machine collaboration fluency, feedback predictability, and error recovery ability based on indicators at each stage. It also supports A / B version comparison of different interaction strategies (A / B versions include different versions of robot systems).
[0051] VI. The automated evaluation report generation and visualization module summarizes the stage indicators, capability scores and version differences, and generates a visualized interactive evaluation report. This report consists of three core dimensions, covering the complete interaction chain of "understanding → execution → learning". It aims to quantify, comprehensively and interpretably evaluate the interaction quality of embodied intelligence from both the system and user perspectives, and support the rapid evaluation and iterative optimization of interaction design in the R&D cycle.
[0052] Based on the functions of each module of the system, the following describes the process of the embodied intelligent human-computer interaction evaluation method based on multimodal data disclosed in this disclosure, starting with how to obtain interaction event data.
[0053] In some embodiments of this disclosure, interactive event data is acquired, for example, firstly, robot motion event data and robot voice data are acquired; then, user motion event data and user voice data are acquired; and finally, process event data of the robot performing interactive tasks is acquired.
[0054] Specifically, acquiring robot motion event data involves continuously recording the robot's actions during interaction through the embodied interaction scene and event management module. This includes the start and end times of actions, the motion trajectories of key joints (such as 3D position sequences, joint angle changes, and end effector path points), motion dynamic parameters (velocity, acceleration, and angular velocity), and pose changes (six-dimensional pose sequences). Simultaneously, the system records the action execution status, such as task success, grasping failure, path deviation, and obstacle avoidance triggering. This event information is used to reconstruct the robot's behavioral trajectory, providing core evidence for subsequent human-robot collaboration identification and motion smoothness analysis.
[0055] Specifically, acquiring robot voice data involves, for example, collecting the robot's voice output through the embodied interaction scenario and event management module. This includes prompts, instructions, confirmations, error messages, etc., and recording the semantic type, output text, temporal position, and feedback intent (such as guidance, inquiry, or correction). Simultaneously, the system records the robot's recognition results of user input, such as "understood," "not recognized," or "repeated is needed," to reflect the quality of the robot's intent presentation and feedback chain, providing a foundation for subsequent analysis of user comprehension load and interaction clarity.
[0056] Specifically, user action event data can be acquired, for example, by using the embodied interaction scenario and event management module to collect user actions during interaction through hand tracking, posture recognition, or input device logs. These actions include catching objects, pointing, reaching out, obstacle avoidance, and button confirmation, and the start and end times, action paths, posture stability, and action results (such as correct, incorrect, slow, or no response). These events can reflect the user's understanding of the robot's intentions, physical coordination load, and the usability of the operation method.
[0057] Specifically, user voice data is acquired, for example, by using the embodied interaction scenario and event management module to recognize user voice input in real time, recording the start and end times of the voice, the recognized text, semantic category (confirmation, inquiry, correction, rejection, etc.), and features such as voice fluency. This information (data) is used to analyze the user's understanding of the robot's prompts, the difficulty of expressing intent, and the language burden in the interaction, providing a basis for assessing the naturalness and misunderstanding rate of voice interaction.
[0058] Specifically, acquiring process event data of the robot's interactive tasks involves extracting interactive process events from the robot's internal task scheduler or state machine through the embodied interaction scenario and event management module. These events include logical nodes such as the start and end of task steps, conditional branch triggers, error states (e.g., "target does not exist"), and recovery paths (e.g., "instruction retry"). Process event data (or task logic events) constitute the structural boundary of the interactive process, enabling the system to accurately identify the stage transitions in the interactive process and providing clear structural markers for staged evaluation and indicator calculation.
[0059] In the technical solution of this disclosure embodiment, robot action event data and robot voice data are first acquired, then user action event data and user voice data are acquired, and finally the process event data of the robot performing interactive tasks is acquired. A complete interactive event flow is constructed by combining the robot-side behavior event data, user-side behavior event data and interactive task logic events, so as to accurately identify the stage boundaries of the interactive process, thereby providing effective support for the refinement and interpretability of subsequent stage-based evaluation.
[0060] Then, it describes how to divide the interaction process into five stages based on the interaction event data acquired during the user-robot interaction process: "intent expression - perception and understanding - embodied execution - human-robot collaboration - feedback and adjustment". This allows for the determination of the target interaction stage based on specific robot behavior events, user behavior events, and task flow events.
[0061] Specifically, the target interaction stage is determined by the interaction stage identification module based on specific interaction event data. The target interaction stage includes at least one of the following: user intention initiation stage, robot perception and intention understanding stage, robot execution and human-machine collaboration stage, user feedback stage, and robot strategy adjustment stage. The interaction event data characteristics corresponding to each interaction stage are described below: (1) User Intention Initiation Stage: The user expresses the interaction intention through voice, gestures, body movements, interface clicks or touches, etc. This stage is triggered by user-side behavioral events. The system records the way, timing and input path of the user's instructions to define the starting point of the interaction and to clarify the trigger time of the robot's perception action. This stage mainly describes the moment when "the person begins to communicate with the robot", which is used for input positioning in the subsequent perception stage.
[0062] (2) Robot Perception & Understanding Stage: The robot receives signals from the user through perception modules such as vision, hearing, touch, and environmental modeling, and performs processing steps such as speech recognition, target recognition, intent parsing, and task planning. This stage defines boundaries through robot-side events (such as recognition start / end, recognition confidence, parsing success / failure, candidate intent generation, etc.). This stage reflects the robot's ability to understand external input and its internal decision-making process, and is a prerequisite for the robot to form its next embodied behavioral strategy.
[0063] (3) Robot Execution & Human Collaboration Stage: Based on the understanding results, the robot executes actions or outputs language, including embodied behaviors such as moving, grasping, delivering, giving instructions, and interactive prompts. This stage also includes the user's cooperative behavior with the robot's actions, such as picking up objects, avoiding obstacles, and cooperating in placement, thus reflecting the collaborative process between the two parties. The stage boundary is jointly determined by robot action events and user action events. This stage is the most action-intensive and state-complex part of the intelligent interaction link, used to describe "how the robot acts and how the user responds".
[0064] (4) User Feedback Stage: After the robot completes an action or outputs feedback, the user evaluates the result of its behavior, including confirmation, denial, correction, supplementation, or expression of confusion. The identification of this stage is based on user voice events, operation actions, and interface behavior logs, emphasizing "how people evaluate the robot's performance," which is an important link in determining whether the interaction has achieved the agreed goal.
[0065] (5) Robot Adaptation & Correction Phase: The robot updates its strategy based on user feedback, including correcting erroneous actions, adjusting path planning, enhancing language prompts, or reinterpreting user intent. This phase is defined by robot-side state machine events (strategy update, replanning, secondary execution, etc.) and represents the robot's ability to adapt to feedback. This phase is an important part of the entire interaction cycle to form a closed loop and is also a key feature that distinguishes embodied intelligence from traditional passive systems.
[0066] Therefore, this disclosure proposes a phased interaction structure covering "intent expression - perception and understanding - embodied execution - human-machine collaboration - feedback adjustment - strategy self-learning", which divides the complex and continuous behavioral flow of embodied intelligent systems into identifiable, alignable and quantifiable evaluation units. This fills the technical gap that traditional UX / HMI evaluation cannot cover the multi-stage behavioral quality of embodied intelligent agents, and provides a unified structural foundation for multimodal data mounting and phased indicator construction.
[0067] In the technical solution of this disclosure, based on the division and corresponding determination of target interaction stages using interactive event data, and through the automatic identification of the above five stages, embodied intelligent interaction is structured from a continuous behavioral flow into a resolvable stage chain, giving the complex interaction process clear logical boundaries and temporal structure. This stage division framework can not only adapt to different types of intelligent robot tasks, but also provide a unified structural foundation for subsequent multimodal electrophysiological data alignment, stage index selection, and interaction performance quantification.
[0068] Next, it describes how to acquire the user's multimodal signals and align them to the target interaction stage.
[0069] In some embodiments of this disclosure, after determining the target interaction stage based on interaction event data, the acquired multimodal signals are uniformly time-aligned, for example, the multimodal signals are aligned to the target interaction stage.
[0070] Specifically, the acquisition of user multimodal signals is achieved by using the system's multimodal electrophysiological and behavioral data acquisition synchronization module to simultaneously acquire EEG, ECG, EMG, EDA, eye tracking, and user's visible actions and speech during the embodied interaction process. The system also uses NTP reference clock and high-precision crystal oscillator calibration technology to achieve unified time alignment and build a consistent time reference across devices and modalities to align the multimodal signals to each interaction stage (target interaction stage). Specifically, the following is done: (1) EEG signal acquisition and synchronization: The EEG system acquires the user's brain activity through a multi-channel EEG device and performs EEG signal analysis based on this. The analysis data is used to align and synchronize the signal to the specific interaction stage. The analysis focuses on: mismatch negative wave (MMN): used to evaluate the user's ability to automatically detect abnormalities or deviations in the robot's feedback and to determine whether the user has identified feedback errors or interaction inconsistencies.
[0071] Early visual components (P1 / N1): reflect the user's initial perceptual processing and attention capture of the robot's visual cues.
[0072] Early auditory components (N1 / P2): used to assess the user's processing speed and level of attention to robot voice prompts.
[0073] Energy characteristics in frequency bands such as α / θ / β: used to identify changes in a user's cognitive load, attention deficit, and emotional inhibition.
[0074] θ / β ratio: used to assess short-term cognitive load and decline in attentional resources.
[0075] Alpha power variation: used to identify visual processing inhibition effects and fatigue trends.
[0076] Functional Connectivity: Used to analyze the intensity of coordinated activity between different brain regions during the interaction phase, and to assess the quality of neural activation, information integration ability, and stability of interaction-related neural networks.
[0077] The aforementioned EEG indicators are mainly extracted during the robot execution phase (including action execution and voice output) and the user feedback presentation phase. They are used to assess the user's understanding of the robot's behavior, error sensitivity, and changes in neural load and brain network activation during the interaction process.
[0078] (2) The ECG / PPG system collects the user's ECG signal through ECG or PPG equipment and analyzes the changes in the user's heart rate activity based on this. The analysis data is used to align and synchronize the signal to the specific interaction stage. The analysis focuses on: SDNN (Standard Deviation of Heart Rate): used to assess the user's overall heart rate variability, reflecting the long-term autonomic nervous system regulation ability and stress tolerance level.
[0079] RMSSD (a heart rate metric): used to identify short-term heart rate fluctuations, reflecting a user's immediate stress, attentional engagement, and rapid recovery ability during interactive tasks.
[0080] LF / HF ratio (a heart rate indicator): used to assess changes in the sympathetic-parasympathetic balance and identify autonomic nervous system responses caused by interactive conflict, waiting frustration, or decision-making stress.
[0081] pNN20: Used to assess the proportion of short, rapid changes in the heart rate interval, in order to identify the user's immediate physiological stress and tension during interaction.
[0082] The aforementioned electrocardiogram indicators can be used to assess users' physiological stress changes, tension fluctuations, and physiological adaptability to the interaction rhythm and information load during embodied interaction.
[0083] (3) The electromyography (EMG) system collects the activity of the user's upper limb or hand muscles through surface electromyography sensors and performs electromyography signal analysis based on this. The analysis focuses on: root mean square value (RMS): used to assess the overall activation level of muscles and identify whether there is continuous force or tension in the interactive movement.
[0084] Integral electromyography (iEMG): Used to quantify the total muscle load throughout the entire movement process, in order to compare the physical cost of different interaction designs.
[0085] Median frequency (MF): Used to identify local muscle fatigue trends and determine whether interactive postures lead to fatigue accumulation.
[0086] Coefficient of variation of electromyography (CV-EMG): used to assess the stability, smoothness and consistency of maneuvering.
[0087] Co-contraction Index: Used to identify tension, poor posture, or compensatory muscle exertion.
[0088] The aforementioned EMG indicators can be used to evaluate the rationality of the interaction posture, operational load, motion smoothness, and the impact of the interaction method on the motion execution process.
[0089] (4) The skin conductance signal acquisition (EDA) system collects the skin conductance activity of the user during the interaction process through skin conductance sensors and performs skin conductance signal analysis based on this. The analysis focuses on: SCL rising slope: used to determine the speed at which the user's tension level rises when there is interaction conflict, waiting delay or increased uncertainty.
[0090] SCL Recovery Slope: Used to assess how quickly a user recovers emotionally after completing a task, receiving positive feedback, or resolving confusion.
[0091] SCL Acceleration: Used to identify the intensity and reactivity of a user's emotional fluctuations during interaction.
[0092] SCL Stability Duration: Used to identify the user's emotional stability and adaptability to the rhythm of the interaction.
[0093] SCR peak frequency: used to capture instantaneous emotional fluctuations caused by sudden prompts, feedback errors, or unexpected events.
[0094] The aforementioned EDA indicators are more suitable for assessing emotional tension trends, waiting tolerance, behavioral uncertainty responses, and emotional recovery capabilities during the interaction process, and can reflect the quality of the interaction experience more comprehensively than traditional SCR.
[0095] (5) Eye Tracking system collects visual attention behavior generated by users during robot action execution, voice output and interface prompts through eye tracker, and performs eye tracking data analysis based on this. The analysis focuses on: fixation point position and attention area distribution (Fixation Map): used to determine whether the user effectively pays attention to robot prompts, operation targets or key information areas.
[0096] Fixation Duration: Used to assess the information processing difficulty for users to understand the robot's output information and the cognitive time required for users to understand the prompts.
[0097] Saccade Amplitude & Velocity: Used to identify visual search efficiency and determine whether the interface and prompts are easy to locate.
[0098] Time to First Fixation (TTFF): Used to assess the salience of cues and the predictability of robot behavior.
[0099] Pupil dilation: used to identify changes in cognitive load, such as momentary peaks of load during decision-making, waiting, or dealing with conflict.
[0100] Revisit Count: Used to identify interface uncertainty, difficulty in understanding prompts, or repeated confirmation behaviors.
[0101] Scanpath Entropy: Used to evaluate the efficiency of a user's visual strategies during free exploration or information searching.
[0102] The aforementioned eye-tracking metrics can be used to analyze the clarity of interactive cues, the accessibility of information structures, the predictability of robot behavior, and the efficiency of user attention resource allocation at different interaction stages.
[0103] (6) The Speech and Emotion feature acquisition and analysis system collects the user's speech features during the interaction process through the microphone and performs speech and emotion analysis based on this. The analysis focuses on: Speech Rate: which reflects the user's information processing speed and tension during the interaction. When the interaction is difficult to understand or when pressure is generated, the speech rate usually slows down.
[0104] Pause Duration: Used to identify a user's mental load, level of hesitation, or obstruction of understanding. Longer pauses often occur when there is interaction conflict or increased uncertainty.
[0105] Speech Emotion: Based on fundamental frequency (F0), energy changes, and emotion classification models, it is used to identify the user's emotional state during interaction, such as calm, tension, frustration, and impatience.
[0106] The aforementioned voice and emotion indicators can be used to assess a user's emotional changes, comprehension burden, and the quality of their subjective response to robot prompts or feedback during embodied interaction.
[0107] In the technical solution of this disclosure embodiment, based on a unified timestamp mechanism and interactive event labeling strategy, the collected EEG, ECG, EMG, ESC, eye movement, as well as user's visible actions and speech, multimodal signals are uniformly aligned to the target interaction stage with millisecond-level precision, thereby achieving accurate correlation analysis between physiological reactions and intelligent agent behavior and improving the credibility and precision of the evaluation.
[0108] In some embodiments of this disclosure, the target interaction phase includes at least one of the following: user intent issuance phase, robot perception and intent understanding phase, robot execution and human-robot collaboration phase, user feedback phase, and robot strategy adjustment phase.
[0109] Interactive evaluation metrics include, for example, the Robot Perception and Understanding Quality Index (PUQI), the Behavior Planning and Design Rationality Index (SBPRI), the User Physiological Load and Collaborative Comfort Index (UPLC), the Intent Clarification Strategy Quality Index (CSQI), and the Feedback-Driven Self-Learning Performance Index (FLPI).
[0110] In one example, the interaction evaluation metrics for any two different target interaction stages are at least partially different. For example, the user intent emission stage includes at least the Intent Clarification Policy Quality Index (CSQI); the robot perception and intent understanding stage includes at least the Robot Perception and Understanding Quality Index (PUQI); the robot execution and human-robot collaboration stage includes at least the Behavior Planning and Design Rationality Index (SBPRI) and the User Physiological Load and Collaboration Comfort Index (UPLC); the user feedback stage includes at least the Intent Clarification Policy Quality Index (CSQI) and the Feedback-Driven Self-Learning Performance Index (FLPI); and the robot strategy adjustment stage includes at least the Feedback-Driven Self-Learning Performance Index (FLPI). It is understood that this is done to illustrate that the interaction evaluation metrics for any two different target interaction stages are at least partially different; each target interaction stage may include more types of metrics, and it is possible that some different target interaction stages may have the same interaction evaluation metrics. See below for details.
[0111] Determine the interaction evaluation indicators corresponding to the target interaction stage. For example, for at least one of the user intent issuance stage and the robot perception and intent understanding stage, determine the corresponding interaction evaluation indicators, including the Robot Perception and Understanding Quality Index (PUQI); or for the robot execution and human-robot collaboration stage, determine the corresponding interaction evaluation indicators, including at least one of the Behavior Planning and Design Rationality Index (SBPRI) and the User Physiological Load and Collaboration Comfort Index (UPLC); or for at least one of the user intent issuance stage, the user feedback stage, and the robot strategy adjustment stage, determine the corresponding interaction evaluation indicators, including at least one of the Intent Clarification Strategy Quality Index (CSQI) and the Feedback-Driven Self-Learning Performance Index (FLPI).
[0112] The following section provides a detailed explanation of the interaction evaluation indicators and the calculation process for the indicator values corresponding to each target interaction stage.
[0113] In some embodiments of this disclosure, for the robot's perception and understanding quality index, the index value corresponding to the interaction evaluation index (PUQI) is determined based on interaction event data. For example, firstly, based on the interaction event data, the robot's intent recognition result is determined, and the consistency of the intent recognition result and the prediction results of multiple models is compared to determine the multi-model intent consistency index value; then, based on the interaction event data, the time difference between the user's intent input time and the robot's intent parsing completion time is determined, and the intent parsing delay index value is determined based on the time difference; next, based on the interaction event data, the number of times the robot triggers clarification and the number of times it successfully clarifies the user's ambiguous intent is determined, and the ambiguous intent re-determination capability index value is determined based on the ratio between the number of successful clarifications and the number of triggers; finally, the multi-model intent consistency index value, the intent parsing delay index value, and the ambiguous intent re-determination capability index value are weighted to determine the robot's perception and understanding quality index (PUQI) value.
[0114] Specifically, through the phased comprehensive indicator system construction module, when the target interaction stage includes at least one of the user intent issuance stage and the robot perception and intent understanding stage, the Robot Perception and Understanding Quality Index (PUQI) is determined as the test indicator. This indicator includes three core dimensions: multi-model intent consistency index, intent parsing delay index, and fuzzy intent re-determination capability index. It is used to evaluate the embodied intelligent robot's ability to perceive, parse, and clarify input signals after the user's intent is issued. Specifically, the multi-model intent consistency index (multi-model intent consistency index value) is used to measure whether multiple recognition models give consistent judgments for the same user input; and whether the robot's intent is consistent with the judgments of most models.
[0115] The evaluation system first obtains the predictive intent of a set of models, namely, the prediction results of multiple models M1(I), M2(I), ..., M n (I), and then calculate the average similarity (semantic vector similarity or label consistency rate) between these model results and the robot recognition result R(I), to form a consistency index, as shown in formula (1): C=avg(sim(M i (I), R(I)))(1)wherein, the higher the consistency index C value, the clearer the user’s expression and the more accurate the robot’s understanding.
[0116] The intent parsing delay metric (intent parsing delay metric value) is used to quantify the speed at which the robot understands user input. The evaluation system records the user input completion time (user intent input time) t. inThe time t for the robot to complete semantic parsing out Then the delay is defined as: L = t out -t in At the same time, latency is mapped to an efficiency score of 0 to 1 for unified measurement, for example, The longer the delay, the lower the score (i.e., the rating result).
[0117] The Fuzzy Intent Re-determination Capability Index (Fuzzy Intent Re-determination Capability Index Value) is used to evaluate the robot's processing ability when the input is unclear, ambiguous, or lacks information. The evaluation system first counts whether the robot successfully completes the "clarification → confirmation → execution" link in a fuzzy input scenario. The capability can be represented by the success rate of clarification, as shown in formula (2): A = (2) Among them, the higher the clarification ratio (the index value of the ability to reconfirm fuzzy intentions) A value, the stronger the robustness and recovery ability of the robot under non-standard input conditions.
[0118] Finally, the three sub-indicators are weighted and combined to form the total score for this stage, which is the index value of the robot's perception and understanding quality index, as shown in formula (3): PUQI=w1C+w2 + w3A (3) Wherein, PUQI is the robot perception and understanding quality index, and w1, w2, w3 are weights configured according to the type of interaction task (e.g., consistency weight can be increased for voice-dominated tasks, and clarification ability weight can be increased for safety tasks).
[0119] In the technical solution of this disclosure embodiment, firstly, based on interaction event data, the robot intent recognition result is determined, and the consistency of the intent recognition result and the prediction results of multiple models is compared to determine the multi-model intent consistency index value. Then, based on the interaction event data, the time difference between the user intent input time and the robot intent parsing completion time is determined, and the intent parsing delay index value is determined based on the time difference. Next, based on the interaction event data, the number of times the robot triggers clarification and the number of times it successfully clarifies the user's ambiguous intent is determined, and the ambiguous intent re-determination capability index value is determined based on the ratio between the number of successful clarifications and the number of triggers. Finally, the multi-model intent consistency index value, the intent parsing delay index value, and the ambiguous intent re-determination capability index value are weighted to determine the index value of the robot's perception and understanding quality index. A multimodal, multi-model parallel recognition strategy is adopted to compare the system's internal recognition results with the robot's actual parsing results to construct a robot perception and understanding quality index system to evaluate the embodied intelligent robot's ability to perceive, analyze, and clarify input signals after the user's intent is issued.
[0120] In some embodiments of this disclosure, for the Behavior Planning and Design Rationality Index (SBPRI), multimodal signals include electroencephalogram (EEG) signals and eye-tracking signals. The index value corresponding to the interaction evaluation index is determined based on interaction event data and multimodal signals. For example, firstly, robot behavior is determined based on interaction event data, user expected behavior is determined based on EEG signals, and user neural mismatch response index value is determined based on the behavioral deviation between robot behavior and user expected behavior; then, expected visual data is determined based on interaction event data, user visual data is determined based on eye-tracking signals, and visual attention shift index value is determined based on the visual deviation between expected visual data and user visual data; finally, the index value of the Behavior Planning and Design Rationality Index (SBPRI) is determined based on the user neural mismatch response index value and the visual attention shift index value.
[0121] Specifically, in the phased comprehensive indicator system construction module, when the target interaction stage includes the robot execution and human-machine collaboration stage, at least one of the behavior planning and design rationality indicators and user physiological load and collaboration comfort indicators is determined as the test indicator. Among them, the behavior planning and design rationality indicator is used to quantify whether the robot's behavior planning (robot behavior, including trajectory, triggering timing, feedback logic, action speed, etc.) in the execution stage is consistent with the user's expectations (user expected behavior). An automated scoring model is established through objective signals such as neural mismatch response (EEG) and attention compensation behavior (eye movement) to realize the quantitative representation of the naturalness and interpretability of system behavior, thereby calculating and determining the indicator value of the behavior planning and design rationality indicator. The specific contents are as follows: (1) Neural Mismatch Response When the robot behavior deviates from the user's psychological expectations (e.g., abrupt action, unnatural path, mismatched rhythm or delayed feedback), the user's brain will generate a series of automated neural responses, including: MMN (Mismatch Negativity, a type of EEG signal), which represents the early automatic detection of abnormal events and usually appears in a time window of about 100 to 250 ms after the stimulus is presented.
[0122] ERN (Error-Related Negativity, an electroencephalogram signal) characterizes the error correction monitoring of behavioral errors or sudden deviations, and usually appears within a time window of about 0 to 150 ms after the error response occurs.
[0123] An enhancement of the P3 component (P300 wave, a type of EEG signal) indicates the need for additional interpretation of system behavior and reconstruction of the predictive model. It typically occurs within a time window of approximately 250–500 ms after feedback or result presentation.
[0124] When a peak value of the user's EEG signal is successfully detected in a trial, and the latency of the peak value falls within the preset MMN, ERN, or P3 time window range, it is recorded as a "neural mismatch event" (i.e., determining the user's neural mismatch response index value) for subsequent statistics and evaluation. The specific boundaries of the above time window can be fine-tuned based on this in different task scenarios to adapt to different stimulus types and interaction rhythms.
[0125] (2) Visual Attention Shift Indicators The system generates the Expected Attention Region and Expected Scanpath based on the robot’s current action and feedback position, i.e., expected visual data, and aligns and compares the user’s real eye movement data (user visual data) with it. It adopts a “two-condition” offset criterion. When both the first fixation deviation and the trajectory complexity deviation occur at the same time, it is judged as a visual mismatch event (i.e., the visual attention shift index value is determined). The specific contents are as follows: First Fixation Deviation, i.e., the user’s first fixation point does not fall into the expected attention region, indicating that he cannot immediately predict the system’s action or needs to confirm the target position.
[0126] Scanpath Complexity Excess (SCL) refers to the fact that the actual visual trajectory complexity of a user is more than 30% higher than the expected visual route, indicating that the user is engaging in additional search or monitoring behavior when tracking system actions.
[0127] Therefore, after obtaining the index value of the behavior planning and design rationality index, based on the index, we evaluate whether the robot system's action logic is natural, predictable, and in line with human expectations, starting from the user's neural response and attention dynamics.
[0128] In the technical solution of this disclosure, robot behavior is first determined based on interactive event data, user expected behavior is determined based on EEG signals, and user neural mismatch response index value is determined based on the behavioral deviation between robot behavior and user expected behavior. Then, expected visual data is determined based on interactive event data, user visual data is determined based on eye-tracking signals, and visual attention shift index value is determined based on the visual deviation between expected visual data and user visual data. Finally, the index value of the rationality index of behavior planning and design is determined based on user neural mismatch response index value and visual attention shift index value. In addition to traditional behavior evaluation methods such as success rate, trajectory deviation, and reaction time, a dual-channel fusion evaluation mechanism of "robot behavior planning - user physiological response" is innovatively introduced. It utilizes multimodal physiological signals such as user neural mismatch response, attention shift behavior, and muscle activation to conduct a phased quantitative evaluation of the rationality of system action planning, behavior naturalness, and human-machine collaboration fluency.
[0129] In some embodiments of this disclosure, for the User Physical Load and Collaborative Comfort Index (UPLC), multimodal signals include electroencephalogram (EEG) signals, heart rate variability (HRV) signals, electrodermal transfer (EDT) signals, electromyography (EMG) signals, speech information, and facial expression information. The index values corresponding to the interaction evaluation index are determined based on the multimodal signals. For example, firstly, user load status is identified based on EEG signals, HRV signals, and EDT signals to determine the load status index value; then, user emotion is identified based on EEG signals, HRV signals, EDT signals, speech information, and facial expression information to determine the emotion index value; next, the comfort level of user collaboration with the robot is determined based on EMG signals to determine the collaborative comfort index value; finally, the index value of the User Physical Load and Collaborative Comfort Index (UPLC) is determined based on the load status index value, emotion index value, and collaborative comfort index value.
[0130] Specifically, during the target interaction phase, including the robot execution and human-machine collaboration phases, a phased comprehensive indicator system construction module determines at least one of the following as test indicators: the rationality indicator of behavior planning and design, and the user's physiological load and collaboration comfort indicator. The user's physiological load and collaboration comfort indicator is based on previously extracted features such as electroencephalogram (EEG), heart rate variability (HRV), electrical conductance analysis (EDA), electromyography (EMG), speech information, and facial expression information. A multi-model fusion structure is used to classify and determine the user's physiological stress, emotional state, and physical comfort during the collaboration process. Using feature vectors as input, three recognition models targeting different dimensions are constructed: a load state recognition model, an emotional state recognition model, and a collaboration comfort model. These models uniformly output three levels—high, medium, and low—to determine the indicator values for the user's physiological load and collaboration comfort.
[0131] (1) Input to the load state identification model: load-related feature combination of EEG+HRV+EDA (i.e., EEG signal, heart rate variability signal and skin conductance signal features).
[0132] Objective: To determine the overall physiological workload level of users during the collaboration phase.
[0133] Output levels: High load, requires continuous monitoring or extra attention; Moderate load, normal operation but slightly higher load; Low load, overall relaxed and natural state.
[0134] Among them, the load status index value is used to reflect whether the collaborative task causes excessive cognitive stress or continuous tension to the user.
[0135] (2) Input to the emotion recognition model: EEG+HRV+EDA+related dynamic features of speech and facial expressions (i.e., EEG signals, heart rate variability signals, skin conductance signals, speech information and facial expression information features).
[0136] Objective: To identify users' emotional states during the collaboration process.
[0137] Output levels: Positive emotions: stable, relaxed, confident; Neutral emotions: no significant emotional fluctuations; Negative emotions: tense, frustrated, irritable, impatient.
[0138] Among them, the emotion index value is used to identify whether the robot system's behavior triggers a positive experience for the user or causes emotional tension and dissatisfaction, serving as an important basis for the adaptability of intelligent collaboration.
[0139] (3) Input to the collaborative comfort model: electromyography (EMG) features (i.e., electromyographic signals) are used to assess the user’s force patterns and postural stability.
[0140] Objective: To determine the user's physical comfort during collaborative actions.
[0141] Output levels: High Comfort: Natural exertion, smooth movements, and no tension in posture; Moderate Comfort: Some increase in exertion or slight compensation; Low Comfort: Significant increase in exertion, tense posture, and compensatory behavior.
[0142] Among them, the Collaborative Comfort Index is used to characterize the friendliness of embodied interaction at the physical action level.
[0143] Finally, based on the load status index, emotion index, and collaborative comfort index, the index values of the user's physiological load and collaborative comfort index are determined.
[0144] In the technical solution of this disclosure, firstly, user load status is identified based on EEG signals, heart rate variability signals, and skin conductance signals to determine load status index values. Then, user emotion is identified based on EEG signals, heart rate variability signals, skin conductance signals, voice information, and facial expression information to determine emotion index values. Next, based on electromyography signals, the comfort level of user collaboration with the robot is determined to determine collaboration comfort index values. Finally, based on load status index values, emotion index values, and collaboration comfort index values, the index values of user physiological load and collaboration comfort index are determined. Multimodal electrophysiological signals and behavioral performance data are fused to construct comprehensive physiological indicators such as cognitive load, emotional stress, attentional confusion, and motor tension, and these are jointly modeled with task performance indicators. Through a multi-model fusion structure, the physiological stress, emotional state, and physical comfort of the user during the collaboration process are graded and judged to form a quantitative evaluation system covering multi-dimensional capabilities of intelligent interaction, thereby achieving a comprehensive assessment of interaction quality.
[0145] Therefore, this disclosure integrates neural mismatch response, heart rate variability, skin conductance trend characteristics, electromyographic exertion patterns, and visual attention behavior (gaze path and trajectory complexity) into a unified physiological-behavioral comprehensive index system, which can objectively assess the user's cognitive load, emotional stress, attentional confusion, and motor tension during the interaction process, and realize the leap from single behavior evaluation to a multi-dimensional quantitative evaluation paradigm of "behavior + physiology + neurology".
[0146] Meanwhile, a dual-indicator system of rationality of system behavior planning and user physiological load and collaborative comfort is introduced. The robot's motion planning is evaluated from the perspective of multimodal signals such as user neural mismatch response, visual deviation behavior (eye movement / trajectory complexity), electromyographic load, heart rate variability and skin conductance trend. At the same time, the user's perceived physical stress and collaborative comfort are quantified. For the first time, an interpretable evaluation framework for "evaluating robot design quality with user physiological response" is realized.
[0147] In some embodiments of this disclosure, for the Policy Quality Index for Clarification of Intent (CSQI), the index value corresponding to the interaction evaluation index is determined based on interaction event data. For example, firstly, based on the interaction event data, the actual number of clarifications and the theoretical number of clarifications performed by the robot to clarify the user's ambiguous intent are determined, and the clarification rationality index value is determined based on the ratio between the actual number of clarifications and the theoretical number of clarifications; then, based on the interaction event data, the efficiency of the robot in clarifying the user's ambiguous intent is determined, and the clarification efficiency index value is determined; next, based on the interaction event data, the number of triggered clarifications and the number of successful clarifications performed by the robot to clarify the user's ambiguous intent are determined, and the clarification success index value is determined based on the ratio between the number of successful clarifications and the number of triggered clarifications; finally, based on the clarification rationality index value, the clarification efficiency index value, and the clarification success index value, the index value of the Policy Quality Index for Clarification of Intent (CSQI) is determined.
[0148] Specifically, by constructing a phased comprehensive indicator system, when the target interaction phase includes at least one of the user intent issuance phase, user feedback phase, and robot strategy adjustment phase, at least one of the intent clarification strategy quality indicator and feedback-driven self-learning performance indicator is determined as the test indicator. The intent clarification strategy quality indicator is used to evaluate whether the robot system can adopt a reasonable, effective, and non-disruptive clarification strategy when the user intent is unclear, the confidence level is insufficient, or the multimodal parsing results are conflicting.
[0149] The evaluation system first automatically identifies “low-confidence intent segments” in the interaction log, namely, situations where the speech parsing entropy increases, the candidate intent probabilities are close, the intermodal recognition is inconsistent, or the robot internally marks them as “uncertain / unclear”. In these “fuzzy intent scenarios”, the robot system records whether clarification is initiated, the number of clarification rounds, the clarification time, and whether the user intent is ultimately successfully locked. Based on this information, the following three key sub-dimensions can be constructed: (1) Clarification trigger rationality: The system statistically analyzes the proportion of actual clarifications triggered in low-confidence scenarios that should theoretically be clarified, that is, determining the actual number of clarifications and the theoretical number of clarifications made by the robot for the user’s fuzzy intent, and determining the clarification rationality index value. This dimension simultaneously punishes insufficient triggering (guessing without asking) and excessive triggering (frequent questioning), with the goal of ensuring that the system “only asks when it is truly fuzzy, and does not disturb when it is clear.”
[0150] (2) The clarification efficiency system determines the clarification efficiency index value by statistically analyzing the number of rounds and time required from initiating the first clarification to finally identifying the true intention, in order to measure whether the clarification process is efficient. The fewer the number of clarification rounds and the shorter the time, the more targeted the robot's clarification questions are, and the more ineffective repetitions are avoided.
[0151] (3) Clarification Success Rate: In all scenarios where clarification has been initiated, the proportion of cases in which the user's true intent was ultimately restored is calculated. That is, based on the ratio between the number of successful clarifications and the number of clarifications triggered, a clarification success rate is determined to measure the effectiveness of the robot's clarification strategy. The higher the success rate, the more reasonable the robot's "what to ask" and "how to ask" are under uncertain conditions.
[0152] The quality index of intent clarification strategy is composed of three dimensions: the clarification rationality index, the clarification efficiency index, and the clarification success index. The higher the index value of the intent clarification strategy quality index, the more intelligent and robust the robot is in handling uncertain intents. Conversely, it indicates that there are problems such as misjudgment, excessive questioning, or clarification failure.
[0153] Therefore, based on the clarification strategy quality index and feedback-driven self-learning performance index, the system can automatically identify low-confidence input scenarios, analyze the rationality of clarification triggers, the number of clarification rounds and time efficiency, and the clarification success rate. After the user provides error correction, the system continuously tracks the rate of error reduction, the change in clarification frequency, and the ability to retain error correction memory in similar scenarios, thereby achieving a quantitative evaluation of whether the system becomes smarter with use, and providing a reliable evaluation basis for the long-term adaptability of the agent and the optimization of interaction strategies.
[0154] In the technical solution of this disclosure embodiment, firstly, based on interaction event data, the actual number of clarifications and the theoretical number of clarifications performed by the robot to clarify the user's ambiguous intent are determined. Then, based on the ratio between the actual number of clarifications and the theoretical number of clarifications, a clarification rationality index value is determined. Next, based on the interaction event data, the efficiency of the robot in clarifying the user's ambiguous intent is determined, and a clarification efficiency index value is determined. Then, based on the interaction event data, the number of triggered clarifications and the number of successful clarifications performed by the robot to clarify the user's ambiguous intent are determined. Then, based on the ratio between the number of successful clarifications and the number of triggered clarifications, a clarification success index value is determined. Finally, based on the clarification rationality index value, the clarification efficiency index value, and the clarification success index value, an index value for the quality index of intent clarification strategy is determined. Based on the three dimensions of clarification rationality, clarification efficiency, and clarification success rate, the system evaluates whether the robot can adopt a reasonable, effective, and non-intrusive clarification strategy when the user's intent is unclear, the confidence level is insufficient, or there are conflicts in the multimodal parsing results. A comprehensive evaluation of interaction quality is achieved based on a multi-dimensional index system.
[0155] In some embodiments of this disclosure, the interaction event data includes interaction event data before the robot update and interaction event data after the robot update. For the Feedback-Driven Self-Learning Performance Index (FLPI), the index value corresponding to the interaction evaluation index is determined based on the interaction event data. For example, firstly, based on the interaction event data before and after the robot update, the index values for the improvement of intent recognition error rate, the change in intent clarification frequency and efficiency, and the error correction memory retention capability are determined. Then, based on the index values for the improvement of intent recognition error rate, the change in intent clarification frequency and efficiency, and the error correction memory retention capability, the index value of the Feedback-Driven Self-Learning Performance Index (FLPI) is determined.
[0156] Specifically, by constructing a phased comprehensive indicator system, when the target interaction phase includes at least one of the user intent issuance phase, user feedback phase, and robot strategy adjustment phase, at least one of the intent clarification strategy quality indicator and feedback-driven self-learning performance indicator is determined as the test indicator. The feedback-driven self-learning performance indicator is used to evaluate whether the robot can produce a stable and quantifiable performance improvement in subsequent interactions after the user provides error correction, repetition, or explicit feedback.
[0157] The evaluation system first automatically identifies "user error correction events" and "system policy update events" from the interaction logs, and divides them into "pre-learning window" and "post-learning window" accordingly. Both select fragments with similar intent types and similar interaction patterns for comparative analysis. The feedback-driven self-learning performance index consists of the following three core performance dimensions: (1) Error rate improvement magnitude (error rate improvement index value). This index value is used to compare the accuracy of the robot's intent recognition for similar inputs with the interaction event data before the robot update and the interaction event data after the robot update. If the robot's error rate decreases significantly after learning, it indicates that the parsing boundary has been successfully corrected using user feedback; if the change is not obvious or worsens in the opposite direction, it indicates that the robot has not effectively absorbed the feedback information.
[0158] (2) Changes in clarification frequency and efficiency (Indicator value of change in intention clarification frequency and efficiency) This indicator value is used to measure whether the robot’s understanding ability in the first round after learning is enhanced, including the clarification trigger frequency, the average number of rounds and the average time spent in each clarification. If the number of clarifications and the number of rounds decrease, it indicates that the robot can understand the user’s intention faster under the same expression.
[0159] (3) Error correction memory retention ability (error correction memory retention ability index value) This index value is used to evaluate whether the robot can retain the learned correction results in the future. When the user uses the same or highly similar expression again, will the robot still repeat the previous error? If the repeated error is significantly reduced, it indicates that it has stable error correction memory and long-term retention ability.
[0160] The evaluation system comprehensively scores the above three dimensions and finally outputs the index value of the feedback-driven self-learning performance index as three levels: "excellent", "medium" and "poor", which are used to reflect the robot's self-learning performance in interactive iteration.
[0161] In the technical solution of this disclosure embodiment, firstly, based on the interaction event data before and after the robot update, the indicator values for the improvement of intent recognition error rate, the change in intent clarification frequency and efficiency, and the indicator value for error correction memory retention capability are determined. Then, based on the indicator values for the improvement of intent recognition error rate, the change in intent clarification frequency and efficiency, and the indicator value for error correction memory retention capability, the indicator values for the feedback-driven self-learning performance are determined. This allows for the evaluation of whether the robot can produce a stable and quantifiable performance improvement in subsequent interactions after the user provides error correction, repetition, or explicit feedback. A comprehensive evaluation of interaction quality is achieved based on a multi-dimensional indicator system.
[0162] Finally, the evaluation system uses the intelligent interaction capability assessment and design comparison module to score core interaction capabilities such as intent understanding ability, human-computer collaboration fluency, feedback predictability, and error recovery ability based on indicators at each stage. The automated evaluation report generation and visualization module summarizes the stage indicators, capability scores, and version differences to generate a visualized interaction evaluation report, which will be explained in detail below.
[0163] The Embodied Intelligence Interaction Design Evaluation Report consists of three core dimensions, covering the complete interaction chain of "understanding → execution → learning". It aims to provide a quantitative, comprehensive, and interpretable evaluation of the interaction quality of embodied intelligence from both the system and user perspectives.
[0164] Phase 1: Robot Perception and Understanding Quality (PUQI) This dimension evaluates the system's ability to perceive, interpret, and clarify user intentions. It focuses on three key capabilities: multi-model intent consistency (how accurately the system understands the user); intent interpretation speed (how promptly the system responds); and the ability to redefine ambiguous intents (whether the system can successfully recover the true intent in ambiguous situations). This dimension reflects the robot's ability to "understand the user," laying the foundation for the accuracy of subsequent interactions.
[0165] Phase Two: Robotic Execution and Human-Robot Collaboration Performance (SBPRI / UPLC) This dimension assesses the rationality of the robot's motion planning and the physiological and emotional responses it elicits in the user during collaboration. Its core includes two aspects: System Behavior Planning Rationality (SBPRI): Analyzing whether the robot's movements are natural, predictable, and rhythmic, using EEG neural mismatch responses and eye-tracking attention shifts, to ensure they meet user expectations.
[0166] User physiological workload and collaborative comfort (UPLC): Using signals such as EEG, EDA, HRV, and EMG to identify users' physiological workload, emotional state, and physical comfort.
[0167] This dimension reflects whether the robot "acts naturally and whether the human is comfortable," and is a key indicator of the quality of the robot's behavior layer.
[0168] Phase 3: Uncertainty Intent Handling and Interaction Self-Learning Capability (CSQI / FLPI) This dimension focuses on the quality of the robot's strategies in "uncertain scenarios" and whether it can continuously improve after user corrections. It includes two types of evaluation: Clarification Strategy Quality (CSQI): assessing whether the robot can provide reasonable, effective, and non-disruptive clarification when input is unclear.
[0169] Feedback-Driven Self-Learning (FLPI): This examines whether a robot can reduce its error rate, decrease the number of clarifications, and retain long-term memory after receiving user corrections.
[0170] This dimension reflects whether a robot "becomes smarter with use," and is an important manifestation of its embodied intelligence and adaptive capabilities.
[0171] The above three dimensions together constitute the overall evaluation framework of this report, realizing full-link evaluation from intent understanding, action execution to strategy learning, and providing quantifiable, comparable and traceable quality basis for embodied intelligent interaction design.
[0172] Therefore, based on an automated multimodal signal processing pipeline, feature extraction, comprehensive index calculation, stage scoring, interaction quality assessment, and visualization report generation can be automatically completed, forming an automated evaluation closed loop that runs through data acquisition, analysis, and evaluation. This supports rapid iteration and version comparison of intelligent robot interaction systems. Through the above innovative approach, this disclosure provides a structured, multimodal, and automated interaction evaluation system to meet the interaction design quality assessment needs of embodied intelligent robots. It can accurately quantify key performance aspects of intelligent systems, such as understanding ability, reasoning and decision-making, human-machine collaboration, and feedback mechanisms. Furthermore, it can be used for interaction strategy optimization, A / B version comparison, and continuous improvement of agent interaction performance, demonstrating broad engineering application prospects.
[0173] This disclosure constructs an automated multimodal signal processing pipeline that can automatically complete multi-source physiological signal preprocessing, staged feature extraction, index modeling, interactive capability scoring, and visualization report generation. It forms a fully automated closed loop from data acquisition and analysis to evaluation, significantly reducing manual processing costs and greatly improving the efficiency and repeatability of system interactive evaluation. It provides engineering-level support for the rapid iteration of intelligent robot interaction strategies and A / B version comparison.
[0174] Figure 2 is a block diagram of an intelligent robot evaluation device based on multimodal signals provided in an embodiment of this disclosure.
[0175] As shown in Figure 2, the intelligent robot evaluation device 200 based on multimodal signals provided in this embodiment of the present disclosure includes: an acquisition module 210, used to acquire interaction event data during the interaction between the user and the robot and multimodal signals collected from the user.
[0176] The first determining module 220 is used to determine a target interaction stage from a preset multi-stage interaction based on interaction event data. The target interaction stage includes at least one of the following: user intent issuance stage, robot perception and intent understanding stage, robot execution and human-machine collaboration stage, user feedback stage, and robot strategy adjustment stage.
[0177] The second determining module 230 is used to determine the interaction evaluation indicators corresponding to the target interaction stage.
[0178] The third determining module 240 is used to determine the index value corresponding to the interaction evaluation index based on the interaction event data and / or multimodal signals.
[0179] The fourth determination module 250 is used to determine the evaluation results of human-computer interaction based on the indicator values.
[0180] In some embodiments of this disclosure, the second determining module 230 is further configured to: for at least one of the user intent issuance stage and the robot perception and intent understanding stage, determine corresponding interaction evaluation indicators, including robot perception and understanding quality indicators; for the robot execution and human-machine collaboration stage, determine corresponding interaction evaluation indicators, including at least one of behavior planning and design rationality indicators and user physiological load and collaboration comfort indicators; for at least one of the user intent issuance stage, user feedback stage, and robot strategy adjustment stage, determine corresponding interaction evaluation indicators, including at least one of intent clarification strategy quality indicators and feedback-driven self-learning performance indicators.
[0181] In some embodiments of this disclosure, for robot perception and understanding quality indicators, determining the corresponding indicator values for interaction evaluation indicators based on interaction event data includes: determining the robot intent recognition result based on interaction event data, and comparing the consistency of the intent recognition result with the prediction results of multiple models to determine a multi-model intent consistency indicator value; determining the time difference between the user intent input time and the robot intent parsing completion time based on interaction event data, and determining an intent parsing delay indicator value based on the time difference; determining the number of times the robot triggers clarification and the number of times it successfully clarifies the user's ambiguous intent based on interaction event data, and determining an ambiguous intent re-determination capability indicator value based on the ratio between the number of successful clarifications and the number of triggers for clarification; and weighting the multi-model intent consistency indicator value, the intent parsing delay indicator value, and the ambiguous intent re-determination capability indicator value to determine the indicator value of the robot perception and understanding quality indicators.
[0182] In some embodiments of this disclosure, for the behavior planning and design rationality index, the multimodal signals include electroencephalogram (EEG) signals and eye-tracking signals. Determining the index value corresponding to the interaction evaluation index based on interaction event data and multimodal signals includes: determining robot behavior based on interaction event data, determining the user's expected behavior based on EEG signals, and determining the user's neural mismatch response index value based on the behavioral deviation between the robot behavior and the user's expected behavior; determining expected visual data based on interaction event data, determining the user's visual data based on eye-tracking signals, and determining the visual attention shift index value based on the visual deviation between the expected visual data and the user's visual data; and determining the index value of the behavior planning and design rationality index based on the user's neural mismatch response index value and the visual attention shift index value.
[0183] In some embodiments of this disclosure, for user physiological load and collaborative comfort indicators, multimodal signals include electroencephalogram (EEG) signals, heart rate variability (HRV) signals, electrodermal transfer (EDT) signals, electromyography (EMG) signals, speech information, and facial expression information. Determining the corresponding indicator values for the interaction evaluation indicators based on the multimodal signals includes: identifying user load status based on EEG signals, HRV signals, and EDT signals to determine load status indicator values; identifying user emotions based on EEG signals, HRV signals, EDT signals, speech information, and facial expression information to determine emotion indicator values; determining the comfort level of user-robot collaboration based on EMG signals to determine collaborative comfort indicator values; and determining the indicator values for user physiological load and collaborative comfort based on load status indicator values, emotion indicator values, and collaborative comfort indicator values.
[0184] In some embodiments of this disclosure, for the quality index of intent clarification strategy, the index value corresponding to the interaction evaluation index is determined based on interaction event data, including: determining the actual number of times the robot clarifies the user's ambiguous intent and the theoretical number of times the robot clarifies the user's ambiguous intent based on the interaction event data, and determining the clarification rationality index value based on the ratio between the actual number of clarifications and the theoretical number of clarifications; determining the efficiency of the robot in clarifying the user's ambiguous intent based on the interaction event data, and determining the clarification efficiency index value; determining the number of times the robot triggers clarification and the number of times the robot successfully clarifies the user's ambiguous intent based on the interaction event data, and determining the clarification success index value based on the ratio between the number of successful clarifications and the number of times the robot triggers clarification; and determining the index value of the intent clarification strategy quality index based on the clarification rationality index value, the clarification efficiency index value, and the clarification success index value.
[0185] In some embodiments of this disclosure, the interaction event data includes interaction event data before the robot update and interaction event data after the robot update. For feedback-driven self-learning performance indicators, the indicator values corresponding to the interaction evaluation indicators are determined based on the interaction event data, including: determining the indicator values for the improvement in intent recognition error rate, the change in intent clarification frequency and efficiency, and the indicator value for error correction memory retention capability before and after the robot update based on the interaction event data before and after the robot update; and determining the indicator values for the feedback-driven self-learning performance indicators based on the indicator values for the improvement in intent recognition error rate, the change in intent clarification frequency and efficiency, and the error correction memory retention capability.
[0186] In some embodiments of this disclosure, acquiring interaction event data includes: acquiring robot motion event data and robot voice data; acquiring user motion event data and user voice data; and acquiring process event data of the robot performing interactive tasks.
[0187] In some embodiments of this disclosure, after determining the target interaction stage based on interaction event data, the device 200 further includes an alignment module for: aligning the multimodal signals to the target interaction stage.
[0188] Figure 3 is a block diagram of an edge computing device provided in an embodiment of this disclosure.
[0189] As shown in FIG3, this embodiment of the present disclosure provides an edge computing device 300, including a memory 302 and a processor 301. The memory 302 stores a computer program, and the processor 301 executes the computer program to implement the steps of the method of any of the above embodiments.
[0190] This disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in any of the above embodiments.
[0191] The embodiments of this disclosure have now been described in detail. To avoid obscuring the concept of this disclosure, some details known in the art have not been described. Those skilled in the art can fully understand how to implement the technical solutions disclosed herein based on the above description.
[0192] While specific embodiments of this disclosure have been described in detail by way of examples, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments or equivalent substitutions can be made to some technical features without departing from the scope and spirit of this disclosure. In particular, as long as there is no structural conflict, the technical features mentioned in the various embodiments can be combined in any manner.
Claims
1. A method for evaluating embodied intelligent human-computer interaction based on multimodal data, characterized in that, include: Acquire interaction event data and multimodal signals collected from the user during the interaction process between the user and the robot; Based on the interaction event data, a target interaction stage is determined from the preset multi-stage interaction, wherein the target interaction stage includes at least one of the following: user intent issuance stage, robot perception and intent understanding stage, robot execution and human-machine collaboration stage, user feedback stage, and robot strategy adjustment stage; an interaction evaluation index corresponding to the target interaction stage is determined; based on the interaction event data and / or the multimodal signal, the index value corresponding to the interaction evaluation index is determined; and the evaluation result of human-machine interaction is determined based on the index value.
2. The method according to claim 1, characterized in that, The determination of the interaction evaluation indicators corresponding to the target interaction stage includes: for at least one of the user intent issuance stage and the robot perception and intent understanding stage, determining the corresponding interaction evaluation indicators to include a robot perception and understanding quality indicator; for the robot execution and human-machine collaboration stage, determining the corresponding interaction evaluation indicators to include at least one of a behavior planning and design rationality indicator and a user physiological load and collaboration comfort indicator; for at least one of the user intent issuance stage, the user feedback stage, and the robot strategy adjustment stage, determining the corresponding interaction evaluation indicators to include at least one of an intent clarification strategy quality indicator and a feedback-driven self-learning performance indicator.
3. The method according to claim 2, characterized in that, For the robot perception and understanding quality index, the index value corresponding to the interaction evaluation index is determined based on the interaction event data, including: determining the robot intent recognition result based on the interaction event data, and comparing the intent recognition result with the prediction results of multiple models to determine the multi-model intent consistency index value; determining the time difference between the user intent input time and the robot intent parsing completion time based on the interaction event data, and determining the intent parsing delay index value based on the time difference; determining the number of times the robot triggers clarification and the number of times it successfully clarifies the user's ambiguous intent based on the interaction event data, and determining the ambiguous intent re-determination capability index value based on the ratio between the number of successful clarifications and the number of triggered clarifications; and weighting the multi-model intent consistency index value, the intent parsing delay index value, and the ambiguous intent re-determination capability index value to determine the index value of the robot perception and understanding quality index.
4. The method according to claim 2, characterized in that, For the behavior planning and design rationality index, the multimodal signals include electroencephalogram (EEG) signals and eye-tracking signals. Determining the index value corresponding to the interaction evaluation index based on the interaction event data and the multimodal signals includes: determining robot behavior based on the interaction event data, determining the user's expected behavior based on the EEG signals, and determining the user's neural mismatch response index value based on the behavioral deviation between the robot behavior and the user's expected behavior; determining expected visual data based on the interaction event data, determining the user's visual data based on the eye-tracking signals, and determining the visual attention shift index value based on the visual deviation between the expected visual data and the user's visual data; and determining the index value of the behavior planning and design rationality index based on the user's neural mismatch response index value and the visual attention shift index value.
5. The method according to claim 2, characterized in that, For the user physiological load and collaborative comfort index, the multimodal signals include electroencephalogram (EEG) signals, heart rate variability (HRV) signals, electrodermal transfer (EDT) signals, electromyography (EMG) signals, speech information, and facial expression information. Determining the corresponding index value for the interaction evaluation index based on the multimodal signals includes: identifying the user's load state based on the EEG signals, HRV signals, and EDT signals to determine the load state index value; identifying the user's emotion based on the EEG signals, HRV signals, EDT signals, speech information, and facial expression information to determine the emotion index value; determining the user's comfort level in collaborating with the robot based on the EMG signals to determine the collaborative comfort index value; and determining the index value for the user physiological load and collaborative comfort index based on the load state index value, the emotion index value, and the collaborative comfort index value.
6. The method according to claim 2, characterized in that, For the quality index of the intent clarification strategy, the index value corresponding to the interaction evaluation index is determined based on the interaction event data, including: determining the actual number of times the robot clarified the user's ambiguous intent and the theoretical number of times the robot clarified the user's ambiguous intent based on the interaction event data, and determining the clarification rationality index value based on the ratio between the actual number of clarifications and the theoretical number of clarifications; determining the efficiency of the robot in clarifying the user's ambiguous intent based on the interaction event data, and determining the clarification efficiency index value; determining the number of times the robot triggered clarification and the number of times the robot successfully clarified the user's ambiguous intent based on the interaction event data, and determining the clarification success index value based on the ratio between the number of successful clarifications and the number of times the robot triggered clarification; and determining the index value of the quality index of the intent clarification strategy based on the clarification rationality index value, the clarification efficiency index value, and the clarification success index value.
7. The method according to claim 2, characterized in that, The interaction event data includes interaction event data before and after the robot update. For the feedback-driven self-learning performance index, the index values corresponding to the interaction evaluation index are determined based on the interaction event data, including: determining the intention recognition error rate improvement index value, the intention clarification frequency and efficiency change index value, and the error correction memory retention capability index value before and after the robot update based on the interaction event data before and after the robot update; and determining the index value of the feedback-driven self-learning performance index based on the intention recognition error rate improvement index value, the intention clarification frequency and efficiency change index value, and the error correction memory retention capability index value.
8. The method according to any one of claims 1-7, characterized in that, Acquiring interaction event data includes: acquiring robot motion event data and robot voice data; acquiring user motion event data and user voice data; acquiring process event data of the robot performing interactive tasks; and / or, after determining the target interaction stage based on the interaction event data, the method further includes: aligning the multimodal signals to the target interaction stage.
9. An edge computing device, characterized in that, The method includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method according to any one of claims 1-8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-8.