Method and apparatus for testing human-computer interaction of intelligent device
By collecting and processing multimodal data to generate fused semantic features, and combining them with a virtual environment to construct a high-fidelity interactive scenario, the problem of multimodal data fusion in human-computer interaction testing of intelligent devices is solved, and efficient and reliable multidimensional evaluation is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 启元实验室
- Filing Date
- 2026-01-06
- Publication Date
- 2026-05-12
AI Technical Summary
现有技术在智能设备的人机交互测试中缺乏多模态数据的统一理解与融合能力,无法满足高复杂度任务的需求。
采集原始多模态数据,通过统一处理生成多模态融合语义特征,结合虚拟环境生成高保真交互场景,进行多任务预测并输出关键评估指标的综合量化结果和可解释文本说明。
It enables real-time, multi-dimensional evaluation of interactive behavior, improves the systematic nature and efficiency of testing, enhances the reliability and generalization ability of test results, and reduces physical testing costs.
Smart Images

Figure CN121455345B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-computer interaction technology, for example to a testing method and apparatus for human-computer interaction of a smart device. Background Technology
[0002] Intelligent devices refer to electronic devices that are connected to the internet, possess data computing and processing capabilities, and can sense their environment and interact intelligently with people or other devices. In recent years, the rapid development of artificial intelligence technology in intelligent devices, intelligent algorithm systems, and intelligent command and control systems has placed higher demands on the real-time performance, accuracy, and reliability of human-computer interaction. Especially in tasks such as intelligent algorithm verification, equipment operation evaluation, and command decision analysis, testing technology not only needs to have a high-precision understanding of complex multimodal information but also needs to be able to conduct realistic and effective verification in dynamically changing scenarios. Currently, research and applications of human-computer interaction testing are involved in various industries, including the control and operation testing of intelligent devices, the operation evaluation of industrial robots, and the verification of intelligent traffic management systems. These systems often rely on multiple sensory channels such as voice, vision, gestures, and physiological signals for information interaction.
[0003] In related technologies, existing testing technologies lack the ability to uniformly understand and integrate multimodal data during multimodal data processing, which makes it impossible to fully meet the interactive testing needs of highly complex tasks. Summary of the Invention
[0004] This application aims to provide a testing method and apparatus for human-computer interaction of smart devices.
[0005] According to one aspect of this application, a testing method for human-computer interaction of intelligent devices is proposed, comprising:
[0006] Collect raw multimodal data of the preset test targets and process them uniformly to generate multimodal fusion semantic features;
[0007] Based on the multimodal fusion semantic features, determine the intermediate results of multi-task prediction;
[0008] Based on multimodal fusion semantic features and test objectives, a scene description of a high-fidelity interactive scenario corresponding to the test objectives is generated in a virtual environment.
[0009] Based on the intermediate results of multi-task prediction, scenario description, and operational data during scenario execution, determine and output the comprehensive quantitative results of key evaluation indicators and corresponding interpretable text descriptions.
[0010] According to one aspect of this application, a testing apparatus for human-computer interaction of intelligent devices is provided, comprising:
[0011] The data acquisition and processing module is used to acquire the raw multimodal data of the preset test target and process it in a unified manner to generate multimodal fusion semantic features;
[0012] The intermediate result determination module is used to determine the intermediate results of multi-task prediction based on multimodal fusion semantic features;
[0013] The scene generation module is used to generate a scene description of a high-fidelity interactive scene corresponding to the test target in a virtual environment based on multimodal fusion semantic features and test targets;
[0014] The evaluation module is used to determine and output the comprehensive quantitative results of key evaluation indicators and corresponding interpretable text descriptions based on the intermediate results of multi-task prediction, scenario descriptions, and operational data during scenario operation.
[0015] According to one aspect of this application, an electronic device is provided, comprising: a processor; and a memory storing a computer program that, when executed by the processor, causes the processor to perform the method described above.
[0016] According to one aspect of this application, a non-transitory computer-readable medium is proposed, on which readable instructions are stored, which, when executed by a processor, cause the processor to perform the method described above.
[0017] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this application.
[0018] Beneficial effects:
[0019] The embodiments provided in this application, by collecting and uniformly processing the raw multimodal data during the interaction process, generate fused semantic features, overcoming the limitations of traditional single-modal testing, and more comprehensively and accurately reflecting the complex information of real interaction scenarios. Based on the fused features, multi-task prediction intermediate results are generated simultaneously, and combined with a high-fidelity interaction scenario constructed in a virtual environment, real-time, multi-dimensional evaluation of interaction behavior is achieved, improving the systematic nature and efficiency of testing. By integrating multi-task prediction results, scenario descriptions, and operational data, comprehensive quantitative results and interpretable textual explanations of key evaluation indicators are output, ensuring that test conclusions not only have objective data support but also clear business logic explanations, facilitating R&D personnel in locating problems and optimizing interaction design. Utilizing a virtual environment to generate high-fidelity interaction scenarios allows for flexible simulation of various edge cases and complex working conditions, reducing physical testing costs while ensuring that the test scenario closely resembles the real user environment, enhancing the reliability and generalization ability of test results. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings, without exceeding the scope of protection claimed by this application.
[0021] Figure 1 The system architecture diagram based on the native multimodal unified large model provided in the embodiments of this application is as follows:
[0022] Figure 2 This is a schematic diagram of the internal structure of the native multimodal unified large model provided in the embodiments of this application;
[0023] Figure 3 A flowchart illustrating a testing method for human-computer interaction of an intelligent device provided in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram illustrating the process of acquiring and preprocessing raw multimodal data provided in the embodiments of this application;
[0025] Figure 5 A schematic diagram of the intelligent adaptive testing process provided in the embodiments of this application;
[0026] Figure 6 A flowchart illustrating the implementation details of generating a high-fidelity interactive scenario provided in the embodiments of this application;
[0027] Figure 7 A flowchart illustrating the implementation of multi-dimensional evaluation in the embodiments of this application;
[0028] Figure 8 A schematic diagram of the system deployment and integration provided for embodiments of this application;
[0029] Figure 9 A block diagram of a testing apparatus for human-computer interaction of intelligent devices provided in an embodiment of this application;
[0030] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0031] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this application will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.
[0032] Furthermore, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this application. However, those skilled in the art will recognize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this application.
[0033] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0034] The flowcharts shown in the accompanying drawings are merely illustrative and do not necessarily include all content and operations / steps, nor do they necessarily have to be performed in the described order. For example, some operations / steps can be broken down, while others can be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.
[0035] It should be understood that although the terms first, second, third, etc., may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component discussed below may be referred to as the second component without departing from the teachings of this application. As used herein, the term "and / or" includes all combinations of any one and more of the associated listed items.
[0036] Figure 1 This is a system architecture diagram based on a native multimodal unified large model provided in the embodiments of this application. Figure 2 This is a schematic diagram of the internal structure of a native multimodal unified large model provided for an embodiment of this application. (See attached diagram.) Figure 1As shown, the system in this application runs on a high-performance computing platform, with a preferred configuration including: a GPU cluster supporting native multimodal unified large-scale model inference, a host with at least 64GB of memory, a data acquisition card / synchronization controller supporting multi-channel high-speed input, a high-definition camera / depth camera / directional microphone / pressure touchpad, physiological signal devices (eye tracker, heart rate monitor, EEG acquisition device), and VR (Virtual Reality) or AR (Augmented Reality) headsets or large-screen display devices. The system unifies the timestamps of each sensor channel through a time synchronization module to ensure low-latency alignment of multimodal data (typically ≤20ms). The operating system is a dual-platform Linux / Windows; the native multimodal unified large-scale model is developed based on the Transformer architecture (implemented in PyTorch / TensorFlow); the simulation uses a 3D graphics engine that supports real-time rendering; data management uses a hybrid storage of "time-series database + relational database"; the front end includes an immersive interactive interface and a test management console.
[0037] The system adopts a layered architecture, including: a multimodal data acquisition layer, a native multimodal processing layer, an intelligent simulation engine layer, an intelligent test control layer, a comprehensive evaluation and analysis layer, and a user interface layer. Data exchange and control command transmission between these layers occur through standardized interfaces. The internal structure of the native multimodal unified large model is referenced from [reference needed]. Figure 2 .
[0038] For modality-specific encoder groups, dedicated encoders are constructed for speech, vision, text, gesture, and physiological signals, respectively. The speech encoder extracts temporal spectral features and performs endpoint detection; the vision encoder extracts hierarchical features from frame-level images and supports ROI (Region of Interest) alignment; the text encoder performs word segmentation / sub-word level embedding; the gesture encoder models the trajectory based on skeleton points or hand key points; and the physiological encoder extracts frequency domain / time domain statistical features from eye movements, heart rate, and EEG.
[0039] For the cross-modal alignment module, a cross-modal attention mechanism is adopted to perform semantic alignment of multi-source signals on a unified time axis: the feature sequences of each modality (i.e., multimodal fusion semantic features) are mapped to a shared time index, and the correlation weight between modalities is calculated through the alignment head to eliminate sampling rate differences and hysteresis, thereby improving cross-modal semantic consistency.
[0040] For the unified representation learning layer, a unified Transformer is used as the core to fuse aligned multimodal features end-to-end, resulting in a shared semantic representation. This unified semantic representation (i.e., the shared semantic representation) serves as input to the multi-task learning head, used for tasks such as interactive intent recognition, sentiment state recognition, quality assessment, and scene understanding. These tasks collectively constitute the core functional modules of the system, with interactive intent recognition and sentiment state recognition being the main technological innovations, and quality assessment and scene understanding serving as auxiliary modules to form a complete testing and evaluation loop.
[0041] For the multi-task learning heads, the parameter settings of each task head follow the general principles of the Transformer multi-task learning framework, achieved by adding independent task output layers on top of a shared encoding layer. The intent recognition head uses a classification structure, the sentiment analysis head uses a bidirectional attention network structure, and the interaction quality assessment head combines regression and classification dual-objective functions. This module enables parallel inference across multiple tasks within a unified semantic space, significantly reducing inter-task interference. A joint loss function is used to share underlying representations and improve overall performance. To meet real-time requirements, model pruning and quantization strategies can be combined to achieve low-latency inference. The shared underlying representations are generated by a unified Transformer encoding layer and are shared by all task heads. Overall performance includes three comprehensive metrics: recognition accuracy, processing latency, and model robustness. After latency optimization, intermediate multi-task prediction results, including intent category, sentiment label, task state score, and interaction quality level, can be output within 20ms.
[0042] For specific implementation details, please refer to the following examples.
[0043] Figure 3 A flowchart illustrating a testing method for human-computer interaction of a smart device provided in an embodiment of this application. Figure 3 As shown, the method includes steps S30, S31, S32 and S33.
[0044] In step S30, the original multimodal data corresponding to the preset test target is collected and processed uniformly to generate multimodal fusion semantic features.
[0045] In this application, the test objective is used to characterize the test purpose of the human-computer interaction that needs to be achieved.
[0046] When users interact with VR, AR, or large-screen environments using voice, gestures, touch, and eye tracking, a multimodal data acquisition layer can be utilized. This layer activates cameras, depth sensors, microphones, pressure touchpads, and other physiological sensors at the same level, synchronously acquiring raw multimodal data based on a unified trigger signal or master clock. This layer can include intelligent sensor arrays, environmental sensing devices, and physiological signal acquisition devices. Specifically, the intelligent sensor array collects external interaction data such as vision, voice, and gestures; the environmental sensing devices collect environmental parameters such as scene background, lighting, and temperature; and the physiological signal acquisition devices acquire user physiological state information such as eye movements, heart rate, and electroencephalogram (EEG).
[0047] According to the preset preprocessing methods, such as time synchronization and noise filtering, different types of original multimodal data are uniformly processed to obtain multimodal fusion semantic features.
[0048] In step S31, intermediate results of multi-task prediction are determined based on multimodal fusion semantic features.
[0049] In this application, multimodal fusion semantic features can be input into a pre-set... Figure 1 The intelligent simulation engine layer in the system shown obtains intermediate results from multi-task prediction.
[0050] In some implementations, multimodal fusion semantic features and test targets can be input into a pre-set intelligent scene generator. This intelligent scene generator can determine the task type and interaction model, obtain intermediate results, and use them as intermediate results for multi-task prediction.
[0051] In step S32, a scene description of a high-fidelity interactive scene corresponding to the test target is generated in the virtual environment based on the multimodal fusion semantic features and the test target.
[0052] This application can further match the scene template library and task parameters corresponding to the test target of multimodal fusion semantic features, such as environmental complexity, interaction frequency, and interference level, to establish a high-fidelity interaction scene. The high-fidelity interaction scene of this application includes environment, objects, task targets, and interference factors, and supports parameterized adjustment and custom difficulty adjustment. The components included in this scene can be a specific scene description.
[0053] In some implementations, refer to Figure 1The system utilizes an intelligent simulation engine layer to receive a unified semantic representation—a multimodal fusion semantic feature—generated from the native multimodal processing layer. Based on this semantic information, it generates high-fidelity interactive scenes relevant to the task within a virtual environment. The mixed reality rendering module takes visual, spatial, and environmental parameters as input and outputs immersive images and spatial feedback; the physical interaction simulation module takes task commands and device models as input and outputs environmental responses, collisions, and physical behavior results. Together, they generate a visualized, dynamic testing environment. Mixed reality rendering constructs an immersive, high-fidelity interactive environment; physical interaction simulation simulates realistic effects such as device operation, environmental changes, and physical collisions.
[0054] In step S33, based on the intermediate results of multi-task prediction, the scenario description, and the running data during the scenario operation, the comprehensive quantitative results of key evaluation indicators and the corresponding interpretable text descriptions are determined and output.
[0055] This application can construct high-fidelity interactive scenarios that match the testing objectives. These scenarios can automatically run based on multimodal fusion semantic features, and the data generated during the process can be used as runtime data. This data may include objective achievement rate, speech pauses, speech error correction ratio, and standardized eye-tracking regression counts, etc.
[0056] The calculation methods for key evaluation indicators in the current scenario can be pre-set based on the scenario description. Intermediate results and operational data from multi-task predictions can be substituted into the calculation methods to obtain comprehensive quantitative results. This enables multi-dimensional evaluation of task completion, interaction naturalness, and collaboration efficiency. Furthermore, an automatic improvement suggestion can be generated based on the comprehensive quantitative results using a pre-set language generation module. In some implementations, a multi-dimensional evaluation engine can be used to calculate intent recognition accuracy, emotion recognition rate, task completion, collaboration efficiency, and cognitive load, and generate interpretable text descriptions.
[0057] The data sources for intent recognition accuracy, emotion recognition rate, task completion rate, collaboration efficiency, and cognitive load are as follows: the recognition results and their labels / standard answers are derived from task configuration and manual annotation, and the system performs comparative statistics on a test round basis. The calculation method is as follows:
[0058] (1) Intent / emotion recognition accuracy = (number of correct recognitions / total number of samples), and weighted according to sample weight;
[0059] (2) Completion rate = (Number of targets achieved / Number of targets to be achieved) × Time efficiency coefficient;
[0060] (3) Collaboration efficiency is based on a comprehensive score of communication delay and information consistency;
[0061] (4) Cognitive load is normalized based on heart rate variability and the number of eye movements and regressions;
[0062] (5) Text generation: The report generator reads the above indicators and key events, and calls the language generation module to form a natural language description;
[0063] (6) Text composition: includes indicator values and classifications, main problems and possible causes, comparison with historical baselines, suggestions for further optimization and corresponding responsibility modules.
[0064] The test report outputs a complete set of information, including scenario descriptions, key evaluation metrics, operational data, and improvement suggestions, all presented as interpretable text. Specifically, performance metrics are derived from the calculation results of the evaluation module, scenario descriptions are automatically generated by the scenario generation module by extracting task parameters, interaction logs originate from the data acquisition and logging module, and improvement suggestions are automatically generated by the language generation module based on the evaluation results.
[0065] In some implementations, test reports can be stored for historical comparison and analysis, thereby enabling optimization. Based on the evaluation results, subsequent test scenarios and interaction strategies can be adjusted to achieve closed-loop optimization.
[0066] In the specific implementation process, refer to Figure 1 The system can utilize an intelligent test control layer to receive scene state information and evaluation feedback from the simulation engine as input. The adaptive test scheduler automatically adjusts the test flow and task order based on test objectives, real-time evaluation results, and task completion rates. The intelligent scene generator dynamically adjusts the scene difficulty and complexity based on the semantic understanding results of the multimodal model and user performance, achieving adaptive adjustment of the test content. The adaptive test scheduler dynamically adjusts the test flow and tasks based on test objectives and evaluation feedback. The intelligent scene generator, relying on the scene understanding capabilities of the native multimodal large model, automatically generates complex interactive scenes that match the task objectives. When the system detects changes in user performance during task execution, such as reaction time, error rate, or cognitive load, the intelligent scene generator automatically adjusts the scene difficulty and complexity, achieving adaptive adjustment of the test content.
[0067] Can be used Figure 1 The user interface layer shown presents and provides feedback on high-fidelity interactive scenarios, comprehensive quantitative results, and corresponding interpretable text descriptions through VR / AR headsets or large screens.
[0068] This application overcomes the limitations of traditional single-modal testing by collecting and uniformly processing raw multimodal data during the interaction process to generate fused semantic features, thus reflecting the complex information of real-world interaction scenarios more comprehensively and accurately. Based on these fused features, it simultaneously generates intermediate prediction results for multiple tasks and combines them with high-fidelity interaction scenarios constructed in a virtual environment to achieve real-time, multi-dimensional evaluation of interaction behavior, improving the systematic nature and efficiency of testing. By integrating multi-task prediction results, scenario descriptions, and operational data, it outputs comprehensive quantitative results and interpretable textual explanations for key evaluation indicators, ensuring that test conclusions are not only supported by objective data but also have clear business logic explanations, facilitating developers in identifying problems and optimizing interaction design. Utilizing a virtual environment to generate high-fidelity interaction scenarios allows for flexible simulation of various edge cases and complex working conditions, reducing physical testing costs while ensuring that the test scenarios closely resemble real-world user environments, enhancing the reliability and generalization ability of test results.
[0069] According to some embodiments, in the process of determining intermediate results, specifically: multimodal fusion semantic features are input into a preset native multimodal unified large model to output multi-task prediction intermediate results, wherein the multi-task prediction intermediate results include intent category, sentiment label, task state score and interaction quality level.
[0070] This application can input multimodal fused semantic features into... Figure 2 The native multimodal unified large model shown outputs intermediate results for multi-task prediction, which include the intent category of the interaction intent, the sentiment label representing the emotional state, and indicators such as the task completion degree and interaction quality level representing the task state score.
[0071] This application achieves parallel and collaborative prediction of multiple tasks such as intent recognition, sentiment analysis, task status judgment, and interaction quality assessment by directly inputting the fused multimodal semantic features into the native unified multimodal model. This avoids the error accumulation and semantic fragmentation problems caused by the serial processing of modules in traditional methods, and improves the consistency and accuracy of overall perception.
[0072] According to some embodiments, multimodal fusion semantic features can be input into a preset native multimodal unified large model; the multi-task learning head of the native multimodal unified large model can be used to analyze the multimodal fusion semantic features for multiple tasks to determine the intermediate results of multi-task prediction.
[0073] refer to Figure 2 The multi-task learning head includes an intent recognition head, a sentiment analysis head, and a quality assessment head. Different learning heads can perform corresponding analysis on different multimodal fusion semantic features, output corresponding intermediate results, and thus obtain the overall multi-task prediction intermediate results.
[0074] In some implementations, a multi-task learning head can be established by adding an independent task output layer on top of the shared encoding layer, based on a transformer multi-task learning framework.
[0075] This application enables parallel analysis and prediction of multiple tasks on the same set of fused semantic features by calling the pre-defined multi-task learning head in the native multimodal unified large model. This structured task head design avoids the computational redundancy and parameter bloat caused by building a separate model for each task, significantly improving the efficiency of multi-task processing and the overall system performance.
[0076] According to some embodiments, multimodal fusion semantic features can be input into a preset scene generator to determine the corresponding task type and interaction mode; a scene template can be selected and task parameters can be determined according to the task type and interaction mode; a high-fidelity interactive scene can be established in a virtual environment based on the scene template and task parameters; and the task parameters can be dynamically adjusted based on a reinforcement learning mechanism to generate a scene description with adaptive difficulty.
[0077] In this application, multimodal fusion semantic features are input into a pre-defined scene generator, which can output different interaction modes corresponding to task types and special functions. Interaction modes can include voice, vision, gestures, and physiological signals, and the data sources for different modes may differ. Scene templates and basic task parameters corresponding to different task types and interaction modes can be pre-set for direct matching. Then, a high-fidelity interactive scene is built based on the scene templates and task parameters, and a reinforcement learning mechanism is used to dynamically adjust the task parameters to adapt to different difficulty requirements, thereby updating the scene and obtaining a scene description. The task parameters can include environmental complexity, interaction frequency, and interference level.
[0078] This application determines specific task types and interaction patterns by analyzing multimodal fusion semantic features, ensuring that the generated virtual scenarios are highly relevant to the test objectives. Based on this, scene templates are selected and parameters are configured, achieving automated transformation from abstract test requirements to specific, structured virtual scenarios. This allows test scenarios to cover core interaction logic while also being finely customized for specific test objectives. The scene generator, combining selected high-fidelity scene templates and refined task parameters, can construct highly realistic test scenarios in the virtual environment in terms of vision, interaction logic, and physical characteristics. This greatly enhances the "ecological validity" of the test, ensuring that test results effectively reflect the interactive performance of intelligent devices in real, complex environments, overcoming the problems of traditional scripted test scenarios being singular and detached from reality.
[0079] According to some embodiments, a preset multi-dimensional evaluation engine can be used to evaluate and calculate the intermediate results of multi-task prediction based on the running data, so as to determine the key evaluation indicators and comprehensive quantitative results in the intermediate results of multi-task prediction; and an interpretable text description can be generated based on the scenario description and comprehensive quantitative results using a preset large language model.
[0080] This application may refer to Figure 1 The multi-dimensional evaluation engine calculates key evaluation indicators such as task completion, interaction naturalness, cognitive load, collaborative efficiency, and stability. Among these:
[0081] Completion rate = Target achievement rate × Time efficiency coefficient;
[0082] Naturalness = a weighted sum of the language fluency score (based on the pause / correction ratio) and the gesture-speech consistency score;
[0083] Cognitive load = Standardized HRV × weight + Standardized eye movement regression count × weight;
[0084] Collaboration efficiency = average communication delay inverse ratio score × team consistency score;
[0085] Stability = the inverse ratio of the variance of the results from multiple rounds of testing.
[0086] The final comprehensive score, or comprehensive quantitative result, is a weighted average, with weights configured according to task type.
[0087] Among them, the target achievement rate, language fluency score, average communication delay inverse ratio score, and team consistency score can be extracted from the operational data.
[0088] The comprehensive quantization results and scene description obtained from the above calculations are input into the large language model to generate interpretable text descriptions.
[0089] In the specific implementation process, the comprehensive evaluation and analysis layer can be used to quantitatively analyze indicators such as task completion, interaction naturalness, cognitive load, and collaborative efficiency by taking the interaction data (i.e., runtime data) and scenario descriptions output by the model inference as input. The analysis results are output to the test control layer for subsequent adaptive optimization. Multi-dimensional evaluation comprehensively quantifies task completion, interaction naturalness, cognitive load, and collaborative efficiency. Data sources for the evaluation can include process logs, model output, and scenario and device-side parameters, combining these three sources into one. Specifically, the data sources for the comprehensive evaluation and analysis layer include three categories:
[0090] (1) Interaction process data recorded by the monitoring subsystem: latency, click / accidental trigger events, system alarms; (2) Model inference results: intent, emotion, task status score and confidence level; (3) Scene and device operation data: environmental parameters, interference injection records, task configuration. All data are aligned with a unified timestamp and then enter the evaluation engine. Human-machine collaborative analysis combines multimodal data to analyze collaboration patterns, cooperation levels and potential problems, and generates improvement suggestions.
[0091] In some implementations, a multi-dimensional evaluation approach using an indicator library and a calculator can be employed to achieve comprehensive quantitative assessment. Specifically, each indicator is configured with a data source, time window, and calculation formula, which is then executed by the evaluation calculator. Human-machine collaborative analysis identifies common problems based on clustering and correlation analysis, such as the correlation between high cognitive load and false triggering, and the language generation module outputs interpretable conclusions. Furthermore, the evaluation calculator uses a hybrid batch processing and streaming computing mode. Streaming computing is used for real-time indicators (including latency and false triggering rate), while batch processing is used for periodic indicators (including stability and collaborative efficiency). The results are written to a message bus and time-series database in standardized JSON format for the control layer to subscribe to and use in report generation.
[0092] This application utilizes a pre-defined multi-dimensional evaluation engine to automatically correlate and calculate scenario operation data with intermediate results from multi-task predictions, generating standardized key evaluation indicators and comprehensive quantitative results. This avoids the subjectivity and inconsistency of manual evaluation, ensuring the repeatability and objectivity of test results, and shifting the evaluation process from experience-driven to data-driven. The evaluation engine not only relies on operational data and prediction results during calculation but also fully integrates the complete context provided by the scenario description. This ensures that the evaluation indicators accurately reflect the interactive performance in specific test scenarios, distinguishing between general capabilities and scenario-specific performance, making the evaluation conclusions more targeted and practically instructive. The output comprehensive quantitative results provide clear numerical measurements of the interactive performance of intelligent devices. This structured data facilitates version comparison, regression testing, performance baseline management, and bottleneck identification, providing traceable and verifiable data support for R&D decisions. Utilizing a large language model to automatically generate interpretable text descriptions transforms complex quantitative data and technical indicators into easily understandable narrative reports. This not only lowers the professional threshold for test reports but also clearly identifies problems, potential causes, and improvement suggestions, directly driving product optimization.
[0093] According to some embodiments, the original multimodal data corresponding to the test target can be collected; features can be extracted from the original multimodal data to determine initial features; the initial features can be processed for temporal and semantic consistency to determine unified features; and the unified features can be deeply fused to generate multimodal fused semantic features.
[0094] refer to Figure 4This application first initializes the system and equipment, and uses a preset feature extraction method to extract features from the collected raw multimodal data to obtain initial features, which are low-level feature representations of each modality. A cross-modal attention mechanism is then used to semantically align the multi-source signals on a unified time axis to obtain feature sequences for each modality. These feature sequences are mapped to a shared time index, and inter-modal correlation weights are calculated using an alignment head to eliminate sampling rate differences and hysteresis, improve cross-modal semantic consistency, and obtain unified features.
[0095] In some implementations, the acquired raw multimodal data can be preprocessed. This includes noise reduction and endpoint detection for speech, distortion correction and illumination equalization for images, smoothing and imputation of depth and skeleton data, and filtering, noise reduction, and artifact removal for physiological data. Subsequently, feature standardization and timestamp alignment are performed. The data from each modality is written to a circular buffer in (timestamp, feature vector) format, providing a bounded-delay data stream for downstream large-scale model inference.
[0096] In some implementations, the native multimodal unified large model and the test control module of the execution subject of this application can be initialized, a connection with the acquisition device can be established, and test environment parameters can be configured.
[0097] Specifically, initial features can be extracted using a modality-specific encoder group. Then, using a unified Transformer as the core, the aligned unified features can be fused end-to-end to obtain a shared semantic representation, i.e., multimodal fusion semantic features. These multimodal fusion semantic features serve as input to a multi-task learning head, used for tasks such as interactive intent recognition, emotion state recognition, quality assessment, and scene understanding. These tasks collectively constitute the core functional modules of the system, with interactive intent recognition and emotion state recognition being the main technological innovations, and quality assessment and scene understanding serving as auxiliary modules to form a complete testing and evaluation closed loop.
[0098] This application effectively addresses the heterogeneity issues of different modalities in sampling rate, timestamps, and semantic granularity by performing unified feature extraction on raw multimodal data and implementing temporal and semantic consistency processing. This enables heterogeneous data such as visual, speech, tactile, and sensor data to be aligned within the same spatiotemporal and semantic framework, providing structured, high-quality input for subsequent deep fusion. Deep fusion of unified features can mine and utilize complementary information between multimodal data (e.g., speech referential disambiguation can be aided by visual target localization, and tactile force information can assist in confirming intent strength), thereby generating more comprehensive and discriminative fused semantic features that cannot be provided by a single modality. This provides a richer semantic foundation for downstream multi-task understanding and evaluation.
[0099] According to some embodiments, reinforcement learning mechanisms can also be used to adjust custom difficulty information based on comprehensive quantification results. Custom difficulty information includes task density, interference intensity, and information occlusion ratio.
[0100] In this application, based on the changing trends of task completion and error type in the comprehensive quantification results, reinforcement learning algorithms can be used to adjust scene complexity, such as task density and interference level, and interaction strategy parameters, such as response latency and prompt frequency, to achieve real-time adaptive optimization.
[0101] In some implementations, after determining the comprehensive quantitative results, a structured report containing indicators, trajectories, events, and improvement suggestions can be generated and stored in a database for auditing and retesting comparison. The data sources for the report include:
[0102] (1) Evaluation calculator output metrics set; (2) Scene parameters and version number of the scene generator; (3) Interaction events and timeline of the monitoring system; (4) Improvement suggestion text of the language generation module. The report generator summarizes the above four types of data, outputs a structured report and archives it.
[0103] In some implementations, the evaluation report is provided by Figure 1 The system automatically generates the results shown. In the specific implementation process, after calculating the scores for each dimension, the engine calls the language generation module to convert the results into a natural language report, which includes explanations of key indicators, descriptions of anomalies, and targeted optimization suggestions.
[0104] According to other embodiments, reference Figure 5 This paper describes the intelligent adaptive testing process and control process from a system perspective.
[0105] First, the system loads a unified multimodal model and control strategy, completes device enumeration, port detection, and parameter configuration, and achieves initialization. Users interact with the system via voice, gestures, touch, and eye tracking in VR / AR or large-screen environments. The system collects voice, visual, gesture, and physiological data in parallel, performing preprocessing and alignment. During the model inference phase, the system sequentially inputs the collected raw multimodal data into the modal encoder group, cross-modal alignment module, and unified representation learning layer. After calculation by the multi-task learning head, intermediate results are output, including indicators such as interaction intent, emotional state, and task completion rate. Based on semantic representation and testing objectives derived from model understanding, scene elements are constructed from a template library, including environment, objects, tasks, and interference factors, supporting parameterized configuration and adaptive difficulty adjustment. The system provides low-latency feedback and monitors key indicators, including response latency, recognition confidence, and false trigger rate. The multi-dimensional evaluation engine calculates indicators such as task completion, interaction naturalness, cognitive load, collaborative efficiency, and stability, and outputs interpretable conclusions in conjunction with the language generation module. Based on the changing trends of task completion and error types in the evaluation results, the system uses reinforcement learning algorithms to adjust scene complexity and interaction strategy parameters, achieving real-time adaptive optimization. It generates a structured report containing metrics, trajectories, events, and improvement suggestions, and stores it in a database for auditing and retesting comparison.
[0106] According to some embodiments, reference Figure 6 This describes the implementation details of intelligent scenarios, namely high-fidelity interactive scenarios.
[0107] It can pre-configure common interaction modes, including voice control, gesture commands, multi-source information fusion, and anomaly handling, and supports industry-specific template segmentation to form a basic scenario template library. Through parameter sets, including complexity, interaction frequency, interference level, and task priority, it can quickly derive variant scenarios and record parameter versions for repeatability verification. The parameterized configuration engine takes the collected multimodal data analysis results as input, extracts task features, and maps them to parameter sets (complexity, interference level, interaction frequency, etc.). The system quickly derives different variant scenarios based on the parameter sets; these scenarios, with the same task logic, are implemented by changing parameters.
[0108] A reinforcement learning mechanism is introduced to dynamically adjust task density, interference intensity, and information masking ratio based on real-time evaluation feedback (success rate, reaction time, error type). Unreasonable or invalid scenarios are filtered out through large-scale model semantic consistency scoring, realism scoring, and task effectiveness scoring to ensure effective test coverage. The three scores are derived from simulation logs and model evaluation results, respectively. Specifically, the semantic consistency score is calculated from the semantic matching degree, the realism score is derived from the difference between simulation and real-world samples, and the task effectiveness score is generated from the task completion index. If the score is below a threshold (e.g., consistency < 0.6 or effectiveness < 0.5), the scenario is deemed invalid, and the system automatically removes it and regenerates a new scenario to replace it.
[0109] According to some embodiments, reference Figure 7 This paper provides a detailed overall description of the implementation of a multi-dimensional human-machine collaboration effectiveness evaluation.
[0110] The data sources for the evaluation indicator system are as follows:
[0111] Completion level is determined by task execution logs and target configuration; naturalness by voice / gesture processing logs and alignment consistency; cognitive load by physiological signal devices (heart rate variability, eye movement); and collaborative efficiency by multi-person collaborative event streams (message timestamps, shared state consistency). All source data is aligned by the time synchronization module before being used for metric calculation.
[0112] It adopts a multi-dimensional system (15 dimensions and more than 50 indicators) that includes task completion, interaction naturalness, cognitive load, and collaborative efficiency. Specifically, it includes completion time, step redundancy, error rate, voice and gesture recognition accuracy, eye movement regression count, heart rate variability, emotional stability, and team communication latency.
[0113] The system calculates a comprehensive score through multimodal fusion, performs confidence interval and significance analysis on key indicators, and generates cluster reports for common problems. The comprehensive score is calculated using a weighted average model, with weights dynamically adjusted based on task importance and relevance. Confidence intervals are estimated using a bootstrap sampling method, and significance analysis is performed based on t-tests. Common problems refer to recurring interaction anomalies or performance degradation types in multiple tests, such as speech misrecognition, gesture delays, or excessive cognitive load. The cluster reports are generated by analyzing historical test data using the K-means clustering algorithm, summarizing the types and frequencies of common problems, and serving as a supplementary result to the comprehensive evaluation analysis layer. Together with improvement suggestions from natural language generation, they constitute the conclusions of the human-machine collaborative analysis.
[0114] Leveraging the language generation capabilities of large models, the system outputs natural language descriptions of problem causes and improvement suggestions. Clustering reports and interpretable text descriptions belong to different levels of results. The former is a structured statistical output used for quantitative analysis of common problems; the latter is an interpretive description generated from natural language, used for qualitative explanation of problem causes and improvement directions. The two complement each other.
[0115] According to some embodiments, combined with Figure 2 and Figure 6 This paper provides supplementary descriptions of key algorithms and parameter suggestions for generating high-fidelity interactive scenarios within the system.
[0116] The cross-modal attention implementation in this application adopts a query-key-value mechanism to weight the feature sequences of different modalities, and the weights are jointly determined by temporal neighborhood and semantic relevance; the number of alignment heads and hidden dimensions are pruned according to real-time requirements.
[0117] Real-time inference optimization uses methods such as model pruning, weight quantization, and parallel pipelined inference to prioritize ensuring the interaction loop latency of tests performed based on user interaction information.
[0118] The reward function incorporates factors such as improved completion rate, reduced errors, optimized reaction time, and enhanced stability. An online strategy is used to fine-tune the scenario parameters, gradually approaching the boundary conditions to expand the test coverage.
[0119] According to some embodiments, reference Figure 8 This provides a supplementary description of the deployment and integration process.
[0120] A unified driver and acquisition protocol is implemented on the sensor side; model services are provided to the front-end and scheduler via RPC or REST; logs and metrics are written to the time-series database via a message bus. Each functional layer is decoupled through standardized interfaces, allowing replacement of any modal encoder or simulation engine without affecting overall operation; expansion with new scene templates and evaluation metrics is supported. Data acquisition, inference results, and evaluation reports are digitally signed and versioned to ensure consistency and traceability in retesting.
[0121] Combination Figure 5 and Figure 7 The corresponding output and retesting mechanism are described in detail: the report includes scene parameters, interaction trajectory, key indicators, timeline of abnormal events, model attention visualization summary, and improvement suggestions. Retesting can be conducted using either the "same template, same parameters" or "same template, progressive parameters" approach to compare and evaluate score change trends and verify the effectiveness of optimization.
[0122] According to other embodiments, an example of an actual testing process using the methods of this application is provided.
[0123] Example 1: Human-Computer Interaction Test of Intelligent Algorithm
[0124] Scenario setting: Select the "Algorithm Verification" template, and the task is a joint test of "Multimodal Intent Recognition + Emotional State Recognition"; add background noise and lighting change interference (medium intensity).
[0125] Execution process: Follow steps S30 to S33 to collect multimodal data, generate target / interference mixed scenarios through model inference, subject interaction, real-time monitoring, evaluation, difficulty adjustment, and report generation.
[0126] Key takeaways: The end-to-end fusion model avoids information loss during post-fusion; reinforcement learning is used to dynamically increase task density or add anomalous events to test the algorithm's robustness; the evaluation report includes metrics such as completion rate, recognition accuracy, and sentiment stability, and provides suggestions for improvement.
[0127] Expected results: The accuracy of cross-modal understanding, interactive intent recognition, and emotion recognition will reach the preset targets, and the system response latency will be ≤20ms.
[0128] Example 2: Intelligent Device Control Test
[0129] Scenario setting: Select the "Device Operation Evaluation" template and open four types of interaction channels: voice command, gesture control, touch control and eye tracking; inject three types of interference: "device lag", "false alarm" and "visual occlusion", and set three levels of difficulty from easy to difficult.
[0130] Execution process: Implemented according to steps S30 to S33, including multimodal data acquisition and inference, scene generation, adaptive difficulty adjustment, real-time monitoring and evaluation.
[0131] Key points: Establish an "abnormal pattern recognition + early warning" mechanism to provide early warnings for delays, lags, and false triggers; evaluate the coupling relationship between "human-machine-environment" in human-machine collaborative analysis and form optimization suggestions.
[0132] Evaluation and playback: Quantitatively evaluate task efficiency, erroneous actions, and cognitive load; fully record the control trajectory, supporting playback and debriefing analysis.
[0133] Example 3: Intelligent Command and Control Test
[0134] Scenario setting: Select the "Command Decision Analysis" template to construct a multi-source information fusion situation, including images, text briefings, voice broadcasts, map plotting, etc., and set three types of decision-making tasks: routine, emergency, and uncertainty.
[0135] Execution process: Implemented according to steps S30 to S33, presenting situational information in a multimodal manner, supporting natural language queries and gesture annotations; compressing the decision-making time window in emergency tasks, and increasing the frequency and degree of information disturbance.
[0136] Evaluation method: The evaluation will assess indicators such as information acquisition efficiency, analysis depth, collaboration quality, and stress adaptability, and will be compared and analyzed based on historical templates to provide personalized improvement suggestions.
[0137] This application employs an end-to-end native multimodal unified large model to achieve unified representation learning and cross-modal semantic alignment of data from multiple modalities, including speech, vision, gestures, and physiological signals. This avoids information loss in traditional post-fusion methods and significantly improves cross-modal understanding accuracy. Through an intelligent scene generation module, it can automatically construct complex interactive test scenarios closely resembling real-world applications and, combined with reinforcement learning mechanisms, achieve adaptive difficulty adjustment, resulting in significantly improved scene coverage and test relevance compared to existing methods. A multi-dimensional human-machine collaboration performance evaluation system is constructed, encompassing task completion, interaction naturalness, cognitive load, and collaborative efficiency. This system comprehensively reflects the overall performance of intelligent devices in practical applications and generates interpretable improvement suggestions through the large model. Supported by a high-performance computing platform and optimized inference strategies, the system achieves low-latency real-time multimodal processing capabilities, ensuring both high accuracy and low latency to meet the stringent real-time requirements of intelligent devices and command and control scenarios. The overall architecture features a modular design and standardized hardware and software interfaces, possessing excellent scalability and portability, and can be widely adapted to various application scenarios such as intelligent algorithm testing, intelligent device operation, and intelligent command and control.
[0138] The following describes an apparatus embodiment of this application, which can be used to perform the method embodiment of this application. For details not disclosed in the apparatus embodiment of this application, please refer to the method embodiment of this application.
[0139] Figure 9 A block diagram of a testing apparatus for human-computer interaction of a smart device provided in an embodiment of this application. Figure 9 As shown, the testing device 900 for human-computer interaction of intelligent devices includes a data acquisition and processing module 901, an intermediate result determination module 902, a scene generation module 903, and an evaluation module 904.
[0140] The data acquisition and processing module 901 is used to acquire the original multimodal data of the preset test target and perform unified processing to generate multimodal fusion semantic features;
[0141] The intermediate result determination module 902 is used to determine the intermediate results of multi-task prediction based on the multimodal fusion semantic features;
[0142] The scene generation module 903 is used to generate a scene description of a high-fidelity interactive scene corresponding to the test target in a virtual environment based on multimodal fusion semantic features and test targets;
[0143] Evaluation module 904 is used to determine and output the comprehensive quantitative results of key evaluation indicators and corresponding interpretable text descriptions based on the intermediate results of multi-task prediction, scenario descriptions and operational data during scenario operation.
[0144] Optionally, the intermediate result determination module 902 is specifically used for:
[0145] Multimodal fusion semantic features are input into a pre-defined native multimodal unified large model to output intermediate results for multi-task prediction. These intermediate results include intent category, sentiment label, task state score, and interaction quality level.
[0146] Optionally, the intermediate result determination module 902, when inputting multimodal fused semantic features into a preset native multimodal unified large model to output multi-task prediction intermediate results, is specifically used for:
[0147] Input the multimodal fused semantic features into the pre-defined native multimodal unified large model;
[0148] We utilize the multi-task learning head of the native multimodal unified large model to analyze multimodal fused semantic features for various tasks in order to determine the intermediate results of multi-task prediction.
[0149] Optionally, the scene generation module 903 is specifically used for:
[0150] Multimodal fusion semantic features are input into a preset scene generator to determine the corresponding task type and interaction mode;
[0151] Select a scene template and determine the task parameters based on the task type and interaction mode;
[0152] Based on scene templates and task parameters, high-fidelity interactive scenes are created in a virtual environment;
[0153] The task parameters are dynamically adjusted based on the reinforcement learning mechanism to generate scene descriptions with adaptive difficulty.
[0154] Optionally, the evaluation module 904 is specifically used for:
[0155] Using a pre-defined multi-dimensional evaluation engine, the intermediate results of multi-task prediction are evaluated and calculated based on the running data to determine the key evaluation indicators and comprehensive quantitative results in the intermediate results of multi-task prediction.
[0156] Using a pre-defined large language model, interpretable text descriptions are generated based on scene descriptions and comprehensive quantification results.
[0157] Optionally, the data acquisition and processing module 901 is specifically used for:
[0158] Collect the raw multimodal data corresponding to the test target;
[0159] Feature extraction is performed on the original multimodal data to determine the initial features;
[0160] The initial features are processed for temporal and semantic consistency to determine unified features;
[0161] Deeply fuse unified features to generate multimodal fused semantic features.
[0162] Optionally, the testing apparatus 900 for human-computer interaction of intelligent devices also includes a custom module 905 for:
[0163] By utilizing reinforcement learning mechanisms, the custom difficulty information is adjusted based on the comprehensive quantification results. The custom difficulty information includes task density, interference intensity, and information occlusion ratio.
[0164] The device performs functions similar to those described above; other functions are described in the preceding descriptions and will not be repeated here.
[0165] Figure 10 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 10 As shown, the electronic device 1000 of this embodiment may include a memory 1001 and a processor 1002.
[0166] The memory 1001 stores a computer program, which, when executed by the processor 1002, causes the processor 1002 to perform the method described in the above embodiments.
[0167] The processor 1002 and the memory 1001 are connected, for example, via a bus.
[0168] Optionally, the electronic device 1000 may also include a transceiver. It should be noted that in practical applications, the transceiver is not limited to one, and the structure of the electronic device 1000 does not constitute a limitation on the embodiments of this application.
[0169] Processor 1002 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 1002 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0170] A bus can include a pathway for transmitting information between the aforementioned components. The bus can be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one thick line is used in the diagram, but this does not imply that there is only one bus or one type of bus.
[0171] The memory 1001 can be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or it can be EEPROM (Electrically Erasable Programmable Read Only Memory), CD. ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed discs, laser discs, optical discs, digital universal discs, Blu-ray discs, etc.), disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0172] The memory 1001 is used to store application code that executes the solution of this application, and its execution is controlled by the processor 1002. The processor 1002 is used to execute the application code stored in the memory 1001 to implement the content shown in the foregoing method embodiments.
[0173] Electronic devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Servers can also be included. Figure 10 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0174] The electronic device in this embodiment can be used to execute the method of any of the above embodiments, and its implementation principle and technical effect are similar, so they will not be described again here.
[0175] This application also provides a non-transitory computer-readable storage medium storing computer-readable instructions thereon, which, when executed by a processor, cause the processor to perform the method as described in the above embodiments.
[0176] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a non-transitory computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0177] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. Furthermore, any changes or modifications made by those skilled in the art based on the ideas of this application, and on the specific implementation methods and application scope of this application, are all within the scope of protection of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A testing method for human-computer interaction in intelligent devices, characterized in that, include: Collect raw multimodal data of the preset test targets and process them uniformly to generate multimodal fusion semantic features; Based on the multimodal fusion semantic features, determine the intermediate results of multi-task prediction; Based on the multimodal fusion semantic features and the test target, a scene description of a high-fidelity interactive scene corresponding to the test target is generated in a virtual environment; Based on the intermediate results of the multi-task prediction, the scenario description, and the operational data during the scenario operation, determine and output the comprehensive quantitative results of the key evaluation indicators and the corresponding interpretable text descriptions. The step of generating a scene description of a high-fidelity interactive scene corresponding to the test target in a virtual environment based on the multimodal fusion semantic features and the test target includes: The multimodal fusion semantic features are input into a preset scene generator to determine the corresponding task type and interaction mode; Select a scene template and determine task parameters based on the task type and interaction mode; Based on the scene template and the task parameters, the high-fidelity interactive scene is established in the virtual environment; The task parameters are dynamically adjusted based on a reinforcement learning mechanism to generate a scene description with adaptive difficulty. The step of determining and outputting a comprehensive quantitative result of key evaluation indicators and corresponding interpretable text descriptions based on the intermediate results of the multi-task prediction, the scenario description, and the operational data during the scenario operation includes: Using a preset multi-dimensional evaluation engine, the intermediate results of the multi-task prediction are evaluated and calculated based on the running data to determine the key evaluation indicators and the comprehensive quantitative results in the intermediate results of the multi-task prediction. Using a pre-defined large language model, the interpretable text description is generated based on the scenario description and the comprehensive quantization results.
2. The method according to claim 1, characterized in that, The step of determining the intermediate results of multi-task prediction based on the multimodal fusion semantic features includes: The multimodal fusion semantic features are input into a preset native multimodal unified large model to output the multi-task prediction intermediate results, wherein the multi-task prediction intermediate results include intent category, sentiment label, task state score and interaction quality level.
3. The method according to claim 2, characterized in that, The step of inputting the multimodal fused semantic features into a preset native multimodal unified large model to output the multi-task prediction intermediate results includes: The multimodal fused semantic features are input into a preset native multimodal unified large model; The multi-task learning head of the native multimodal unified large model is used to analyze the multimodal fused semantic features for multiple tasks in order to determine the intermediate results of the multi-task prediction.
4. The method according to claim 1, characterized in that, The process of collecting and uniformly processing the raw multimodal data of the preset test target to generate multimodal fusion semantic features includes: Collect the original multimodal data corresponding to the test target; Feature extraction is performed on the original multimodal data to determine initial features; The initial features are subjected to temporal and semantic consistency processing to determine unified features; The unified features are deeply fused to generate the multimodal fused semantic features.
5. The method according to any one of claims 1-4, characterized in that, Also includes: Using a reinforcement learning mechanism, the custom difficulty information is adjusted based on the comprehensive quantization results. The custom difficulty information includes task density, interference intensity, and information occlusion ratio.
6. A testing device for human-computer interaction of intelligent devices, characterized in that, include: The data acquisition and processing module is used to acquire the raw multimodal data of the preset test target and process it in a unified manner to generate multimodal fusion semantic features; The intermediate result determination module is used to determine the intermediate results of multi-task prediction based on the multimodal fusion semantic features; The scene generation module is used to generate a scene description of a high-fidelity interactive scene corresponding to the test target in a virtual environment based on the multimodal fusion semantic features and the test target; The evaluation module is used to determine and output the comprehensive quantitative results of key evaluation indicators and corresponding interpretable text descriptions based on the intermediate results of the multi-task prediction, the scenario description, and the running data during the scenario operation. Specifically, the scene generation module is used for: The multimodal fusion semantic features are input into a preset scene generator to determine the corresponding task type and interaction mode; Select a scene template and determine task parameters based on the task type and interaction mode; Based on the scene template and the task parameters, the high-fidelity interactive scene is established in the virtual environment; The task parameters are dynamically adjusted based on a reinforcement learning mechanism to generate a scene description with adaptive difficulty. Specifically, the evaluation module is used for: Using a preset multi-dimensional evaluation engine, the intermediate results of the multi-task prediction are evaluated and calculated based on the running data to determine the key evaluation indicators and the comprehensive quantitative results in the intermediate results of the multi-task prediction. Using a pre-defined large language model, the interpretable text description is generated based on the scenario description and the comprehensive quantization results.
7. An electronic device, characterized in that, include: processor; A memory storing a computer program that, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-5.
8. A non-transitory computer-readable storage medium, characterized in that, It stores computer-readable instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-5.