Digital Assistant Chat Skill Evaluation Using Automated Test Cases
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for evaluating digital assistants rely heavily on manual testing, which is inefficient, lacks objectivity, and is difficult to standardize, leading to inconsistent evaluation results.
Innovation Solution
An automated evaluation method that uses predefined test cases to assess a digital assistant's chat skills, determining a target evaluation index including a chat skill score, and generating a quality evaluation result based on these assessments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual testing is used to evaluate digital assistants, then evaluation can be performed with current methods, but evaluation efficiency is low and objectivity is lacking
Solution Approach 1:
The evaluation system performs self-service by automatically executing test cases against digital assistants without requiring manual intervention. The system autonomously sends test questions, collects responses, evaluates performance against criteria, and generates reports, eliminating the need for human evaluators to manually test each assistant.
Solution Approach 2:
The patent replaces the mechanical manual testing process with an automated computational system. Instead of human evaluators manually interacting with digital assistants, the system uses automated scripts and evaluation algorithms to perform the same functions, substituting mechanical human actions with automated digital processes.
2Manufacturing precision
If manual testing is used for evaluation, then flexibility in assessment is maintained, but standardization and consistency are difficult to achieve
Solution Approach 1:
The evaluation system segments the assessment process into distinct, standardized modules: test case generation, question sending, response collection, performance evaluation, and report generation. Each module operates independently with defined inputs and outputs, ensuring consistent application of evaluation criteria across all digital assistants while maintaining manageable system complexity through modular architecture.
3Reliability
If automated evaluation with predefined test cases is used, then evaluation efficiency and objectivity are improved, but the complexity of the evaluation system increases
Solution Approach 1:
The evaluation system achieves universality by designing a single platform that can evaluate multiple digital assistants across various skills and domains using the same standardized process. The system handles different types of digital assistants (customer service, education, entertainment) through a unified architecture, reducing overall system complexity while maintaining high reliability and objectivity through consistent automated evaluation procedures.
Data Source
AI summary
The disclosure relates to digital assistant evaluation. In an example method, in response to an evaluation request for a target digital assistant, at least one set of test cases for the target digital assistant is obtained, and each set of test cases includes at least one test question related to a chat skill of the target digital assistant. The at least one set of test cases is provided to the target digital assistant to obtain a reply to the at least one set of test cases by the target digital assistant. A target evaluation index for the target digital assistant is determined based at least on the at least one set of test cases and the reply to the at least one set of test cases by the target digital assistant. A quality evaluation result of the target digital assistant is determined based on the target evaluation index.


