Digital Assistant Chat Skill Evaluation Using Automated Test Cases

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for evaluating digital assistants rely heavily on manual testing, which is inefficient, lacks objectivity, and is difficult to standardize, leading to inconsistent evaluation results.

Innovation Solution

An automated evaluation method that uses predefined test cases to assess a digital assistant's chat skills, determining a target evaluation index including a chat skill score, and generating a quality evaluation result based on these assessments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If manual testing is used to evaluate digital assistants, then evaluation can be performed with current methods, but evaluation efficiency is low and objectivity is lacking

Engineering Contradiction:
Improveevaluation efficiencyVSAvoidautomation level
Core Design Contradiction:
ProductivityVSExtent of automation

Solution Approach 1:

The evaluation system performs self-service by automatically executing test cases against digital assistants without requiring manual intervention. The system autonomously sends test questions, collects responses, evaluates performance against criteria, and generates reports, eliminating the need for human evaluators to manually test each assistant.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual testing process with an automated computational system. Instead of human evaluators manually interacting with digital assistants, the system uses automated scripts and evaluation algorithms to perform the same functions, substituting mechanical human actions with automated digital processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual testing is used for evaluation, then flexibility in assessment is maintained, but standardization and consistency are difficult to achieve

Engineering Contradiction:
Improveevaluation consistencyVSAvoidevaluation system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The evaluation system segments the assessment process into distinct, standardized modules: test case generation, question sending, response collection, performance evaluation, and report generation. Each module operates independently with defined inputs and outputs, ensuring consistent application of evaluation criteria across all digital assistants while maintaining manageable system complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

3Reliability

If automated evaluation with predefined test cases is used, then evaluation efficiency and objectivity are improved, but the complexity of the evaluation system increases

Engineering Contradiction:
Improveevaluation objectivityVSAvoidsystem structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The evaluation system achieves universality by designing a single platform that can evaluate multiple digital assistants across various skills and domains using the same standardized process. The system handles different types of digital assistants (customer service, education, entertainment) through a unified architecture, reducing overall system complexity while maintaining high reliability and objectivity through consistent automated evaluation procedures.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20260050800A1Digital assistant evaluation
Publication Date: 2026.02.19 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260050800A1 patent drawing
  • US20260050800A1 patent drawing
  • US20260050800A1 patent drawing

AI summary

The disclosure relates to digital assistant evaluation. In an example method, in response to an evaluation request for a target digital assistant, at least one set of test cases for the target digital assistant is obtained, and each set of test cases includes at least one test question related to a chat skill of the target digital assistant. The at least one set of test cases is provided to the target digital assistant to obtain a reply to the at least one set of test cases by the target digital assistant. A target evaluation index for the target digital assistant is determined based at least on the at least one set of test cases and the reply to the at least one set of test cases by the target digital assistant. A quality evaluation result of the target digital assistant is determined based on the target evaluation index.