Intelligent call dynamic response method integrating ASR and emotion recognition

By adopting a parallel processing architecture of ASR and emotion recognition in the call center system, combined with dynamic weight allocation and closed-loop learning mechanism, the problems of traditional ASR ignoring emotional information and response delay are solved, and real-time and accurate emotional human-computer interaction is achieved.

CN120913558APending Publication Date: 2025-11-07SHANGHAI ZHAOKUN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511098018.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

In existing call center systems, traditional ASR (Automatic Responsibility Response) focuses only on text conversion and ignores emotional information, leading to inappropriate response strategies. Furthermore, emotion recognition systems have high response latency and cannot meet the needs of real-time interaction.

Method used

A parallel processing architecture for ASR and emotion recognition is adopted. The ASR text stream and emotion feature stream are processed in real time through dual channels. The dynamic weight allocation model and closed-loop learning mechanism are combined to optimize the emotion judgment threshold to achieve real-time response.

Benefits of technology

It achieves simultaneous analysis of ASR and emotion recognition, reducing latency by 40%, and adjusts the weight of emotion factors according to business type, improving the accuracy of response strategies and the real-time performance of responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913558A_ABST
    Figure CN120913558A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent call dynamic response method and system fusing ASR (automatic voice recognition) and emotion recognition, and is applied to an interaction scene of a voice robot and a call center. The robot response strategy is dynamically adjusted by synchronously analyzing the text content (ASR) and emotional characteristics (such as intonation, speech speed and energy) of the user voice in real time. When the negative emotion is detected, a manual seat is automatically triggered to switch over or switch the pacifying verbal skill; and for the positive emotion of the high-value customer, a precision marketing module is started. The method solves the problems that an existing call center is single in response mode and cannot sense the emotion of a user, and the customer satisfaction and the service conversion rate are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent voice interaction, and particularly relates to an emotional human-computer interaction method in a call center scene. BACKGROUND

[0002] The existing call center system has two defects:

[0003] 1. The traditional ASR only focuses on text conversion, and ignores emotional information, resulting in improper response strategies;

[0004] 2. The emotional recognition system has high response delay and cannot meet real-time interaction requirements.

[0005] The application solves the above problems through a parallel processing architecture of ASR and emotion recognition (see Figure 2 ). SUMMARY

[0006] Core innovation points:

[0007] 1. Dual-channel real-time processing: synchronous analysis of ASR text stream and emotional feature stream (technical effect: delay reduction of 40%);

[0008] 2. Dynamic weight distribution model: adjusting the weight of emotional factors according to the type of business (for example, the emotional weight in the complaint scene is greater than or equal to 0.7);

[0009] 3. Closed-loop learning mechanism: optimizing the emotional judgment threshold based on the subsequent processing results of the agent. DETAILED DESCRIPTION

[0010] Example 1: product after-sales service scene

[0011] User input voice: "Your new system is completely useless!"

[0012] Parallel processing results:

[0013] ASR output: {Complaint words: "useless", product words: "new system"} (confidence 0.91)

[0014] Emotional analysis: anger (0.83), anxiety (0.76)

[0015] Decision process:

[0016] Semantic score: 0.95 (complaint type)

[0017] Emotional score: 0.79 (composite value)

[0018] The history record shows that the user is a VIP customer (γ = 0.3)

[0019] Final Urgency = 0.95*0.5 + 0.79*0.4 + 0.3*0.1 = 0.816

[0020] System response:

[0021] Play soothing words immediately: "I'm sorry for the inconvenience, we will prioritize your processing..."

[0022] Push to the agent terminal: ① User profile ② Similar problem solution ③ Emotional relief words prompt BRIEF DESCRIPTION OF DRAWINGS

[0023] Attached Figure 1 : System architecture diagram (showing voice input → ASR / emotion double channel → decision engine → response output)

[0024] Attached Figure 2 : Emotional feature extraction flowchart (acoustic preprocessing → feature extraction → emotion classification).

Claims

1. A method for call dynamic response by fusing ASR and emotion recognition, characterized in that The method comprises the following steps: Real-time receiving user voice stream, parallel executing ASR text conversion and emotion feature extraction; Generating composite decision parameters based on semantic analysis and emotion score, which is calculated by acoustic features (fundamental frequency, short-time energy) and language features (keyword frequency).

2. Dynamically selecting response mode according to decision parameters: regular robot response, emotional dialogue adjustment or manual agent transfer.

3. The method of claim 1, wherein: The emotion feature extraction adopts a CNN-LSTM hybrid model, and the input features include MFCC, F0 and speech speed variance.

4. The method of claim 1, wherein: When the emotion score is below the threshold and the ASR recognizes complaint keywords, manual agent transfer is triggered preferentially and user historical service records are pushed. A system for implementing the method of claims 1-3, comprising: a voice receiving module (101), an ASR engine (102), an emotion analysis module (103); a dynamic decision engine (104) for weighted fusion of semantic intent and emotion label; a multi-modal response generator (105) supporting collaborative output of voice, text and visual interface.

Citation Information

Cited By

  • AI question and answer expert model construction method and system based on psychological counseling and medium

    CN121171262A

  • Psychological counseling-based ai question and answer expert model construction method and system, and medium

    CN121171262B