Intelligent call dynamic response method integrating ASR and emotion recognition
By adopting a parallel processing architecture of ASR and emotion recognition in the call center system, combined with dynamic weight allocation and closed-loop learning mechanism, the problems of traditional ASR ignoring emotional information and response delay are solved, and real-time and accurate emotional human-computer interaction is achieved.
Patent Information
- Application Number
- CN202511098018.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-07
AI Technical Summary
In existing call center systems, traditional ASR (Automatic Responsibility Response) focuses only on text conversion and ignores emotional information, leading to inappropriate response strategies. Furthermore, emotion recognition systems have high response latency and cannot meet the needs of real-time interaction.
A parallel processing architecture for ASR and emotion recognition is adopted. The ASR text stream and emotion feature stream are processed in real time through dual channels. The dynamic weight allocation model and closed-loop learning mechanism are combined to optimize the emotion judgment threshold to achieve real-time response.
It achieves simultaneous analysis of ASR and emotion recognition, reducing latency by 40%, and adjusts the weight of emotion factors according to business type, improving the accuracy of response strategies and the real-time performance of responses.
Smart Images

Figure CN120913558A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of intelligent voice interaction, and particularly relates to an emotional human-computer interaction method in a call center scene. BACKGROUND
[0002] The existing call center system has two defects:
[0003] 1. The traditional ASR only focuses on text conversion, and ignores emotional information, resulting in improper response strategies;
[0004] 2. The emotional recognition system has high response delay and cannot meet real-time interaction requirements.
[0005] The application solves the above problems through a parallel processing architecture of ASR and emotion recognition (see Figure 2 ). SUMMARY
[0006] Core innovation points:
[0007] 1. Dual-channel real-time processing: synchronous analysis of ASR text stream and emotional feature stream (technical effect: delay reduction of 40%);
[0008] 2. Dynamic weight distribution model: adjusting the weight of emotional factors according to the type of business (for example, the emotional weight in the complaint scene is greater than or equal to 0.7);
[0009] 3. Closed-loop learning mechanism: optimizing the emotional judgment threshold based on the subsequent processing results of the agent. DETAILED DESCRIPTION
[0010] Example 1: product after-sales service scene
[0011] User input voice: "Your new system is completely useless!"
[0012] Parallel processing results:
[0013] ASR output: {Complaint words: "useless", product words: "new system"} (confidence 0.91)
[0014] Emotional analysis: anger (0.83), anxiety (0.76)
[0015] Decision process:
[0016] Semantic score: 0.95 (complaint type)
[0017] Emotional score: 0.79 (composite value)
[0018] The history record shows that the user is a VIP customer (γ = 0.3)
[0019] Final Urgency = 0.95*0.5 + 0.79*0.4 + 0.3*0.1 = 0.816
[0020] System response:
[0021] Play soothing words immediately: "I'm sorry for the inconvenience, we will prioritize your processing..."
[0022] Push to the agent terminal: ① User profile ② Similar problem solution ③ Emotional relief words prompt BRIEF DESCRIPTION OF DRAWINGS
[0023] Attached Figure 1 : System architecture diagram (showing voice input → ASR / emotion double channel → decision engine → response output)
[0024] Attached Figure 2 : Emotional feature extraction flowchart (acoustic preprocessing → feature extraction → emotion classification).
Claims
1. A method for call dynamic response by fusing ASR and emotion recognition, characterized in that The method comprises the following steps: Real-time receiving user voice stream, parallel executing ASR text conversion and emotion feature extraction; Generating composite decision parameters based on semantic analysis and emotion score, which is calculated by acoustic features (fundamental frequency, short-time energy) and language features (keyword frequency).
2. Dynamically selecting response mode according to decision parameters: regular robot response, emotional dialogue adjustment or manual agent transfer.
3. The method of claim 1, wherein: The emotion feature extraction adopts a CNN-LSTM hybrid model, and the input features include MFCC, F0 and speech speed variance.
4. The method of claim 1, wherein: When the emotion score is below the threshold and the ASR recognizes complaint keywords, manual agent transfer is triggered preferentially and user historical service records are pushed. A system for implementing the method of claims 1-3, comprising: a voice receiving module (101), an ASR engine (102), an emotion analysis module (103); a dynamic decision engine (104) for weighted fusion of semantic intent and emotion label; a multi-modal response generator (105) supporting collaborative output of voice, text and visual interface.
Citation Information
Cited By
AI question and answer expert model construction method and system based on psychological counseling and medium
CN121171262A
Psychological counseling-based ai question and answer expert model construction method and system, and medium
CN121171262B