Customer service method and system based on AI digital human
Through the cross-modal feature synchronization engine and multi-scale timing convolution network, the strategy decision tree for dynamic emotional state and service needs is constructed, which solves the problem of insufficient multimodal perception in the existing customer service system, and realizes efficient and personalized multimodal response, improving user satisfaction and interaction nature.
Patent Information
- Application Number
- CN202510905765.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-02
AI Technical Summary
The existing customer service system lacks multimodal perception and unified modeling, and it is difficult to dynamically adjust the interaction strategy, resulting in one-sided situational understanding, lack of targeted response behavior, and a reduction in user sentiment escalation or satisfaction.
The cross-modal feature synchronization engine generates space-time-aligned fusion data packets, combines multi-scale time-sequence convolution networks and pre-trained emotion encoders to realize high spatio-temporal consistency processing of multimodal information, builds a strategy decision tree for dynamic emotional states and service needs, and generates multimodal responses to match user emotions and needs.
It realizes refined and low-latency perception of user emotions and service intentions, improves the personalized service capabilities and user satisfaction of AI digital people, and enhances the naturalness and adaptability of interaction.
Smart Images

Figure CN120406893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a customer service method and system based on an AI digital human. Background Art
[0002] With the rapid development of artificial intelligence, natural language processing, and computer vision technologies, AI digital humans have gradually been introduced into the customer service field to replace traditional human customer service and rule-based robots, achieving 7×24-hour automated response, problem guidance, and service execution. Especially with the support of speech recognition, multimodal interaction, and deep learning emotion analysis models, AI digital humans already have a certain degree of emotion recognition and semantic understanding capabilities, providing a new path for improving service efficiency and user experience. On this basis, integrating multimodal information such as speech, images, and physiology to achieve service responses closer to the human perception mechanism has become an important development trend of intelligent customer service systems.
[0003] However, there are still several key problems to be solved in the prior art. First, most customer service systems mainly rely on single-modal input, lacking the collaborative perception and unified modeling of speech, vision, and physiological signals, resulting in one-sided situation understanding and lack of pertinence in response behavior. Second, existing AI customer services often adopt static response logic, ignoring the coupling relationship between the dynamic changes of user emotions and the urgency of services, and it is difficult to flexibly adjust interaction strategies, leading to the escalation of user emotions or a decrease in service satisfaction. In addition, speech and facial generation mechanisms usually rely only on content-driven, lacking a parameter regulation mechanism linked to emotional states, and it is difficult to present a real and natural digital human interaction experience. Summary of the Invention
[0004] Based on the above purposes, the present invention provides a customer service method and system based on an AI digital human, providing an AI digital human customer service method and system with multimodal perception, dynamic emotion modeling, and flexible strategy control capabilities to meet user service needs in a more accurate, efficient, and emotionally affinity way.
[0005] The customer service method based on an AI digital human includes the following steps:
[0006] S1: Receive multimodal interaction data input by the user, and generate a spatio-temporally aligned fusion data packet through a cross-modal feature synchronization engine;
[0007] S2: Perform real-time emotional state tracking on the fusion data packet, output dynamic emotional labels including emotional intensity values and emotional type encodings, and simultaneously extract service demand features to generate a demand vector;
[0008] S3: Input the dynamic emotion label and the demand vector into a policy decision tree to generate a policy identifier including an emotion compensation instruction, where the policy identifier includes an emotion soothing coefficient and a response urgency;
[0009] S4: Drive the AI digital human to generate a multimodal response according to the policy identifier, where the multimodal response includes micro-expression parameters positively correlated with the emotion intensity value and speech prosody features matching the emotion soothing coefficient.
[0010] Optionally, S2 includes:
[0011] S11: Receive multimodal interaction data from a user terminal, where the multimodal interaction data includes a speech stream, a visual stream, and auxiliary sensor data;
[0012] S12: Perform feature decoupling on the speech stream, extract a Mel-frequency cepstral coefficient sequence and a fundamental frequency trajectory, and generate a speech feature sequence;
[0013] S13: Perform dynamic partitioning processing on the visual stream, extract a facial action unit intensity matrix and a line-of-sight focus coordinate, and generate a visual feature sequence.
[0014] Optionally, S1 further includes:
[0015] S14: Analyze physiological indexes of the auxiliary sensor data to generate a physiological feature vector, including a skin conductance response value and a heart rate variability;
[0016] S15: Input the speech feature sequence, the visual feature sequence, and the physiological feature vector into a cross-modal feature synchronization engine, perform time series calibration through a timestamp alignment algorithm, and generate a multimodal feature set after timestamp alignment;
[0017] S16: Perform cross-modal embedding representation on the multimodal feature set after timestamp alignment, fuse features using a gated attention mechanism, and generate a spatio-temporally aligned fusion data packet.
[0018] Optionally, S2 includes:
[0019] S21: Perform real-time emotion state tracking on the spatio-temporally aligned fusion data packet, extract emotion features through a multi-scale time series convolutional network, and generate a primary emotion state vector;
[0020] S22: Quantify the emotion intensity of the primary emotion state vector, calculate the current emotion fluctuation index and map it to the [0,1] interval, and output an emotion intensity value;
[0021] S23: Perform emotion type classification based on the primary emotion state vector, and generate a discretized emotion type code through a pre-trained emotion encoder.
[0022] Optionally, S2 further includes:
[0023] S24: Bind the emotional intensity value with the emotional type code, and generate a dynamic emotional label with a time stamp by combining the time stamp;
[0024] S25: Synchronously analyze the service requirements of the spatio-temporally aligned fusion data packet, extract the intention features, key entities and service urgency indicators, and generate a dimension-normalized requirement vector.
[0025] Optionally, S3 includes:
[0026] S31: Input the dynamic emotional label and the requirement vector into the policy decision tree, perform pattern matching through the emotional state-requirement association matrix, and generate a primary policy instruction set;
[0027] S32: Perform emotional compensation requirement analysis on the primary policy instruction set, calculate the deviation degree between the emotional intensity value and the emotional baseline, and generate an emotional compensation instruction;
[0028] S33: Perform policy optimization based on the emotional compensation instruction, and calculate the emotion soothing coefficient through a fuzzy logic controller.
[0029] Optionally, S3 further includes:
[0030] S34: Combine the service urgency indicator in the requirement vector and the emotional intensity value in the dynamic emotional label, and calculate the response urgency;
[0031] S35: Encode the emotional compensation instruction, the emotion soothing coefficient and the response urgency to generate a policy identifier including the operation type, intensity parameter and time limit constraint.
[0032] Optionally, S4 includes:
[0033] S41: Analyze the policy identifier, and extract the operation type code, the emotion soothing coefficient and the response urgency;
[0034] S42: Generate voice prosody control parameters according to the emotion soothing coefficient, calculate the fundamental frequency adjustment amount and the speech rate adjustment factor, and form voice prosody features;
[0035] S43: Generate a facial muscle movement instruction set based on the operation type code and the emotional intensity value in the policy identifier, and calculate micro-expression parameters positively correlated with the emotional intensity value.
[0036] Optionally, S4 further includes:
[0037] S44: Input the speech prosody features and micro-expression parameters into the AI digital human behavior engine, and generate a temporally aligned multi-modal response through the multi-modal synchronization controller;
[0038] S45: Adjust the output priority of the multi-modal response according to the response urgency. If the response urgency ≥ 0.7, interrupt the current interaction thread and immediately output the response.
[0039] A customer service system based on an AI digital human is used to implement the above-mentioned customer service method based on an AI digital human, and includes the following modules:
[0040] Data acquisition module: used to receive multi-modal interaction data from the user terminal. The multi-modal interaction data includes voice streams, visual streams, and auxiliary sensor data, and generate a preliminary multi-modal feature sequence;
[0041] Cross-modal feature synchronization module: used to perform timestamp alignment and gated attention fusion on the multi-modal feature sequence to generate a spatio-temporally aligned fusion data packet;
[0042] Emotional state tracking and demand analysis module: used to extract a primary emotional state vector, calculate an emotional intensity value and an emotional type code based on the fusion data packet, generate a dynamic emotional label, and extract intention features, key entities, and service urgency indicators to construct a demand vector;
[0043] Policy decision-making generation module: used to input the dynamic emotional label and demand vector into a policy decision tree to generate a primary policy instruction set, calculate an emotional soothing coefficient and a response urgency, and finally output a policy identifier including an operation type, an intensity parameter, and a time limit constraint;
[0044] Multi-modal response generation module: used to drive the AI digital human to generate a multi-modal response according to the policy identifier. The multi-modal response includes micro-expression parameters and speech prosody features to achieve coordinated control of emotional soothing and service response.
[0045] Advantages of the present invention:
[0046] In the present invention, through the constructed cross-modal feature synchronization engine, which fuses voice streams, visual streams, and auxiliary sensor data, and adopts a timestamp alignment algorithm and a gated attention mechanism to achieve high spatio-temporal consistency processing of multi-modal information, effectively alleviating the problem of insufficient situational understanding of traditional single-modal perception. Via the multi-scale temporal convolutional network in S2 to extract emotional features, combined with a pre-trained emotional encoder and an emotional fluctuation quantization model, the system can dynamically generate emotional labels with time stamps and standardized demand vectors, ensuring fine-grained and low-latency perception of user emotions and service intentions.
[0047] In the present invention, through the constructed emotion-demand coupling analysis mechanism, by integrating the emotion intensity value, emotion type coding, and service urgency index, a strategy decision tree is constructed and a strategy identifier is generated, which includes multi-layer parameters such as operation type, emotion soothing coefficient, and response urgency. This mechanism breaks through the rigid constraints of the traditional static response process, realizes the generation of personalized service paths centered on the user's real emotional state, enables the AI digital human to have the ability to actively adjust the rhythm, style, and response mode, and effectively improves user satisfaction and trust.
[0048] In the present invention, based on the strategy identifier, the AI digital human is driven to generate multi-modal responses. The amplitude of micro-expression parameters is controlled by the emotion soothing coefficient, and the speech prosody features (including speech rate, pitch, pause interval, etc.) are dynamically adjusted in a positive correlation with the emotion intensity value, so as to generate an interaction form highly matching the user's current psychological state. Compared with the deficiencies of single response mode and rigid feedback in the prior art, the proposed method can significantly improve the emotional adaptability and interaction naturalness of the AI digital human in complex service scenarios without increasing additional labor costs. Brief Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only those of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a schematic diagram of the S1 process of the embodiment of the present invention;
[0051] Figure 2 It is a schematic diagram of the S2 process of the embodiment of the present invention;
[0052] Figure 3 It is a schematic diagram of the system process of the embodiment of the present invention. Detailed Embodiments
[0053] The following will describe the present invention in detail with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawing part is only for more specifically describing the embodiments, and is not intended to specifically limit the present invention.
[0054] As Figure 1 - Figure 2 shown, the customer service method based on the AI digital human includes the following steps:
[0055] S1: Receive the multi-modal interaction data input by the user, and generate a spatio-temporally aligned fusion data packet through a cross-modal feature synchronization engine;
[0056] S2: Perform real-time emotional state tracking on the fusion data packet, output dynamic emotional labels including emotional intensity values and emotional type encodings, and at the same time extract service demand features to generate a demand vector;
[0057] S3: Input the dynamic emotional label and the demand vector into a policy decision tree to generate a policy identifier including an emotional compensation instruction, and the policy identifier includes an emotion soothing coefficient and a response urgency;
[0058] S4: Drive the AI digital person to generate a multi-modal response according to the policy identifier, and the multi-modal response includes micro-expression parameters positively correlated with the emotional intensity value and speech prosody features matching the emotion soothing coefficient.
[0059] S1 includes:
[0060] S11, Receive multi-modal interaction data: The system first receives multi-modal interaction data from the user terminal, and these data include a speech stream, a visual stream, and auxiliary sensor data.
[0061] Specifically, the speech stream is collected by the user through the terminal microphone, and the visual stream is obtained by collecting facial video data through the camera; the auxiliary sensor data is collected by a heart rate sensor and a skin conductance sensor device to provide physiological information.
[0062] For example, in a certain user interaction scenario, the terminal simultaneously records the audio data when the user is speaking, the captured facial dynamic video, and the real-time heart rate and skin conductance data, and these raw data will be used as the input for subsequent processing.
[0063] S12, Speech stream feature decoupling: Perform feature decoupling processing on the received speech stream, mainly including the extraction of two parts of features:
[0064] Mel-frequency cepstral coefficient sequence extraction: After pre-emphasizing, framing, windowing, and performing a fast Fourier transform (FFT) on the speech stream to obtain spectral information, use a Mel filter bank to convert the spectrum to the Mel scale, and then obtain the Mel-frequency cepstral coefficient (MFCC) sequence through a discrete cosine transform (DCT).
[0065] In addition, the system also extracts the fundamental frequency trajectory from the speech stream, and calculates the fundamental frequency value of each frame using an autocorrelation algorithm or other fundamental frequency detection algorithms;
[0066] Generate a speech feature sequence: After integrating the above Mel-frequency cepstral coefficient sequence and the fundamental frequency trajectory, form a speech feature sequence describing the characteristics of the speech signal to provide sufficient speech information for subsequent fusion;
[0067] S13, Visual Flow Dynamic Partitioning Processing: After the visual flow data is collected, the system performs dynamic partitioning processing on the video data. The specific steps include:
[0068] (1) Facial Region Detection and Partitioning: Use a face detection algorithm to locate the user's facial region, and then perform regional processing on the facial image.
[0069] (2) Facial Action Unit Intensity Matrix Extraction: For each region, use a facial expression recognition algorithm (such as a method based on a convolutional neural network) to extract the intensity information of facial action units (Action Units, AUs), forming a matrix representation:
[0070] ;
[0071] where represents the intensity of the th action unit in the th region;
[0072] (3) Gaze Focus Coordinate Extraction: Use a gaze detection algorithm to extract the two-dimensional coordinates of the user's fixation points in the video , forming a sequence of gaze focus coordinates;
[0073] These two parts of information together constitute the visual feature sequence, providing key visual information for subsequent multimodal fusion;
[0074] S14, Physiological Index Analysis of Auxiliary Sensor Data: For the auxiliary sensor data, the system sequentially preprocesses the collected physiological signals, including filtering, denoising, and normalization. Subsequently, key physiological features are extracted to generate a physiological feature vector. Specifically, it includes:
[0075] Skin Conductance Response Value (EDA): Extract the instantaneous conductance level from the conductance sensor data.
[0076] Heart Rate Variability (HRV): Calculate the difference in the intervals between adjacent heartbeats from the heart rate sensor data, and further analyze short-term and long-term heart rate variability. HRV is calculated as:
[0077] ;
[0078] where represents the th heartbeat interval;
[0079] After combining the skin conductance response value and heart rate variability, a physiological feature vector is formed to describe the user's current physiological state.
[0080] S14, Timestamp Alignment and Temporal Calibration: When collecting voice streams, visual streams, and auxiliary sensor data, each is accompanied by its own timestamp. To achieve spatio-temporal alignment of cross-modal data, the system uses a timestamp alignment algorithm to calibrate each feature sequence. The specific steps are as follows:
[0081] (1)Data Timestamp Annotation: Establish a timestamp index sequence for the voice stream: ;
[0082] Establish a corresponding timestamp index sequence for the visual stream: , where is the cross-device latency compensation amount. The auxiliary sensor data is also marked with timestamps ;
[0083] (2)Timestamp Alignment Condition: When the voice frame timestamp , visual frame timestamp , and physiological data timestamp meet the following conditions:
[0084] ;
[0085] Then this set of data is considered to be consistent in time, and the multi-modal feature set after timestamp alignment can be output.
[0086] Example: Assume that during a data collection process, the timestamp of a certain frame of voice data is 1000ms, and the timestamp of the corresponding visual stream data is (for example, ), and the timestamp recorded by the auxiliary sensor is 1005ms. After calculating:
[0087] Since 10ms is less than or equal to 20ms, this set of data meets the alignment condition.
[0088] S15, Cross-modal Embedding Representation and Feature Fusion: For the multi-modal feature set after timestamp alignment, it is further processed using a cross-modal feature synchronization engine, mainly including:
[0089] (1)Cross-modal Embedding Representation: Embed the voice feature sequence, visual feature sequence, and physiological feature vector into the same feature space through non-linear mapping respectively. Assuming the mapping functions are , , and , then we get respectively:
[0090] Voice feature sequence), (Visual feature sequence), (Physiological feature vector);
[0091] (2) Gated Attention Mechanism for Feature Fusion: To fully integrate information from various modalities, a gated attention mechanism is introduced to fuse the embedding vectors of different modalities. Let the gating coefficients of each modality be , and the following calculation formula can be adopted:
[0092] ;
[0093] where represents the Sigmoid activation function, and are the corresponding trainable parameters, and the finally fused feature representation is:
[0094] ;
[0095] where represents element-wise multiplication.
[0096] S16, Generating Spatiotemporally Aligned Fusion Data Packets: After combining the feature vector fused by the gated attention mechanism with the aligned timestamp information, spatiotemporally aligned fusion data packets are formed, providing unified and accurate multi-modal data support for subsequent emotional state tracking and policy decision-making.
[0097] Through the above process, it can be ensured that data from different acquisition devices are accurately aligned spatiotemporally, and effective fusion of features from each modality is achieved through embedding representation and attention mechanism, thus constructing high-quality fusion data packets and providing a solid data foundation for the downstream AI digital human customer service capabilities.
[0098] S2 includes:
[0099] S21 Real-time Emotional State Tracking: The system first performs real-time emotional state tracking on the input spatiotemporally aligned fusion data packets, and uses a multi-scale temporal convolutional network to model the temporal patterns of the multi-modal feature sequences to extract emotional features with time dependence.
[0100] The multi-scale temporal convolutional network contains multiple convolutional kernels with different receptive field sizes to cover short-term, medium-term, and long-term dependence features, and the network output is the primary emotional state vector:
[0101] where is the multi-modal fusion feature vector in the spatiotemporally aligned fusion data packet, represents the primary emotional state vector.
[0102] Example: When the user's voice is excited, the facial AUs tension increases, and the EDA suddenly rises, the MTCN will capture this set of changes and at is reflected as a high-amplitude fluctuation pattern;
[0103] S22, Emotional intensity quantification: The system estimates the intensity of the primary emotional state vector using a weighted emotional fluctuation index model to calculate the current emotional fluctuation index , and maps it to the standardized interval [0,1] to output the emotional intensity value , expressed as:
[0104] ;
[0105] where, is the weight of the th emotional dimension, is the static mean of the emotional dimension, is the normalization upper limit value, represents the emotional intensity value.
[0106] Example: When the system detects significant voice jitter, sudden changes in AU intensity, and an increase in heart rate, the fluctuation index , if , then , indicating a medium-high intensity emotional state.
[0107] S23, Emotional type classification: After obtaining the primary emotional state vector, the system performs discrete classification of emotional types through a pre-trained emotional encoder. This encoder is a type of deep learning model (BiLSTM classifier) that can map the input vector to a preset emotional category.
[0108] The pre-trained model is , then there is: ;
[0109] where: represents the emotional type encoding, is the total number of emotional categories (such as joy, anger, anxiety, satisfaction).
[0110] The output is a discrete encoding for subsequent policy mapping.
[0111] Example: In a certain interaction, the emotional type encoding output by the system is , indicating that the user is currently in an "anxious" state.
[0112] S24, Dynamic emotional label generation: The system binds the aforementioned emotional intensity value with the emotional type encoding , and combines it with the interaction timestamp to form a dynamic emotional label with a time-effective mark :
[0113] ;
[0114] This dynamic sentiment label is used in subsequent policy generation to real-time regulate the behavior performance of the AI digital human.
[0115] Example output: , indicating that the user is in an anxious state at 10:21:33, and the emotional intensity is 0.72.
[0116] S25, Service demand parsing and demand vector generation: The system simultaneously performs service demand parsing on the input spatio-temporally aligned fusion data packet, mainly including:
[0117] Intention feature extraction: Use a semantic recognition model (such as BERT) to identify service intentions in the user input (such as "query bill", "repair", "cancel order", etc.);
[0118] Key entity recognition: Extract key business information from the input (such as account number, device type, amount);
[0119] Service urgency index calculation: Used to characterize the service response priority, and the specific calculation method is as follows:
[0120] (1) Keyword-triggered emergency weight increase: If keywords in the keyword set "urgent", "immediately", "right away" are detected in the input text, then: ;
[0121] (2) Basic business type assignment: If the service type belongs to high-priority tasks (such as "fault repair", "payment failure"), then the basic urgency: ;
[0122] (3) Conversation turn duration correction factor: If the current conversation turn time is , and the standard response duration is , then the urgency increment is:
[0123] ;
[0124] Example calculation:
[0125] For the initial business type "payment failure", we get ;
[0126] If the current turn takes 18s and the standard duration is 12s, then:
[0127] ;
[0128] Finally, standardize and encode the intention features, key entities, and urgency indicators to generate a demand vector with consistent dimensions: ;
[0129] S3 includes:
[0130] S31, generation of emotion-demand coupling strategy: First, input the dynamic emotion label and the demand vector into the policy decision tree for joint analysis. The policy decision tree adopts a conditional branch structure and formulates a preliminary behavior response path under different emotion type encodings and service types. The core of policy decision-making is to construct an emotion-demand coupling mapping relationship:
[0131] ;
[0132] Among them, the output primary policy instruction set includes the initially set operation types (such as "soothing tone", "active questioning", "process simplification") and their candidate execution parameters.
[0133] Example: If (anxiety), and the service urgency in the demand vector is high (for example ), then generate policy instructions such as: Type: Soothe + Accelerate, candidate speech rate , guiding operation ;
[0134] S32, generation of emotion deviation detection and compensation instructions: The system performs emotion deviation detection on the emotion response part in the primary policy instruction set, judges the difference between the current emotional state and the user's individualized emotional baseline , and then generates emotion compensation instructions ;
[0135] Define emotion deviation as: ;
[0136] If (threshold, preset to 0.2), then the system triggers the compensation strategy and outputs the compensation type (mitigation / enhancement), the compensation target module (voice / expression), and the applicable duration.
[0137] Example: If , while Generate compensation instructions: Mitigation, module: voice ;
[0138] S33, calculation of emotion soothing coefficient: Based on the emotion compensation instructions, the system further calculates a quantified emotion soothing coefficient , as the basis for dynamic parameter adjustment in digital human response regulation.
[0139] The soothing coefficient consists of the following three factors: the current emotional intensity value , the emotional type code , and the emotional trend derivative ;
[0140] The soothing coefficient calculation function is:
[0141] ;
[0142] Among them: Encode( ) is the influence weight of the emotional type (e.g., anxiety is positive, calm is negative), is the weighted coefficient obtained from experience or training.
[0143] Example: If , (anxiety, corresponding encoding value 0.5), , then set , ;
[0144] Get: ;
[0145] This coefficient will be used as the core adjustment parameter to drive the control of the micro-expressions, intonation, etc. of the Al digital human.
[0146] S34, Response urgency calculation: The system further combines the service urgency indicator with the emotional intensity value to calculate the comprehensive response urgency , which is used to adjust control strategies such as the response speed and speech density of the digital human.
[0147] The designed calculation function is:
[0148] ;
[0149] Among them, are the weighted factors of service orientation and emotional orientation respectively (for example ;
[0150] Example: If , then:
[0151] ;
[0152] The system accordingly raises the response level and sets the priority response strategy.
[0153] S35, Policy identifier generation final: Uniformly encode the above-generated emotional compensation instructions, emotional soothing coefficients, and response urgencies to generate a policy identifier ;
[0154] The policy identifier structure includes the following three parts:
[0155] 1. Operation type: such as soothing intonation, task simplification, active guidance;
[0156] 2. Intensity parameter: including intonation rhythm factor, micro-expression amplitude, etc., controlled by the emotion soothing coefficient;
[0157] 3. Time limit constraint: Set the response delay threshold and timeout warning parameter policy identifier according to the response urgency.
[0158] S4 includes:
[0159] S41, Policy identifier parsing: First, perform a parsing operation on the policy identifier to extract three core parameters: operation type code, emotion soothing coefficient, and response urgency. Among them:
[0160] The operation type code is used to indicate the execution method of the digital human behavior (such as voice output, facial expression, or action feedback);
[0161] The emotion soothing coefficient is derived from the calculation result of the previous step and reflects the adjustment intensity in the current emotional state;
[0162] The response urgency is jointly evaluated based on the user demand urgency and the emotion intensity value, and determines the output priority of the interaction response.
[0163] S42, Generation of speech prosody control parameters: According to the extracted emotion soothing coefficient, generate the prosody control parameters required for voice output, mainly including fundamental frequency adjustment amount and speech rate adjustment factor. Its calculation model is as follows:
[0164] According to the extracted emotion soothing coefficient, generate the prosody control parameters required for voice output, mainly including the fundamental frequency adjustment amount and the speech rate adjustment factor , and the calculation model is as follows:
[0165] ;
[0166] Among them, is the prosody regulation coefficient set by the system.
[0167] Example: If the emotion soothing coefficient is 0.6 and the emotion intensity rising trend is 0.3;
[0168] and , then:
[0169] ;
[0170] ;
[0171] It means that the fundamental frequency of the voice is increased by 15 Hz and the speech rate is decreased to 48.7% of the original.
[0172] S43, Generation of Facial Muscle Movement Instruction Set: Based on the operation type code and emotional intensity value in the policy identifier, generate the corresponding facial muscle movement instruction set, and calculate the set of micro-expression parameters that is positively correlated with the emotional intensity value.
[0173] The set of micro-expression parameters is defined as follows:
[0174] ;
[0175] Among them, represents the amount of facial muscle movement in the i-th group, is the emotional weight corresponding to each group of muscles.
[0176] The system determines the type of expression (such as happy, surprised, angry, etc.) according to the operation type code, and then linearly amplifies the muscle movement amplitude by the emotional intensity to form a facial feedback.
[0177] S44, Multimodal Synchronous Control: Input the speech prosody features generated in the above steps and the set of micro-expression parameters into the AI digital human behavior engine together. The multimodal synchronous controller integrates them, and by aligning the speech output time axis and the facial expression control signal, realizes the generation of a time-synchronized multimodal response.
[0178] The controller internally uses a sliding window mechanism to predict and correct the delays of each channel, ensuring the synchronous display of speech changes and facial movements, and enhancing the user interaction immersion.
[0179] S45: Priority scheduling driven by response urgency;
[0180] According to the response urgency η extracted from the policy identifier, dynamically adjust the response priority of the current interaction thread. When the response urgency satisfies ;
[0181] The system immediately interrupts the current low-priority interaction thread, inserts the current response content into the front of the execution queue, and immediately outputs the response through the AI digital human;
[0182] Example: If a voice broadcast of the scenario content is currently in progress and the user suddenly issues a high-emotion intensity + high-urgency request (such as "Stop immediately"), resulting in η = 0.83, the system will immediately interrupt the introduction speech, give priority to outputting an interruption response (such as "Okay, I will stop now") and synchronously adjust the expression to be serious or alert.
[0183] Such as Figure 3As shown, the customer service system based on the AI digital human is used to implement the above-mentioned customer service method based on the AI digital human, and includes the following modules:
[0184] Data acquisition module: It is used to receive multimodal interaction data from user terminals. The multimodal interaction data includes voice streams, visual streams, and auxiliary sensor data, and generates a preliminary multimodal feature sequence;
[0185] Cross-modal feature synchronization module: It is used to perform timestamp alignment and gated attention fusion on the multimodal feature sequence to generate a spatio-temporally aligned fusion data packet;
[0186] Emotional state tracking and demand analysis module: It is used to extract a primary emotional state vector, calculate the emotional intensity value and emotional type coding based on the fusion data packet, generate a dynamic emotional label, and extract intention features, key entities, and service urgency indicators to construct a demand vector;
[0187] Policy decision generation module: It is used to input the dynamic emotional label and the demand vector into the policy decision tree to generate a primary policy instruction set, calculate the emotion soothing coefficient and response urgency, and finally output a policy identifier including the operation type, intensity parameter, and time limit constraint;
[0188] Multimodal response generation module: It is used to drive the AI digital human to generate a multimodal response according to the policy identifier. The multimodal response includes micro-expression parameters and speech prosody features to achieve the coordinated control of emotion soothing and service response.
[0189] The present invention covers any substitutions, modifications, equivalent methods, and solutions made within the essence and scope of the present invention. To enable the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention. However, those skilled in the art can fully understand the present invention without these detailed descriptions. In addition, well-known methods, processes, procedures, components, and circuits are not described in detail to avoid unnecessary confusion to the essence of the present invention.
[0190] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. A customer service method based on an AI digital human, characterized in that, It includes the following steps: S1: Receive the multi-modal interaction data input by the user, and generate a spatio-temporally aligned fusion data packet through a cross-modal feature synchronization engine; S2: Perform real-time emotional state tracking on the fusion data packet, output dynamic emotional labels including emotional intensity values and emotional type encodings, and at the same time extract service demand features to generate a demand vector; S3: Input the dynamic emotional label and the demand vector into a policy decision tree to generate a policy identifier including an emotional compensation instruction, and the policy identifier includes an emotional soothing coefficient and a response urgency; S4: Drive an AI digital person to generate a multi-modal response according to the policy identifier, and the multi-modal response includes micro-expression parameters positively correlated with the emotional intensity value and speech prosody features matching the emotional soothing coefficient.
2. The customer service method based on an AI digital human according to claim 1, wherein The S2 includes: S11: Receive multi-modal interaction data from a user terminal, and the multi-modal interaction data includes a speech stream, a visual stream, and auxiliary sensor data; S12: Perform feature decoupling on the speech stream, extract a Mel-frequency cepstral coefficient sequence and a fundamental frequency trajectory, and generate a speech feature sequence; S13: Perform dynamic partitioning processing on the visual stream, extract a facial action unit intensity matrix and a line-of-sight focus coordinate, and generate a visual feature sequence.
3. The customer service method based on an AI digital human according to claim 2, wherein The S1 further includes: S14: Analyze the physiological indexes of the auxiliary sensor data to generate a physiological feature vector, including a skin conductance response value and a heart rate variability; S15: Input the speech feature sequence, the visual feature sequence, and the physiological feature vector into a cross-modal feature synchronization engine, and perform time series calibration through a timestamp alignment algorithm to generate a multi-modal feature set with timestamp alignment; S16: Perform cross-modal embedding representation on the multi-modal feature set with timestamp alignment, and fuse features using a gated attention mechanism to generate a spatio-temporally aligned fusion data packet.
4. The customer service method based on an AI digital human according to claim 3, wherein The S2 includes: S21: Perform real-time emotional state tracking on the spatio-temporally aligned fusion data packet, extract emotional features through a multi-scale time series convolutional network, and generate a primary emotional state vector; S22: Quantify the emotional intensity of the primary emotional state vector, calculate the current emotional fluctuation index and map it to the [0,1] interval, and output the emotional intensity value; S23: Perform emotional type classification based on the primary emotional state vector, and generate a discretized emotional type encoding through a pre-trained emotional encoder.
5. The customer service method based on an AI digital human according to claim 4, wherein The S2 further includes: S24: Bind the emotional intensity value and the emotional type encoding, and generate a dynamic emotional label with a time stamp in combination with the time stamp; S25: Synchronously perform service demand analysis on the spatio-temporally aligned fusion data packet, extract intention features, key entities, and service urgency indicators, and generate a dimension-normalized demand vector.
6. The customer service method based on an AI digital human according to claim 5, wherein The S3 includes: S31: Input the dynamic emotional label and the demand vector into a policy decision tree, perform pattern matching through an emotional state-demand association matrix, and generate a primary policy instruction set; S32: Analyze the emotional compensation demand of the primary policy instruction set, calculate the deviation degree between the emotional intensity value and the emotional baseline, and generate an emotional compensation instruction; S33: Based on the optimization of the emotional compensation instruction execution strategy, calculate the emotion soothing coefficient through a fuzzy logic controller.
7. The customer service method based on an AI digital human according to claim 6, wherein S3 further includes: S34: Combine the service urgency indicator in the demand vector and the emotion intensity value in the dynamic emotion label to calculate the response urgency; S35: Encode the emotional compensation instruction, emotion soothing coefficient, and response urgency into a strategy to generate a strategy identifier that includes the operation type, intensity parameter, and time limit constraint.
8. The customer service method based on an AI digital human according to claim 7, wherein, S4 includes: S41: Analyze the strategy identifier, extract the operation type code, emotion soothing coefficient, and response urgency; S42: Generate voice prosody control parameters according to the emotion soothing coefficient, calculate the fundamental frequency adjustment amount and speech rate adjustment factor, and form voice prosody features; S43: Based on the operation type code and emotion intensity value in the strategy identifier, generate a facial muscle movement instruction set and calculate micro-expression parameters positively correlated with the emotion intensity value.
9. The customer service method based on an AI digital human according to claim 8, wherein S4 further includes: S44: Input the voice prosody features and micro-expression parameters into the AI digital human behavior engine, and generate a time-aligned multimodal response through a multimodal synchronization controller; S45: Adjust the output priority of the multimodal response according to the response urgency. If the response urgency ≥ 0.7, interrupt the current interaction thread and immediately output the response.
10. An AI digital human-based customer service system for implementing the AI digital human-based customer service method according to any one of claims 1-9, characterized in that, It includes the following modules: Data acquisition module: Used to receive multimodal interaction data from the user terminal. The multimodal interaction data includes voice streams, visual streams, and auxiliary sensor data, and generate a preliminary multimodal feature sequence; Cross-modal feature synchronization module: Used to perform timestamp alignment and gated attention fusion on the multimodal feature sequence to generate a spatio-temporally aligned fusion data packet; Emotional state tracking and demand analysis module: Used to extract the primary emotional state vector, calculate the emotion intensity value and emotion type encoding based on the fusion data packet, generate a dynamic emotion label, and extract intention features, key entities, and service urgency indicators to construct a demand vector; Strategy decision-making and generation module: Used to input the dynamic emotion label and demand vector into a strategy decision tree, generate a primary strategy instruction set, calculate the emotion soothing coefficient and response urgency, and finally output a strategy identifier that includes the operation type, intensity parameter, and time limit constraint; Multimodal response generation module: Used to drive the AI digital human to generate a multimodal response according to the strategy identifier. The multimodal response includes micro-expression parameters and voice prosody features to achieve the coordinated control of emotion soothing and service response.
Citation Information
Patent Citations
Interactive digital human driving method
CN119273818A
Digital human emotion perception and dynamic response method, system and device and storage medium
CN119884327A
DC series arc filure diagnosis apparatus using artificial machine learning
KR1020250120690A
KR20230151162A
Cited By
Digital human live broadcast voice interaction system fused with emotion calculation
CN121393436A
Digital human live voice interaction system fused with affective computing
CN121393436B
Digital customer service information generation method based on multi-modal information, medium and equipment
CN121561823A
Intelligent customer service interaction method and system based on AI large model
CN121903622A