Personalized customer portrait label generating and updating system based on voice interaction behavior
By employing cross-modal deep fusion and dynamic update mechanisms, the problems of single modality and lagging labels in customer profiles have been solved, enabling accurate mapping of customer intentions and emotions and business guidance, thereby improving the accuracy and real-time nature of customer profiles.
Patent Information
- Application Number
- CN202511762876.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-02-06
AI Technical Summary
Existing customer profiling technologies are too simplistic and fail to capture nuances. Emotional tags are disconnected from business logic, and tag updates are slow and lack dynamism, resulting in distorted profiles and an inability to guide business decisions.
We employ cross-modal deep fusion technology, utilize forced alignment and cross-attention mechanisms to analyze speech prosodic features, construct an emotion-business two-layer mapping mechanism, and achieve dynamic label updates through physical decay and Bayesian inference.
It enables the accurate capture of complex customer intentions, generates tags with direct business guidance significance, dynamically tracks the psychological flow of customers, and improves the accuracy of profiles and business guidance capabilities.
Smart Images

Figure CN121483264A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, big data analysis, and human-computer interaction. Specifically, it relates to a system and method that utilizes natural language processing (NLP) and digital signal processing (DSP) technologies to deeply mine multimodal data during voice interaction, deeply integrate voice features with text features, and realize real-time generation of customer profile tags, business logic mapping, dynamic evolution of weights, and long-term and short-term hierarchical management. Background Technology
[0002] In modern customer relationship management (CRM), intelligent contact centers (ICC), and telemarketing scenarios, user profiling is the foundation for achieving personalized services. Traditional profiling techniques mainly rely on static structured data (such as gender, age, and account balance) and non-real-time historical text records (such as work order records and chat logs).
[0003] However, the applicant's research has revealed the following significant technical shortcomings in existing portrait technologies: 1. Limited Modality, Inability to Perceive Subtext: Most existing solutions rely solely on text transcribed by Automatic Speech Recognition (ASR) for label extraction. However, psychological research (such as the Mehrabian rule) indicates that over 38% of the information in human communication is contained in paralinguistics such as tone, volume, speech rate, and pauses. Analyzing only the text leads to serious misjudgments of intent. For example, a customer saying "okay" in a short, rising tone indicates pleasant agreement; while a low, drawn-out tone accompanied by a sigh suggests reluctant compromise or underlying dissatisfaction. Existing systems cannot distinguish between these two distinct states, resulting in distorted user profile labels.
[0004] 2. Disconnect between emotion labels and business logic: Existing emotion recognition systems often only output general emotion labels such as "happy," "angry," and "neutral," failing to link these emotion characteristics with specific business contexts. For example, "fast speaking speed + loud volume" represents "anger" in a complaint scenario, but may represent "high purchase intention" in a limited-time flash sale scenario. The lack of this context-based mapping mechanism means that emotion labels cannot directly guide business decisions.
[0005] 3. Delayed and dynamistic tag updates: Traditional profiling typically uses an offline "T+1" update model and lacks a dynamic weight decay mechanism based on real-time interaction intensity. It cannot capture the psychological shifts of customers during a single call (such as from "calm" to "anxious"), nor can it distinguish between "instantaneous state" and "inherent attributes," resulting in "noise pollution" in the profiling database.
[0006] Therefore, there is an urgent need for a dynamic profiling system that can deeply integrate acoustic and semantic features and accurately map multimodal features to real business tags. Summary of the Invention
[0007] The purpose of this invention is to provide a personalized customer profile tag generation and update system based on voice interaction behavior, aiming to solve the problems of "listening to the text but not the voice" and the lack of business relevance of tags in the existing technology.
[0008] The core technical means of this invention are as follows: 1. Cross-modal deep fusion technology: Unlike simple feature splicing, this invention utilizes forced alignment and cross-attention mechanisms to treat the prosodic features of speech as an "interpreter" of text semantics. The system can accurately analyze "how each word is expressed" and use a gate mechanism to decide whether to believe the text or the tone, thereby accurately capturing complex intentions such as "irony" and "hesitation".
[0009] 2. Emotion-Business Dual-Layer Mapping Mechanism: This invention constructs a mapping logic from "physical characteristics" to "psychological state" and then to "business tags." The system no longer simply outputs "emotionally agitated," but combines business scenarios (such as debt collection, sales, and after-sales service) to output tags with direct business guidance significance, such as "high-risk resistance," "urgent purchase," and "potential complaints."
[0010] 3. Dynamic updates based on physical decay and Bayesian inference: An exponential decay function similar to Newton's law of cooling is introduced, which allows labels representing instantaneous states to automatically decay over time; a Bayesian evidence accumulation mechanism is introduced to ensure the robustness of long-term attribute labels. Attached Figure Description
[0011] Figure 1 The diagram shows the overall logical architecture of the system of this invention, illustrating the dual-stream feature extraction and fusion process.
[0012] Figure 2 This is a schematic diagram of the cross-modal attention fusion and gating unit structure.
[0013] Figure 3 A logical association diagram mapping emotional features to business tags.
[0014] Figure 4 This is a schematic diagram illustrating the dynamic decay and accumulation curves of label weights as they change with interactive behavior. Detailed Implementation
[0015] The implementation methods of this invention will be described in detail below with reference to specific mathematical models and application scenarios.
[0016] I. Feature Extraction and Deep Fusion Mechanism of Audio-Text Integration This system adopts a dual-stream parallel and interactive fusion architecture, focusing on solving the problems of "feature extraction" and "feature combination".
[0017] 1. Dual-stream feature extraction: Acoustic Flow: The system not only extracts basic features such as MFCC (Multiple-Cost Combination of Voice Control) but also focuses on capturing microprosodic features. For example, the system calculates the "turn-taking overlap" and "response latency jitter" during conversation turn-by-turn switching. For each syllable, the system extracts the slope of its F0 (fundamental frequency) trajectory. If the F0 slope rises sharply and the energy is concentrated in the 2-4kHz frequency band when a negative word (such as "no" or "not") appears, the system determines it as a "firm rejection".
[0018] Semantic Stream: Speech is transcribed into text using ASR, and semantic vectors are extracted using the BERT model.
[0019] 2. Cross-modal forced alignment and gating fusion: The system introduces a timestamp mapping mechanism. For each word w(i) output by ASR, the system traces back the original audio based on its start and end times [t_start, t_end] and extracts the corresponding acoustic segment.
[0020] Subsequently, a cross-attention model is used, with text vectors as queries and acoustic vectors as keys and values. This means that the system will "query" the corresponding tone intensity based on the semantic content.
[0021] Fusion Example: When a customer says "This service is great" (semantic positive), if the acoustic features show an extremely slow speech rate, monotonous tone, and accompanying sighs (acoustic negative), the gating unit z will capture this modal conflict. The fused feature vector will be biased towards the acoustic features, thus generating a "sarcasm / dissatisfaction" fused feature instead of a simple "satisfaction".
[0022] II. Correlation Mapping Between Emotional Characteristics and Real Business Tags Another major innovation of this invention lies in business logic mapping. The system has multiple built-in "feature-business" mapping templates, which transform abstract signal features into executable business actions.
[0023] 1. Feature mapping in sales scenarios (identifying purchase intent): Feature combination: [Speeding up speech] + [Interrupting machine speech (Barge-in)] + [Mentioning keywords such as "price" or "how much"].
[0024] Traditional labels: excitement, inquiry.
[0025] This system maps to: "Price-sensitive - Urgent buyers".
[0026] Business value: This tag instantly becomes the top priority, triggering the sales script system to jump to the "direct quotation" stage, reducing small talk and increasing the closing rate.
[0027] 2. Feature mapping (risk identification) in risk control and anti-fraud scenarios: Feature combination: [pause for more than 2 seconds before answering key identity questions] + [unstable intonation (high jitter value)] + [frequent use of filler words "um... um..."].
[0028] Traditional labels: hesitation, tension.
[0029] This system is mapped to either "potential fraud risk" or "operation not performed by the user".
[0030] Business value: The system generates high-risk warnings and automatically triggers the "facial recognition" or "human verification" process to prevent business risks.
[0031] 3. Feature mapping in customer care scenarios (churn identification): Feature combination: [volume decreases sentence by sentence] + [falling tone at the end of sentence] + [short response ("Oh", "Okay")].
[0032] Traditional labels: sadness, depression.
[0033] This system maps "service fatigue" or "potential churn risk".
[0034] Business value: Generate retention tags and recommend that customer service immediately use empathy scripts to reassure the customer.
[0035] III. Dynamic Evolution Algorithm for Label Weights 1. Exponential decay of short-term state labels For labels representing a momentary state (such as "emotion: impatient"), their weight should decay rapidly over time to reflect the customer's true current state.
[0036] Calculation formula: W(t) = W(t-1) * exp[-(t - t_last) / τ] + [ A * S_event ] / [ 1 + exp(-k * (S_event - θ)) ] Here, exp[-(t - t_last) / τ] simulates the forgetting curve of human emotions. If the customer calms down in the next few seconds, the label weight quickly returns to zero, and the system will not mistakenly believe that the customer is still "angry".
[0037] 2. Bayesian accumulation of long-term attribute tags Labels representing inherent attributes (such as "personality: decisive") cannot be established based on a single behavior; they require accumulated evidence.
[0038] After each call, the posterior probability is updated using Bayes' theorem: P(H|D_new) = [ P(D_new|H) * P(H_old) ] / P(D_new) The denominator P(D_new) is the expansion of the total probability formula.
[0039] The system only writes the attribute into the long-term profile database when P(H|D_new) has been consistently above the threshold for N consecutive interactions. This mechanism effectively prevents occasional emotions (such as speaking loudly due to a noisy environment) from being misjudged as customer personality traits (such as a loud voice).
[0040] Through the above mechanism, this invention effectively solves the two major pain points of traditional profiling systems: "not understanding the implied meaning" and "labels cannot guide business," and achieves accurate capture and dynamic tracking of customers' deep intentions.
Claims
1. A personalized customer profile tag generation and updating system based on voice interaction behavior, characterized in that, The system includes: The multi-channel high-fidelity speech capture and preprocessing unit is used to access the speech data stream of the client terminal in real time, use adaptive filters for echo cancellation (AEC) and beamforming processing, perform endpoint detection (VAD) based on energy and spectral entropy dual thresholds, and segment the continuous speech stream into a speech frame sequence containing effective information. The dual-stream multimodal feature deep extraction network is configured with a dual-path parallel architecture: the first path is the acoustic feature extraction branch, which is used to extract paralinguistic acoustic feature vectors including fundamental frequency trajectory, energy envelope, Mel frequency cepstral coefficients (MFCC), jitter, amplitude flicker, and formants; the second path is the semantic understanding branch, which transcribes speech into text sequences through automatic speech recognition (ASR) and uses a pre-trained language model to extract semantic embedding vectors containing contextual information. The cross-modal attention fusion and feature enhancement module is used to force alignment and weighted fusion of the acoustic feature vector and the semantic embedding vector in the time dimension through a multi-head cross-attention mechanism to construct a joint audio-text representation tensor. This module identifies and amplifies the "implied meaning" features by calculating the gating weights of acoustic features to semantic features, and generates an enhanced semantic vector that incorporates emotional coloring. A business-oriented emotion and intent mapping engine is used to establish a mapping relationship from underlying audio-text fusion features to upper-level business logic, directly mapping the identified deep psychological state of customers to specific business attribute tags. The business attribute tags include at least: purchase intention intensity, potential complaint risk level, fraud tendency index, and price sensitivity. The dual-layer profile tag lifecycle dynamic management module is used to build a short-term state tag library and a long-term attribute tag library. Based on the persistence, frequency and business impact of the interaction behavior, it controls the generation, migration, freezing and destruction of tags between the two libraries. The weighted adaptive evolution module based on Bayesian inference is used to update the activity of short-term state labels in real time using a nonlinear time decay algorithm and to correct the posterior confidence of long-term attribute labels in real time using a Bayesian probability update algorithm, based on the intensity of the current interaction behavior and the feedback of business results.
2. The system according to claim 1, characterized in that, The cross-modal attention fusion and feature enhancement module is specifically configured as follows: A timestamp mapping mechanism is adopted. For each word w(i) in the text sequence, the corresponding feature segment A(i) is extracted from the acoustic feature stream based on its start and end timestamps [t_start, t_end]. The aligned features are processed using a gated multimodal unit (GMU), and the calculation formula is as follows: H_fused = z * tanh(W_a * A(i) + b_a) + (1-z) * tanh(W_t * T(i) + b_t) Where A(i) is the acoustic feature vector, T(i) is the text semantic vector, W and b are the weight and bias matrices, and z is the adaptive gating coefficient learned based on the current context; When a significant conflict is detected between the emotional valence in acoustic features and the emotional polarity in text semantics (i.e., "irony" or "insincerity"), the system automatically increases the weight z of acoustic features and prioritizes generating labels based on speech intonation features.
3. The system according to claim 1, characterized in that, When generating business attribute tags, the business-oriented sentiment and intent mapping engine executes the following specific mapping logic: Complaint risk mapping: When the "speech speed acceleration" exceeds the threshold and the "average pitch" is significantly higher than the baseline (acoustic feature), accompanied by "repetitive rhetorical questions" or "high density of negative words" (semantic feature), the system does not directly output the "anger" emotion label, but instead converts it into a "high risk - complaint about to escalate" business warning label; Fraud risk mapping: When it is detected that the "response delay" exceeds the average and is accompanied by an abnormally high "jitter" when answering key authentication questions, the system generates a "potential fraud risk" or "unauthorized operation" label. Purchase intention mapping: When the system detects that a customer is "barge-in" and speaks faster when discussing product prices or offers, it generates a "price-sensitive - urgent purchase" tag.
4. The system according to claim 1, characterized in that, The weight adaptive evolution module based on Bayesian inference calculates the real-time weight S(t) of the short-term state label using the following formula: S(t) = S(t-1) * exp(-Δt / τ) + α * [ I_curr / (1 + exp(-β * (I_curr - I_th))) ] in: S(t-1) is the weight of the label at the previous time step; Δt is the time interval since the last tag activation; τ (tau) is the time decay constant of the tag, used to control the memory decay rate; I_curr is the currently detected audio-text fusion interaction intensity value, which is composed of the weighted sum of the acoustic energy normalization value and the semantic keyword TF-IDF value; α (alpha) is the behavioral influence coefficient, and β (beta) is the slope parameter of the Sigmoid activation function.
5. The system according to claim 1, characterized in that, The weight adaptive evolution module based on Bayesian inference calculates the posterior confidence P(L|E) of the long-term attribute label using the following Bayesian formula: P(L|E) = [ P(E|L) * P(L) ] / [ P(E|L) * P(L) + P(E|¬L) * (1 - P(L)) ] in: P(L) is the prior probability of the existence of this long-term attribute label (such as "high net worth preference") at the current moment; P(E|L) is the likelihood probability, which represents the statistical probability that a customer with this business attribute will exhibit the current specific voice interaction mode E (such as "low speech speed + long pauses + use of professional terminology"). P(E|¬L) is the false alarm probability, representing the probability that a customer who does not possess this attribute will exhibit pattern E; The system will store the label in the image database only if the posterior confidence P(L|E) remains above the locking threshold Γ (Gamma) for N consecutive interactions.
6. The system according to claim 1, characterized in that, The system also includes an abnormal interaction resistance and retention calculation unit, used to calculate the customer's resistance index R_index: R_index = w1 * (T_delay / T_avg) + w2 * log(1 + N_interrupt) + w3 * V_variance + w4 * S_neg in: T_delay is the average customer response delay time; T_avg is the system's historical average response time; N_interrupt is the number of times the client interrupts the machine's broadcast (interruption frequency); V_variance is the variance of the customer's speech rate (representing emotional instability). S_neg represents the semantic distance of the negative rejection word extracted based on semantic analysis; w1, w2, w3, w4 are normalized weighting coefficients; When R_index exceeds the preset warning value, the system generates a "strong emotional rejection" business label and triggers a script reduction or transfer to a human agent process.
7. A method for generating and updating personalized customer profile tags based on voice interaction behavior, characterized in that, Includes the following steps: Step S1: Initialize the session environment and load the base voiceprint features and historical preference tags of the customer's historical profile data; Step S2: Real-time acquisition of speech stream and processing of noise reduction and echo cancellation, and parallel input to dual-stream network to extract MFCC acoustic features and BERT semantic features; Step S3: Perform cross-modal feature fusion, calculate the correction coefficient of acoustic features to semantic features based on time alignment algorithm and attention mechanism, detect "lexical level" micro-prosodic changes, and identify implicit intentions that do not match words. Step S4: Perform business scenario mapping, input the fused features into the classifier, and generate specific business scenario labels (such as "hesitant type" and "complaint type") according to the preset business logic rules. Step S5: Calculate the intensity index of the current interaction behavior, determine the label type, and if it is a short-term state label, perform a weight update based on exponential decay; if it is a long-term attribute label, perform a probability update based on Bayesian evidence accumulation. Step S6: Monitor the gradient of label weight changes in real time. If a surge in the weight of a key negative business label is detected, immediately output a strategy adjustment instruction to the dialogue management system.