Conversation-based user satisfaction analysis method

By extracting multi-dimensional features from dialogue records and modeling the context, combined with machine learning models, the complex language recognition and data fusion problems in user satisfaction analysis were solved, enabling more accurate satisfaction prediction and service optimization decisions.

CN120951086AInactive Publication Date: 2025-11-14SOUTHWEAT UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511065743.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify complex linguistic phenomena and emotional changes in user satisfaction analysis. They fail to fully integrate textual, acoustic features, and behavioral data, and ignore the differential impact of event timing on evaluation results, leading to inaccurate satisfaction predictions.

Method used

By acquiring complete dialogue records between target users and service providers, multi-dimensional features such as text semantics, interaction patterns, and voice features are extracted. A context-aware model is used to fuse time-series information, and a machine learning model is used to predict satisfaction. A time decay factor is introduced to quantify the impact of events.

Benefits of technology

It improves the accuracy of satisfaction prediction, analyzes ironic metaphors, reveals the temporal correlation between response fluctuations and satisfaction, integrates the complementary value of multimodal data, and provides real-time and actionable decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120951086A_ABST
    Figure CN120951086A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of dialogue data analysis, and discloses a dialogue-based user satisfaction analysis method. According to the method, through context emotion evolution tracking, an anti-explanation metaphor and emotion turning are analyzed; modeling based on interaction index dynamic distribution, and revealing time sequence association of response fluctuation and satisfaction; a cross-modal collaborative analysis technology is utilized, and text-voice-behavior data complementary values are fused; and introducing an aging attenuation factor, and quantifying the influence intensity of the event time sequence on evaluation. And finally, while the real-time performance is ensured, the prediction accuracy is improved, and an executable decision basis is provided for service optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of dialogue data analysis technology, specifically to a method for analyzing user satisfaction based on dialogue. Background Technology

[0002] In the current field of user satisfaction analysis, existing technologies mainly employ two types of methods: sentiment analysis methods based on dialogue text and statistical methods based on interaction behavior indicators. Text analysis methods typically rely on keyword matching or general sentiment models, judging satisfaction by detecting the frequency of sentiment words in user statements or the matching degree of a pre-set satisfaction word library. Behavioral analysis methods, on the other hand, evaluate service efficiency through statistical metrics such as response time and dialogue rounds.

[0003] However, traditional text analysis methods lack the ability to effectively identify complex linguistic phenomena such as irony and metaphor, as well as emotional changes during dialogue; behavioral statistics methods do not fully capture the dynamic distribution characteristics of interaction indicators (such as response delay) and their temporal correlation with satisfaction; in multimodal voice interaction scenarios, existing technologies have failed to effectively integrate the complementary relationship between text, acoustic features and behavioral data; at the same time, they generally ignore the differential impact of the timing of events on user evaluation results, making it difficult to accurately reflect the characteristics of satisfaction evolution over time. Therefore, a dialogue-based user satisfaction analysis method is proposed. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a dialogue-based user satisfaction analysis method to solve the problems described in the background.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for analyzing user satisfaction based on dialogue, comprising the following steps:

[0006] S1, Dialogue Data Acquisition: Acquire the complete dialogue record between the target user and the service provider, wherein the dialogue record contains at least a sequence of interactive text arranged in chronological order;

[0007] S2, Multi-dimensional Feature Extraction:

[0008] S21, Text feature extraction: Perform natural language processing on the user statements in the interactive text sequence to extract text semantic features. The text semantic features include at least the intensity of sentiment tendency, the degree of matching and frequency with the preset satisfaction keyword library, and the clarity of the core appeal.

[0009] S22, Interaction pattern feature extraction: Analyze the temporal structure of the interaction text sequence and extract the interaction pattern features that represent the dynamic process of the dialogue. The interaction pattern features include at least the number of rounds in which the user's question is first resolved, the number of times the user asks the question repeatedly, the distribution of the service provider's response delay duration, the number of times and duration of the user's active silence, and the proportion of the service provider's guiding questions.

[0010] S3, Feature Fusion and Context Modeling: The text semantic features and interaction pattern features extracted in step S2 are fused according to the temporal sequence of the dialogue rounds and input into the context-aware model to generate a comprehensive feature vector that incorporates temporal contextual information.

[0011] S4, Satisfaction Prediction: Input the comprehensive feature vector generated in step S3 into the pre-trained user satisfaction prediction model, and output the target user's quantitative score or classification result of the current dialogue satisfaction; the user satisfaction prediction model is a machine learning model trained based on historical dialogue data labeled with real satisfaction.

[0012] Preferably, in step S21, the extraction of the intensity of sentiment tendency adopts a sentiment analysis model based on a domain sentiment dictionary and context awareness; the matching degree and frequency with the preset satisfaction keyword library, wherein the satisfaction keyword library includes hierarchical positive keywords, negative keywords and intensity indicators; the clarity of the core demand is determined by an intent recognition model to determine whether the user's statement clearly expresses the specific problem or need.

[0013] Preferably, in step S22, the number of times and duration of user-initiated silence refer to the number and average duration of pause time segments exceeding a preset threshold after the user's statement ends and before the service provider responds; the proportion of guiding questions from the service provider refers to the proportion of guiding questions raised by the service provider to the total number of questions raised by the service provider.

[0014] Preferably, the dialogue record is a voice dialogue, and step S2 further includes:

[0015] S23, speech feature extraction: extract acoustic features of user speech segments from the original speech signal of the dialogue record, wherein the acoustic features include at least average speech rate, pitch variation range, energy fluctuation intensity, and frequency of non-fluency phenomena.

[0016] Preferably, in step S3, the feature fusion and context modeling step also incorporates the acoustic features extracted in step S23, and the context-aware model simultaneously processes the temporal information of text semantic features, interaction pattern features and acoustic features.

[0017] Preferably, in step S3, the context-aware model is one of a recurrent neural network (RNN), a long short-term memory network (LSTM), a gated recurrent unit (GRU), or a Transformer encoder; the model introduces an attention mechanism during training to automatically learn the feature weights of different rounds of dialogue for predicting the final satisfaction level.

[0018] Preferably, in step S3, the feature fusion and context modeling step further includes: introducing a time decay factor into the time-sensitive features of the interaction pattern features, so that the interaction event closer to the end of the dialogue has a greater weight in the satisfaction prediction.

[0019] Preferably, the time decay factor is calculated using an exponential decay function, and the decay strength is negatively correlated with the time difference between the event occurrence and the end of the dialogue, specifically:

[0020] weight t =exp(-λ·(Tt))

[0021] Where T is the total duration of the dialogue, t is the time point when the event corresponding to this feature occurs, and λ is an adjustable attenuation coefficient.

[0022] Preferably, the user satisfaction prediction model is one of Support Vector Machine (SVM), Random Forest, Gradient Boosting Decision Tree, or Deep Neural Network; the output of the user satisfaction prediction model is a continuous satisfaction score or a discrete satisfaction classification result.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] This invention analyzes ironic metaphors and emotional shifts through contextual sentiment evolution tracking; reveals the temporal correlation between response fluctuations and satisfaction based on dynamic distribution modeling of interaction indicators; integrates the complementary value of text-voice-behavioral data using cross-modal collaborative analysis technology; and introduces a time-effect decay factor to quantify the impact of event timing on evaluation. Ultimately, while ensuring real-time performance, it improves prediction accuracy and provides actionable decision-making basis for service optimization.

[0025] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0026] Figure 1 This is a flowchart of the dialogue data acquisition process of the present invention;

[0027] Figure 2This is a diagram of the multi-dimensional feature extraction architecture of the present invention;

[0028] Figure 3 This is the feature fusion and modeling process of the present invention;

[0029] Figure 4 This invention is applied to the prediction of satisfaction levels.

[0030] Figure 5 This is a flowchart of the end-to-end system of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Please see Figure 1-5 The present invention provides a dialogue-based user satisfaction analysis method, which achieves accurate quantitative analysis of user satisfaction through multi-dimensional feature extraction, temporal context modeling, and machine learning prediction.

[0033] S1. Dialogue Data Acquisition

[0034] We construct a standardized, time-aligned dialogue data structure to lay the foundation for subsequent multi-dimensional feature extraction. The core solution addresses three key issues: ensuring the integrity of dialogue records (avoiding feature distortion caused by truncation), temporal accuracy (supporting high-precision calculation of key indicators such as response latency), and multi-source adaptability (compatible with heterogeneous data sources such as text and voice dialogues).

[0035] enter:

[0036] Original dialogue source:

[0037] Text conversations: Structured log files exported from the online customer service system (containing the text content of the interaction between the user and the customer service representative, a unique user identifier, and a statement start timestamp accurate to milliseconds);

[0038] Voice conversation: The original telephone recording file (such as WAV format) and its corresponding time-stamped text sequence generated by automatic speech recognition (ASR);

[0039] Auxiliary parameters:

[0040] Timestamp parsing rules: Specify the standard format of the timestamp (e.g., YYYY-MM-DDHH:MM:SS.fff);

[0041] Silence threshold: A standard value used to define the duration of a user's voluntary silence (denoted as τ). silence This is used as a global configuration parameter input to guide subsequent feature extraction.

[0042] 1. Multi-source data acquisition and analysis:

[0043] Text dialogue processing:

[0044] Parse the dialogue sequence field in the log file and extract the core information triples for each interaction record:

[0045] (u i ,s j ,t k )

[0046] Among them, u i This represents the text content of the user's i-th statement; s j This indicates the content of the j-th response text from the service provider (customer service); t k Indicates the statement start timestamp (accurate to milliseconds).

[0047] Automatically complete missing timestamps:

[0048] When timestamps are missing in the log, an interpolation algorithm based on the historical interval mean is used:

[0049]

[0050] Among them, t missing Indicates the missing timestamp to be completed; t prev Indicates the previous valid timestamp; n represents the average interval between adjacent rounds; t k and t k-1 These are the timestamps of the k-th and k-1-th valid records, respectively.

[0051] Voice dialogue processing:

[0052] The raw audio is processed by an Automatic Speech Recognition (ASR) engine to generate a text sequence with precise time stamps:

[0053]

[0054] Among them, text m This represents the transcribed text content of the m-th speech segment; and These represent the start and end timestamps of the segment, respectively; M is the total number of audio segments.

[0055] Speaker identification: A voiceprint recognition model is used to label the speaker role of each speech segment, generating a role identifier of u (user) or s (service provider).

[0056] 2. Data structuring:

[0057] Generate a temporal dialogue matrix:

[0058]

[0059] Among them, D struct A matrix representing the entire structured dialogue data; Role K This indicates the speaker role in the k-th round of dialogue, with values ​​of u (user) or s (service provider); Content K Represents the text content of the k-th round of dialogue (after abstract representation, such as "user inquiry statements", "customer service response statements", etc.); Indicates the start timestamp of the k-th round of dialogue; This represents the end timestamp of the k-th round of dialogue; K represents the total number of rounds of dialogue.

[0060] Key validation: Perform round-based continuity validation to ensure that the service provider's response timestamps meet timing constraints.

[0061]

[0062] Among them, s j Indicates the start time of the j-th response from the service provider; u i This indicates the end time of the user's i-th statement.

[0063] Output: Structured dialogue record (K represents the total number of rounds).

[0064] By constructing a temporally accurate dialogue structure matrix, the challenge of integrating multi-source heterogeneous data was solved, providing standardized input for subsequent feature extraction. Voiceprint recognition technology was employed to accurately distinguish speaker roles in voice dialogues, and an automatic timestamp completion algorithm ensured data integrity. A temporal continuity verification mechanism effectively prevented dialogue turn misalignment, laying a solid foundation for the accurate calculation of interactive dynamic features.

[0065] S2. Multidimensional Feature Extraction

[0066] This study deconstructs the dialogue process from three orthogonal dimensions: textual semantics, interactive dynamics, and voice features, quantifying the core factors influencing user satisfaction. The textual dimension breaks through the limitations of keyword matching, enabling context-aware emotional intensity analysis; the interactive dimension pioneers a service efficiency quantification system; and the voice dimension captures blind spots in textual analysis through paralinguistic cues.

[0067] Input: The structured dialogue record D output by S1 struct .

[0068] S21. Text Feature Extraction

[0069] 1. Calculation of emotional tendency intensity:

[0070] Employing a domain-adaptive deep semantic model:

[0071] s i =BERT finetuned (u i ;θ domain )×α context

[0072] Where: s i This represents the sentiment intensity value of the user's i-th statement; BERT finetuned This refers to a pre-trained language model that has been fine-tuned on a vertical domain corpus; u i For user statement text; θ domain Domain-specific adaptation parameters; α context This is a context correction factor.

[0073] Context correction mechanism: When an ironic expression is detected, polarity reversal is automatically triggered.

[0074]

[0075] 2. Satisfaction keyword matching:

[0076] Building a hierarchical weighted keyword library Includes positive / negative keywords and their strength weights w k ;

[0077] Calculate keyword contribution:

[0078]

[0079] in, Here, tf(k) is the indicator function (1 if the keyword appears, 0 otherwise); idf(k) is the term frequency; idf(k) is the inverse document frequency (calculated across dialogue corpora); w is the keyword database. A single keyword in the dictionary (i.e., the iteration variable when traversing the dictionary).

[0080] 3. Assessing the clarity of core demands:

[0081] Intent clarity assessment based on sequence labeling model:

[0082] c i =Softmax(W c ·CRF(h BiLSTM ))

[0083] Among them, c i Indicates the score for the clarity of the appeal; h BiLSTMThe hidden state vector is the output of a bidirectional long short-term memory network; CRF is the conditional random field decoding layer; W c This is the classification weight matrix.

[0084] Output: Text feature vector

[0085] A domain-adaptive sentiment analysis model effectively addresses the problem of quantifying the sentiment intensity of ambiguous expressions. Hierarchical weighted keyword contribution calculation significantly enhances the decision weight of high-intensity words. Combining deep learning with a clearer assessment of demands accurately distinguishes between valid complaints and vague feedback, providing fine-grained textual semantic features for satisfaction prediction.

[0086] S22. Interaction Pattern Feature Extraction

[0087] 1. Initial solution for round-robin positioning:

[0088] Define the problem-solving detection function:

[0089]

[0090] Among them, isSolved(s j ) indicates whether the j-th response statement from the service provider contains an explicit instruction to resolve the problem; NER indicates Named Entity Recognition Model.

[0091] Calculate the efficiency index for solving the problem:

[0092]

[0093] Where Δr is the first resolution round metric, representing the number of dialogue rounds from when a user raises a question to when the question is first resolved. A smaller value indicates higher service efficiency; i u The starting round index for the user's question, i.e., the dialogue round number in which the user first raised the issue that needs to be resolved; The minimum value operator finds the earliest occurrence among all server responses that meet the given conditions; j is the index number of the server response statement, representing the j-th round of server response in the dialogue; isSolved(s j ) is the problem-solving detection function, which determines the service provider's statement s in the j-th round. j Does it include a solution to the problem? Returns a boolean value (true / false).

[0094] 2. Response Delay Analysis:

[0095] Single-round delay calculation:

[0096]

[0097] Where, d iThis represents the response delay duration of the i-th round of interaction;

[0098] Global statistical characteristics:

[0099]

[0100] Where, μ d This indicates the average response time (a service stability indicator); This represents the response delay variance (volatility indicator); N represents the total number of effective dialogue rounds.

[0101] 3. User silence detection:

[0102] Effective silence identification:

[0103]

[0104] Among them, t silence (i) represents the silence duration during the i-th round of interaction.

[0105] Statistical indicators:

[0106]

[0107] Where, N silence This indicates that the silence threshold τ has been exceeded. silence The effective number of silences; K represents the total number of dialogue rounds; This represents the average duration of silence.

[0108] 4. Percentage of leading questions:

[0109] Define a guiding question decision function:

[0110]

[0111] Among them, isGuiding(s j ) represents the function that determines whether the j-th statement of the service provider is a leading question; SyntaxTree represents the syntax tree analysis result of the statement; contains WH-question indicates that the syntax tree contains question structures led by WH question words (such as "what", "why", "how" etc.).

[0112] Calculate the percentage:

[0113]

[0114] Among them, R guide Indicates the percentage of guiding questions; s j |isGuiding(s j ) = true represents the set of service provider statements that satisfy the guiding question conditions; |sj |isGuiding(s j ) = true | represents the total number of guiding questions; s j |s j `isquestion` represents the set of all queries made by the service provider (including guided and non-guided queries); |s j |s j is question| indicates the total number of questions asked by the service provider.

[0115] Output: Interactive feature vector

[0116] By quantifying service efficiency through the first-ever implementation of round-based metrics, and revealing service stability through response latency distribution, user silence pattern analysis effectively captures frustration during the conversation, while the proportion of leading questions assesses the proactiveness of the service. These innovative metrics collectively construct a complete quantitative system for the dynamic process of conversation.

[0117] S23. Speech Feature Extraction

[0118] 1. Speech rate quantification:

[0119]

[0120] Among them, v speech N represents the average speech rate, which is the number of syllables spoken per unit of time. syllable This indicates the total number of syllables detected in a speech segment.

[0121] 2. Pitch dynamic range:

[0122]

[0123] Where Δf0 represents the dynamic range of the fundamental frequency (i.e., pitch), which is the difference between the maximum and minimum values ​​of the fundamental frequency in the speech segment; f0(t) is a function of the fundamental frequency changing with time; τ represents the time variable, which is defined in the interval [t...]. start ,t end Internal changes; Indicates the time period [t] start ,t end The maximum value of the internal fundamental frequency; Indicates the time period [t] start ,t end The minimum value of the internal fundamental frequency.

[0124] 3. Frequency of non-fluent phenomena:

[0125]

[0126] Among them, freq disfluency N represents the frequency of non-fluent phenomena.fillers Indicates the number of times filler words (such as "uh" or "ah") appear; N repeats This represents the number of times a word is repeated.

[0127] Output: Acoustic feature vector

[0128] By analyzing speech rate variations, the system captured users' anxiety states; the dynamic range of pitch quantified fluctuations in emotional intensity; and the frequency of non-fluency phenomena identified moments of hesitation and uncertainty in the conversation. These acoustic features effectively supplemented the blind spots of text analysis, providing additional dimensions for sentiment analysis.

[0129] S3. Feature Fusion and Context Modeling

[0130] This step aims to construct a dynamic fusion framework for multi-source heterogeneous features. By addressing three core issues—spatiotemporal scale mismatch (differences in text / speech feature sampling rates), long-range dependency modeling (satisfaction evolution across dozens of dialogue rounds), and key event reinforcement (optimizing the weight of recency events on final satisfaction)—it achieves temporal alignment and semantic synergy of multi-dimensional features. Specifically, it quantifies the psychological recency effect using a time decay function, captures state transition patterns in the dialogue process using a bidirectional recurrent neural network, and introduces an attention mechanism to automatically identify key dialogue rounds (such as problem-solving moments and emotional turning points) that have a decisive impact on satisfaction prediction. This elevates discrete local features into dynamic representations that integrate global contextual information.

[0131] enter:

[0132] Text feature sequence: {T1,T2,...,T} K},in Represents the text features of the i-th round;

[0133] Interactive feature vector: This indicates the interaction pattern characteristics of the entire dialogue;

[0134] Acoustic feature sequence: {A1,A2,...,A K},in Represents the speech features of the i-th round;

[0135] Time decay coefficient: λ (configurable parameter).

[0136] 1. Feature alignment and tensor construction:

[0137] Multimodal feature stitching:

[0138]

[0139] in, Let represent the feature vector fused in the i-th round.

[0140] Generate time series feature matrix:

[0141]

[0142] Where K represents the total number of dialogue rounds.

[0143] 2. Time decay weighted

[0144] Time-sensitive feature decay:

[0145]

[0146] in, Represents the weighted feature vector of the i-th round; ⊙ represents the Hadamard product (element-by-element multiplication); ω i For the decay weight vector:

[0147]

[0148] Among them, J time-sensitive is the time-sensitive feature index set (corresponding to response delay, silence duration, etc.), context; K is the total dialogue duration, i is the current round index, and λ is an adjustable decay coefficient.

[0149] A bidirectional recurrent neural network is used:

[0150]

[0151] in: and These represent the forward and backward hidden states, respectively, where d is the hidden layer dimension;

[0152] Context vector generation:

[0153]

[0154] Attention weight calculation:

[0155]

[0156] Where, α i Let be the attention weight for the i-th round. This represents the learnable weight matrix for the attention mechanism; d represents the learnable parameter vector for the attention mechanism. a For the attention dimension.

[0157] Generate context vectors:

[0158]

[0159] Output: Context fusion vector

[0160] The time decay mechanism significantly enhances the influence of events in the second half of the dialogue and effectively suppresses the interference of redundant information in the early stage; the attention weight distribution can interpretably locate the key dialogue stages for the formation of satisfaction (such as high-weight rounds corresponding to the moment of problem-solving or conflict outbreak); the bidirectional network structure fully models the bidirectional interaction process of user emotions and service response, and finally achieves deep integration of text semantics, interaction dynamics and acoustic features in the temporal dimension, providing a unified feature representation with time awareness for satisfaction prediction.

[0161] S4. Satisfaction Prediction

[0162] This step aims to build a robust and interpretable satisfaction quantification model. Through a multi-granularity output architecture (parallel operation of continuous rating and discrete classification), a hybrid learning strategy (integrating the complementary advantages of tree models and neural networks), and a decision-interpretable engine (measuring feature contribution), it addresses industry pain points such as insufficient prediction stability (single models are sensitive to noise), weak business adaptability (different output formats are required for different scenarios), and lack of attribution analysis (inability to pinpoint the root causes of service defects). It focuses on utilizing gradient boosting trees to capture nonlinear interactions and threshold effects between features, combining this with neural networks to uncover higher-order latent patterns, and using feature contribution to inversely deconstruct the decision-making basis of the prediction results.

[0163] Input: The context fusion vector output by S3

[0164] Hybrid prediction model

[0165] Neural network regression branch:

[0166]

[0167] Where σ(·) represents the Sigmoid activation function; b represents the weight matrix vector of the neural network; r This indicates the bias term.

[0168] Synthetic tree classification branches:

[0169]

[0170] Where, p k This represents the probability of satisfaction for the k-th class. Let M represent the classification function of the m-th decision tree; M represents the total number of decision trees.

[0171] Final prediction output:

[0172]

[0173] 2. Interpretability Engine

[0174] Feature contribution metric:

[0175]

[0176] in, j z is the SHAP value of the j-th feature; \j denoted as occluding the j-th dimension feature; f represents the trained prediction model (i.e., the hybrid prediction model in S4); z represents the complete feature vector (i.e., the context fusion vector output by S3).

[0177] Output:

[0178] Main output: Satisfaction prediction results Dynamically generate satisfaction heatmaps when When the threshold is reached, an early warning system is triggered, notifying superiors to intervene and simultaneously generating a KPI report for the staff involved. As a core performance indicator, it drives team optimization.

[0179] Auxiliary output: Identify the features with the highest negative contribution (such as latency > 0), automatically generate service improvement suggestions, and guide customer service training focus (e.g., strengthen script training when the number of guiding questions is consistently low).

[0180] The integrated model significantly improves prediction robustness and effectively overcomes the performance degradation of a single model under data distribution shifts; the dual-output mode simultaneously supports refined service monitoring (continuous scoring) and automated service grading (discrete classification), meeting the business needs of multiple scenarios; feature contribution analysis clearly quantifies the influence of each dimension of features on the prediction results (such as the negative contribution of response latency variance or the positive contribution of the proportion of guided questions), providing quantifiable and traceable decision-making basis for service optimization and driving continuous improvement in service quality.

[0181] This embodiment provides a dialogue-based user satisfaction analysis method. It constructs a multi-source heterogeneous dialogue foundation by acquiring standardized time-series data, extracts deep features from three orthogonal dimensions: text semantics, interaction dynamics, and voice features, and integrates time-sensitive information using a time decay mechanism and a context-aware model. Finally, it outputs a quantitative satisfaction score and interpretable attribution through a hybrid prediction model, forming a closed-loop management of service quality consisting of "monitoring-early warning-diagnosis-optimization".

Claims

1. A method for analyzing user satisfaction based on dialogue, characterized in that, Includes the following steps: S1, Dialogue Data Acquisition: Acquire the complete dialogue record between the target user and the service provider, wherein the dialogue record contains at least a sequence of interactive text arranged in chronological order; S2, Multi-dimensional Feature Extraction: S21, Text feature extraction: Perform natural language processing on the user statements in the interactive text sequence to extract text semantic features. The text semantic features include at least the intensity of sentiment tendency, the degree of matching and frequency with the preset satisfaction keyword library, and the clarity of the core appeal. S22, Interaction pattern feature extraction: Analyze the temporal structure of the interaction text sequence and extract the interaction pattern features that represent the dynamic process of the dialogue. The interaction pattern features include at least the number of rounds in which the user's question is first resolved, the number of times the user asks the question repeatedly, the distribution of the service provider's response delay duration, the number of times and duration of the user's active silence, and the proportion of the service provider's guiding questions. S3, Feature Fusion and Context Modeling: The text semantic features and interaction pattern features extracted in step S2 are fused according to the temporal sequence of the dialogue rounds and input into the context-aware model to generate a comprehensive feature vector that incorporates temporal contextual information. S4, Satisfaction Prediction: Input the comprehensive feature vector generated in step S3 into the pre-trained user satisfaction prediction model, and output the target user's quantitative score or classification result of the current dialogue satisfaction; the user satisfaction prediction model is a machine learning model trained based on historical dialogue data labeled with real satisfaction.

2. The method for analyzing user satisfaction based on dialogue according to claim 1, characterized in that, In step S21, the extraction of the intensity of sentiment tendency adopts a sentiment analysis model based on a domain sentiment dictionary and context awareness; the matching degree and frequency with the preset satisfaction keyword library, the satisfaction keyword library includes hierarchical positive keywords, negative keywords and intensity indicators; the clarity of the core demand is determined by an intent recognition model to determine whether the user's statement clearly expresses the specific problem or need.

3. The method for analyzing user satisfaction based on dialogue according to claim 1, characterized in that, In step S22, the number of times and duration of user-initiated silence refer to the number and average duration of pause time segments exceeding a preset threshold after the user's statement ends and before the service provider responds; the proportion of guiding questions from the service provider refers to the proportion of guiding questions raised by the service provider to the total number of questions raised by the service provider.

4. The method for analyzing user satisfaction based on dialogue according to claim 1, characterized in that, The dialogue record is a voice dialogue, and step S2 further includes: S23, speech feature extraction: extract acoustic features of user speech segments from the original speech signal of the dialogue record, wherein the acoustic features include at least average speech rate, pitch variation range, energy fluctuation intensity, and frequency of non-fluency phenomena.

5. The method for analyzing user satisfaction based on dialogue according to claim 4, characterized in that, In step S3, the feature fusion and context modeling step also incorporates the acoustic features extracted in step S23. The context-aware model simultaneously processes the temporal information of text semantic features, interaction pattern features, and acoustic features.

6. The method for analyzing user satisfaction based on dialogue according to claim 1, characterized in that, In step S3, the context-aware model is one of the following: Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), or Transformer Encoder; the model introduces an attention mechanism during training to automatically learn the feature weights of different rounds of dialogue for predicting the final satisfaction level.

7. The method for analyzing user satisfaction based on dialogue according to claim 1, characterized in that, In step S3, the feature fusion and context modeling step further includes: introducing a time decay factor into the time-sensitive features in the interaction pattern features, so that the interaction event closer to the end of the dialogue has a greater weight in the satisfaction prediction.

8. The method for analyzing user satisfaction based on dialogue according to claim 7, characterized in that, The time decay factor is calculated using an exponential decay function, and the decay intensity is negatively correlated with the time difference between the event occurrence and the end of the dialogue. ω i =exp(-λ·(K-i)) Where i is the round index of the event corresponding to the feature, K is the total number of rounds of dialogue, and λ is an adjustable decay coefficient.

9. The method for analyzing user satisfaction based on dialogue according to claim 1, characterized in that, The user satisfaction prediction model is one of Support Vector Machine (SVM), Random Forest, Gradient Boosting Decision Tree, or Deep Neural Network; the output of the user satisfaction prediction model is a continuous satisfaction score or a discrete satisfaction classification result.