A spoken dialogue data processing method based on an agent
By combining sentiment analysis and prosodic features in a conversational dialogue system, a path graph and propagation matrix are constructed, which solves the problem of insufficient sentiment understanding, achieves the naturalness and consistency of the dialogue system, and improves user experience and task processing efficiency.
Patent Information
- Application Number
- CN202511160428.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing conversational systems are inadequate in terms of emotion understanding and context processing. They cannot accurately capture users' emotional fluctuations and changes in voice, resulting in a poor user experience. Furthermore, they cannot flexibly respond and adapt to users' needs and emotional states in real time.
By receiving voice input, converting it into text data and extracting emotional features, generating emotional tags and speech prosody features, constructing a path graph, calculating edge weights, combining emotional and semantic information for intelligent path reasoning, and using a propagation matrix for real-time correction to ensure the naturalness and consistency of the dialogue content.
It improves user experience and dialogue fluency, enhances the deep understanding of dialogue context, achieves continuity and consistency in multi-turn dialogues, improves the efficiency and accuracy of task processing, and can adjust the dialogue path in real time according to user emotions.
Smart Images

Figure CN120726996B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of artificial intelligence big data processing, and particularly relates to a spoken dialogue data processing method based on an agent. BACKGROUND
[0002] With the development of artificial intelligence technology, spoken dialogue systems have been widely applied in various intelligent customer service, virtual assistants and smart home fields. However, the existing spoken dialogue systems usually have the problems of insufficient emotional understanding and limited context processing capability, which makes it difficult to accurately capture the emotional fluctuations and emotional changes in the user's voice. Traditional speech recognition technology mainly focuses on converting speech into text, while ignoring the emotional features, speech rate changes and intonation fluctuations in the speech, which makes the system unable to adjust the dialogue content and task execution path according to the emotional state of the user in multi-round dialogue, resulting in poor user experience.
[0003] In addition, the existing agent system mostly relies on fixed rules or models when processing tasks, and cannot flexibly respond and adapt to the user's needs and emotional state in real time. The existing technology fails to fully consider the emotional changes and intonation features in the speech, resulting in inaccurate emotional adaptation and inability to dynamically respond to the emotional fluctuations of the user in multi-round dialogue. The decision of each round of dialogue cannot be effectively connected, and lacks continuity and consistency. It is unable to flexibly adjust the path selection according to the emotional state and semantic content of the user, resulting in low efficiency of task execution. In order to improve the naturalness of the dialogue and the intelligent level of the system, a new method is needed to integrate speech recognition, emotional analysis and intelligent reasoning capability to realize more accurate multi-round dialogue and task processing. SUMMARY
[0004] The purpose of the present application is to provide a spoken dialogue data processing method based on an agent to solve the problems raised in the background art.
[0005] The spoken dialogue data processing method based on an agent of the present application comprises the following steps:
[0006] S1, receiving a voice input file sent by a user terminal, converting the voice signal into text data, extracting emotional features, generating emotional labels and speech prosody features, combining the text and emotional information to form dialogue anchor point information, constructing a path graph and calculating edge weights;
[0007] S2, constructing a propagation matrix according to the similarity of the current dialogue anchor point and the historical path node, performing intelligent path reasoning, calculating path scores, ensuring the influence of emotional fluctuations on task path selection, and selecting the path with the highest score for subsequent task processing according to the path score by the agent, and performing real-time correction after user feedback.
[0008] Further, the S1 voice input file is a WAV or PCM format voice input file sent by the receiving user end, the sampling rate is 16 kHz, the intelligent agent is initiated to call, the audio input calls Whisper for streaming analysis, that is, the audio signal is converted into text data data through the voice2wenzi method, and the processing, word segmentation and denoising steps of the audio data are included; the current user's emotional color is analyzed by using the emotion analysis, and is stored to the current user mark, the time, the text data and the emotional color are substituted into the large language model LSTM, the pitch and the rhythm characteristics of the speech speed are extracted, the emotional information is analyzed, and the emotional label in the speech rhythm is obtained.
[0009] Further, the information extracted in the S1 includes the text expression of the voice and the emotional color of the i-th round of dialogue in the audio, the text expression of the voice is the semantic information, the emotional color is the emotional label and the rhythm characteristics related to the voice, the emotional label , is the emotional main classification label, is the change range of the pitch, is the change of the speech speed; used to reflect the emotional fluctuation of the user in different dialogue rounds, affect the emotional adaptation and path selection of the subsequent dialogue.
[0010] Further, by comprehensively processing the emotional and semantic information, a dialogue anchor point is generated for each round of dialogue, the dialogue anchor point information in the S1, and the calculation method of the dialogue anchor point is as follows:
[0011] ;
[0012] wherein, is the anchor point information of the i-th round of dialogue, which contains the semantic and emotional information of the i-th round of voice input; is the semantic vector of the i-th round of dialogue input, representing the semantic information of the current round of text; is the change range of the pitch in the i-th round of dialogue; is the pitch sensitivity coefficient, used to adjust the influence degree of the pitch change on the path reasoning; is the change of the speech speed in the i-th round; is the timestamp of the i-th round of dialogue, representing the processing time of the i-th round of voice input.
[0013] Further, the S1 constructs a path graph and calculates the edge weight, specifically, a weighted directed graph is constructed, wherein each node represents a function execution point, representing a function or task executed by the user in the process of interacting with the system, the nodes are connected through the edge weight, the weight of the edge calculates the relationship strength between the semantic content and the emotional characteristics of the current dialogue and the historical dialogue, and the edge weight is calculated by the following expression:
[0014] ;
[0015] wherein, denotes the edge weight between node i and node j; is the path hop sensitive coefficient, aiming to balance the coupling relationship between "anchor changing rate" and "behavior staying time", and control the sensitivity of path weight; is the anchor information of the i-th round of dialogue, is the anchor information of the j-th round of dialogue; is the functional staying time corresponding to node i; is the functional staying time corresponding to node j.
[0016] Further, each node in the weighted directed graph is represented as wherein, denotes the function name defined by the system, derived from the function_name field of the user_function_record table; is the average duration of user staying in the current function, used to evaluate the weight of the task.
[0017] Further, the propagation matrix formula in S2 is:
[0018] ;
[0019] wherein, is the emotion-semantic similarity between node i and historical path node j, indicating the matching degree of the current user input and the historical path node; is the cosine similarity index, used to measure the semantic structure similarity between anchors, is the L2 norm; is the adjustment function, aiming to further amplify the path deviation signal caused by emotion state resonance by analyzing the alignment degree of current tone change and speech speed amplitude, is a perturbation factor to prevent zero division setting.
[0020] Further, the path score in S2 is calculated according to the similarity between each path and the current dialogue anchor, as well as the correlation between the historical path and the emotion evolution, to score all possible paths:
[0021] ;
[0022] wherein, denotes the task or function comprehensive score of the current input, based on the weighted calculation results of the current input emotion, semantic, time difference and historical path, evaluate each candidate function in the current conversation, select the highest score task or function as the next step operation; is the total number of nodes in the path graph, that is, the total number of dialogue turns; is the edge weight between node i and node j; is the emotion-semantic similarity between node i and historical path node j; is the time decay adjustment coefficient, which controls the influence of time difference on task path selection; is the standardized value for calculating the influence of the time interval between two task nodes i and j on path decay.
[0023] Further, each value forms a function ranking table, which is pushed to the agent as the basis for the next workflow branch call, and the path is automatically selected based on the configured Prompt and tool chain node rules for workflow process planning; Based on the coupling strength of the current conversation emotion context and the path context, the agent autonomously calls the API module, database query, RAG knowledge base or executes the Python code module in the workflow, realizes the combined multi-step decision of cross-tool chain, and records the emotional feedback in the final response result, which is used to update the emotional memory chain.
[0024] The beneficial effects of the technical solutions of the present application are:
[0025] 1. By combining emotional information with semantic information, the system can adapt to the emotional state of the user more naturally in different dialogue scenarios, improve user experience and dialogue fluency, and dynamically adjust according to the emotional fluctuations and voice prosody changes of the user; By introducing emotion analysis technology and audio feature extraction, the emotional changes (such as pitch changes and speech rate changes) in the voice signal are accurately extracted, so that the system can understand and respond to the emotional state of the user, thereby improving the accuracy of voice recognition and the sensitivity of emotion recognition;
[0026] 2. The present application generates a sentiment-semantic composite feature vector for each round of conversation, i.e. a dialogue anchor point, to capture the semantic information of the current conversation, and also combines the emotional color and voice prosody features of the user, thereby enhancing the depth of understanding of the dialogue context by the agent, supporting the continuity and consistency of multi-round dialogue; By constructing a weighted directed graph and a propagation matrix, the present application can realize path reasoning based on historical dialogue data, combine the emotional changes of the user and the semantic similarity, and the agent can more accurately select the most suitable task path for execution, greatly improving the efficiency and accuracy of task processing;
[0027] 3. When users are dissatisfied with the system's response, this invention can adjust the anchor information of the current dialogue in real time through sentiment analysis technology, and adjust the reasoning process according to feedback, forming a closed-loop adaptive adjustment mechanism, thereby optimizing the system's response quality and enhancing users' trust in the system. Attached Figure Description
[0028] Figure 1 This is a flowchart of the present invention. Detailed Implementation
[0029] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0031] The following description, in conjunction with the accompanying drawings, details a specific scheme for a speech-based dialogue data processing method based on intelligent agents provided by this invention.
[0032] See attached document Figure 1 This illustrates an agent-based spoken dialogue data processing method provided by an embodiment of the present invention, comprising the following steps:
[0033] S1. Receive the voice input file sent by the user terminal, convert the voice signal into text data, extract emotional features, generate emotional tags and voice prosody features, combine text and emotional information to form dialogue anchor information, construct a path graph, and calculate edge weights.
[0034] The system receives WAV or PCM format voice input files from the user client, with a sampling rate of 16kHz. It then initiates an intelligent agent call, invoking Whisper to perform streaming parsing of the input audio. This involves converting the audio signal into text data (data) using the voice2wenzi method, including audio data processing, word segmentation, and noise reduction. Sentiment analysis is used to determine the current user's emotional tone and store it in the current user's identifier. The time, text data, and emotional tone are then fed into a large language model LSTM (Long Short-Term Memory) to extract prosodic features of pitch and speech rate. Emotional information is analyzed to obtain emotional labels in the speech rhythm (such as pitch variance, speech rate, and the main emotion classification label "emotion").
[0035] The information extracted in S1 includes the textual representation of the voice, i.e. semantic information, and the emotional color of the i-th round of dialogue contained in the audio, i.e. emotional tags and prosodic features related to the voice, the emotional tags , are emotional primary classification tags (emotion, used to distinguish basic emotional types such as joy, anger, sadness, and pleasure), are the change amplitudes of the pitch (pitch variance, used to capture the expression of the speaker's emotional intensity or emotional fluctuations), and are the changes in speech rate (speech rate, related to the emotional state of the speaker such as anxiety, tension, etc.). Not only can it reflect the emotional fluctuations of the user in different dialogue rounds, but also can affect the emotional adaptation and path selection of subsequent dialogue.
[0036] By comprehensive processing of emotional and semantic information, a dialogue anchor point is generated for each round of dialogue, which not only contains the semantic information of the current round, but also combines the emotional state and voice prosodic features, fully reflecting the dynamic changes of dialogue content, emotional color and tone. As a "emotion-semantic" composite feature vector, the dialogue anchor point becomes the core basis for task selection, path planning and emotional adaptation in the subsequent reasoning process. Specifically, by comprehensive processing of emotional and semantic information, a dialogue anchor point is generated for each round of dialogue, and the calculation method of the dialogue anchor point in S1 is as follows:
[0037] ;
[0038] wherein, is the anchor point information of the i-th round of dialogue, which contains the semantic and emotional information of the i-th round of voice input; is the semantic vector of the i-th round of dialogue input, representing the semantic information of the current round of text; is the change amplitude of the pitch in the i-th round of dialogue; is the tone sensitivity coefficient, used to adjust the influence degree of the change of the pitch on the path reasoning; is the change of the i-th speech rate; is the timestamp of the i-th round of dialogue, representing the processing time of the i-th round of voice input. The calculation of the dialogue anchor point converts the original text semantic and audio emotional information into anchor point information with consistent structure through unified mapping, which can accurately encode the emotional changes and semantic information in the voice.
[0039] Each round of conversation input of the user is converted into emotion and semantic anchor information, to build a weighted directed graph for understanding the relationship between the current conversation and the historical conversation, wherein each node represents a function execution point, indicating a function or task executed by the user in the process of interacting with the system, and each node contains not only the basic description of the current function, but also semantic information, emotion label and execution time related to the current function and the like. For example, in the process of a user operating the system, each user request (such as querying the weather or checking the account information) can be regarded as a function execution point. Each time the user triggers a function, the system records the operation as a node and associates the corresponding emotion and voice features, to form a complete path graph. The nodes are connected through edge weights, and the edge weight calculates the relationship strength between the semantic content and the emotion features of the current conversation (the current node) and the historical conversation (the historical node), to ensure that the emotion change of the user in the voice expression and the influence on the path selection can be captured.
[0040] Specifically, each node in the weighted directed graph is represented as , wherein represents the function name (for example, function A, function B) defined by the system, which is derived from the function_name field of the user_function_record table; is the average duration of the user staying in the current function, used to evaluate the weight of the task. In S1, the path graph is constructed, and the edge weight is calculated, specifically, a weighted directed graph is constructed, wherein each node represents a function execution point, indicating a function or task executed by the user in the process of interacting with the system, and the nodes are connected through edge weights, and the edge weight calculates the relationship strength between the semantic content and the emotion features of the current conversation and the historical conversation. The edge weight is calculated through the following expression:
[0041] ;
[0042] , wherein represents the edge weight between node i and node j; is a path jump sensitive coefficient, aiming to balance the coupling relationship between the “anchor change rate” and the “behavior residence time”, and control the sensitivity of the path weight; is the function residence time corresponding to node i; is the function residence time corresponding to node j. is the anchor information of the i-th round of conversation, is the anchor information of the j-th round of conversation.
[0043] S2, construct a propagation matrix according to the similarity between the current dialogue anchor point and the historical path node, perform intelligent path reasoning, calculate the path score, ensure the influence of emotional fluctuations on task path selection, and the agent selects the path with the highest score for subsequent task processing according to the path score, and makes real-time correction after user feedback. The path with the highest score is the optimal path.
[0044] In order to realize intelligent path reasoning based on user emotion and semantic information, it is necessary to match the anchor point information of the current dialogue with the nodes in the historical path, calculate the association strength between them, determine which historical path best meets the current user intent, and form a propagation matrix.
[0045] Specifically, each item of the propagation matrix represents the similarity between the anchor point information of the current dialogue and the historical path node. The cosine similarity calculation is used to reflect the similarity of the semantic structure, and an adjustment factor of emotional change is introduced to consider the influence of pitch change on path prediction, so as to ensure that emotional transition has a necessary influence on path selection, and to select a path that is more consistent with the current emotional state. The agent can flexibly adjust the path according to the double factors of emotion-semantic change in a complex multi-round dialogue scene, improve the naturalness and efficiency of multi-round dialogue, and better serve the user's needs. The formula of the propagation matrix in S2 is:
[0046] ;
[0047] wherein, is the emotion-semantic similarity between node i and historical path node j, indicating the matching degree between the current user input and the historical path node; is the cosine similarity index, used to measure the semantic structure similarity between the anchor points, is the L2 norm; is the adjustment function, which is intended to further amplify the path deviation signal caused by emotional state resonance by analyzing the alignment degree of the current intonation change (pitch) and speech speed amplitude, is a perturbation factor to prevent zero division setting.
[0048] After constructing the propagation matrix, all possible paths need to be scored according to the similarity between each path and the current dialogue anchor point, as well as the relevance of historical paths to the evolution of emotions. The calculation of each path score integrates three factors: the edge weight represents the frequency and intensity of the user's execution of the current path in history, that is, it reflects the actual execution importance of the current path; the propagation matrix measures the semantic and emotional matching degree between the current dialogue anchor point and the historical path; the time decay factor adjusts the time sensitivity of the path, reducing the influence of historical paths with large time differences on the current path recommendation. Thus, the agent platform selects the appropriate workflow for automated processing, scoring all paths.
[0049] ;
[0050] wherein, is the comprehensive score representing the task or function based on the weighted calculation results of the current input emotions, semantics, time differences and historical paths, evaluating each candidate function in the current dialogue, and selecting the highest scoring task or function as the next step operation; is the total number of nodes in the path graph, that is, the total number of dialogue turns; is the edge weight between node i and node j; is the emotion-semantic similarity between node i and historical path node j; is the time decay adjustment coefficient, which controls the influence of time difference on task path selection; is the standardized value used to calculate the influence of the time interval between two task nodes i and j on path decay.
[0051] The final value forms a function ranking table, which is pushed to the agent as the basis for the next workflow branch call, and the agent automatically selects the best path based on the configured Prompt and tool chain node rules to plan the workflow process. Based on the coupling strength of the current dialogue emotion context and the path context, the agent can autonomously call API modules, database queries, RAG knowledge bases or execute Python code modules in the workflow, implement combined multi-step decision-making across tool chains, and record the emotional feedback in the final response result to update the emotional memory chain.
[0052] When the user is not satisfied with the output result of this round, the emotion analysis technology is called to extract the emotional features of the user feedback, automatically update the anchor point information of the current dialogue turn, and perform the propagation matrix and path scoring steps again to realize the inference correction closed loop based on real-time feedback.
[0053] In summary, a kind of oral conversation data processing method based on agent is completed.
[0054] The order of the embodiments is merely for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0055] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other, and each embodiment mainly explains the difference from other embodiments.
[0056] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit it; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that the technical solutions recorded in the foregoing embodiments can still be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for processing spoken dialogue data based on an agent, characterized by, Comprise the following steps: S1, receiving the voice input file sent by the user end, converting the voice signal into text data, extracting the emotional features, generating emotional labels and voice prosody features, combining the text and emotional information to form the dialogue anchor point information, constructing the path graph and calculating the edge weight; the construction of the path graph and the calculation of the edge weight are as follows: a weighted directed graph is constructed, each node represents a function execution point, which represents a function or task executed by the user during interaction with the system, the nodes are connected by edge weight, and the weight of the edge calculates the relationship strength between the semantic content and the emotional features of the current dialogue and the historical dialogue; S2, according to the similarity between the current dialogue anchor point and the historical path node, a propagation matrix is constructed, intelligent path reasoning is carried out, path score is calculated, the influence of emotional fluctuation on task path selection is ensured, and the agent selects the path with the highest score for subsequent task processing according to the path score, and real-time correction is carried out after user feedback; the propagation matrix is the emotional-semantic similarity between the current node and the historical path node, which represents the matching degree between the current user input and the historical path node; The cosine similarity index and the adjustment function are used to calculate the propagation matrix, the cosine similarity index is used to measure the semantic structure similarity between anchor points, and the adjustment function analyzes the alignment degree of current intonation change and speech speed amplitude, further amplifying the path deviation signal caused by emotional state resonance. 2.The method of claim 1, wherein, The voice input file in S1 is a WAV or PCM format voice input file sent by the receiving user end, the sampling rate is 16kHz, the intelligent agent is initiated, the input audio calls Whisper for stream parsing, that is, the audio signal is converted into text data data through the voice2wenzi method, and includes the processing, word segmentation and denoising steps of audio data; the current user's emotional color is analyzed by the emotion analysis and stored in the current user flag, the time, text data and emotional color are substituted into the large language model LSTM, the prosody features of pitch and speech speed are extracted, the emotional information is analyzed, and the emotional label in the voice rhythm is obtained. 3.The method of claim 2, wherein, The information extracted in the S1 includes the textual representation of the voice, i.e. semantic information, and the emotional color of the i-th round of dialogue in the audio, i.e. emotional tags and prosodic features related to the voice , emotional main classification tags, the change range of the pitch, the change of the speech rate; used to reflect the emotional fluctuations of the user in different dialogue rounds, and affect the emotional adaptation and path selection of the subsequent dialogue.
4. The method of claim 1, wherein, Through comprehensive processing of emotional and semantic information, a dialogue anchor point is generated for each round of dialogue, and the calculation method of the dialogue anchor point in S1 is as follows: ; wherein, is the anchor information of the i-th round of dialogue, containing the semantic and emotional information of the i-th round of voice input; is the semantic vector of the i-th round of dialogue input, representing the semantic information of the current round of text; is the change amplitude of the i-th round of dialogue; is the tone sensitivity coefficient, used to adjust the influence degree of the change of the pitch on the path reasoning; is the change of the i-th round of speech speed; is the timestamp of the i-th round of dialogue, indicating the processing time of the i-th round of voice input.
5. The method of claim 4, wherein the method further comprises: The edge weight in S1 is expressed by the following formula: ; wherein, represents the edge weight between node i and node j; is the path hop sensitive coefficient, aiming to balance the coupling relationship between "anchor change rate" and "behavior stay time", and control the sensitivity of path weight; is the anchor information of the i-th round of dialogue, is the anchor information of the j-th round of dialogue; is the function stay time corresponding to node i; is the function stay time corresponding to node j.
6. The method of claim 5, wherein the method further comprises: Each node in the weighted directed graph is represented as where represents the function name defined by the system, derived from the function_name field of the user_function_record table; is the average duration of the user staying in the current function, used to evaluate the weight of the task.
7. The method of claim 6, wherein the method further comprises: The formula of the propagation matrix in S2 is: ; wherein, is the emotion-semantic similarity between node i and historical path node j, representing the matching degree of the current user input and the historical path node; is the cosine similarity index, used to measure the semantic structure similarity between anchor points, is the L2 norm; is the adjustment function, which is intended to further amplify the path deviation signal caused by the emotion state resonance by analyzing the alignment degree of the current tone change and the speech speed amplitude, is a perturbation factor to prevent zero exclusion settings.
8. The method of claim 7, wherein the method further comprises: In S2, the path score is calculated, the similarity between each path and the current dialogue anchor point, and the correlation between the historical path and the emotional evolution are used to score all possible paths: ; wherein, is a comprehensive score representing a task or function , based on the weighted calculation results of the current input emotion, semantics, time difference and historical path, each candidate function in the current dialogue is evaluated, and the task or function with the highest score is selected as the next step operation; is the total number of nodes in the path graph, i.e. the total number of dialogue turns; is the edge weight between node i and node j; is the emotion-semantics similarity between node i and historical path node j; is a time decay adjustment coefficient that controls the influence of time difference on task path selection; is a standardized value for calculating the influence of the time interval between two task nodes i and j on path decay.
9. The method of claim 8, wherein, Each The values form a function ranking table, which is pushed to the agent as the basis for the next workflow branch call. The path is automatically selected by the Prompt and tool chain node rules based on the configuration to plan the workflow process. Based on the coupling strength of the current dialogue emotion context and the path context, the agent autonomously calls the API module in the workflow, database query, RAG knowledge base or executes the Python code module, realizes the combined multi-step decision of cross-tool chain, and records the emotional feedback in the final response result, which is used to update the emotional memory chain.
Citation Information
Patent Citations
Digital human interaction control method and device fusing emotional semantics and logical reasoning and storage medium
CN120316321A
Cited By
Spoken language dialogue dynamic generation method and system based on speech emotion feedback
CN122157703A