Dialogue interaction state recognition method and system, electronic equipment and storage medium

By combining temporal and semantic features to identify the user's dialogue interaction status, and using preset rules and decision tree models, the problem that fixed thresholds are difficult to adapt to users' personalized expressions is solved, and high-precision recognition of hesitation and interruption states is achieved, thereby improving the naturalness and coherence of human-computer interaction.

CN120656484AActive Publication Date: 2025-09-16SHANGHAI XULU INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511134763.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-09-16
Estimated Expiration
2045-08-14

AI Technical Summary

Technical Problem

In existing human-computer voice interactions, the fixed threshold silence duration is difficult to adapt to the personalized expression habits and diverse interaction scenarios of different users, resulting in insufficient recognition accuracy for complex interaction states such as user hesitation and interruption.

Method used

By combining temporal features and semantic features to identify the user's conversation interaction status, and using preset rules and decision tree models to analyze voice rhythm, voice content and voice intensity, it can identify silence, interruption and hesitation states, thereby improving recognition accuracy.

Benefits of technology

It improves the recognition accuracy of complex interaction states such as user hesitation and interruption, and enhances the naturalness and consistency of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656484A_ABST
    Figure CN120656484A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a dialogue interaction state recognition method and system, electronic equipment and a storage medium, and relates to the technical field of voice interaction, and the method comprises the steps: extracting time sequence features and semantic features from obtained voice information; the time sequence feature represents the rhythm change condition of the user in the dialogue process, and the semantic feature represents the voice content and the voice intensity of the user in the dialogue process; recognizing the current dialogue interaction state of the user according to the time sequence features and the semantic features; the dialogue interaction state comprises one of a silent state, an interrupted state and a hesitant state. Therefore, the conversation interaction state of the user is recognized by combining the time sequence features and the semantic features, voice signal level analysis is considered, and changes of voice rhythm, voice content and voice intensity are also considered, so that the conversation interaction state of the user is accurately recognized, and the recognition precision of complex interaction states such as hesitation and interruption is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction technology, and in particular to a method, system, electronic device and storage medium for identifying a dialogue interaction state. Background Art

[0002] With the rapid development of voice interaction technology, intelligent voice assistants have been widely used in many fields such as smart homes, car systems and mobile terminals, and have gradually become one of the core ways of human-computer interaction.

[0003] In traditional human-computer voice interaction, the determination of user status mainly relies on the silence timeout mechanism, which sets a fixed threshold, such as 1.5-3 seconds of continuous silence to determine the end of speaking, or 25 seconds of no input to trigger a system prompt. Although this mechanism is simple to implement, it has significant limitations. The fixed silence duration threshold is difficult to adapt to the personalized expression habits and diverse interaction scenarios of different users; it also lacks the ability to recognize subtle behavioral characteristics during user interaction. For example, when the user hesitates or interrupts in speech, misjudgment is likely to occur. Therefore, most existing methods focus only on the analysis of the voice signal level with a fixed threshold, resulting in the recognition accuracy of complex interaction states such as hesitation and interruption still needs to be improved. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method, system, electronic device and storage medium for identifying conversation interaction states, which identify the user's conversation interaction state by combining temporal features and semantic features, taking into account not only the voice signal level analysis, but also the changes in voice rhythm, voice content and voice intensity, so as to accurately identify the user's conversation interaction state, thereby improving the recognition accuracy of complex interaction states such as hesitation and interruption.

[0005] In order to achieve the above-mentioned objectives, in a first aspect, an embodiment of the present invention provides a method for identifying a conversation interaction state, the method comprising: extracting temporal features and semantic features from the acquired voice information; the temporal features characterize the changes in the user's rhythm during the conversation, and the semantic features characterize the user's voice content and voice intensity during the conversation; identifying the user's current conversation interaction state based on the temporal features and the semantic features; the conversation interaction state includes one of a silent state, an interrupted state, and a hesitant state.

[0006] In this embodiment, by acquiring timing features and semantic features, it is easy to know the changes in voice rhythm, voice content and voice intensity in the user's voice information, so as to identify the user's current dialogue interaction state as one of a silent state, an interruption state and a hesitant state based on the changes in voice rhythm, voice content and voice intensity; in this way, by combining timing features and semantic features to identify the user's dialogue interaction state, not only the analysis at the voice signal level is considered, but also the changes in voice rhythm, voice content and voice intensity, so as to accurately identify the user's dialogue interaction state, thereby improving the recognition accuracy of complex interaction states such as hesitation and interruption.

[0007] In some embodiments, identifying the user's current conversation interaction state based on the timing features and the semantic features includes: when the pause duration in the timing features is greater than or equal to a first preset pause duration, identifying the user's current conversation interaction state as a silent state; when the pause duration is greater than a second preset pause duration and less than the first preset pause duration, inputting the timing features and the semantic features into a preset decision tree model to identify whether the user's current conversation interaction state is a silent state.

[0008] With this setting, since there is a clear time threshold difference between the first preset pause duration and the second preset pause duration, the former is used to directly determine the silent state, and the latter is used to trigger the decision tree model to assist in judgment, thereby forming a hybrid recognition mechanism of rules plus models, so that a high recognition accuracy can still be maintained when facing boundary samples, thereby improving the accuracy of identifying the user's dialogue interaction state as a silent state.

[0009] In some embodiments, identifying the user's current dialogue interaction state based on the timing features and the semantic features includes: when the overlapping duration in the timing features is greater than or equal to a first preset overlapping duration and the semantic features include continuous speech, identifying the user's current dialogue interaction state as an interruption state; the continuous speech is the voice content continuously input by the user in a single voice input; when the overlapping duration is greater than a second preset overlapping duration and less than the first preset overlapping duration and the voice intensity in the semantic features indicates that the voice intensity is unstable, inputting the timing features and the semantic features into a preset decision tree model to identify whether the user's current dialogue interaction state is an interruption state.

[0010] With this setting, since there is a clear time threshold difference between the first preset overlapping time length and the second preset overlapping time length, the former is used in conjunction with continuous speech to directly determine the silent state, and the latter is used in conjunction with voice intensity to trigger the decision tree model to assist in judgment, thereby forming a hybrid recognition mechanism of rules plus models, so that a high recognition accuracy can still be maintained when facing boundary samples, thereby improving the accuracy of identifying the user dialogue interaction state as an interruption state.

[0011] In some embodiments, the timing features include voice activity signals, voice energy and signal-to-noise ratio, timing context and historical behavior, and the semantic features include semantic keywords. Identifying the user's current conversation interaction state based on the timing features and the semantic features includes: based on the voice activity signal, voice energy and signal-to-noise ratio, semantic keywords, timing context and historical behavior, identifying whether the user's current conversation interaction state is an interruption state.

[0012] This setting forms a multi-dimensional combined judgment through voice activity signals, voice energy and signal-to-noise ratio, semantic keywords, temporal context and historical behavior, avoiding misjudgments caused by simple judgment based on time overlap thresholds, and improving the ability to perceive the user's true interruption intention and the rationality of the response.

[0013] In some embodiments, identifying the user's current dialogue interaction state based on the timing features and the semantic features includes: identifying the user's current dialogue interaction state as a hesitant state when the speech rate decrease rate in the timing features is greater than or equal to a first preset speech rate decrease rate, the number of pauses in the timing features is greater than or equal to a first preset number of pauses, and the semantic features include hesitant keywords; inputting the timing features and the semantic features into a preset decision tree model to identify whether the user's current dialogue interaction state is a hesitant state when the speech rate decrease rate is greater than a second preset speech rate decrease rate and less than the first preset speech rate decrease rate, the number of pauses is not less than a second preset number of pauses and less than the first preset number of pauses, and the semantic features include hesitant keywords.

[0014] With this setting, since there is a clear threshold difference between the first preset speech speed decrease rate and the second preset speech speed decrease rate, and there is a clear threshold difference between the first preset number of pauses and the second preset number of pauses, the former is used in conjunction with hesitation keywords to directly determine the hesitation state, and the latter is used in conjunction with hesitation keywords to trigger the decision tree model to assist in judgment, so as to form a hybrid recognition mechanism of rules plus models, so that a high recognition accuracy can still be maintained when facing boundary samples, thereby improving the accuracy of identifying the user dialogue interaction state as a hesitation state.

[0015] In some embodiments, the method further includes: when the user's current dialogue interaction state is a silent state, triggering a first response strategy, the first response strategy representing entering the next round of prompts; when the user's current dialogue interaction state is an interruption state, triggering a second response strategy, the second response strategy representing cutting into user voice input; when the user's current dialogue interaction state is a hesitation state, triggering a third response strategy, the third response strategy representing a delayed response and providing candidate sentences.

[0016] This setting makes it easier to respond to silent states, interruption states, and hesitation states, significantly improving the naturalness and consistency of human-computer interaction.

[0017] In a second aspect, an embodiment of the present invention provides a conversation interaction state recognition system, the system comprising: an acquisition module for extracting temporal features and semantic features from the acquired voice information; a state recognition module for identifying the user's current silent state based on the temporal features according to a first preset rule; identifying the user's current interruption state based on the temporal features or the semantic features according to a second preset rule; and identifying the user's current hesitation state based on the temporal features and the semantic features according to a third preset rule.

[0018] In some embodiments, the system includes a response module, which is used to: trigger a first response strategy when the user's current dialogue interaction state is a silent state, and the first response strategy represents entering the next round of prompts; trigger a second response strategy when the user's current dialogue interaction state is an interruption state, and the second response strategy represents cutting into the user's voice input; trigger a third response strategy when the user's current dialogue interaction state is a hesitation state, and the third response strategy represents delaying the response and providing candidate sentences.

[0019] In a third aspect, an embodiment of the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the method for identifying a conversation interaction state as described in the first aspect.

[0020] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for identifying a conversation interaction state as described in the first aspect.

[0021] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0023] Figure 1 A flow chart of a method for identifying a conversation interaction state provided by an embodiment of the present invention; Figure 2 for Figure 1 Flowchart of sub-steps S202 to S201 of step S200; Figure 3 for Figure 1 Flowchart of sub-steps S203 to S204 of step S200; Figure 4 for Figure 1 Flowchart of sub-steps S205-S206 of step S200; Figure 5 A schematic diagram of the functional modules of a conversation interaction state recognition system provided by an embodiment of the present invention; Figure 6 A schematic diagram of an electronic device provided by an embodiment of the present invention.

[0024] Icon: 1000 - dialogue interaction state recognition system; 1100 - acquisition module; 1200 - state recognition module; 1300 - response module; 2000 - electronic device; 2100 - processor; 2200 - memory; 2300 - bus; 2400 - communication interface. DETAILED DESCRIPTION

[0025] The following will be combined with the accompanying drawings to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations.

[0026] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention as claimed, but is merely intended to represent selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative work are within the scope of protection of the present invention.

[0027] It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

[0028] The following is a brief introduction to some concepts that may be involved in the embodiments of the present invention.

[0029] VAD: Voice Activity Detection.

[0030] ASR: Automatic Speech Recognition.

[0031] TTS: Text-to-Speech.

[0032] Fuzzy words: words that express uncertainty, such as "maybe", "seem", and "should".

[0033] Repetitive words: Words or phrases that appear repeatedly when the user is hesitant or thinking.

[0034] As described in the background technology, traditional human-computer voice interaction mostly focuses on the analysis of voice signals at a fixed threshold, resulting in the recognition accuracy of complex interaction states such as hesitation and interruption still needing to be improved.

[0035] To this end, the embodiment of the present invention provides a method for identifying a dialog interaction state. Figure 1 , Figure 1 This is a flow chart of a method for identifying a user's conversational interaction state, provided by an embodiment of the present invention. This method identifies the user's conversational interaction state by combining temporal and semantic features. It not only considers voice signal analysis but also focuses on changes in voice rhythm, voice content, and voice intensity to accurately identify the user's conversational interaction state, thereby improving the recognition accuracy of complex interaction states such as hesitation and interruption. The method includes steps S100 to S200: S100. Extract the temporal features and semantic features from the obtained speech information; the temporal features characterize the rhythm change of the user during the conversation, and the semantic features characterize the speech content and speech intensity of the user during the conversation.

[0036] In this embodiment, the speech information specifically includes the original speech stream emitted during the interaction with the user, with a sampling rate of 16 kHz, and its start and end times are determined by voice activity detection (VAD). The temporal features are defined as characterizing the rhythm change of the user during the conversation, specifically including the start and end times of speech segments, speech rate (in words per second), volume change, pause duration, overlap time, etc. The semantic features are defined as characterizing the speech content and speech intensity of the user during the conversation, specifically including semantic-level information such as keywords, fuzzy words, repeated structures, semantic offsets, etc. in the text content transcribed by automatic speech recognition (ASR), as well as audio features such as speech energy, signal-to-noise ratio, and speech intensity change.

[0037] In the process of extracting the temporal features and semantic features, it is necessary to perform time alignment analysis by combining the start and end timestamps of TTS with the start and end timestamps of the user's speech to determine whether there is a speech overlap situation. Specifically, TTSStart represents the absolute time when TTS starts playing, TTSEnd represents the absolute time when TTS finishes playing; SpeechStart represents the absolute time when the user starts speaking, and SpeechEnd represents the absolute time when the user finishes speaking. If SpeechStart < TTSEnd and SpeechEnd > TTSStart, it is considered that there is a time overlap between the user's speech and the speech, thus providing basic data support for subsequent judgment of interruption behavior.

[0038] S200. Identify the current dialogue interaction state of the user according to the temporal features and semantic features; the dialogue interaction state includes one of the silent state, interruption state, and hesitation state.

[0039] In this embodiment, a preliminary judgment is made according to the set rule thresholds. When the temporal feature and semantic feature data fall within the preset boundary range, it is directly determined as one of the silent state, interruption state, and hesitation state; when the temporal feature and semantic feature data are in a critical state, the decision tree model is triggered to intervene to improve the recognition accuracy and further identify the dialogue interaction state as one of the silent state, interruption state, and hesitation state.

[0040] Exemplarily, when the current dialogue interaction state of the user is the silent state, refer to Figure 2 , Figure 2 is Figure 1 the flowchart of sub-steps S201~S202 of step S200 in S201: When a pause duration in the time series feature is greater than or equal to a first preset pause duration, identify that the user's current dialogue interaction state is a silent state.

[0041] In this embodiment, the pause duration specifically refers to the time interval between the end of one user's voice input and the start of the next voice input, which is detected and recorded by voice activity detection (VAD). It can be understood that the first preset pause duration is set as a threshold for directly determining that the user is in a silent state, typically 2.5 seconds. If the pause duration between the user's voice segments is greater than or equal to the first preset pause duration, the user's current dialogue interaction state is directly identified as a silent state. Furthermore, this judgment logic can be combined with the TTS playback status. For example, when content is being broadcast, if the user does not issue a voice signal for more than 2.5 seconds, it can be determined that the user has no intention of interrupting the current content, and the next round of prompting can be entered accordingly.

[0042] S202: When the pause duration is greater than the second preset pause duration and less than the first preset pause duration, input the temporal features and the semantic features into a preset decision tree model to identify whether the user's current dialogue interaction state is a silent state.

[0043] In this embodiment, considering the diversity of user speech behavior, simply using a fixed threshold for judgment may not cover all situations. For example, some users may experience long pauses of nearly 2.5 seconds while thinking, but they are not truly silent. In another example, environmental noise or speech recognition errors may cause VAD misjudgment, thereby mistakenly mistaking valid speech for pauses. To this end, an auxiliary judgment mechanism based on a decision tree model is introduced. Specifically, when it is detected that the pause duration between user speech segments is greater than the second preset pause duration (e.g., 2.2 seconds) and less than the first preset pause duration (e.g., 2.5 seconds), the temporal features and semantic features corresponding to the pause duration are input into the preset decision tree model to further determine whether the user is truly silent.

[0044] At this time, the temporal features include but are not limited to pause time, speech rate change rate, overlapping time, etc., while the semantic features include the frequency and position of ambiguous words and repeated words, speech energy fluctuations, noise ratio, etc. These features serve as input variables of the decision tree model to assist the decision tree model in making more precise classification judgments on boundary samples. For example, when the user pauses for 2.3 seconds, and it is detected that he has made hesitant expressions such as "um..." and "should..." many times before the pause, accompanied by a decrease in speech energy and a slowdown in speech speed, the decision tree model is more inclined to judge that the sample is not in a silent state, but in a hesitant state; on the contrary, if the user's speech content before and after the pause is complete and the semantics are clear, and there is no semantic variation or speech intensity fluctuation, the decision tree model is more likely to support it as a silent state.

[0045] It should be noted that the decision tree model in the embodiments of the present invention is designed as a lightweight model, such as a single tree model trained using the CART (Classification and Regression Tree) algorithm (this applies to all decision tree models used in the embodiments below). This model has the advantages of simple structure, fast inference speed, and low deployment cost, and is particularly suitable for edge computing devices or resource-constrained voice interaction terminals. In addition, the model is only called when the rule is judged to be in a critical state, and normal samples are still judged by the rule, thereby effectively reducing the frequency of model calls and improving overall processing efficiency and throughput.

[0046] For example, when the user's current dialogue interaction state is interrupted, refer to Figure 3 , Figure 3 for Figure 1 Flowchart of sub-steps S203-S204 of step S200, step S200 further includes steps S203-S204: S203. When the overlapping duration in the temporal features is greater than or equal to a first preset overlapping duration and the semantic features include continuous speech, identify that the user's current dialogue interaction state is an interruption state; continuous speech is voice content continuously input by the user in one voice input.

[0047] In this embodiment, the overlapping duration specifically refers to the overlapping part between the start and end time of the user's voice and the start and end time of the TTS playback, and its calculation is based on a unified time reference. It can be understood that the first preset overlapping duration is set as a threshold for directly determining that the user is in an interruption state, which is usually 200 milliseconds. If it is detected that the overlapping duration of the user's voice and the TTS playback is greater than or equal to the first preset overlapping duration, and the semantic features contain continuous speech content, the user's current dialogue interaction state is identified as an interruption state. At this time, continuous speech is defined as the voice content continuously input by the user in a voice input, that is, there is no obvious pause or voice interruption during a speech, and the voice content has strong semantic coherence.

[0048] In some other embodiments, the timing features include voice activity signals, voice energy and signal-to-noise ratio, timing context and historical behavior. Based on the voice activity signals, voice energy and signal-to-noise ratio, semantic keywords, timing context and historical behavior, it is identified whether the user's current conversation interaction state is an interruption state.

[0049] In this embodiment, the voice activity signal, generated by the Voice Activity Detection (VAD) module, is first analyzed to determine whether the user is uttering valid speech at the current time. If a user voice activity signal is detected during TTS playback, it indicates that the user may be attempting to interrupt. Furthermore, combining speech energy and signal-to-noise ratio analysis can effectively distinguish valid speech from ambient noise. For example, when the user's speech energy is significantly higher than the background noise and the signal-to-noise ratio is high, it can be determined as valid speech input, thus supporting the recognition of interruptions.

[0050] Based on this, the system further analyzes semantic keywords in the speech content, such as "wait a moment," "stop talking," and "let me do it," typical expressions of interruption intent. Detecting these keywords further strengthens the confidence level in the user's interruption intent. Furthermore, the temporal context analysis module assesses whether the current speech behavior aligns with the logical flow of the conversation. For example, if a user interrupts a long sentence during playback, the overlapping speech in this context is more likely to be a valid interruption.

[0051] In addition, the system incorporates historical user behavior patterns as a basis for judgment. For example, some users tend to interrupt frequently during past interactions, while others do so less frequently. By recording users' historical interactions in the same or similar scenarios, a personalized behavior model is established to more accurately determine whether they actually intend to interrupt in the current state. For example, for users who habitually interrupt, if the overlap time is slightly shorter than the threshold but the speech content is complete and the speech energy is stable, they may still be judged as interrupting. However, for users who rarely interrupt, a clearer overlap time and speech content are required before they can be judged as interrupting.

[0052] This demonstrates that by combining multiple dimensions of information, including voice activity signals, voice energy and signal-to-noise ratio, semantic keywords, temporal context, and historical behavior, a comprehensive judgment of interruption status is achieved. This judgment mechanism not only improves the accuracy of identifying user interruption intent but also enhances its adaptability to diverse user behavior patterns, resulting in excellent recognition performance and user experience value in practical applications.

[0053] S204. When the overlap duration is greater than the second preset overlap duration and less than the first preset overlap duration, and the speech intensity in the semantic feature indicates that the speech intensity is unstable, the timing feature and the semantic feature are input into a preset decision tree model to identify whether the user's current dialogue interaction state is an interruption state.

[0054] In this embodiment, taking into account the diversity of user interruption behaviors and the uncertainty of voice recognition, there is still a risk of misjudgment based solely on the combined judgment of overlapping duration and voice content integrity. For this reason, an auxiliary judgment mechanism based on a decision tree model is introduced. Specifically, when it is detected that the overlapping duration of the user's voice and TTS playback is greater than the second preset overlapping duration (for example, 180 milliseconds) and less than the first preset overlapping duration (for example, 200 milliseconds), and the voice intensity in the semantic feature indicates that the voice intensity is unstable (for example, the voice energy fluctuates greatly, there is background noise interference, the voice intensity fluctuates, etc.), the temporal features and semantic features corresponding to the overlapping time are input into the preset decision tree model to further determine whether the user is actually in an interruption state.

[0055] Temporal features include, but are not limited to, the start and end times of the speech segments, speech rate, overlap time, and voice activity detection signals, while semantic features include speech intensity, speech energy, signal-to-noise ratio, keywords (such as "wait a moment," "let me talk"), and semantic completeness. These features serve as input variables for the decision tree model, assisting it in making more refined classification decisions for boundary samples. For example, if the overlap between the user's speech and the TTS playback is 190 milliseconds, and the speech content is not completely coherent, but the speech intensity suddenly increases and contains interruption keywords such as "stop talking," the model tends to identify it as an interruption. Conversely, if the speech intensity change is small and the speech content does not clearly indicate an interruption, it may be identified as a non-interruption.

[0056] It should be noted that measuring speech intensity instability primarily relies on energy analysis of the speech signal. Specifically, after acquiring the user's speech stream, it is first divided into several short time frames (typically 10ms to 30ms in length). The energy of each frame is then calculated to generate time series data of speech intensity. The speech intensity mentioned here specifically includes, but is not limited to, the following metrics: Frame energy, for example, calculates the energy value of each speech frame after windowing (e.g., a Hamming window) the speech signal. This value is the mean of the squared amplitude of the speech signal. Frame energy reflects the intensity of the speech frame and is commonly used to measure the loudness of speech. Short-term energy, based on frame energy, calculates the sum or mean of the frame energies within a time window (e.g., 200ms) to characterize the overall intensity of the speech segment. Energy fluctuation, for example, calculates the energy variation between adjacent speech frames, or the energy variance within a specific time window, to measure speech intensity stability. A large energy fluctuation indicates unstable speech intensity. For example, the signal-to-noise ratio (SNR) measures the energy ratio of the speech component to the background noise component in the speech signal to determine whether the speech intensity is affected by background noise. A low SNR indicates that the speech intensity may be unstable due to noise. Another example is the peak energy to average energy ratio, which compares the ratio of the maximum energy frame to the average energy frame in a speech segment to determine whether the speech intensity fluctuates. If this ratio exceeds a preset threshold, the speech intensity is considered unstable.

[0057] Furthermore, in actual applications, voice activity detection (VAD) results are combined to perform contextual filtering on speech intensity measurements. For example, energy calculation is performed only on speech frames marked as "voice activity" by VAD, thereby avoiding misinterpreting background noise as valid speech intensity changes. Furthermore, the time window length for energy calculation can be dynamically adjusted to accommodate users with varying speaking speeds and rhythms. For example, when a user says "Wait a minute!", their speech intensity is typically high and the energy distribution is concentrated. By calculating their short-term energy and peak energy, we can determine that the speech intensity is stable. However, when a user expresses hesitation, using speech such as "Hmm... Hmm... Loan... Loan," their speech intensity may fluctuate greatly. By calculating the frame energy variance and signal-to-noise ratio, we can identify unstable speech intensity, thereby supporting the determination of hesitant or borderline states and further determining whether to invoke a decision tree model for auxiliary recognition.

[0058] It can be seen that the measurement of speech intensity is achieved through a comprehensive analysis of multiple dimensions of the speech signal, such as short-time energy, energy fluctuation rate, signal-to-noise ratio, etc. It is not only used to determine the intensity of the speech content, but also to identify whether the speech intensity is stable, thereby providing important audio-level support for judging the user's current interaction status.

[0059] For example, when the user's current dialogue interaction state is in a hesitant state, refer to Figure 4 , Figure 4 for Figure 1 Flowchart of sub-steps S205-S206 of step S200, step S200 further includes steps S205-S206: S205. When the speech rate decrease rate in the time series feature is greater than or equal to the first preset speech rate decrease rate, the number of pauses in the time series feature is greater than or equal to the first preset number of pauses, and the semantic feature includes hesitation keywords, identify the user's current dialogue interaction state as a hesitation state.

[0060] In this embodiment, the speech rate decrease rate specifically refers to the rate of change of the user's speech rate in the current speech segment compared to the previous speech segment, and its calculation method can be evaluated based on the rate of change of the number of words expressed per unit time. It can be understood that the first preset speech rate decrease rate is set as a threshold for directly determining that the user is in a hesitant state, for example, the user's current speech rate decreases by more than 50%. If it is detected that the speech rate decrease rate is greater than or equal to the first preset speech rate decrease rate, and the number of pauses in the time series feature is greater than or equal to the first preset number of pauses (for example, 3 times), and the semantic feature contains hesitant keywords (such as "maybe", "should", "almost", "not very clear", "I think", "is it", etc. uncertain expressions), then the user's current dialogue interaction state is directly identified as a hesitant state.

[0061] S206. When the speech rate decrease rate is greater than the second preset speech rate decrease rate and less than the first preset speech rate decrease rate, the number of pauses is not less than the second preset number of pauses and less than the first preset number of pauses, and the semantic features include hesitation keywords, the temporal features and the semantic features are input into the preset decision tree model to identify whether the user's current dialogue interaction state is a hesitation state.

[0062] In this embodiment, taking into account the diversity of user expressions, the hesitation behavior of some users may not fully reach the threshold set by the above rules, but still have obvious hesitation characteristics. For this reason, an auxiliary judgment mechanism based on a decision tree model is introduced. Specifically, when it is detected that the speech speed drop rate is greater than the second preset speech speed drop rate (for example, 45%) and less than the first preset speech speed drop rate (for example, 50%), the number of pauses is not less than the second preset number of pauses (for example, 2 times) and less than the first preset number of pauses (for example, 3 times), and the semantic features contain hesitation keywords, the speech speed drop rate, number of pauses and hesitation keywords are used as input features and input into the preset decision tree model to further determine whether the user is actually in a hesitant state.

[0063] The temporal features at this time include but are not limited to the rate of speech speed drop, the number of pauses, the trend of speech speed change, the length of the speech segment, etc., while the semantic features include the frequency and position of hesitant keywords, the use of repeated words, the degree of semantic deviation, etc. These features are used as input variables of the decision tree model to assist the model in making more precise classification judgments on boundary samples. For example, when the user's speech speed drops by 48%, the number of pauses is 2, and the voice content contains repeated or hesitant expressions such as "um...um...loan...loan" or "I...I want to check...", the model will tend to judge the sample as being in a hesitant state; on the contrary, if there is a drop in speech speed but no obvious hesitant keywords or a small number of pauses, it may be judged as a non-hesitant state.

[0064] In some embodiments, in order to respond to identifying the user's current dialogue interaction state, embodiments of the present invention provide a first response strategy corresponding to the silent state, a second response strategy corresponding to the interruption state, and a third response strategy corresponding to the hesitation state.

[0065] Exemplarily, when the user's current dialogue interaction state is silent, the first response strategy is triggered, and the first response strategy represents entering the next round of prompts.

[0066] In this embodiment, silence specifically refers to a situation where the user does not provide valid voice input after being asked a question or making a statement, and the pause duration reaches or exceeds a preset threshold (e.g., 2.5 seconds). In this case, the user is determined to have no intention of continuing the current topic or may not be ready to respond, so the next prompting process is proactively entered. For example, in an automotive finance scenario, during the loan application process, the question may be: "Are you applying for a new car loan or a used car loan?" If the user does not respond within the specified time, an automatic prompt will be displayed: "If you need a new car loan, please say 'new car loan'; if you need a used car loan, please say 'used car loan'; you can also say 'return' to exit." In addition, other auxiliary guidance information may be provided, such as: "If you feel uncomfortable speaking, you can complete the application later on the mobile app." This response strategy is designed to avoid interruptions in interaction and improve the consistency of the user experience.

[0067] Exemplarily, when the user's current dialogue interaction state is an interruption state, the second response strategy is triggered, and the second response strategy represents switching to the user's voice input.

[0068] In this embodiment, an interruption occurs when a user inserts voice input during TTS playback and meets the interruption criteria (e.g., the voice and TTS playback overlap for more than 200 milliseconds, or the user is accompanied by continuous speech). If this judgment is met, the current TTS playback is immediately stopped and the voice recognition process is switched to capture the user's voice input. For example, if the announcement is: "Hello, your current bill is 1,500 yuan. Please pay it by the 25th of this month." If the user suddenly says "Wait a moment!" at the part "Please pay this month...", the system detects the interruption, quickly stops TTS playback, switches to voice recognition mode, and waits for the user to continue speaking. This response strategy enables more timely response to user interruptions, avoiding the unnatural phenomenon of traditional voice announcements continuing after the user interrupts, thereby improving the real-time and responsive nature of the interaction.

[0069] Exemplarily, when the user's current dialogue interaction state is a hesitant state, the third response strategy is triggered, and the third response strategy represents a delayed response and provides candidate sentences.

[0070] In this embodiment, hesitation specifically refers to a situation where the user's speech expresses characteristics such as a decrease in speech speed, frequent pauses, and the use of ambiguous or repetitive words, and also meets preset conditions (e.g., a decrease in speech speed exceeding 50%, more than three pauses, and the inclusion of ambiguous expressions such as "maybe," "should," and "almost"). If these conditions are met, the user will not respond immediately, but will instead delay the response by 500 to 1000 milliseconds to simulate a humane "listening and waiting" interaction. After the delayed response, the user will proactively offer alternative sentences or clarifications to assist the user in completing their expression. For example, in an auto finance scenario, if a user hesitantly expresses, "Hmm... I might be looking for a loan... probably for a new car...", the user can respond with, "Are you applying for a new or used car loan?" or provide further guidance, "If you're unsure, you can start by saying 'new car' or 'used'." Furthermore, the user can introduce relevant business details, such as, "I can explain the difference between new and used car loans. Do you need one?" This response strategy aims to enhance interaction guidance and smooth user expression, reducing interruptions or misunderstandings caused by hesitation.

[0071] In some embodiments, after making a corresponding response strategy, voice playback needs to be performed, and fade-in and fade-out processing is adopted based on the two situations of normal playback start or end and "quick but not abrupt" end when an interruption is detected.

[0072] In this embodiment, when TTS playback begins, the audio volume is gradually increased linearly or nonlinearly to achieve a fade-in effect, for example, gradually increasing the volume from 0 to the target volume, with the duration controlled within 500 milliseconds. When TTS playback ends or is interrupted, the audio volume is gradually decreased linearly or nonlinearly to achieve a fade-out effect, for example, gradually decreasing the volume from the target volume to 0, to avoid auditory discomfort caused by sudden audio changes. After the user finishes speaking, playback is resumed from the starting position or interruption position of TTS playback to ensure complete output of the TTS content.

[0073] Exemplarily, resuming playback from the starting position or interruption position of TTS playback corresponds to a scenario where the user interrupts the speech during TTS playback. For example, when "Please complete the repayment before the 25th of this month" is playing, the user suddenly says "Wait a minute!" After the interruption is detected, the TTS playback interruption is immediately triggered and fade-out processing is performed. After the user finishes speaking, you can choose to replay from the starting position of TTS or resume playback from the last interruption position. This mechanism records the TTS playback progress (such as audio frame index or timestamp), saves the playback status when interrupted, and reloads it from the cached audio data when resuming, thereby achieving semantic integrity of the context and playback continuity.

[0074] For example, let's take a car finance assistant. The assistant says, "Hello, your current bill is 1,500 yuan. Please pay it back by the 25th of this month." In the first scenario, normal playback occurs, with the TTS fade-in starting at 500ms and fading out at 500ms. In the second scenario, an interruption occurs. When the assistant reaches "Please pay back this month..." and the user suddenly says, "Wait a minute!", the assistant detects the interruption, quickly fades out, stops the TTS, and switches to user speech recognition. In the third scenario, the assistant plays again after the interruption. After the user finishes speaking, the assistant continues playing, "Please pay back by the 25th of this month," completing the unfinished content.

[0075] Based on the above method, the embodiment of the present invention also provides a system corresponding to the above method, such as Figure 5 As shown, Figure 5 This is a functional module diagram of a conversation interaction state recognition system 1000 according to an embodiment of the present invention. It should be noted that the basic principles and technical effects of the conversation interaction state recognition system 1000 provided in this embodiment are the same as those of the aforementioned method embodiment. For the sake of brevity, reference is made to the corresponding content in the method embodiment for any parts not mentioned in this embodiment.

[0076] In this embodiment, the dialog interaction state recognition system 1000 includes an acquisition module 1100 and a state recognition module 1200. The acquisition module 1100 is used to extract temporal features and semantic features from the acquired speech information. It can be understood that the acquisition module 1100 is used to perform the above step S100.

[0077] State recognition module 1200 is configured to identify the user's current silent state based on temporal features according to a first preset rule; identify the user's current interruption state based on temporal features or semantic features according to a second preset rule; and identify the user's current hesitation state based on temporal features and semantic features according to a third preset rule. It is understood that state recognition module 1200 is configured to perform step S200 above.

[0078] In some embodiments, state identification module 1200 is further configured to identify the user's current conversational interaction state as silent if the pause duration in the temporal features is greater than or equal to a first preset pause duration. If the pause duration is greater than a second preset pause duration but less than the first preset pause duration, the temporal features and semantic features are input into a preset decision tree model to identify whether the user's current conversational interaction state is silent. It is understood that state identification module 1200 is also configured to perform steps S201 and S202 above.

[0079] In some embodiments, state recognition module 1200 is further configured to identify the user's current conversational interaction state as interrupting when the overlap duration in the temporal features is greater than or equal to a first preset overlap duration and the semantic features include continuous speech; continuous speech refers to speech content continuously input by the user during a single speech input. If the overlap duration is greater than a second preset overlap duration and less than the first preset overlap duration, and the speech intensity in the semantic features indicates unstable speech intensity, the temporal features and semantic features are input into a preset decision tree model to identify whether the user's current conversational interaction state is interrupting. It is understood that state recognition module 1200 is also configured to perform steps S203 and S204 above.

[0080] In some embodiments, the temporal features include voice activity signals, voice energy and signal-to-noise ratio, temporal context and historical behavior, and the semantic features include semantic keywords. The state recognition module 1200 is also used to identify whether the user's current dialogue interaction state is an interruption state based on the voice activity signal, voice energy and signal-to-noise ratio, semantic keywords, temporal context and historical behavior.

[0081] In some embodiments, state identification module 1200 is further configured to identify the user's current conversational interaction state as a hesitant state if the speech rate decrease rate in the temporal features is greater than or equal to a first preset speech rate decrease rate, the number of pauses in the temporal features is greater than or equal to a first preset number of pauses, and the semantic features include a hesitant keyword. Furthermore, if the speech rate decrease rate is greater than a second preset speech rate decrease rate and less than the first preset speech rate decrease rate, the number of pauses is not less than the second preset number of pauses and less than the first preset number of pauses, and the semantic features include a hesitant keyword, the temporal features and semantic features are input into a preset decision tree model to identify whether the user's current conversational interaction state is a hesitant state. It is understood that state identification module 1200 is further configured to perform steps S205-S206 above.

[0082] In some embodiments, the dialogue interaction state recognition system 1000 also includes a response module 1300, which is used to: trigger a first response strategy when the user's current dialogue interaction state is a silent state, and the first response strategy represents entering the next round of prompts; trigger a second response strategy when the user's current dialogue interaction state is an interruption state, and the second response strategy represents cutting into the user's voice input; trigger a third response strategy when the user's current dialogue interaction state is a hesitation state, and the third response strategy represents delaying the response and providing candidate sentences.

[0083] Based on the same inventive concept as disclosed above, correspondingly, the embodiment of the present invention further provides a block diagram of an electronic device 2000 for executing the above method, please refer to Figure 6 , Figure 6 This is a schematic diagram of an electronic device 2000 provided in an embodiment of the present invention. The electronic device 2000 includes a processor 2100, a memory 2200, a bus 2300, and a communication interface 2400. The processor 2100 and the memory 2200 are connected via the bus 2300, and the processor 2100 communicates with external devices via the communication interface 2400.

[0084] The processor 2100 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits in the processor 2100 or by software instructions. The above processor 2100 may be a general-purpose processor 2100, including a central processing unit 2100 (CPU), a network processor 2100 (NP), etc.; it may also be a digital signal processor 2100 (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0085] The memory 2200 is used to store computer programs. For example, the conversation interaction state recognition system 1000 in the embodiment of the present invention includes at least one software function module that can be stored in the memory 2200 in the form of software or firmware. After receiving an execution instruction, the processor 2100 executes the program to implement the conversation interaction state recognition method in the embodiment of the present invention.

[0086] The memory 2200 may include a high-speed random access memory 2200 (RAM) or a non-volatile memory 2200. Alternatively, the memory 2200 may be a storage device built into the processor 2100 or a storage device independent of the processor 2100.

[0087] The bus 2300 may be an ISA bus 2300, a PCI bus 2300, an EISA bus 2300, or the like. Figure 6 The use of only one bidirectional arrow does not mean that there is only one bus 2300 or only one type of bus 2300 .

[0088] The electronic device 2000 may be a computer device such as a mobile phone, a tablet computer, a laptop computer, or a desktop computer.

[0089] Based on the same inventive concept, an embodiment of the present invention further provides a computer-readable storage medium storing a computer program that, when executed by the processor 2100, implements the above-described method for identifying a dialog interaction state. The computer-readable storage medium may include a USB flash drive, a mobile hard drive, a read-only memory 2200 (ROM), a random access memory 2200 (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0090] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for identifying a dialogue interaction state, characterized in that: The method comprises: Extracting temporal features and semantic features from the acquired voice information; the temporal features characterize changes in the user's rhythm during the conversation, and the semantic features characterize the user's voice content and voice intensity during the conversation; The user's current dialogue interaction state is identified based on the temporal features and the semantic features; the dialogue interaction state includes one of a silent state, an interruption state, and a hesitation state, wherein: If the speech rate decrease rate in the time series feature is greater than or equal to a first preset speech rate decrease rate, the number of pauses in the time series feature is greater than or equal to a first preset number of pauses, and the semantic feature includes a hesitation keyword, identifying the user's current dialogue interaction state as a hesitation state; When the speech rate decrease rate is greater than the second preset speech rate decrease rate and less than the first preset speech rate decrease rate, the number of pauses is not less than the second preset number of pauses and less than the first preset number of pauses, and the semantic features include hesitant keywords, the temporal features and the semantic features are input into a preset decision tree model to identify whether the user's current dialogue interaction state is a hesitant state.

2. The method according to claim 1, characterized in that The identifying the user's current dialogue interaction state according to the temporal features and the semantic features includes: When the pause duration in the time series feature is greater than or equal to a first preset pause duration, identifying that the user's current dialogue interaction state is a silent state; When the pause duration is greater than a second preset pause duration and less than the first preset pause duration, the temporal features and the semantic features are input into a preset decision tree model to identify whether the user's current dialogue interaction state is a silent state.

3. The method according to claim 1, characterized in that The identifying the user's current dialogue interaction state according to the temporal features and the semantic features includes: When the overlapping duration in the temporal features is greater than or equal to a first preset overlapping duration and the semantic features include continuous speech, identifying the user's current dialogue interaction state as an interruption state; the continuous speech is speech content continuously input by the user in one speech input; When the overlap duration is greater than the second preset overlap duration and less than the first preset overlap duration, and the speech intensity in the semantic feature indicates that the speech intensity is unstable, the timing feature and the semantic feature are input into a preset decision tree model to identify whether the user's current dialogue interaction state is an interruption state.

4. The method according to claim 1, wherein The temporal features include voice activity signals, voice energy and signal-to-noise ratio, temporal context, and historical behavior; the semantic features include semantic keywords; and identifying the user's current conversation interaction state based on the temporal features and the semantic features includes: Based on the voice activity signal, voice energy and signal-to-noise ratio, semantic keywords, temporal context and historical behavior, it is identified whether the user's current dialogue interaction state is an interruption state.

5. The method according to claim 1, wherein The method further comprises: When the user's current dialogue interaction state is silent, triggering a first response strategy, wherein the first response strategy indicates entering a next round of prompts; When the current dialogue interaction state of the user is an interruption state, triggering a second response strategy, wherein the second response strategy represents switching to user voice input; When the user's current dialogue interaction state is a hesitant state, a third response strategy is triggered, wherein the third response strategy represents a delayed response and provides candidate sentences.

6. A conversation interaction state recognition system, characterized in that: The system comprises: An acquisition module is used to extract temporal features and semantic features from the acquired voice information, wherein the temporal features represent the rhythm changes of the user during the conversation, and the semantic features represent the voice content and voice intensity of the user during the conversation; A state recognition module is configured to recognize the user's current conversation interaction state based on the temporal features and the semantic features; the conversation interaction state includes one of a silent state, an interruption state, and a hesitation state; the state recognition module is further configured to: If the speech rate decrease rate in the time series feature is greater than or equal to a first preset speech rate decrease rate, the number of pauses in the time series feature is greater than or equal to a first preset number of pauses, and the semantic feature includes a hesitation keyword, identifying the user's current dialogue interaction state as a hesitation state; When the speech rate decrease rate is greater than the second preset speech rate decrease rate and less than the first preset speech rate decrease rate, the number of pauses is not less than the second preset number of pauses and less than the first preset number of pauses, and the semantic features include hesitant keywords, the temporal features and the semantic features are input into a preset decision tree model to identify whether the user's current dialogue interaction state is a hesitant state.

7. The conversation interaction state recognition system according to claim 6, characterized in that: The system includes a response module configured to: When the user's current dialogue interaction state is silent, triggering a first response strategy, wherein the first response strategy indicates entering a next round of prompts; When the current dialogue interaction state of the user is an interruption state, triggering a second response strategy, wherein the second response strategy represents switching to user voice input; When the user's current dialogue interaction state is a hesitant state, a third response strategy is triggered, wherein the third response strategy represents a delayed response and provides candidate sentences.

8. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the method for identifying a conversation interaction state according to any one of claims 1 to 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying a dialog interaction state according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • All-media interactive intelligent customer service interaction method and system

    CN118428343A

  • Dialogue round end judgment method and device, electronic equipment, medium and vehicle

    CN119207485A

  • AI voice call method and system, program product and storage medium

    CN119583716A

  • Telephone answering system based on AI

    CN119854414A

  • Speaker'S situation adaptive voice interactive device and ticket issuing device

    JP2000244609A

Cited By

  • AI interview method, electronic equipment, storage medium and program product

    CN121212135A

  • AI interviewing methods, electronic devices, storage media, and software products

    CN121212135B

  • Breaking management method and device for voice large model dialogue, medium and program product

    CN121214935A