Dialogue robot interaction method, system, medium, product and terminal
By combining automatic speech recognition and speech activity detection algorithms, the chatbot can simultaneously support dual detection of DTMF and ASR, solving the problem of limited interaction flexibility and adaptability in existing technologies, and improving user experience and dialogue efficiency.
Patent Information
- Application Number
- CN202510165719.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-02-14
AI Technical Summary
Existing chatbots only support one of DTMF or ASR, which limits the flexibility and adaptability of user interaction and affects user experience, especially in complex scenarios where it is difficult to meet multiple interaction needs.
By combining automatic speech recognition algorithms and voice activity detection algorithms, the current dialogue state of the interactive object is determined, and an appropriate recognition method is selected based on DTMF key signals or voice data to generate interactive recognition results, supporting dual detection of DTMF and ASR.
It improves the response speed and semantic understanding capabilities of chatbots, enhances their adaptability and flexibility, reduces the possibility of operational errors, and improves user interaction experience and dialogue efficiency.
Smart Images

Figure CN120071923B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of dialogue robots, and in particular to a dialogue robot interaction method, system, medium, product and terminal. BACKGROUND
[0002] In the prior art, when configuring a dialogue robot, only one of DTMF (Dual Tone Multi Frequency) or ASR (Automatic Speech Recognition) is usually supported. If the dialogue robot is configured to only handle DTMF type nodes, it will ignore any ASR input, which means that if a user tries to interact through voice instructions, the system will not be able to recognize and process accordingly. Similarly, if the dialogue robot is configured to only support ASR, information entered by the user through the keypad will also be ignored.
[0003] Firstly, this single configuration limits the flexibility of user interaction. In various scenarios, users may prefer to use one of voice or keypad to interact with the dialogue robot due to personal habits, environmental restrictions, or device conditions. However, when the dialogue robot only supports one of these methods, users have to adapt to this single interaction mode, leading to inconvenience or dissatisfaction, affecting the interaction experience. Especially in scenarios that require quick and accurate confirmation of information, the flexibility of user interaction is particularly important.
[0004] Secondly, with the continuous development of technology and the changing needs of users, dialogue robots need to continuously adapt to new interaction methods and scenarios. However, when the dialogue robot only supports one of DTMF or ASR, its system adaptability will be limited, making it difficult to meet the needs of different scenarios.
[0005] Finally, when the dialogue robot only supports ASR, users need to interact with the robot through voice instructions. However, in the absence of proper speech environment, such as noisy background, unclear pronunciation or fast speech speed, voice instructions may not be accurately recognized by the robot. This can lead to interaction failure, and users may need to repeat the instructions or use other methods (such as manual input) to complete the interaction, thereby reducing the user experience. On the other hand, when the dialogue robot only supports DTMF, users need to input instructions through the keypad. However, in situations where it is not convenient to press the keys, such as when the user's hands are occupied, has poor eyesight or is not convenient to operate, the user cannot effectively confirm the requirements or other instructions, which also leads to interaction failure and reduces user experience. SUMMARY
[0006] In view of the above-mentioned disadvantages of the prior art, the purpose of the present application is to provide a dialogue robot interaction method, system, medium, product and terminal, which are used to solve the technical problems of single configuration mode of the existing dialogue robot, limited interaction flexibility of the user, limited adaptability, and influence on the interaction experience.
[0007] To achieve the above-mentioned purposes and other related purposes, the first aspect of the present application provides a dialogue robot interaction method, comprising: determining a current dialogue state of an interaction object based on an automatic speech recognition algorithm and a voice activity detection algorithm after a dialogue robot sends interaction information; the current dialogue state includes an interaction state; if the current dialogue state is in the interaction state, determining whether a DTMF key signal is received, and acquiring to-be-recognized data input by the interaction object according to the determination result of whether the DTMF key signal is received; generating an interaction recognition result according to the to-be-recognized data; wherein the way of acquiring the to-be-recognized data input by the interaction object according to the determination result of whether the DTMF key signal is received includes: if the DTMF key signal is received, taking the DTMF key signal as the to-be-recognized data; if the DTMF key signal is not received, acquiring voice data of the interaction object, and taking the voice data as the to-be-recognized data.
[0008] In some embodiments of the first aspect of the present application, the way of determining the current dialogue state of the interaction object includes: determining whether there is a first ASR text based on the automatic speech recognition algorithm; if there is the first ASR text, determining that the current dialogue state is the interaction state; if there is no first ASR text, acquiring a current voice continuous time length based on the voice activity detection algorithm, and determining the current dialogue state according to the current voice continuous time length and a preset voice continuous time length.
[0009] In some embodiments of the first aspect of the present application, the current dialogue state further includes a waiting state; the way of determining the current dialogue state according to the current voice continuous time length and the preset voice continuous time length includes: determining whether the current voice continuous time length is greater than or equal to the preset voice continuous time length; if the current voice continuous time length is greater than or equal to the preset voice continuous time length, determining that the current dialogue state is the interaction state; if the current voice continuous time length is less than the preset voice continuous time length, determining that the current dialogue state is the waiting state.
[0010] In some embodiments of the first aspect of the present application, the manner of determining the current dialogue state of the interactive object comprises: upon receiving the DTMF key signal, obtaining a current key waiting duration, and determining whether the current dialogue state is the waiting state based on a preset key waiting duration; if the current key waiting duration is greater than or equal to the preset key waiting duration, determining that the current dialogue state is the waiting state, and taking the received DTMF key signal as the to-be-recognized data; if the current key waiting duration is less than the preset key waiting duration, determining that the current dialogue state is the interactive state.
[0011] In some embodiments of the first aspect of the present application, the manner of generating an interactive recognition result according to the to-be-recognized data comprises: when the voice data of the interactive object is taken as the to-be-recognized data of the interactive object input, generating a second ASR text according to the voice data of the interactive object; and generating an interactive recognition result according to the second ASR text.
[0012] In some embodiments of the first aspect of the present application, the manner of generating an interactive recognition result according to the to-be-recognized data comprises: when the received DTMF key signal is taken as the to-be-recognized data of the interactive object input, performing a value assignment operation on the DTMF key signal to generate a key ASR text; and generating an interactive recognition result according to the key ASR text.
[0013] To achieve the above object and other related objects, the second aspect of the present application provides a dialogue robot interactive system, comprising: an interactive state determination module, configured to determine a current dialogue state of an interactive object based on an automatic speech recognition algorithm and a voice activity detection algorithm after a dialogue robot sends interactive information; the current dialogue state comprises an interactive state; a data acquisition module, configured to determine whether a DTMF key signal is received if the current dialogue state is in the interactive state, and acquire to-be-recognized data input by the interactive object according to a determination result of whether the DTMF key signal is received; the manner comprises: taking the DTMF key signal as the to-be-recognized data if the DTMF key signal is received; and acquiring voice data of the interactive object and taking the voice data as the to-be-recognized data input by the interactive object if the DTMF key signal is not received; and an recognition result generation module, configured to generate an interactive recognition result according to the to-be-recognized data.
[0014] To achieve the above object and other related objects, the third aspect of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the dialogue robot interactive method as described above.
[0015] To achieve the above object and other related objects, the fourth aspect of the present application provides a computer program product, which comprises computer program codes, and when the computer program codes are run on a computer, the computer is caused to implement the dialogue robot interaction method as described above.
[0016] To achieve the above object and other related objects, the fifth aspect of the present application provides an electronic terminal, which comprises a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to implement the dialogue robot interaction method as described above.
[0017] As described above, the dialogue robot interaction method, system, medium, product and terminal of the present application have the following beneficial effects: simultaneously supporting the DTMF and ASR dialogue types, the dialogue robot can simultaneously detect and respond to the two kinds of input signals, when the interactive object expresses the intention through the key or voice, the robot can quickly and accurately identify and process these signals, this double detection mechanism not only improves the response speed of the robot, but also can more quickly identify and respond to the intention of the interactive object, significantly enhances the semantic understanding ability to generate more accurate interaction recognition results, thereby providing a reply more in line with the needs of the interactive object. Further, simultaneously supporting the DTMF and ASR dialogue types makes the dialogue robot have stronger adaptability and flexibility to cope with various complex interaction scenarios, improves the interaction experience of the dialogue robot. In the scene where the key reply and voice reply coexist, more selection space is provided for the interactive object, the possibility of operation failure is reduced, the interactive object experience is improved, the dialogue period is shortened, and the dialogue efficiency is improved. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 A flowchart of the voice interaction is shown.
[0019] Figure 2 A flowchart of the dialogue robot interaction method in an embodiment of the present application is shown.
[0020] Figure 3 A flowchart of the dialogue robot interaction method in an embodiment of the present application is shown.
[0021] Figure 4 A flowchart of the dialogue robot interaction method in an embodiment of the present application is shown.
[0022] Figure 5 A flowchart of the dialogue robot interaction method in an embodiment of the present application is shown.
[0023] Figure 6Fig. 1 shows a schematic block diagram of a dialogue robot interaction system according to an embodiment of the present application.
[0024] Figure 7 Fig. 2 shows a schematic block diagram of an electronic terminal according to an embodiment of the present application. DETAILED DESCRIPTION
[0025] The present application is described in detail below by specific embodiments. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the specification. The present application can also be implemented or applied by other different specific embodiments, and the details in the specification can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0026] Before the present application is further described, the nomenclature and terminology used in the embodiments of the present application are explained, which are applicable to the following explanations:
[0027] <1>DTMF (Dual tone multi frequency): dual tone multi frequency, which is a general name in key telephone signaling. It is most familiar to users as the sound produced when a number key is pressed. It is equivalent to the button dialing system used internally by the Bell System. In telephone communication, when a user presses a key on the telephone keypad, a combination of two different frequencies of sound is produced to transmit information such as numbers.
[0028] <2>ASR (Automatic Speech Recognition): automatic speech recognition algorithm, which is a technology for converting human speech into text. The goal is to convert the content of the words in human language into computer-readable input.
[0029] <3>VAD (Voice Activity Detection): voice activity detection algorithm, also known as speech breakpoint detection or speech boundary detection. The purpose is to determine and eliminate long periods of silence from the sound signal stream to save valuable bandwidth resources and reduce the end-to-end delay perceived by users without degrading service quality.
[0030] In the current dialogue scenarios supported by dialogue robots, such as scenarios of confirming the delivery intention (acceptance or rejection) of users by dialogue robots in the logistics cash on delivery scenarios, confirming the complaint reason of users in the logistics complaint scenarios, confirming the consultation intention (query service, mailing service, information update service or other services, etc.) of users in the logistics customer service scenarios, confirming the purchase intention of users in the retail scenarios, confirming the purchase category of users in the retail scenarios, confirming the payment method of users in the retail scenarios, confirming the purchase of insurance types of users in the insurance service scenarios, confirming the basic information of users in the insurance service scenarios, confirming the repayment intention of users in the financial scenarios, etc., the dialogue nodes requiring information input of users usually only support one of DTMF or ASR, so as to reduce the system configuration depth and shorten the dialogue interaction time by identifying the user intention through a single mode.
[0031] If the dialogue robot is configured to only process the DTMF type node, it will ignore any ASR input, which means that if the user tries to interact through voice instructions, the system will not be able to identify and make corresponding processing. Similarly, if the dialogue robot is configured to only support ASR, the information input by the user through the key input will also be ignored. This single configuration mode, although it can adapt to most simple call interaction scenarios, limits the interaction flexibility of users, affects the user experience, and the system adaptability will be limited, making it difficult to meet the needs of the above dialogue scenarios in which the user input contains a large amount of complex information.
[0032] Taking a dialogue scenario configured with single ASR recognition as an example, as shown in FIG. 1, the dialogue robot is configured to support ASR, and the dialogue node is configured to support DTMF. Figure 1As shown, a flowchart of voice interaction is shown. (1) audio play: the robot starts playing voice. (2) robot announces time: determine whether the robot has announced the corresponding speech. If not, continue to wait; if yes, determine whether the user is speaking. (3) determine whether the user is speaking: if the user is not speaking, the robot enters the "waiting for the user to speak" state; if the user is speaking, determine whether it can be interrupted. (4) whether it can be interrupted: determine whether the interruption condition is met at this time, if not, the robot continues to wait; if yes, determine whether the maximum number of interruptions is exceeded. (5) whether the maximum number of interruptions is exceeded: check whether the interruption operation has exceeded the maximum number of pre-set times. If exceeded, the robot does not interrupt; if not exceeded, determine whether the number of continuous interruptions is exceeded. (6) whether the number of continuous interruptions is exceeded: determine whether the number of continuous interruptions exceeds the set value, if exceeded, the robot does not interrupt; if not exceeded, determine whether it is a modal word. (7) whether it is a modal word: determine whether the content spoken by the user is a modal word, if not, execute the interruption; if it is a modal word, determine whether the modal word interruption is allowed. (8) whether the modal word interruption is allowed (high): determine whether the modal word interruption is allowed, if allowed, execute the interruption operation; if not allowed, do not execute the interruption.
[0033] In Figure 1 In the voice interaction scenario shown, only the ASR voice recognition mode is configured to interact with the user, generate an ASR text based on the user's voice input content, and perform semantic understanding based on the generated ASR text; common semantic understanding schemes include intent recognition through a natural language understanding algorithm (NLU-Natural Language Understanding) and named entity recognition through a named entity recognition algorithm (NER-Named Entity Recognition); a dialog management module of a dialog robot is called to generate a reply content based on the semantic understanding result and a pre-set dialog strategy and feed back to the user. In actual scenarios, taking the example of confirming the delivery intention of the user in the logistics cash on delivery scenario, the user intent that the dialog robot needs to obtain only includes receiving or refusing, but the user may state the receiving or refusing reason or other service requests during the dialog process, which increases the call time and reduces the service quality and efficiency; similarly, if only DTMF is configured, the user may be mistaken, and the user may not be convenient to press the keys; therefore, the ASR and DTMF composite recognition mode is a more suitable dialog scheme for the scenario.
[0034] But directly superimposing the ASR recognition and DTMF recognition scheme, the total time of detecting ASR signal + detecting DTMF signal is too long, which leads to too long call time, affects user interaction experience and increases call cost; in addition, the ASR signal and the DTMF signal interrupt mechanism are different, if the user speaks and presses the key at the same time, or speaks and presses the key alternately, or the user mis-touches, etc., the ASR recognition and DTMF recognition process will further increase the user waiting time of single interaction, which leads to the lack of dialogue solution of the existing dialogue robot in this scene.
[0035] To solve the above problems, the present application provides a dialogue robot interaction method, system, medium, product and terminal, which is used to solve the technical problems of single configuration mode of the existing dialogue robot, limiting the interaction flexibility of the user, limited adaptability, affecting the interaction experience, etc.
[0036] In order to facilitate the understanding of the embodiments of the present application, the embodiments of the present application will be described in combination with Figure 2 The detailed description is as follows. Figure 2 The flowchart of the dialogue robot interaction method in the embodiment of the present application is shown. The dialogue robot interaction method in the embodiment mainly includes the following steps:
[0037] S201: After the dialogue robot sends interaction information, the current dialogue state of the interaction object is determined based on automatic speech recognition algorithm and voice activity detection algorithm; the current dialogue state includes interaction state.
[0038] In the embodiment, the interaction scenario provided includes the case of voice dialogue interaction or text dialogue interaction between the dialogue robot and the interaction object (user or user terminal or third party service platform), and the complete interaction process between the dialogue robot and the user includes several interaction nodes, and the interaction process between the dialogue robot and the user at each interaction node is that the user and the dialogue robot alternately dialogue. Therefore, when configuring the corresponding interaction scheme of the embodiment for one or more interaction nodes, the interaction recognition result is generated by the ASR and DTMF common recognition mode after the dialogue robot sends interaction information to the interaction object.
[0039] In the embodiment, the dialogue robot and the interaction object interaction process adopts the ASR and DTMF common recognition mode, which is specifically to determine which recognition mode of ASR and DTMF to use for information recognition according to two interaction parameters in the interaction process: "current dialogue state" and "whether a DTMF signal is received". When judging the current dialogue state, the voice recognition mode is used by default for the voice dialogue robot to determine whether the interaction object is in the interaction state or in the waiting state after the interaction ends or before the interaction.
[0040] S202: If the current dialogue state is in an interactive state, determine whether a DTMF key signal has been received, and obtain the data to be recognized input by the interactive object based on the determination result of whether a DTMF key signal has been received.
[0041] In this embodiment, the method for obtaining the data to be recognized input by the interactive object based on the determination result of whether a DTMF key signal has been received includes:
[0042] (1) If a DTMF key signal is received, the DTMF key signal is used as the data to be identified.
[0043] (2) If no DTMF key signal is received, the voice data of the interactive object is obtained and the voice data is used as the data to be recognized.
[0044] In this embodiment, the data to be recognized is only acquired from the input of the interactive object during the interactive state; if the chatbot receives voice or text input from the user during the waiting state after the current round of interaction, the currently received user input is ignored; when the interactive object and the chatbot are in an interactive state, there are two situations when inputting interactive content to the chatbot via voice, text, or keystrokes:
[0045] In the first scenario, if only the user's voice input is received but no DTMF key input is received, then the voice data of the interactive object is acquired and used as the data to be recognized.
[0046] The second scenario is when a DTMF key signal is received from the user. In this case, all voice data input by the user during this round of interaction is ignored, and only the DTMF key signal is used as the data to be recognized.
[0047] S203: Generate interactive recognition results based on the data to be recognized.
[0048] Based on the dialogue robot interaction scheme constructed in steps S210 to S203, the default method is to use speech recognition to identify the current dialogue state. When the user uses speech recognition to interact throughout the process, the speech data input by the interaction object is obtained as the data to be recognized and a second ASR text is generated. When the interaction is conducted in the interaction state using DTMF, the current dialogue state judgment method and the data type to be recognized input by the interaction object are switched from speech recognition to DTMF to achieve the coupling of ASR and DTMF recognition schemes.
[0049] In this embodiment, as Figure 3 The diagram illustrates a process for determining the current dialogue state using speech recognition in an embodiment of the present invention. The methods for determining the current dialogue state of the interacting object include:
[0050] S2011: determining whether the first ASR text exists based on the automatic speech recognition algorithm.
[0051] In this embodiment, the automatic speech recognition algorithm (ASR) is a technology for converting human speech into text, aiming to convert the lexical content in human speech into readable text format. After the dialogue robot sends the interactive information, within the preset silent waiting time, i.e. during the stage of waiting for the interactive object to speak, it is first attempted to determine whether the interactive object starts speaking, i.e. whether the voice input starts, through the automatic speech recognition algorithm.
[0052] S2012: if the first ASR text exists, determining that the current dialogue state is the interactive state.
[0053] In this embodiment, if the first ASR text is recognized and the recognized ASR text is not the text before the end of the dialogue robot voice broadcast, it is determined that the interactive object starts speaking and the current dialogue state is the interactive state, and the first ASR text is the process file for detecting whether the interactive object starts speaking. After the dialogue robot sends the interactive information, the interactive information includes the content of the dialogue robot voice broadcast. Due to the physical characteristics of the audio signal (such as echo, reverberation, etc.) or processing delay, the ASR may recognize some residual information related to the broadcast content. These residual information are not the speaking content actively issued by the interactive object, and if these residual information are directly regarded as the speaking content of the interactive object, it will lead to misjudgment. Based on this, after the dialogue robot sends the interactive information, a preset silent waiting time is set to wait for the audio signal to completely decay before starting the ASR recognition, so as to reduce the influence of residual information on the ASR recognition. Further, it is determined whether there is a first ASR text meeting the preset requirement. If the recognized first ASR text is not the text before the end of the dialogue robot voice broadcast, it is proved that the recognized first ASR text is not such residual information, but the first ASR text corresponding to the new voice issued by the interactive object, so as to determine that the interactive object starts speaking. By using the recognition ability of ASR on voice content, it can be more directly judged whether the interactive object has issued meaningful voice information, so as to avoid misjudging the residual information of the voice broadcast as the speaking content of the interactive object, and improve the accuracy of the judgment.
[0054] S2013: if the first ASR text does not exist, obtaining the current voice continuous time based on the voice activity detection algorithm, and determining the current dialogue state according to the current voice continuous time and the preset voice continuous time.
[0055] In this embodiment, if the first ASR text does not exist, the current voice continuous time is obtained based on the voice activity detection algorithm, and the current dialogue state is determined according to the current voice continuous time and the preset voice continuous time. Figure 4As shown, a flowchart of determining the current conversation state by using the voice activity detection algorithm is shown. The current conversation state further includes a waiting state; and the manner of determining the current conversation state includes:
[0056] S2013a: determining whether the current voice continuous time length is greater than or equal to the preset voice continuous time length.
[0057] In this embodiment, if there is no first ASR text meeting the preset requirement, a voice activity detection algorithm (VAD) is used to monitor and analyze the voice signal. The voice activity detection algorithm can distinguish the voice signal and non-voice signal (such as background noise, silence, etc.), and mark the valid voice activity data packet, and continuously record the number of valid voice activity data packets to obtain the current voice continuous time length of the interactive object.
[0058] S2013b: if the current voice continuous time length is greater than or equal to the preset voice continuous time length, determining that the current conversation state is an interactive state.
[0059] In this embodiment, the preset voice continuous time length (usually expressed by the number of data packets or time length) is compared with the current voice continuous time length. Exemplarily, the preset voice continuous time length is 300 ms or 15 voice data packets, and the time length of each voice data packet is 20 ms. The preset voice continuous time length of 300 ms or 15 voice data packets is only an example of the present application and is not used to limit the present application. In actual application, the preset voice continuous time length is determined according to the system requirement, and is used to determine whether the interactive object has issued a continuous voice for a long enough time.
[0060] In this embodiment, if the state of the received voice data packet is 1, it indicates that the voice signal exists and is marked as a valid voice activity data packet. According to the number of valid voice activity data packets, the current voice continuous time length of the interactive object is obtained. If the current voice continuous time length of the interactive object is greater than or equal to the preset voice continuous time length, it is considered that the interactive object has issued a continuous voice for a long enough time, and thus it is determined that the current conversation state of the interactive object is an interactive state.
[0061] It is worth noting that when the automatic speech recognition algorithm cannot accurately recognize the voice, the real-time detection ability of the voice activity detection algorithm is used to judge whether the voice signal exists or not based on the characteristics of the voice signal, which can provide a bottom judgment mechanism when the automatic speech recognition algorithm fails, and improve the interactive experience and the accuracy of voice recognition.
[0062] S2013c: if the current voice continuous time length is less than the preset voice continuous time length, determining that the current conversation state is a waiting state.
[0063] In this embodiment, if the current speech continuous duration of the interactive object is less than the preset speech continuous duration, it is considered that the interactive object does not utter continuous speech for a long enough time, and thus it is determined that the current conversation state of the interactive object is the waiting state, at which time the interactive object is either at the end of the call or has not started to speak.
[0064] In this embodiment, to avoid the user rejecting interaction with the dialogue robot for a long time, the dialogue robot interaction method further includes: if it is determined that the current conversation state of the interactive object is the waiting state, determining whether the interactive object needs to reply based on a preset silent waiting duration; and triggering a waiting speech response label when it is determined that the interactive object needs to reply.
[0065] In this embodiment, after it is determined that the current conversation state of the interactive object is the waiting state, the duration for which the interactive object has not spoken is started to be timed. If the duration for which the interactive object has not started to speak does not exceed the preset silent waiting duration, the interactive object is continuously listened to and the interactive object is waited to speak. If the duration for which the interactive object has not started to speak exceeds the preset silent waiting duration, it is determined whether the interactive object needs to reply. When the robot plays some speech such as "hello", the interactive object does not need to reply, and the next round of dialogue is directly entered or another round of speech recognition is performed. When the robot plays some speech requesting the interactive object to provide specific information, such as "please tell me your name" or "please enter your password", it is determined that the interactive object needs to reply, and a waiting speech response label is triggered. The waiting speech response label is an internal label or signal used to indicate the current state or behavior, to trigger the robot to provide feedback to the interactive object, such as playing a prompt tone, displaying prompt information, or automatically performing some operations (such as repeating the request, providing additional information, etc.) according to the context.
[0066] In this embodiment, after the determination manner of the current conversation state and the data type input by the interactive object are both switched from the speech recognition manner to the DTMF manner, the determination result of the current conversation state based on speech recognition needs to be reset, and the current conversation state is determined in the DTMF manner, as shown in FIG. 8, which shows a flowchart of determining the current conversation state in the DTMF manner in an embodiment of the present application. The manner of determining the current conversation state of the interactive object includes: Figure 5
[0067] S2021: When the DTMF key signal is received, the current key waiting duration is obtained, and it is determined whether the current conversation state is the waiting state based on the preset key waiting duration.
[0068] S2022: If the current key waiting duration is greater than or equal to the preset key waiting duration, it is determined that the current conversation state is the waiting state, and the received DTMF key signal is taken as the to-be-recognized data.
[0069] S2023: If the current key waiting duration is less than the preset key waiting duration, it is determined that the current dialogue state of the interactive object is the interactive state.
[0070] In this embodiment, after determining that the current dialogue state of the interactive object is the interactive state, it is determined whether there is a DTMF key signal based on the preset key waiting duration, and the method includes:
[0071] (1) When the current dialogue state of the interactive object is the interactive state, the pre-detection state of the DTMF key signal is started based on the preset key detection duration.
[0072] (2) From the end of the preset key detection duration to the end of the preset key waiting duration, the detection of the DTMF key signal is continuously performed.
[0073] (3) If the DTMF key signal is detected, it is determined whether the current dialogue state of the interactive object is the waiting state based on the preset key waiting duration.
[0074] In this embodiment, the duration of the voice broadcast of the dialogue robot is 10 ms, the preset key detection duration is 6 ms, that is, the DTMF key signal is detected from the 4th ms. The preset key waiting duration is 20 ms. After the 6 ms countdown ends, the detection of the DTMF key signal is continuously performed until the 20 ms ends, and it is determined whether there is a DTMF key signal. If the DTMF key signal is detected and the DTMF key signal occurs before the end of the voice input of the interactive object, that is, before the end of the speech of the interactive object, and the current key waiting duration is greater than or equal to the preset key waiting duration, it is determined that the current dialogue state of the interactive object is the waiting state, the current dialogue is ended, and the determination timing is restarted, and the ASR signal is ignored.
[0075] In this embodiment, during the interaction, the voice data of the interactive object is continuously acquired, and the voice data of the interactive object contains valid voice information, which can be used for subsequent voice recognition and processing. If no DTMF key signal is received during the interaction, the voice data of the interactive object is taken as the to-be-recognized data of the voice input of the interactive object, and it is determined whether the current dialogue state of the interactive object is the waiting state, that is, whether the voice input of the interactive object is ended or whether the voice input of the interactive object is waited again.
[0076] In the embodiment, when it is detected that the sound signal in the voice activity data packet of the interactive object is lower than the preset sound threshold, and the duration that the sound signal is lower than the preset sound threshold exceeds the preset voice pause duration, and the sound signal is still lower than the preset sound threshold within the preset tail misjudgment duration, it is determined that the current conversation state of the interactive object is the waiting state, the current round of conversation is ended, and the subsequent DTMF key signal is ignored. If the voice signal of the interactive object can still be detected within the preset voice pause duration, it is determined that the current conversation state of the interactive object is the interactive state, and the voice input of the interactive object is not ended. In this way, after the duration that the sound signal is lower than the preset sound threshold exceeds the preset voice pause duration, in order to avoid misjudgment (for example, due to environmental noise or temporary pause of the interactive object), the preset tail misjudgment duration is set to further determine whether there is new voice signal within the preset tail misjudgment duration. If there is no new voice signal during this period, it is finally determined that the current conversation state of the interactive object is the waiting state, and the voice input of the interactive object is ended. The normal pause and possible misjudgment are comprehensively considered, so that the determination of the waiting state is more accurate.
[0077] In the embodiment, for example, the preset voice pause duration is 500 ms or 25 voice data packets, and the duration of each voice data packet is 20 ms. It is indicated that when it is detected that the voice signal of the interactive object appears 500 ms pause, it is possible to consider that the voice input of the interactive object is ended. In order to avoid tail misjudgment (for example, the interactive object continues to speak after a short pause for thinking), the preset tail misjudgment duration is 200 ms or 10 voice data packets, and the duration of each voice data packet is 20 ms. That is, when 500 ms pause is detected, it is not immediately determined that the voice input of the interactive object is ended, but 200 ms is waited. If there is no new voice signal during this period, it is finally determined that the voice input of the interactive object is ended.
[0078] In the embodiment, the way of generating the interactive recognition result according to the to-be-recognized data includes: when the voice data of the interactive object is taken as the to-be-recognized data input by the interactive object, generating a second ASR text according to the voice data of the interactive object; and generating the interactive recognition result according to the second ASR text.
[0079] In the embodiment, the second ASR text is generated based on the voice data of the interactive object, and the second ASR text is output to a text processing interface through an ASR text output interface to generate the interactive recognition result.
[0080] In the embodiment, the way of generating the interactive recognition result according to the to-be-recognized data includes: when the received DTMF key signal is taken as the to-be-recognized data input by the interactive object, performing value assignment operation on the DTMF key signal to generate a key ASR text; and generating the interactive recognition result according to the key ASR text.
[0081] In this embodiment, the DTMF key signal is assigned as a key ASR text, and the key ASR text is output to a text processing interface through an API interface to generate an interactive recognition result.
[0082] In this embodiment, the manner of generating an interactive recognition result according to the to-be-recognized data input by the interactive object includes a named entity recognition algorithm or a natural language understanding algorithm.
[0083] In this embodiment, the named entity recognition algorithm (NER-Named Entity Recognition) is a natural language processing technology that aims to identify and classify specific types of entities from text. These entities usually include names of people, places, organizations, dates and times, and other words with specific meanings.
[0084] In this embodiment, the natural language understanding algorithm (NLU-Natural Language Understanding) is a branch of computer science that focuses on enabling computers to understand and process the intent and meaning of human language. NLU attempts to understand the user's intent and context, so as to be able to respond and understand complex human language. This includes parsing sentence structure, word sense disambiguation, sentiment analysis, etc.
[0085] In this embodiment, the named entity recognition algorithm is used to generate an interactive recognition result. If the DTMF key signal is the to-be-recognized data input by the interactive object, the named entity recognition algorithm is used to parse the DTMF key signal to obtain the interactive recognition result. If the to-be-recognized data input by the interactive object is the second ASR text obtained by using ASR parsing, the named entity recognition algorithm is used to parse the second ASR text to obtain the interactive recognition result.
[0086] In this embodiment, the natural language understanding algorithm is used to generate an interactive recognition result. If both the DTMF key signal and the text data obtained by using ASR parsing exist, the text data obtained by using ASR parsing is parsed to obtain the interactive recognition result. In the node where both the DTMF key signal and the text data obtained by using ASR parsing exist, the request text of the natural language understanding algorithm is the DTMF key signal, but because the text data obtained by using ASR parsing is also given, the response text of the natural language understanding algorithm will be processed as the text data obtained by using ASR, and within the time interval of assignment and request, if the forwarded ASR text is received, the previously assigned content will be replaced, and finally it seems that the NLU request is made using ASR.
[0087] It is worth noting that the dialogue robot interaction method of the present application simultaneously supports DTMF and ASR dialogue types, and the dialogue robot can simultaneously detect and respond to both input signals. When the interaction object expresses an intention through key pressing or voice, the robot can quickly and accurately identify and process these signals. This dual detection mechanism not only improves the response speed of the robot, but also enables the robot to more quickly identify and respond to the intention of the interaction object, significantly enhancing its semantic understanding ability to generate more accurate interaction recognition results, thereby providing more personalized responses to the needs of the interaction object. Further, the simultaneous support of DTMF and ASR dialogue types enables the dialogue robot to have stronger adaptability and flexibility to cope with various complex interaction scenarios, improving the interaction experience of the dialogue robot. In scenarios where key responses and voice responses coexist, the interaction object is provided with more choice space, reducing the likelihood of operational errors, improving the experience of the interaction object, shortening the dialogue period, and improving dialogue efficiency.
[0088] In the embodiments of the present application, the terms "first", "second", etc. are used to distinguish the same or similar items with basically the same function and role. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the number and execution order, and the terms "first", "second", etc. do not necessarily mean different.
[0089] It should be noted that in the embodiments of the present application, the words "exemplary" or "for example" indicate an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the use of the words "exemplary" or "for example" is intended to present the relevant concept in a specific manner.
[0090] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. The association relationship between the associated objects is described, which means that there can be three relationships, for example, A and / or B, which can represent the following cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c can represent a, b, c, a-b, a-c, b-c or a-b-c, where a, b and c can be single or multiple.
[0091] Figure 6 is a schematic block diagram of the dialogue robot interaction system provided by the embodiments of the present application. As shown in Figure 6As shown, the dialog robot interaction system 600 includes:
[0092] An interaction state determining module 601 is configured to determine a current dialog state of the interaction object based on an automatic speech recognition algorithm and a voice activity detection algorithm after the dialog robot sends the interaction information; the current dialog state includes an interaction state.
[0093] A data obtaining module 602 is configured to determine whether a DTMF key signal is received if the current dialog state is in the interaction state, and obtain the to-be-recognized data input by the interaction object according to the determination result of whether the DTMF key signal is received; the manner includes: if the DTMF key signal is received, taking the DTMF key signal as the to-be-recognized data; if the DTMF key signal is not received, obtaining voice data of the interaction object, and taking the voice data as the to-be-recognized data input by the interaction object.
[0094] An identification result generating module 603 is configured to generate an interaction identification result according to the to-be-recognized data.
[0095] It should be understood that the specific process in which each module performs the corresponding steps described above has been described in detail in the method embodiments, and thus will not be described here again for the sake of brevity.
[0096] It should also be understood that the division of the modules in the embodiments of the present application is illustrative, and is merely a logical functional division. In actual implementation, another division manner can be used. In addition, each functional module in each embodiment of the present application can be integrated in one processor, or can be physically separated, or two or more modules can be integrated in one module. The integrated module can be implemented in the form of hardware or in the form of a software functional module.
[0097] Figure 7 is a schematic block diagram of an electronic terminal provided by the embodiments of the present application. The electronic terminal includes a memory, a processor and a computer program stored in the memory, and the processor executes the computer program to implement the dialog robot interaction method as described above. As shown, Figure 7 the electronic terminal 700 includes at least one processor 701, a memory 702, at least one network interface 703 and a user interface 705. Each component in the apparatus is coupled together through a bus system 704. It can be understood that the bus system 704 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 704 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, all kinds of buses are marked as the bus system in Figure 7 .
[0098] The user interface 705 can include a microphone, a display, a keyboard, a mouse, a trackball, a pointing gun, a key, a button, a touchpad, a touch screen, or the like.
[0099] It can be understood that the memory 702 can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), and the like, which is used as an external cache. By way of example but not limitation, many forms of RAM can be used, such as static random access memory (SRAM), synchronous static random access memory (SSRAM). The memory described in the embodiments of the present application is intended to include but not limited to these and any other suitable categories of memory.
[0100] The memory 702 in the embodiments of the present application is used to store various categories of data to support the operation of the electronic terminal 700. Examples of these data include: any executable program for operating on the electronic terminal 700, such as an operating system 7021 and an application program 7022; the operating system 7021 contains various system programs, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks. The application program 7022 can contain various application programs, such as a media player (Media Player), a browser (Browser), and the like, for implementing various application services. The implementation of the chat robot interaction method provided by the embodiments of the present application can be included in the application program 7022.
[0101] The method disclosed in the embodiments of the present application can be applied to the processor 701 or implemented by the processor 701. The processor 701 can be an integrated circuit chip having a signal processing capability. In the implementation process, the steps of the method can be completed by hardware integrated logic circuits in the processor 701 or by instructions in the form of software. The processor 701 described above can be a general processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The processor 701 can implement or execute the methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general processor 701 can be a microprocessor or any conventional processor, etc. In combination with the steps of the accessory optimization method provided in the embodiments of the present application, the steps can be directly embodied as hardware decoding processor for execution, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium, which is located in the memory, and the processor reads the information in the memory and combines the hardware to complete the steps of the foregoing method.
[0102] In the exemplary embodiments, the electronic terminal 700 can be one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), or the like for executing the foregoing method.
[0103] According to the method provided in the embodiments of the present application, the present application further provides a computer program product, which comprises computer program code, when the computer program code runs on a computer, so that the computer executes Figures 2 to 5 the method of any of the embodiments shown.
[0104] According to the method provided in the embodiments of the present application, the present application further provides a computer readable storage medium, which stores program code, when the program code runs on a computer, so that the computer executes Figures 2 to 5 the method of any of the embodiments shown.
[0105] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0106] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0107] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0108] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0109] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e., may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0110] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.
[0111] In the above embodiments, the functions of each functional unit can be implemented by software, hardware, firmware, or any combination thereof, in whole or in part. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, the flow or function according to the embodiments of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. Computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another, for example, computer instructions can be transferred from one website, computer, server or data center to another website, computer, server or data center through wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. integrated with one or more available media. The available media can be magnetic media (such as floppy disk, hard disk, magnetic tape), optical media (such as high-density digital video disc (digital video disc, DVD), or semiconductor media (such as solid state disk (solid state disk, SSD), etc.
[0112] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0113] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0114] In summary, this application provides a conversational robot interaction method, system, medium, product, and terminal that simultaneously supports DTMF and ASR speech types. The conversational robot can simultaneously detect and respond to both types of input signals. When the interacting object expresses its intention through button presses or voice, the robot can quickly and accurately identify and process these signals. This dual detection mechanism not only improves the robot's response speed but also enables it to more rapidly identify and respond to the interacting object's intention, significantly enhancing its semantic understanding ability to generate more accurate interaction recognition results, thereby providing responses that better meet the interacting object's needs. Furthermore, the simultaneous support for DTMF and ASR speech types gives the conversational robot greater adaptability and flexibility to cope with various complex interaction scenarios, improving the conversational robot's interactive experience. In scenarios where button press responses and voice responses coexist, it provides the interacting object with more choices, reduces the possibility of operational errors, improves the interacting object's experience, shortens the dialogue cycle, and improves dialogue efficiency. Therefore, this application effectively overcomes the various shortcomings of the prior art and has high industrial application value.
[0115] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.
Claims
1. A conversational bot interaction method, characterized by, The method comprises the following steps: After the dialogue robot sends the interactive information, the current dialogue state of the interactive object is determined based on an automatic speech recognition algorithm and a voice activity detection algorithm; The current dialogue state comprises an interactive state and a waiting state; The manner of determining the current dialogue state of the interactive object comprises: determining whether there is a first ASR text based on the automatic speech recognition algorithm; if there is the first ASR text, determining that the current dialogue state is the interactive state; if there is no first ASR text, obtaining a current voice continuous time length based on the voice activity detection algorithm, and determining the current dialogue state according to the current voice continuous time length and a preset voice continuous time length; If the current dialogue state is in the interactive state, it is determined whether a DTMF key signal is received, and the to-be-recognized data input by the interactive object is obtained according to the determination result of whether the DTMF key signal is received; after determining that the current dialogue state of the interactive object is the interactive state, it is determined whether there is a DTMF key signal based on a preset key waiting time length, which comprises: when the current dialogue state of the interactive object is the interactive state, starting to enter a pre-detection state of the DTMF key signal based on a preset pre-detection key time length; from the end of the preset pre-detection key time length to the end of the preset key waiting time length, continuously detecting the DTMF key signal; if the DTMF key signal is detected, determining whether the current dialogue state of the interactive object is the waiting state based on the preset key waiting time length; According to the to-be-recognized data, an interactive recognition result is generated; The manner of obtaining the to-be-recognized data input by the interactive object according to the determination result of whether the DTMF key signal is received comprises: If the DTMF key signal is received, the DTMF key signal is taken as the to-be-recognized data; If the DTMF key signal is not received, voice data of the interactive object is obtained, and the voice data is taken as the to-be-recognized data. 2.The conversational robot interaction method of claim 1, wherein, The manner of determining the current dialogue state according to the current voice continuous time length and the preset voice continuous time length comprises: Determining whether the current voice continuous time length is greater than or equal to the preset voice continuous time length; If the current voice continuous time length is greater than or equal to the preset voice continuous time length, determining that the current dialogue state is the interactive state; If the current voice continuous time length is less than the preset voice continuous time length, determining that the current dialogue state is the waiting state. 3.The conversational robot interaction method of claim 1, wherein, The manner of determining the current dialogue state of the interactive object comprises: When the DTMF key signal is received, a current key waiting time length is obtained, and it is determined whether the current dialogue state is the waiting state based on a preset key waiting time length; If the current key waiting time length is greater than or equal to the preset key waiting time length, it is determined that the current dialogue state is the waiting state, and the received DTMF key signal is taken as the to-be-recognized data; If the current key waiting time length is less than the preset key waiting time length, it is determined that the current dialogue state is the interactive state. 4.The conversational robot interaction method of claim 1, wherein, According to the to-be-recognized data, a manner of generating an interaction recognition result comprises: When the voice data of the interaction object is taken as the to-be-recognized data of the interaction object input, a second ASR text is generated according to the voice data of the interaction object; According to the second ASR text, an interaction recognition result is generated. 5.The conversational robot interaction method of claim 1, wherein, According to the to-be-recognized data, a manner of generating an interaction recognition result comprises: When the received DTMF key signal is taken as the to-be-recognized data of the interaction object input, a key ASR text is generated by performing a value assignment operation on the DTMF key signal; According to the key ASR text, the interaction recognition result is generated.
6. A conversational robot interaction system, characterized in that, Comprise: An interaction state determination module is configured to determine a current dialogue state of an interaction object based on an automatic speech recognition algorithm and a voice activity detection algorithm after a dialogue robot sends interaction information; The current dialogue state comprises an interaction state and a waiting state; The manner of determining the current dialogue state of the interaction object comprises: determining whether there is a first ASR text based on the automatic speech recognition algorithm; if the first ASR text exists, the current dialogue state is determined to be the interaction state; if the first ASR text does not exist, the current voice continuous time length is obtained based on the voice activity detection algorithm, and the current dialogue state is determined according to the current voice continuous time length and a preset voice continuous time length; A data acquisition module is configured to determine whether a DTMF key signal is received if the current dialogue state is in the interaction state, and to acquire to-be-recognized data input by the interaction object according to the determination result of whether the DTMF key signal is received; the manner comprises: if the DTMF key signal is received, the DTMF key signal is taken as the to-be-recognized data; if the DTMF key signal is not received, voice data of the interaction object is acquired, and the voice data is taken as the to-be-recognized data; after determining that the current dialogue state of the interaction object is the interaction state, it is determined whether there is a DTMF key signal based on a preset key waiting time length; the manner comprises: when the current dialogue state of the interaction object is the interaction state, a preset key detection advance time length is started to enter a pre-detection state of the DTMF key signal; from the end of the preset key detection advance time length to the end of the preset key waiting time length, the detection of the DTMF key signal is continuously performed; if the DTMF key signal is detected, it is determined whether the current dialogue state of the interaction object is the waiting state based on the preset key waiting time length; An interaction recognition result is generated according to the to-be-recognized data.
7. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the dialogue robot interaction method of any one of claims 1 to 5.
8. A computer program product, characterised in that, The computer program product comprises computer program code, which, when executed on a computer, causes the computer to implement the dialogue robot interaction method of any one of claims 1 to 5.
9. An electronic terminal comprising a memory, a processor and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the dialogue robot interaction method of any one of claims 1 to 5.
Citation Information
Patent Citations
Processing dual tone multi-frequency signals for use with a natural language understanding system
US6845356B1