System and method for realizing conversation state of telephone robot on engine level
By implementing the conversation status of the telephone robot at the engine level, the coordinated work of the conversation management service and the voice recognition engine are used to pass dynamic parameters, the problem of poor interaction pause duration control in the existing technology is solved, and a better human-computer interaction experience is achieved.
Patent Information
- Application Number
- CN202510323107.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
AI Technical Summary
The existing telephone robot system cannot effectively control the interaction pause duration in long voice input and second voice answer scenarios, resulting in poor human-computer interaction experience. Due to the limitations of the call platform, it is impossible to realize the dynamic control function by node.
The system that realizes the conversation status of the telephone robot at the engine level, through the coordinated work of the conversation management service and the voice recognition engine, communicates directly and coordinates with each other, and passes dynamic parameters of the conversation process node, such as the node pause duration, hot word parameters and dynamic models, so as to realize the ability of the voice recognition engine to dynamically control these parameters according to the conversation process node.
It allows longer pause thinking time in long voice input scenarios, shortens pause time in shorter pause time in shorter pause answer scenarios, improves refined control of human-computer interaction, and improves interactive experience.
Smart Images

Figure CN120238609A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent telephone voice robots, and more specifically, to a system and method for implementing the conversation state of a telephone robot at the engine level. Background Art
[0002] During the human-machine interaction process of a telephone robot system, in the customer voice input stage, such as in the long voice input scenario, the customer's pause and thinking time is relatively long. Since the session process transfer parameter duration of the traditional dialogue system is set and cannot be extended, the customer cannot complete a high-quality conversation. In the short voice response scenario, the interaction pause duration cannot be shortened, and fast interaction cannot be achieved, resulting in a poor human-machine interaction experience. Therefore, the industry usually introduces dynamic control functions such as controlling the interaction pause duration according to session process nodes, hot words, and dynamic recognition model switching to achieve a better interaction experience between the user and the telephone robot.
[0003] The existing dynamic control function of the telephone robot system according to nodes is usually realized by the call platform transmitting node-related parameters. However, in actual projects, it is often impossible to modify the call platform system to support these dynamic control functions due to various reasons, resulting in the inability to implement the dynamic control function according to nodes. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to propose a system and method for implementing the conversation state of a telephone robot at the engine level. When the intelligent telephone voice robot makes an outbound call, it does not require the support of the call platform, bypasses the call platform, and uses the session management service and the speech recognition engine to cooperate and communicate directly with each other, coordinating the transfer of dynamic parameters of the session process nodes, such as node pause duration, hot word parameters, and dynamic models, etc., to enable the speech recognition engine to dynamically control these parameters according to the session process nodes, and to achieve fine-grained control of the interaction pause duration, hot words, or switching of the dynamic recognition model according to the session process nodes, thereby achieving a better human-machine interaction experience.
[0005] The present invention provides a system for implementing the conversation state of a telephone robot at the engine level, including:
[0006] Session Management Service DM: used to implement the design of the human-machine conversation process, configure the prompt words, pause duration, hot words, and dynamic model parameters of each node in the process, and configure the rules and models for semantic understanding, so as to achieve the process control of human-machine multi-round interaction;
[0007] Specifically, the Session Management DM is the core of the dialogue system, responsible for maintaining the dialogue state and deciding the next action.
[0008] The technical selection of the Session Management DM includes:
[0009] Rule - based system: Uses predefined rules and state machines.
[0010] Machine - learning - based system: Uses reinforcement learning or policy networks.
[0011] The implementation steps of the session management DM include:
[0012] State maintenance: Tracks the conversation history and the current state.
[0013] Decision - making: Decides the system response based on the user's intention and context information.
[0014] Action execution: Invokes the corresponding services or APIs to generate an answer.
[0015] Automatic Speech Recognition (ASR) engine: Used to convert human speech into machine - understandable text;
[0016] Specifically, the ASR engine is the first step in converting the user's speech input into text information. Based on deep - learning technology, the ASR can achieve high - accuracy speech - to - text conversion.
[0017] The technology selection of the ASR engine includes:
[0018] End - to - end models: Such as CTC (Connectionist Temporal Classification) and attention - mechanism models.
[0019] Open - source tools: Mozilla DeepSpeech, Kaldi, etc.
[0020] The implementation steps of the ASR engine include:
[0021] Audio acquisition: Uses a microphone or other audio input devices.
[0022] Pre - processing: Noise reduction, gain control, feature extraction.
[0023] Model training: Trains the model using a large amount of labeled speech data.
[0024] Recognition: Converts speech into text in real - time or non - real - time.
[0025] The session management service DM is connected to the ASR engine.
[0026] Furthermore, the system for implementing the dialogue state of the telephone robot at the engine level further includes:
[0027] Call Platform IVR: It is used for automated communication services via telephone, controls telephone communication in the telephone robot system, and realizes telephone operations such as telephone answering, outbound calls, transfers, voice playback and number collection, hanging up, and transfers. By docking with the speech recognition engine, speech synthesis engine, and session management system, a complete telephone robot system is realized;
[0028] The Call Platform IVR is respectively connected to the Session Management Service DM and the Speech Recognition Engine ASR.
[0029] Furthermore, the system for realizing the dialogue state of the telephone robot at the engine level further includes:
[0030] Voice Activity Detection (VAD) module: It is used for voice activity detection, to identify whether there is human voice activity in the audio signal, and separate the voice segment and non-voice segment from the input audio stream for more effective processing and transmission;
[0031] The Voice Activity Detection (VAD) module is connected to the Call Platform IVR, and the Call Platform IVR is respectively connected to the Speech Recognition Engine ASR and the Session Management Service DM.
[0032] Preferably, the system for realizing the dialogue state of the telephone robot at the engine level further includes:
[0033] Text-to-Speech (TTS) engine: It is used to convert text information into natural-sounding speech;
[0034] Specifically, the Text-to-Speech (TTS) engine converts the text answer of the dialogue system into voice output, enabling users to receive information through hearing.
[0035] The technical options of the Text-to-Speech (TTS) engine include:
[0036] Concatenation-based TTS: It uses pre-recorded voice segments for splicing.
[0037] Parameter-based TTS: Such as neural network models like WaveNet and Tacotron.
[0038] The implementation steps of the Text-to-Speech (TTS) engine include:
[0039] Text analysis: Perform preprocessing such as word segmentation and prosody prediction on the text.
[0040] Speech synthesis: Generate speech waveforms according to text features.
[0041] Speech optimization: Adjust the speech rate, pitch, etc. to improve naturalness.
[0042] The call platform IVR is connected to the text-to-speech engine TTS.
[0043] The present invention also provides a method for implementing the dialogue state of a telephone robot at the engine level, which is applied to the system for implementing the dialogue state of a telephone robot at the engine level as described above, and includes: pause dynamic parameter control, hot word dynamic parameter control, and recognition model parameter control.
[0044] Among them, the method for pause dynamic parameter control includes:
[0045] While returning the call result of the call platform IVR, the dialogue management service DM sends the custom VAD pause duration parameter to the automatic speech recognition engine ASR through HTTP; when the call platform IVR calls the automatic speech recognition engine ASR, the automatic speech recognition engine ASR uses the VAD pause duration parameter to control the pause duration; if the automatic speech recognition engine ASR does not receive the VAD pause duration parameter (or other exceptions), the system default pause parameter is used (preferably, the default pause parameter is 800 ms).
[0046] Specifically, in the speech recognition request, the pause duration allowed by VAD during the customer's speech can dynamically transfer parameters according to the needs of each node, which is convenient for allowing the customer a longer pause time for thinking in the long speech input scenario; in the short speech response scenario, the pause duration can be shortened to achieve fast interaction.
[0047] Furthermore, the method for hot word dynamic parameter control includes:
[0048] While returning the call result of the call platform IVR, the dialogue management service DM sends the hot word parameter to the automatic speech recognition engine ASR through HTTP; when the call platform IVR calls the automatic speech recognition engine ASR, the call platform IVR also transfers the current hot word parameter to the automatic speech recognition engine ASR at the same time, and the automatic speech recognition engine ASR uses the current hot word content provided by the call platform IVR to improve the recognition accuracy of relevant keywords.
[0049] Specifically, in the speech recognition request, hot words can be transferred at certain nodes in the conversation process to enable the automatic speech recognition engine ASR to improve the recognition rate of hot words. For example, for the customer input at the satisfaction scoring node, if this node can be switched to a model with higher recognition accuracy for digital parameters, it will bring a better user experience.
[0050] The hot word parameter can be the complete content of the hot word, such as a proper noun like the name of a certain building, and multiple hot words can be separated by commas, with no limit on the length of the number of characters.
[0051] Furthermore, the method for recognition model parameter control includes:
[0052] While the session management service DM returns the call platform IVR call result, it sends model parameters (such as Cantonese parameters, digit recognition parameters, etc.) to the automatic speech recognition (ASR) engine through HTTP; when the call platform IVR calls the ASR engine, the ASR engine dynamically switches to the recognition model corresponding to the model parameters (such as switching to the Cantonese recognition model, digit recognition model, etc.).
[0053] Specifically, in the speech recognition request, the required recognition model can be switched at certain nodes in the session process to improve the recognition rate of the ASR at that node. For example: dedicated address recognition models, digit recognition, or Chinese-English-Cantonese models, etc.
[0054] In the actual outbound robot project of enterprises, customers often already have a call platform for telephone robots. However, in the telephone robot systems of some enterprises, due to the older version, the functions of adjusting pause parameters, hotword parameters, and dynamic recognition model switching according to nodes have not been realized. There are also some customers who only plan to purchase an automatic speech recognition engine, a text-to-speech engine, and a session management service, but hope that the new system can complete the functions of controlling the interactive pause duration, hotwords, and dynamic recognition model switching according to nodes to achieve a better interaction experience between the customer and the robot. The technical solution of the present invention is particularly applicable to these application requirements above.
[0055] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method for implementing the telephone robot dialogue state at the engine level as described above are realized.
[0056] The present invention also provides a computer device. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for implementing the telephone robot dialogue state at the engine level as described above are realized.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] When the system and method for implementing the telephone robot dialogue state at the engine level provided by the present invention are used for outbound calls of intelligent telephone voice robots, they do not require the support of the call platform, bypass the call platform, and use the session management service and the automatic speech recognition engine to cooperate and communicate directly with each other, coordinating and transmitting the dynamic parameters of the session process nodes, including parameters such as node pause duration, hotword parameters, and dynamic models, etc., to enable the automatic speech recognition engine to dynamically control these parameters according to the session process nodes, effectively realizing the fine control of the interactive pause duration, hotwords, or switching of the dynamic recognition model according to the session process nodes, and achieving a better human-computer interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention.
[0060] In the drawings:
[0061] Figure 1 is the implementation flowchart of the pause dynamic parameter control in the embodiment of the present invention;
[0062] Figure 2 is the schematic diagram of the composition of the computer device in the embodiment of the present invention. Detailed Embodiments
[0063] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and products consistent with some aspects of the present disclosure as detailed in the appended claims.
[0064] The terms used in the present disclosure are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0065] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0066] The following further details the embodiments of the present invention.
[0067] The embodiment of the present invention provides a system for implementing the conversation state of a telephone robot at the engine level, including:
[0068] Dialogue Management (DM) Service: It is used to implement the design of the human-machine dialogue process, configure the prompt words, pause durations, hot words, and parameters of the dynamic model for each node in the process, as well as the rules and model configurations for semantic understanding, so as to achieve the process control of multi-round human-machine interaction;
[0069] Automatic Speech Recognition (ASR) Engine: It is used to convert human speech into text that can be understood by machines;
[0070] The Dialogue Management (DM) Service is connected to the Automatic Speech Recognition (ASR) Engine.
[0071] Interactive Voice Response (IVR) Call Platform: It is used to provide automated communication services via telephone, control telephone communication in a telephone robot system, and implement telephone operations such as call answering, outbound calls, transfers, playing tones and receiving numbers, hanging up, and transferring. By docking with the automatic speech recognition engine, text-to-speech engine, and dialogue management system, a complete telephone robot system can be achieved;
[0072] The Interactive Voice Response (IVR) Call Platform is respectively connected to the Dialogue Management (DM) Service and the Automatic Speech Recognition (ASR) Engine.
[0073] Voice Activity Detection (VAD) Module: It is used for voice activity detection, identifying whether there is human voice activity in the audio signal, and separating the voice segment and non-voice segment from the input audio stream for more effective processing and transmission;
[0074] The Voice Activity Detection (VAD) Module is connected to the Interactive Voice Response (IVR) Call Platform, and the Interactive Voice Response (IVR) Call Platform is respectively connected to the Automatic Speech Recognition (ASR) Engine and the Dialogue Management (DM) Service.
[0075] Text-to-Speech (TTS) Engine: It is used to convert text information into natural-sounding speech;
[0076] The Interactive Voice Response (IVR) Call Platform is connected to the Text-to-Speech (TTS) Engine.
[0077] An embodiment of the present invention also provides a method for implementing the dialogue state of a telephone robot at the engine level, which is applied to the system for implementing the dialogue state of a telephone robot at the engine level as described above, and includes: pause dynamic parameter control, hot word dynamic parameter control, and recognition model parameter control;
[0078] Among them, the method for pause dynamic parameter control includes:
[0079] While returning the call platform IVR call result, the session management service DM sends the custom VAD pause duration parameter to the speech recognition engine ASR via HTTP; when the call platform IVR calls the speech recognition engine ASR, the speech recognition engine ASR uses the VAD pause duration parameter to control the pause duration; if the speech recognition engine ASR does not receive the VAD pause duration parameter (or other exceptions), it uses the system default pause parameter of 800ms.
[0080] In the speech recognition request, the pause duration allowed by VAD during the customer's speech can dynamically pass parameters according to the needs of each node, which is convenient for allowing the customer a longer pause time for thinking in the long speech input scenario; in the short speech response scenario, the pause duration can be shortened to achieve fast interaction.
[0081] Figure 1 Shows the implementation process of the pause dynamic parameter control of this embodiment.
[0082] The method for controlling the hot word dynamic parameters includes:
[0083] While returning the call platform IVR call result, the session management service DM sends the hot word parameter to the speech recognition engine ASR via HTTP; when the call platform IVR calls the speech recognition engine ASR, the call platform IVR also passes the current hot word parameter to the speech recognition engine ASR at the same time, and the speech recognition engine ASR uses the current hot word content provided by the call platform IVR to improve the recognition accuracy of relevant keywords.
[0084] In this embodiment, the hot word parameter is the complete content of the hot word, including proper nouns such as the name of a certain building. Multiple hot words are separated by commas, and the length of the words is not limited.
[0085] In the speech recognition request, hot words can be passed at certain nodes in the session process to enable the speech recognition engine ASR to improve the recognition rate of hot words. For example, for the customer input at the satisfaction scoring node, if this node can be switched to a model with higher recognition accuracy for digital parameters, it will bring a better user experience.
[0086] The method for controlling the recognition model parameters includes:
[0087] While returning the call platform IVR call result, the session management service DM sends the model parameter to the speech recognition engine ASR via HTTP; when the call platform IVR calls the speech recognition engine ASR, the speech recognition engine ASR dynamically switches to the recognition model corresponding to the model parameter.
[0088] In a voice recognition request, the required recognition model can be switched at certain nodes in the conversation process to improve the recognition rate of the ASR at that node. The recognition models include: a dedicated address recognition model, a number recognition model, or a Chinese-English-Cantonese model, etc.
[0089] In this embodiment, the model parameters include: Cantonese parameters, number recognition parameters, etc.; switching to the recognition model corresponding to the model parameters includes: switching to a Cantonese recognition model, a number recognition model, etc.
[0090] This embodiment is particularly applicable to the actual application requirements in the outbound robot project of enterprises. For customers who already have a call platform for phone robots, in some enterprises, the phone robot system is too old and has not yet implemented functions such as adjusting pause parameters, hotword parameters, and dynamic recognition model switching by node. There are also some customers who only plan to purchase a voice recognition engine, a voice synthesis engine, and a session management service, but hope that the new system can complete functions such as controlling the interaction pause duration, hotwords, and dynamic recognition model switching by node to achieve a better interaction experience between the customer and the robot.
[0091] The system and method for implementing the phone robot dialogue state at the engine level in this embodiment, when the intelligent phone voice robot makes an outbound call, does not require the support of a call platform. It bypasses the call platform and uses the session management service and the voice recognition engine to cooperate and communicate directly with each other, coordinating and transmitting the dynamic parameters of the conversation process nodes, including parameters such as the node pause duration, hotword parameters, and dynamic models, to enable the voice recognition engine to dynamically control these parameters according to the conversation process nodes, achieving fine-grained control of the interaction pause duration, hotwords, or switching of the dynamic recognition model according to the conversation process nodes, and providing a better human-machine interaction experience.
[0092] The embodiment of the present invention also provides a computer device, Figure 2 which is a schematic structural diagram of a computer device provided by the embodiment of the present invention; see the attached drawing Figure 2 As shown, the computer device includes: an input system 23, an output system 24, a memory 22, and a processor 21; the memory 22 is used to store one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the method for implementing the phone robot dialogue state at the engine level as provided in the above embodiment; where the input system 23, the output system 24, the memory 22, and the processor 21 can be connected through a bus or other means, Figure 2 taking connection through a bus as an example.
[0093] The memory 22 is a computable device-readable and writable storage medium, and can be used to store software programs and computer-executable programs, such as program instructions corresponding to the method for implementing the dialogue state of the telephone robot at the engine level as described in the embodiments of the present invention; the memory 22 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the device, etc.; in addition, the memory 22 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices; in some instances, the memory 22 may further include a memory remotely set relative to the processor 21, and these remote memories may be connected to the device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0094] The input system 23 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the device; the output system 24 may include display devices such as a display screen.
[0095] The processor 21 executes various functional applications and data processing of the device by running software programs, instructions, and modules stored in the memory 22, that is, implements the above-mentioned method for implementing the dialogue state of the telephone robot at the engine level.
[0096] The above-provided computer device can be used to execute the method for implementing the dialogue state of the telephone robot at the engine level provided in the above embodiments, and has corresponding functions and beneficial effects.
[0097] An embodiment of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for implementing the conversation state of the telephone robot at the engine level provided in the above embodiment when executed by a computer processor. The storage medium is any of various types of memory devices or storage devices, and the storage medium includes: installation media, such as CD-ROMs, floppy disks or tape systems; computer system memories or random access memories, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memories, such as flash memories, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc.; the storage medium may also include other types of memories or combinations thereof; in addition, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system, and the second computer system is connected to the first computer system through a network (such as the Internet); the second computer system may provide program instructions to the first computer for execution. The storage medium includes two or more storage media that may reside in different locations (such as in different computer systems connected through a network). The storage medium may store program instructions (such as specifically implemented as a computer program) executable by one or more processors.
[0098] Of course, for a storage medium containing computer-executable instructions provided in an embodiment of the present invention, the computer-executable instructions are not limited to the method for implementing the conversation state of the telephone robot at the engine level as described in the above embodiment, and may also execute related operations in the method for implementing the conversation state of the telephone robot at the engine level provided in any embodiment of the present invention.
[0099] So far, the technical solution of the present invention has been described in combination with the preferred embodiments. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
[0100] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention may have various changes and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A system for realizing the conversation status of a telephone robot at the engine level, characterized in that: include: Conversation management service DM: used to implement the design of human-computer dialogue process, configuration of prompts, pause duration, hot words and dynamic model parameters of each node in the process, semantic understanding rules and model configuration, so as to realize process control of multi-round human-computer interaction; Speech recognition engine ASR: used to convert human speech into machine-understandable text; The session management service DM is connected to the speech recognition engine ASR.
2. The system for realizing telephone robot conversation status at the engine level according to claim 1, characterized in that: Also includes: Calling platform IVR: used for automated communication services via telephone, controlling telephone communications in the telephone robot system, and realizing telephone operations such as answering calls, outbound calls, transfers, playing and collecting numbers, hanging up, and transfers. It is connected with the speech recognition engine, speech synthesis engine, and conversation management system to realize a complete telephone robot system; The call platform IVR is connected to the session management service DM and the speech recognition engine ASR respectively.
3. The system for realizing telephone robot conversation status at the engine level according to claim 2, characterized in that: Also includes: Voice detection module VAD: used for voice activity detection, identifying whether there is human voice activity in the audio signal, and separating the voice segment and non-voice segment from the input audio stream for more efficient processing and transmission; The voice detection module VAD is connected to the call platform IVR, and the call platform IVR is connected to the voice recognition engine ASR and the session management service DM respectively.
4. A method for realizing a telephone robot conversation state at an engine level, applied to a system for realizing a telephone robot conversation state at an engine level as claimed in any one of claims 1 to 3, characterized in that: include: Pause dynamic parameter control, hot word dynamic parameter control, recognition model parameter control; Wherein, the method for controlling the pause dynamic parameters includes: When returning the call result of the call platform IVR, the session management service DM sends the customized VAD pause duration parameter to the speech recognition engine ASR via HTTP; when the call platform IVR calls the speech recognition engine ASR, the speech recognition engine ASR uses the VAD pause duration parameter to control the pause duration; if the speech recognition engine ASR does not receive the VAD pause duration parameter, the system default pause parameter is used.
5. The method for realizing the telephone robot conversation state at the engine level according to claim 4, characterized in that: The method for controlling the hot word dynamic parameters includes: When the session management service DM returns the call result of the call platform IVR, it sends the hot word parameters to the speech recognition engine ASR through HTTP; when the call platform IVR calls the speech recognition engine ASR, the call platform IVR also passes the current hot word parameters to the speech recognition engine ASR. The speech recognition engine ASR uses the current hot word content provided by the call platform IVR to improve the accuracy of identifying related keywords.
6. The method for realizing the telephone robot conversation state at the engine level according to claim 4, characterized in that: The method for identifying model parameter control comprises: When returning the call result of the call platform IVR, the session management service DM sends the model parameters to the speech recognition engine ASR via HTTP; when the call platform IVR calls the speech recognition engine ASR, the speech recognition engine ASR dynamically switches to the recognition model corresponding to the model parameters.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, the steps of the method for realizing the telephone robot conversation status at the engine level as described in any one of claims 4-6 are implemented.
8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for realizing the telephone robot conversation status at the engine level as described in any one of claims 4-6 are implemented.