Intelligent interruption and interaction method based on semantics under IVR (Interactive Voice Response) and medium

By adopting semantic-based intelligent interruption and interaction methods in the IVR system, combining the integration of fixed sound files and TTS synthesized sounds, and the use of energy VAD technology, the blunt and limitations of the existing IVR system in terms of voice interruption and interaction are solved, and a more efficient and user-friendly interactive experience is achieved.

CN120108397APending Publication Date: 2025-06-06KEXUN JIALIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375573.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing IVR system has blunt and limitations in terms of voice interruption and interaction, and cannot effectively skip redundant processes, affecting the user experience.

Method used

Semantic-based intelligent interrupt and interaction method is adopted, and the fixed sound file of the current node is integrated with the TTS synthesized sound, and the MRCP connection recognition engine is used to perform long recognition connections, and combined with energy VAD technology to detect front and back end points, intelligent interrupt and process optimization are achieved.

Benefits of technology

It realizes more natural and intelligent user interaction, can effectively skip redundant processes, reduce user waiting time, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108397A_ABST
    Figure CN120108397A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent interruption and interaction method based on semantics under IVR (Interactive Voice Response) and a medium, relates to the technical field of voice recognition processing, and solves the technical problems that an interruption mode of an existing automatic voice response system is relatively stiff and has limitation, a part of redundant processes cannot be skipped, and the user experience is influenced. According to the invention, the audio file is generated according to the fixed audio file, and the audio file is broadcasted by using the IVR system; starting long identification connection by using an MRCP connection identification engine; a front end point in the audio file is detected and recognized to generate a recognition starting event, and the IVR system is controlled to conduct intelligent interruption; detecting a back end point of the user voice to obtain a recognition result, and judging whether the recognition result has service significance or not; if yes, closing all broadcast contents and identifying the long connection; the IVR node broadcasts a new prompt, and the steps are repeated; if not, normal playing of the IVR system is recovered; redundant processes can be reduced, and user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of speech recognition processing, and relates to speech intelligent interruption and interaction technology, in particular to a semantic-based intelligent interruption and interaction method under IVR, and a medium. Background Art

[0002] Semantic intelligent interruption and interaction is an interactive technology based on natural language understanding and context awareness. It analyzes the user's input content in real time, actively identifies key semantic nodes during the conversation, provides supplementary information in a non-invasive way, corrects errors or optimizes the interaction path. It identifies the user's core needs in real time through semantic analysis, actively interrupts redundant processes, and directly provides solutions, which can reduce user waiting time and improve user experience.

[0003] In the existing automated voice response system, users often need to follow the system prompts to press keys or input voice step by step to complete the required operations. Although the existing IVR system supports voice interruption, users can only interrupt voice input at certain specific nodes of the prompt. The interruption method is relatively blunt and has certain limitations. It is also impossible to skip some redundant processes, which affects the user experience.

[0004] The present invention provides a semantic-based intelligent interruption and interaction method and medium under IVR to solve the above technical problems. Summary of the invention

[0005] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a semantic-based intelligent interruption and interaction method under IVR, a medium, which is used to solve the technical problems that the existing automated voice response system has a relatively rigid interruption method, has certain limitations, and cannot skip some redundant processes, affecting the user experience.

[0006] To achieve the above object, the first aspect of the present invention provides a semantic-based intelligent interruption and interaction method under IVR, comprising:

[0007] Step S1: Obtain a fixed audio file of the current node; generate an audio file according to the fixed audio file, and use the IVR system to broadcast the audio file; when the audio file is being broadcast, use the MRCP connection recognition engine to open a long recognition connection;

[0008] Step S2: using the energy VAD technology to detect and identify the front end point in the audio file to generate a start recognition event; when the IVR system receives the start recognition event, the IVR system is controlled to perform intelligent interruption;

[0009] Step S3: Perform long connection recognition on the user's voice, detect the back-end point of the user's voice to obtain the recognition result; determine whether the recognition result has business significance; if yes, execute step S4; if no, execute step S5;

[0010] Step S4: Stop all broadcast contents in step S1 and close the long recognition connection in step S1; according to the recognition result returned in step S3, perform the broadcast of new prompts by the IVR node and repeat step S1;

[0011] Step S5: Restore the normal playback of the IVR system and execute step S2.

[0012] Preferably, generating an audio file according to a fixed audio file includes:

[0013] Retrieve the fixed audio file of the current node, filter the fixed audio file to obtain the fixed audio file that needs to be broadcast, integrate the fixed audio file that needs to be broadcast and the TTS synthesized audio into a text file; pass the text file to the TTS engine through the MRCP protocol;

[0014] The TTS engine downloads and caches the audio files, and sends them to the IVR system in sequence and at a packet sending frequency. The IVR system converts the text files into audio files.

[0015] The present invention obtains the fixed sound of the current node and integrates the fixed sound with the TTS synthesized sound, which can retain the professional recording sound quality and realize personalized content output. In addition, the TTS engine downloads and caches the audio file in advance, reducing the server pressure of real-time synthesis, especially ensuring the response speed in high-concurrency scenarios.

[0016] Preferably, the method of using the energy VAD technology to detect and identify the front end point in the audio file to generate a start recognition event includes:

[0017] Retrieve the audio signal of the audio file, divide the audio signal into several short-time frames; apply the Hamming window to each of the several short-time frames; use the formula Calculate the energy values ​​of several short-time frames; wherein N represents the frame length; x(n) represents the nth sampling value of the short-time frame;

[0018] When the energy value of the short-time frame is greater than the threshold value, it is marked as active; when the energy value of the short-time frame is less than the threshold value, it is marked as inactive;

[0019] If the short time frame changes from an inactive state to an active state and the duration of the active state exceeds a time threshold, the starting point of the active state is analyzed and the starting point is marked as a front end point;

[0020] If the short time frame changes from the active state to the inactive state and the duration of the inactive state exceeds the time threshold, the starting point of the inactive state is analyzed and marked as the rear endpoint.

[0021] It should be noted that the formula for applying the Hamming window is W(n)=0.54-0.46cos[2πn / (N-1)]; wherein n represents the nth sampling value; n∈[0,(N-1)].

[0022] The present invention divides the audio signal into several short-time frames and applies a Hamming window to each frame in the short-time frames, which can reduce spectrum leakage and make subsequent analysis more accurate; and calculates the energy values ​​of several short-time frames and marks the corresponding states, identifies the front end point and the back end point according to the marked states and durations, avoids misjudgment of burst noise, and truncates the speaking gap.

[0023] Preferably, the method for obtaining the threshold value includes:

[0024] The audio signal of the audio file is retrieved, and when the initial set frame length of the audio signal is a silent segment, the average energy of the corresponding silent segment is calculated to obtain the background noise;

[0025] Otherwise, the noise of the current short-time frame is calculated by the formula En=α×En-1+(1-α)Ec; wherein En is the noise of the corresponding short-time frame, Ec is the energy value of the current short-time frame; α is the smoothing coefficient;

[0026] The energy peaks in several short time frames are identified to obtain the maximum energy value Ep; the energy threshold value is calculated by the formula MZn=En+β×(Ep-En); wherein β is a proportional coefficient greater than 0.

[0027] It should be noted that the core function of the smoothing coefficient α is to balance the weights of historical data and new observations, and its value range is [0,1]. When α is closer to 1, it is more dependent on historical smoothing values, and the smoothing effect is stronger. When α is closer to 0, it is more dependent on new observations, and is more sensitive to changes, but the smoothing effect is weak. In this application, α is usually set to 0.4 to 0.6. The role of β is to dynamically adjust the degree of dependence of the energy threshold value on the current noise energy and the maximum energy value. It is set according to the actual situation. In speech endpoint detection, β is usually set to 0.15 to 0.25.

[0028] The present invention calculates background noise based on silent segments of audio signals. When there are no silent segments, the noise is dynamically analyzed based on the collected values ​​of short-time frames, and a threshold value of energy is calculated based on the noise. The threshold value can be dynamically adjusted according to actual conditions, which is beneficial to improving the rationality of threshold value setting.

[0029] Preferably, the controlling the IVR system to perform intelligent interruption includes:

[0030] Get the user's command information; the command information includes stop, pause and adjust; when receiving the stop command, the IVR system sends a STOP request command to the TTS engine through the MCRP protocol, and the TTS engine immediately stops synthesis according to the STOP request command, releases the audio buffer and returns; the IVR system closes the current RTP stream channel and discards the unplayed audio data packets; jumps to the next node according to the business logic;

[0031] When receiving a pause command, the IVR system sends a PAUSE request command to the TTS engine through the MRCP protocol. The TTS engine obtains the position of the text currently being synthesized according to the PAUSE request command, calls the GetWordBoundary function to identify the word boundary information before and after the current position, and selects the next word boundary of the current position as the pause point. When the playback reaches the pause point, the IVR system's broadcast is paused.

[0032] When an adjustment instruction is received, volume adjustment data is obtained; wherein the volume adjustment data includes: an original volume, a target minimum volume, and a fade-out duration; a quotient of a difference between the original volume and the target minimum volume and the fade-out duration is calculated to obtain an adjustment step; and the volume is gradually adjusted according to the volume adjustment step.

[0033] After the STOP instruction is sent through the MRCP protocol, the TTS engine immediately stops synthesis and releases the audio buffer to avoid buffer accumulation; the IVR closes the current RTP stream channel and discards unplayed data packets to prevent residual data from interfering with subsequent processes, and releases resources, which is beneficial to reducing server load; pausing at word boundaries can ensure context coherence when resuming playback; and adjusting the volume through the calculated volume adjustment step can achieve linear or nonlinear volume attenuation to avoid sudden changes.

[0034] Preferably, the detecting the back-end point of the user's voice to obtain the recognition result includes:

[0035] Acquire the user's voice audio and use the energy VAD technology to detect the back-end point in the user's voice audio; when the back-end point is detected, immediately intercept the voice segment of the back-end point;

[0036] The IVR system sends the intercepted voice segment to the NLP engine, performs semantic analysis on the intercepted voice segment, and obtains voice intent information; and analyzes the integrity of the recognition result based on the voice intent information.

[0037] Preferably, analyzing the integrity of the recognition result according to the speech intention information includes:

[0038] Retrieving speech intention information; obtaining a semantic library; matching the speech intention information with the semantic library to obtain a matching result; wherein the words stored in the semantic library are words without actual meaning; the matching result is success or failure;

[0039] When the matching result is a successful match, the recognition result output is not returned; when the matching result is a failed match, the voice intent information is determined to be complete and the recognition result is returned.

[0040] It should be noted that if clear semantic information or meaningless information appears, the recognition result will not be returned and the voice input will continue to be waited for; the setting of the semantic library is integrated by historical data and meaningless data queried from several channels; when new meaningless words appear, the semantic library will be updated.

[0041] The present invention identifies the user's voice intention information and matches the voice intention information with the semantic library. When the match is successful, it means that the voice intention information is an actual meaningless word, and the recognition result is not returned, and the user's voice input is continued to be waited for. This can avoid the system receiving invalid results, reduce the useless information received by the system, and improve the accuracy of system analysis.

[0042] Preferably, the determining whether the recognition result has business significance includes:

[0043] Retrieve the recognition result; send the complete sentence to the business semantic analysis service and match the business logic of the recognition result;

[0044] When the current node of the IVR system is configured with corresponding business logic, the corresponding business operation is executed and the next node of the IVR system is identified.

[0045] The present invention retrieves the original speech and converts it into a complete sentence, avoiding the traditional keyword matching due to sentence fragmentation or colloquial expressions; directly associates the semantic analysis results with predefined business rules to ensure that the execution operation meets the actual business needs; can quickly obtain the user's actual needs and execute them; it is conducive to accurately executing the user's needs and improving the user experience.

[0046] Preferably, the restoring normal playback of the IVR system includes:

[0047] Retrieve the type of smart interruption; the types of smart interruption include stop, pause and volume adjustment; when the type of smart interruption is stop, jump to the next node of the IVR system for loop operation;

[0048] When the type of intelligent interruption is pause, the word boundary information of the pause is identified and playback continues from the word boundary position;

[0049] When the type of smart interruption is volume adjustment, the volume is gradually adjusted according to the adjustment step until the volume reaches the original volume.

[0050] The present invention resumes playback according to the type of intelligent interruption. When the user clearly expresses the intention to end the current link, the system immediately jumps to the next node to avoid invalid waiting for the user. By identifying the word boundary of the voice pause, the system can accurately locate the interruption position, and there is no need to start from the beginning when resuming playback. This is especially important for long voice broadcasts, reducing the user's annoyance from repeated listening. By gradually increasing the volume with a fixed step length, it ensures that the user can clearly receive the information and avoids the discomfort caused by sudden changes in volume. It is especially suitable for noisy environments or hearing-impaired users, balancing information transmission and auditory comfort.

[0051] A second aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, the program being executable by a processor to implement the above method.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] 1. The present invention obtains the fixed sound of the current node and integrates the fixed sound with the TTS synthesized sound, which can retain the professional recording sound quality and realize personalized content output. In addition, the TTS engine downloads and caches the audio file in advance, reducing the pressure on the server of real-time synthesis, especially ensuring the response speed in high-concurrency scenarios; the audio signal is divided into several short-time frames, and the Hamming window is applied to each frame in the short-time frame, which can reduce spectrum leakage and make subsequent analysis more accurate; and the energy values ​​of several short-time frames are calculated, and the corresponding states are marked, and the front end point and the rear end point are identified according to the marked state and duration, so as to avoid misjudgment of sudden noise and truncation of speaking intervals; the background noise is calculated according to the silent segment of the audio signal, and when there is no During silent segments, the noise is dynamically analyzed based on the collected values ​​of the short-time frames, and the energy threshold is calculated based on the noise; the threshold can be dynamically adjusted based on the actual situation, which is beneficial to improving the rationality of the threshold setting; after sending the STOP command through the MRCP protocol, the TTS engine immediately stops synthesis and releases the audio buffer to avoid buffer accumulation; the IVR closes the current RTP stream channel and discards unplayed data packets to prevent residual data from interfering with subsequent processes, and releases resources, which is beneficial to reducing server load; pausing at word boundaries can ensure context coherence when resuming playback; adjusting the volume through the calculated volume adjustment step can achieve linear or nonlinear volume attenuation to avoid sudden changes.

[0054] 2. The present invention identifies the user's voice intention information, matches the voice intention information with the semantic library, and when the match is successful, it means that the voice intention information is actually meaningless vocabulary, then the recognition result is not returned, and the user's voice input is continued to be waited for, which can avoid the system receiving invalid results, reduce the useless information received by the system, and improve the accuracy of system analysis; retrieve the original voice and convert it into a complete sentence to avoid traditional keyword matching due to sentence breaks or colloquial expressions; directly associate the semantic analysis results with predefined business rules to ensure that the execution operation meets the actual business needs; can quickly obtain the user's actual needs and execute them; is conducive to accurately executing the user's needs and improving the user's experience; resumes the playback according to the type of intelligent interruption, when the user clearly expresses the intention to end the current link; the system immediately jumps to the next node to avoid the user from waiting in vain; by identifying the word boundary of the voice pause, the system can accurately locate the interruption position, and there is no need to start from the beginning when resuming playback. This is particularly important for long voice broadcasts, reducing the user's irritability of listening repeatedly; gradually increasing the volume by a fixed step length, ensuring that the user clearly receives the information and avoiding the discomfort caused by sudden changes in volume. It is especially suitable for noisy environments or hearing-impaired users, balancing information communication and auditory comfort. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0056] Figure 1 It is a schematic diagram of the overall step flow of the present invention;

[0057] Figure 2 A schematic diagram of the intelligent interruption step flow of the present invention;

[0058] Figure 3 A schematic diagram of the user intent identification steps of the present invention. DETAILED DESCRIPTION

[0059] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] See also Figure 1 The first aspect of the present invention provides a semantic-based intelligent interruption and interaction method under IVR, including:

[0061] Step S1: Obtain a fixed audio file of the current node; generate an audio file according to the fixed audio file, and use the IVR system to broadcast the audio file; when the audio file is being broadcast, use the MRCP connection recognition engine to open a long recognition connection;

[0062] Step S2: using the energy VAD technology to detect and identify the front end point in the audio file to generate a start recognition event; when the IVR system receives the start recognition event, the IVR system is controlled to perform intelligent interruption;

[0063] Step S3: Perform long connection recognition on the user's voice, detect the back-end point of the user's voice to obtain the recognition result; determine whether the recognition result has business significance; if yes, execute step S4; if no, execute step S5;

[0064] Step S4: Stop all broadcast contents in step S1 and close the long recognition connection in step S1; according to the recognition result returned in step S3, perform the broadcast of new prompts by the IVR node and repeat step S1;

[0065] Step S5: Restore the normal playback of the IVR system and execute step S2.

[0066] See also Figure 2 , obtain the fixed audio file of the current node; filter the fixed audio file to obtain the fixed audio file that needs to be broadcast, and integrate the fixed audio file that needs to be broadcast with the TTS synthesized audio into a text file; pass the text file to the TTS engine through the MRCP protocol; the TTS engine downloads and caches the audio file, and sends the audio file to the IVR system in sequence and packet sending frequency, and the IVR system converts the text file into an audio file; use the IVR system to broadcast the audio file; when the audio file is being broadcast, use the MRCP connection recognition engine to open a long recognition connection.

[0067] For example, the IVR system combines fixed audio files (such as pre-recorded prompts) and dynamically generated TTS texts into structured data in sequence; uses SSML (Speech Synthesis Markup Language) to encapsulate the content, and <audio>The tag references the fixed audio URL, and the dynamic text is directly embedded;

[0068] The IVR sends a SPEAK request to the TTS engine through the MRCP protocol, passing the SSML content as the request body; specifies audio parameters (such as encoding format audio / L16; rate=8000) and transmission mode (such as RTP stream); uses SIP or RTSP to establish an MRCP control channel;

[0069] The TTS engine parses SSML and extracts all <audio>URL in the tag; transcode the fixed audio into a format consistent with TTS output (such as PCM 16kHz 16bit) to ensure seamless splicing; divide the merged audio stream into RTP packets (such as one packet every 20ms), and transmit it to the IVR's media receiving port in real time via UDP;

[0070] The packet sending interval is adjusted according to the audio sampling rate, for example, 160 samples / packet corresponds to a 20ms interval (16000Hz sampling rate); the IVR receives the RTP stream and decodes and plays it, and monitors the status through MRCP events (such as SPEAK-COMPLETE).

[0071] Retrieve the audio signal of the audio file, divide the audio signal into several short-time frames; apply the Hamming window to each of the several short-time frames; use the formula Calculate the energy values ​​of several short-time frames; wherein N represents the frame length; x(n) represents the nth sampling value of the short-time frame; when the energy value of the short-time frame is greater than the threshold value, it is marked as an active state; when the energy value of the short-time frame is less than the threshold value, it is marked as an inactive state; if the short-time frame changes from an inactive state to an active state and the duration of the active state exceeds the time threshold, analyze the starting point of the active state and mark the starting point as the front end point; if the short-time frame changes from an active state to an inactive state and the duration of the inactive state exceeds the time threshold, analyze the starting point of the inactive state and mark the starting point as the back end point.

[0072] It is worth noting that the methods for obtaining the threshold value include:

[0073] The audio signal of the audio file is retrieved, and when the initial set frame length of the audio signal is a silent segment, the average energy of the corresponding silent segment is calculated to obtain the background noise;

[0074] Otherwise, the noise of the current short-time frame is calculated by the formula En=α×En-1+(1-α)Ec; wherein En is the noise of the corresponding short-time frame, Ec is the energy value of the current short-time frame; α is the smoothing coefficient;

[0075] The energy peaks in several short time frames are identified to obtain the maximum energy value Ep; the energy threshold value is calculated by the formula MZn=En+β×(Ep-En); wherein β is a proportional coefficient greater than 0.

[0076] It should be noted that the core function of the smoothing coefficient α is to balance the weights of historical data and new observations, and its value range is [0,1]. When α is closer to 1, it is more dependent on historical smoothing values, and the smoothing effect is stronger. When α is closer to 0, it is more dependent on new observations, and is more sensitive to changes, but the smoothing effect is weak. In this application, α is usually set to 0.4 to 0.6. The role of β is to dynamically adjust the degree of dependence of the energy threshold value on the current noise energy and the maximum energy value. It is set according to the actual situation. In speech endpoint detection, β is usually set to 0.15 to 0.25.

[0077] parameter Typical Value illustrate Frame length 20-40ms Balanced time-frequency resolution Frame Shift 10-20ms 50% overlap Set frame length of silent segments 100-500ms Noise confirmation time Duration 5-7 frames Avoid misjudgment of sudden noise or speech gaps

[0078] Table 1: Settings of key parameters

[0079] For example, according to the parameter settings in Table 1, the audio signal is divided into short time frames, usually 20-40ms per frame; adjacent frames overlap 50% (such as 10ms) to avoid information loss;

[0080] Apply a Hamming window to each of several short time frames; by the formula Calculate the energy values ​​of several short time frames;

[0081] Take the silence segment of 100-500ms before the audio to calculate the average energy; if there is no silence segment, calculate the noise of the current short-time frame by the formula En=α×En-1+(1-α)Ec; and calculate the energy threshold by the formula MZn=En+β×(Ep-En);

[0082] The energy values ​​of several short-time frames are compared with the threshold value to determine the state of the short-time frame; when the state of the short-time frame changes from an inactive state to an active state, and the active state lasts for 5-7 frames, the starting point is marked as the front end point; when the state of the short-time frame changes from an active state to an inactive state, and the inactive state lasts for 5-7 frames, the starting point is marked as the back end point.

[0083] When the IVR system receives the start recognition event, it obtains the user's instruction information; the instruction information includes stop, pause and adjust; when receiving the stop instruction, the IVR system sends a STOP request instruction to the TTS engine through the MCRP protocol, and the TTS engine immediately stops synthesis according to the STOP request instruction, releases the audio buffer and returns; the IVR system closes the current RTP stream channel and discards the unplayed audio data packets; jumps to the next node according to the business logic;

[0084] When receiving a pause command, the IVR system sends a PAUSE request command to the TTS engine through the MRCP protocol. The TTS engine obtains the position of the text currently being synthesized according to the PAUSE request command, calls the GetWordBoundary function to identify the word boundary information before and after the current position, and selects the next word boundary of the current position as the pause point. When the playback reaches the pause point, the IVR system's broadcast is paused.

[0085] When an adjustment instruction is received, volume adjustment data is obtained; wherein the volume adjustment data includes: an original volume, a target minimum volume, and a fade-out duration; a quotient of a difference between the original volume and the target minimum volume and the fade-out duration is calculated to obtain an adjustment step; and the volume is gradually adjusted according to the volume adjustment step.

[0086] It should be noted that the calculation formula of the adjustment step length is: adjustment step length = (original volume - target minimum volume) / fade-out duration.

[0087] It should be noted that after sending the STOP command through the MRCP protocol, the TTS engine immediately stops synthesis and releases the audio buffer to avoid buffer accumulation; the IVR closes the current RTP stream channel and discards unplayed data packets to prevent residual data from interfering with subsequent processes and releases resources, which is beneficial to reducing server load; pausing at word boundaries can ensure context coherence when resuming playback; adjusting the volume through the calculated volume adjustment step can achieve linear or nonlinear volume attenuation to avoid sudden changes.

[0088] See also Figure 3 , perform long connection recognition on the user's voice, obtain the user's voice audio, and use energy VAD technology to detect the back-end point in the user's voice audio; when the back-end point is detected, the voice segment of the back-end point is immediately intercepted; the IVR system sends the intercepted voice segment to the NLP engine, performs semantic analysis on the intercepted voice segment, and obtains voice intent information; analyzes the integrity of the recognition result based on the voice intent information.

[0089] Retrieve the speech intent information; obtain the semantic library; match the speech intent information with the semantic library to obtain a matching result; wherein the words stored in the semantic library are words without actual meaning; the matching result is success or failure; when the matching result is a successful match, no recognition result output is returned; when the matching result is a failed match, the speech intent information is determined to be complete, and the output recognition result is returned.

[0090] Retrieve the recognition result; send the complete sentence to the business semantic analysis service, and match the business logic of the recognition result; when the current node of the IVR system is configured with the corresponding business logic, execute the corresponding business operation, and the next node of the IVR will broadcast the new prompt and repeat the steps; otherwise, resume the normal playback of the IVR system.

[0091] It should be noted that the purpose of conducting integrity and business semantic service analysis on the user's voice intention information is to accurately understand the user's needs and make corresponding processing based on the user's needs; it is conducive to accurately identifying the user's intentions, making corresponding operations, and improving the user's experience.

[0092] It is worth mentioning that restoring the normal playback of the IVR system includes:

[0093] Retrieve the type of smart interruption; the types of smart interruption include stop, pause and volume adjustment; when the type of smart interruption is stop, jump to the next node of the IVR system for loop operation;

[0094] When the type of intelligent interruption is pause, the word boundary information of the pause is identified and playback continues from the word boundary position;

[0095] When the type of smart interruption is volume adjustment, the volume is gradually adjusted according to the adjustment step until the volume reaches the original volume.

[0096] The second aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which can be executed by a processor to implement the above method.

[0097] Part of the data in the above formula is calculated by removing the dimension and taking its numerical value. The formula is a formula closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.

[0098] The working principle of the present invention is as follows: the present invention comprises step S1: obtaining a fixed audio file of the current node; generating an audio file according to the fixed audio file, and broadcasting the audio file by using an IVR system; when the audio file is being broadcasted, using an MRCP connection recognition engine to start a long recognition connection; step S2: using energy VAD technology to detect and recognize the front end point in the audio file to generate a start recognition event; when the IVR system receives the start recognition event, controlling the IVR system to perform intelligent interruption; step S3: performing long connection recognition on the user's voice, detecting the rear end point of the user's voice to obtain a recognition result; judging whether the recognition result has business significance; if yes, executing step S4; if not, executing step S5; step S4: stopping all broadcast contents in step S1, and closing the recognition long connection in step S1; according to the recognition result returned by step S3, executing the broadcast of a new prompt by the next node of the IVR and repeating step S1; step S5: restoring the normal playback of the IVR system, and executing step S2.

[0099] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.< / audio> < / audio>

Claims

1. A semantic-based intelligent interruption and interaction method under IVR, characterized in that: include: Step S1: Get the fixed audio file of the current node; Generate an audio file based on the fixed audio file, and use the IVR system to broadcast the audio file; when the audio file is being broadcast, use the MRCP connection recognition engine to start a long recognition connection; Step S2: using the energy VAD technology to detect and identify the front end point in the audio file to generate a start recognition event; when the IVR system receives the start recognition event, the IVR system is controlled to perform intelligent interruption; Step S3: Perform long connection recognition on the user's voice, detect the back-end point of the user's voice to obtain the recognition result; and determine whether the recognition result has business significance; If yes, proceed to step S4; If no, proceed to step S5; Step S4: Stop all broadcast contents in step S1 and close the long recognition connection in step S1; according to the recognition result returned in step S3, perform the broadcast of new prompts by the IVR node and repeat step S1; Step S5: Restore the normal playback of the IVR system and execute step S2.

2. According to claim 1, a semantic-based intelligent interruption and interaction method under IVR, characterized in that: The step of generating an audio file according to a fixed audio file comprises: Retrieve the fixed audio file of the current node, filter the fixed audio file to obtain the fixed audio file that needs to be broadcast, integrate the fixed audio file that needs to be broadcast and the TTS synthesized audio into a text file; pass the text file to the TTS engine through the MRCP protocol; The TTS engine downloads and caches the audio files, and sends them to the IVR system in sequence and at a packet sending frequency. The IVR system converts the text files into audio files.

3. According to claim 1, a semantic-based intelligent interruption and interaction method under IVR, characterized in that: The method of using the energy VAD technology to detect and recognize the front endpoint in the audio file to generate a start recognition event includes: Retrieve the audio signal of the audio file, divide the audio signal into several short-time frames; apply the Hamming window to each of the several short-time frames; use the formula Calculate the energy values ​​of several short-time frames; wherein N represents the frame length; x(n) represents the nth sampling value of the short-time frame; When the energy value of the short-time frame is greater than the threshold value, it is marked as active; when the energy value of the short-time frame is less than the threshold value, it is marked as inactive; If the short time frame changes from an inactive state to an active state and the duration of the active state exceeds a time threshold, the starting point of the active state is analyzed and the starting point is marked as a front end point; If the short time frame changes from the active state to the inactive state and the duration of the inactive state exceeds the time threshold, the starting point of the inactive state is analyzed and marked as the rear endpoint.

4. The semantic-based intelligent interruption and interaction method under IVR according to claim 3, characterized in that: The method for obtaining the threshold value includes: The audio signal of the audio file is retrieved, and when the initial set frame length of the audio signal is a silent segment, the average energy of the corresponding silent segment is calculated to obtain the background noise; Otherwise, the noise of the current short-time frame is calculated by the formula En=α×En-1+(1-α)Ec; wherein En is the noise of the corresponding short-time frame, Ec is the energy value of the current short-time frame; α is the smoothing coefficient; The energy peaks in several short time frames are identified to obtain the maximum energy value Ep; the energy threshold value is calculated by the formula MZn=En+β×(Ep-En); wherein β is a proportional coefficient greater than 0.

5. The semantic-based intelligent interruption and interaction method under IVR according to claim 1, characterized in that: The controlling IVR system to perform intelligent interruption includes: Get the user's command information; the command information includes stop, pause and adjust; when receiving the stop command, the IVR system sends a STOP request command to the TTS engine through the MCRP protocol, and the TTS engine immediately stops synthesis according to the STOP request command, releases the audio buffer and returns; the IVR system closes the current RTP stream channel and discards the unplayed audio data packets; jumps to the next node according to the business logic; When receiving a pause command, the IVR system sends a PAUSE request command to the TTS engine through the MRCP protocol. The TTS engine obtains the position of the text currently being synthesized according to the PAUSE request command, calls the GetWordBoundary function to identify the word boundary information before and after the current position, and selects the next word boundary of the current position as the pause point. When the playback reaches the pause point, the IVR system's broadcast is paused. When an adjustment instruction is received, volume adjustment data is obtained; wherein the volume adjustment data includes: an original volume, a target minimum volume, and a fade-out duration; a quotient of a difference between the original volume and the target minimum volume and the fade-out duration is calculated to obtain an adjustment step; and the volume is gradually adjusted according to the volume adjustment step.

6. The semantic-based intelligent interruption and interaction method under IVR according to claim 1, characterized in that: The detection of the user's voice at the back-end to obtain the recognition result includes: Acquire the user's voice audio and use the energy VAD technology to detect the back-end point in the user's voice audio; when the back-end point is detected, immediately intercept the voice segment of the back-end point; The IVR system sends the intercepted voice segment to the NLP engine, performs semantic analysis on the intercepted voice segment, and obtains voice intent information; and analyzes the integrity of the recognition result based on the voice intent information.

7. A semantic-based intelligent interruption and interaction method under IVR according to claim 6, characterized in that: The analyzing the integrity of the recognition result according to the speech intention information includes: Retrieving speech intention information; obtaining a semantic library; matching the speech intention information with the semantic library to obtain a matching result; wherein the words stored in the semantic library are words without actual meaning; the matching result is success or failure; When the matching result is a successful match, the recognition result output is not returned; when the matching result is a failed match, the voice intent information is determined to be complete and the recognition result is returned.

8. The semantic-based intelligent interruption and interaction method under IVR according to claim 1, characterized in that: The determining whether the recognition result has business significance includes: Retrieve the recognition result; send the complete sentence to the business semantic analysis service and match the business logic of the recognition result; When the current node of the IVR system is configured with corresponding business logic, the corresponding business operation is executed and the next node of the IVR system is identified.

9. The semantic-based intelligent interruption and interaction method under IVR according to claim 1, characterized in that: The restoring the normal playback of the IVR system includes: Retrieve the type of smart interruption; the types of smart interruption include stop, pause and volume adjustment; when the type of smart interruption is stop, jump to the next node of the IVR system for loop operation; When the type of intelligent interruption is pause, the word boundary information of the pause is identified and playback continues from the word boundary position; When the type of smart interruption is volume adjustment, the volume is gradually adjusted according to the adjustment step until the volume reaches the original volume.

10. A computer-readable storage medium having a computer program stored thereon, wherein the program can be executed by a processor to implement the semantic-based intelligent interruption and interaction method under IVR as described in any one of claims 1 to 9.