Voice call method and system capable of flexibly interrupting broadcast based on FreeSWITCH, and storage medium
By developing the connection between the module mod_ai and the DUI platform in FreeSWITCH, real-time two-way voice interaction and flexible audio playback control are realized, which solves the problem that the voice interaction system in the existing technology cannot achieve real-time and two-way dialogue, and improves the user experience.
Patent Information
- Application Number
- CN202411948682.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-06
AI Technical Summary
The existing intelligent voice dialogue system based on FreeSWITCH is not able to perform voice recognition and broadcast functions simultaneously, so it can only run in half-duplex form, which cannot meet the needs of real-time and two-way dialogue.
By developing the FreeSWITCH module mod_ai, the module mod_ai is connected to the DUI platform through the WebSocket protocol to realize full-link services of speech recognition (ASR), semantic understanding (NLU) and speech synthesis (TTS), support real-time two-way voice interaction, and parse control commands sent from the DUI platform to adjust the audio playback status.
Real-time two-way interaction between voice input and voice output is realized, and flexible audio playback control is supported, such as pause, recovery, interruption and clearance, which significantly improves the user experience and solves the half-duplex limitation of traditional voice interaction methods.
Smart Images

Figure CN119943048A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice interaction technology, and in particular to a voice call method, system and storage medium based on FreeSWITCH that can flexibly interrupt broadcasting. Background Art
[0002] In recent years, voice interaction technology has developed rapidly. As core technologies, automatic speech recognition (ASR) and text-to-speech synthesis (TTS) have been widely used in voice assistants, smart homes, customer service and other fields. In communication systems, the demand for application scenarios based on real-time voice dialogue is increasing. Users expect the interaction with the system to be smoother and more efficient, and to achieve a natural communication experience similar to human dialogue. As an open source communication platform, FreeSWITCH has become one of the important tools for building intelligent voice interaction systems with its high performance and modular design.
[0003] The existing FreeSWITCH-based intelligent voice dialogue system usually completes a round of voice interaction by calling external ASR and TTS engines. The specific implementation method is: FreeSWITCH first obtains the audio data input by the user and sends the audio data to the ASR module for voice recognition; then, the voice recognition result is sent to the dialogue control system to generate a response text; finally, the TTS module is called to convert the generated text into audio data and return it to the user. In this way, the full-process dialogue capability from voice input to voice output can be realized, which is widely used in call centers, voice navigation and other scenarios.
[0004] However, in the prior art, ASR and TTS processing is performed in a serial manner, and it is impossible to perform speech recognition and broadcasting functions at the same time, resulting in the voice interaction system being able to operate only in half-duplex mode. This half-duplex interaction mode cannot meet the needs of real-time, two-way dialogue, and lacks the ability to flexibly interrupt and resume broadcasting, which is not conducive to improving user experience. Summary of the invention
[0005] This application provides a voice call method, system and storage medium based on FreeSWITCH that can flexibly interrupt broadcasting, which solves the half-duplex limitation of traditional voice interaction methods, improves user experience, and supports flexible, real-time, two-way voice dialogue interaction. This application provides the following technical solutions:
[0006] In a first aspect, the present application provides a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting, the method comprising:
[0007] Based on FreeSWITCH development module mod_ai, the FreeSWITCH is connected to the DUI platform through the module mod_ai communication;
[0008] The module mod_ai reads the user's real-time audio stream from the FreeSWITCH and sends it to the DUI platform for voice recognition;
[0009] The DUI platform calls the ASR engine to perform speech recognition and generate a response text;
[0010] The DUI platform calls the TTS engine to convert the answer text into an audio stream and returns it to the module mod_ai;
[0011] The module mod_ai writes the audio stream into the FreeSWITCH and sends it to the user;
[0012] During the data interaction process, the module mod_ai parses the control command received from the DUI platform and adjusts the state of the audio playback according to the control command.
[0013] In a specific implementation scheme, the FreeSWITCH-based development module mod_ai, the FreeSWITCH is connected to the DUI platform through the module mod_ai, including:
[0014] The module mod_ai is connected to the read interface and write interface of FreeSWITCH, the read interface is used to obtain the audio data sent by the calling user from FreeSWITCH, and the write interface is used to receive the audio data processed by the DUI platform and return the audio data to the calling user through FreeSWITCH;
[0015] The module mod_ai establishes a connection with the DUI platform via the WebSocket protocol.
[0016] In a specific implementation scheme, the module mod_ai reads the user's real-time audio stream from the FreeSWITCH and sends it to the DUI platform for voice recognition, including:
[0017] The module mod_ai reads the real-time audio stream of the calling user through the read interface of FreeSWITCH;
[0018] The module mod_ai transmits the real-time audio stream in the form of a stream to the full-link service interface of the DUI platform.
[0019] In a specific implementation scheme, the DUI platform calls the ASR engine to perform speech recognition and generate a response text including:
[0020] The DUI platform transmits the real-time audio stream to the ASR speech recognition engine, extracts the speech content and converts it into text;
[0021] The DUI platform processes the recognition results according to pre-configured semantics and speech logic to generate corresponding response text.
[0022] In a specific possible implementation scheme, during the data interaction process, the module mod_ai parses the control command received from the DUI platform, and adjusts the state of the audio playback according to the control command, including:
[0023] The module mod_ai divides the data stream read from the DUI platform into two categories according to the content:
[0024] Audio data: audio stream, used for voice broadcast;
[0025] Control commands: control the audio playback process, including pausing, resuming playback, and clearing the buffer;
[0026] For audio data, the module mod_ai stores it in an internal audio buffer.
[0027] In a specific possible implementation scheme, during the data interaction process, the module mod_ai parses the control command received from the DUI platform, and adjusting the state of the audio playback according to the control command further includes:
[0028] If a clear command is received, the module mod_ai will clear the data in the current audio buffer to prepare space for new audio content;
[0029] If a playback-related command is received, the module mod_ai adjusts the state of the audio playback according to the command instruction.
[0030] In a specific possible implementation scheme, during the data interaction process, the module mod_ai parses the control command received from the DUI platform, and adjusting the state of the audio playback according to the control command further includes:
[0031] The specific playback process is based on the audio data in the buffer and the instructions of the control command, including:
[0032] If there is audio data in the buffer and the playback condition is play, the module mod_ai starts playing the audio until the audio in the buffer is finished playing;
[0033] During the playback process, if a pause command is received, the module mod_ai will pause the current audio playback;
[0034] If the audio playback is paused and a resume command is received, the module mod_ai will continue to play the remaining audio data in the buffer and resume the voice broadcast;
[0035] If a clear command is received, the module mod_ai will clear the data in the current audio buffer to prepare for the playback of new audio data, thus interrupting the current audio playback;
[0036] After the audio resumes playing, if a pause command or a clear command is received, the module mod_ai can pause, resume, and interrupt the audio multiple times before the audio ends.
[0037] In the second aspect, the present application provides a voice call system based on FreeSWITCH that can flexibly interrupt broadcasts, using the following technical solutions:
[0038] A voice call system based on FreeSWITCH that can flexibly interrupt broadcasts, including:
[0039] A two-way communication module, used for developing a module mod_ai based on FreeSWITCH, wherein the FreeSWITCH is connected to the DUI platform through the module mod_ai;
[0040] A data reading module, used for the module mod_ai to read the user's real-time audio stream from the FreeSWITCH and send it to the DUI platform for voice recognition;
[0041] A speech recognition module, used for the DUI platform to call the ASR engine to perform speech recognition and generate a response text;
[0042] An audio conversion module, used for the DUI platform to call the TTS engine to convert the answer text into an audio stream and return it to the module mod_ai;
[0043] An audio feedback module, used for the module mod_ai to write the audio stream into the FreeSWITCH and send it to the user;
[0044] The audio control module is used to parse the control commands received from the DUI platform during the data interaction process, and adjust the state of the audio playback according to the control commands.
[0045] In a third aspect, the present application provides an electronic device, comprising a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting as described in the first aspect.
[0046] In a fourth aspect, the present application provides a computer-readable storage medium, wherein a program is stored in the storage medium, and when the program is executed by a processor, it is used to implement a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting as described in the first aspect.
[0047] In summary, the beneficial effects of this application include at least:
[0048] 1) Through the collaborative work of the FreeSWITCH-based module mod_ai and the DUI platform, the system can achieve real-time two-way interaction between voice input and voice output. Different from the traditional half-duplex voice interaction method, the technical solution of this application allows users to issue new instructions during voice broadcasting, and the system can respond immediately, thereby greatly improving the fluency and interaction efficiency of voice dialogue. Audio playback and voice recognition can be carried out in parallel, ensuring the real-time nature of voice interaction.
[0049] 2) This application provides precise control over audio playback, supporting operations such as pausing, resuming, interrupting, and clearing audio. Through the audio buffer and control command processing mechanism of the module mod_ai, the system can flexibly adjust the voice broadcast status according to user needs. For example, the user can interrupt and issue new instructions at any time during the voice broadcast process, and the system can pause the current playback and respond immediately. This flexible control capability significantly improves the user experience, especially in complex interactive scenarios.
[0050] By developing the FreeSWITCH module mod_ai, the system can achieve two-way real-time transmission and processing of audio data with the DUI platform. The module mod_ai is connected to the DUI platform via WebSocket, supporting full-link services for speech recognition (ASR), semantic understanding (NLU) and speech synthesis (TTS), and realizing real-time interaction from voice input to voice output. At the same time, mod_ai can process control commands in audio playback, such as pause, resume and interrupt, to ensure that during voice broadcast, users can issue new instructions and get immediate response from the system. This solution solves the half-duplex limitation of traditional voice interaction methods, improves user experience, and supports flexible, real-time, two-way voice dialogue interaction.
[0051] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application and implement it according to the contents of the specification, the following is a detailed description of the preferred embodiments of the present application in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a flowchart of a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting in an embodiment of the present application.
[0053] Figure 2 It is a flowchart of the internal implementation of the module mod_ai in the embodiment of the present application.
[0054] Figure 3 It is a schematic diagram of the overall process of a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting in an embodiment of the present application.
[0055] Figure 4 It is a block diagram of an electronic device that can flexibly interrupt broadcast voice calls based on FreeSWITCH in an embodiment of the present application. DETAILED DESCRIPTION
[0056] The specific implementation methods of the present application are further described in detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application but are not intended to limit the scope of the present application.
[0057] Optionally, the present application uses the FreeSWITCH-based voice call method with flexible interruption broadcast provided in each embodiment as an example for use in an electronic device, where the electronic device is a terminal or a server. The terminal can be a mobile phone, a computer, a tablet computer, etc. This embodiment does not limit the type of electronic device.
[0058] Reference Figure 1 , is a flow chart of a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting provided by an embodiment of the present application, the method comprising at least the following steps:
[0059] Step S101, developing module mod_ai based on FreeSWITCH, FreeSWITCH is connected to the DUI platform through module mod_ai communication.
[0060] Among them, the DUI platform is an intelligent voice interaction development platform that provides full-link services such as speech recognition (ASR), semantic understanding (NLU), and speech synthesis (TTS), and supports the construction of voice-based human-computer interaction systems. Through the DUI platform, developers can quickly build efficient and flexible voice dialogue capabilities for scenarios such as smart hardware, call centers, and smart customer service.
[0061] Optionally, the DUI platform in this application is the AISpiech DUI platform.
[0062] Step S101: First, develop a new module mod_ai based on FreeSWITCH. The module is responsible for realizing the audio data interaction between FreeSWITCH and the AISpeech DUI platform. The module mod_ai ensures the reception and transmission of real-time audio data by connecting with the audio read interface and write interface of FreeSWITCH. At the same time, it communicates with the DUI platform through the WebSocket network protocol, thereby realizing full-link services such as speech recognition (ASR), semantic processing (NLU) and speech synthesis (TTS).
[0063] Specifically, the module mod_ai runs as a custom module in the FreeSWITCH system. The read interface is used to obtain the audio data sent by the calling user from FreeSWITCH, and send this data to the DUI platform in real time for speech recognition processing. The write interface is used to receive the audio data processed by the DUI platform, and return the audio data to the calling user through FreeSWITCH. Through the WebSocket protocol, the module mod_ai establishes a continuous connection with the DUI platform to achieve two-way transmission of audio streams and control commands. The module mod_ai cyclically reads data from the DUI platform. If it is an audio stream, it is transmitted to FreeSWITCH for playback; if it is a control command, the playback status of the current audio is controlled according to the command, such as pausing, resuming or clearing the buffer, to achieve flexible audio interaction control.
[0064] Step S102, module mod_ai reads the user's real-time audio stream from FreeSWITCH and sends it to the DUI platform for voice recognition.
[0065] The full-link interface of the DUI platform is a standardized service interface that supports developers to interact with it through network protocols and complete the entire process from user voice input to voice output: ASR converts user voice into text. NLU performs semantic analysis on the text and generates response content. TTS synthesizes the response content into a voice stream.
[0066] In step S102, the module mod_ai reads the real-time audio stream of the calling user through the read interface of FreeSWITCH. The audio stream comes from the user's voice input, and the module mod_ai transmits the audio data to the full-link service interface of the DUI platform. Through this interface, the audio stream enters the DUI platform and is passed to the speech recognition engine (ASR) in the platform for processing.
[0067] Specifically, the module mod_ai first obtains audio data from the read interface of the FreeSWITCH system. The audio data can be a voice signal captured by a microphone or other device. Then, the module mod_ai transmits this data in the form of a stream to the full-link service interface of the DUI platform. At this time, the service of the DUI platform receives the audio stream and passes it to its internal ASR engine.
[0068] Step S103: The DUI platform calls the ASR engine to perform speech recognition and generate a response text.
[0069] In step S103, after receiving the audio stream sent by the module mod_ai, the DUI platform full-link service first passes the audio data to its internal ASR speech recognition engine. The engine is responsible for converting the voice signal in the audio stream into text data, that is, completing the speech-to-text (STT) conversion process.
[0070] Specifically, the full-link service of the DUI platform first receives the audio stream data from the module mod_ai, which contains the user's voice input. The DUI platform transmits the audio data to the ASR speech recognition engine, analyzes the audio signal, extracts the voice content and converts it into text. After the ASR engine completes speech recognition and outputs text, the DUI platform then processes the recognition results according to the pre-configured semantics and speech logic. The logical rules include semantic understanding and user intent recognition, and the dialogue management system determines user needs and interaction scenarios to generate corresponding response text.
[0071] Step S104: The DUI platform calls the TTS engine to convert the response text into an audio stream and returns it to the module mod_ai.
[0072] In step S104, after the DUI platform full-link service completes the generation of the response text, it will continue to call the TTS (text-to-speech) speech synthesis engine to convert the generated response text into a voice audio stream. The audio stream will then be returned to the module mod_ai to be played to the calling user through FreeSWITCH.
[0073] Specifically, once the response text is ready, the DUI platform will pass it to the TTS engine. The TTS engine will convert the text into a natural and smooth voice audio stream through synthesis technology based on the content of the response text. This process uses speech synthesis models and audio synthesis technology to ensure that the synthesized speech sounds real, natural, and conforms to the expected voice style. After completing the speech synthesis, the TTS engine will return the generated audio stream to the DUI platform full-link service. The DUI platform then passes the audio stream to the module mod_ai for subsequent processing.
[0074] Step S105, module mod_ai writes the audio stream into FreeSWITCH and sends it to the user.
[0075] In step S105, after receiving the audio stream generated by the TTS engine from the DUI platform, the module mod_ai is responsible for transmitting the audio data to the audio write interface of FreeSWITCH, and then sending the audio stream to the user who is talking through the interface, thereby realizing voice broadcast.
[0076] Specifically, the module mod_ai transmits the voice data to the user end by writing the audio stream to the write interface of FreeSWITCH, so that the calling user can hear the voice response of the system. This process ensures the real-time nature of the voice response and completes the two-way voice interaction with the user.
[0077] It should be noted that the module mod_ai establishes a stable network connection with the DUI platform and begins to enter a real-time, two-way voice dialogue mode. Steps S102 to S105 will be executed in a continuous loop to support real-time, two-way voice interaction. Since the execution of each step does not block other steps, the module mod_ai can recognize when the user is speaking and can also play the system's voice response. When the user issues a new instruction during the system broadcast, the system can respond in real time to pause, resume or interrupt the audio broadcast, forming a true two-way voice interaction.
[0078] In implementation, refer to Figure 2 , the internal implementation of the module mod_ai is as follows:
[0079] First, the module mod_ai establishes a connection with the DUI platform through the WebSocket protocol, and continuously loops to read the audio stream data and control commands from the DUI platform. The module mod_ai divides the data into two categories according to the content of the data stream read from the DUI interface: Audio data: This is an audio stream generated by AI for voice broadcasting. Control commands: These commands control the audio playback process. Specific commands include pausing, resuming playback, clearing the buffer, etc. For audio data, the module mod_ai stores it in the internal audio buffer to ensure that the audio data can be smoothly sent to the calling user for playback. The use of the audio buffer helps to handle various interruptions and recovery operations in audio playback.
[0080] Secondly, the module mod_ai will parse the control commands received from the DUI platform and perform the following operations according to the type of command: If a clear command is received, mod_ai will clear the data in the current audio buffer to prepare space for new audio content. If a playback-related command (such as play, pause, resume) is received, mod_ai will adjust the state of audio playback according to the command instructions to ensure that it can flexibly respond to user needs during a voice call.
[0081] Finally, through the FreeSWITCH write interface, the module mod_ai sends the data in the audio buffer to the calling user. The specific playback process is based on the audio data in the buffer and the instructions of the control command, mainly including the following situations:
[0082] 1) If there is audio data in the buffer and the play condition is "play", the module mod_ai starts playing the audio until the audio in the buffer is finished.
[0083] 2) During the playback process, if a "pause" command is received, the module mod_ai will pause the current audio playback. At this time, the audio stream stops outputting.
[0084] 3) If the audio playback is paused and a "resume" command is received, the module mod_ai will continue to play the remaining audio data in the buffer and resume voice broadcast.
[0085] 4) If a "clear" command is received, the module mod_ai will clear the data in the current audio buffer to prepare for the playback of new audio data, thereby interrupting the current audio playback.
[0086] 5) After the audio resumes playing, if a pause command or a clear command is received, the module mod_ai can pause, resume, interrupt, and other operations multiple times before the audio playback ends until the audio playback ends.
[0087] During the entire process (from step S101 to step S105), mod_ai maintains a continuous connection with the DUI platform full-link service interface. Audio data and control commands are continuously read and processed from the DUI platform, and audio data is delivered to the calling user in real time. At the same time, the control commands implement flexible play, pause and resume operations. This process will continue to loop to ensure a real-time, two-way voice interaction experience.
[0088] In summary, refer to Figure 3, this application provides a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting, aiming to solve the problem that the voice interaction system in the prior art cannot achieve real-time, two-way dialogue, flexible interruption and continuation. By developing the FreeSWITCH module mod_ai, the system can realize two-way real-time transmission and processing of audio data with the DUI platform. The module mod_ai is connected to the DUI platform through WebSocket, supports full-link services of speech recognition (ASR), semantic understanding (NLU) and speech synthesis (TTS), and realizes real-time interaction from voice input to voice output. At the same time, mod_ai can process control commands in audio playback, such as pause, resume and interrupt, to ensure that during voice broadcasting, users can issue new instructions and get immediate response from the system. This solution solves the half-duplex limitation of traditional voice interaction methods, improves user experience, and supports flexible, real-time, two-way voice dialogue interaction.
[0089] An embodiment of the present application also provides a voice call system based on FreeSWITCH that can flexibly interrupt broadcasting, and the system includes at least the following modules:
[0090] Bidirectional communication module, used to develop module mod_ai based on FreeSWITCH. FreeSWITCH is connected to the DUI platform through module mod_ai communication;
[0091] Data reading module, used by module mod_ai to read the user's real-time audio stream from FreeSWITCH and send it to the DUI platform for speech recognition;
[0092] Speech recognition module, used by the DUI platform to call the ASR engine for speech recognition and generate response text;
[0093] The audio conversion module is used by the DUI platform to call the TTS engine to convert the answer text into an audio stream and return it to the module mod_ai;
[0094] Audio feedback module, used by module mod_ai to write audio streams to FreeSWITCH and send them to the user;
[0095] The audio control module is used to parse the control commands received from the DUI platform during data interaction and adjust the audio playback status according to the control commands.
[0096] For relevant details, refer to the above method embodiment.
[0097] Figure 4 4 is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.
[0098] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may also include an AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.
[0099] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction, which is used to be executed by the processor 401 to implement the voice call method based on FreeSWITCH that can flexibly interrupt the broadcast provided in the method embodiment of the present application.
[0100] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402 and the peripheral device interface may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface via a bus, a signal line or a circuit board. Schematically, the peripheral devices include but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply.
[0101] Of course, the electronic device may also include fewer or more components, which is not limited in this embodiment.
[0102] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the voice call method based on FreeSWITCH that can flexibly interrupt the broadcast of the above method embodiment.
[0103] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the voice call method based on FreeSWITCH that can flexibly interrupt the broadcast of the above method embodiment.
[0104] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0105] The above embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A voice call method based on FreeSWITCH that can flexibly interrupt broadcasting, characterized in that: The method comprises: Based on FreeSWITCH development module mod_ai, the FreeSWITCH is connected to the DUI platform through the module mod_ai communication; The module mod_ai reads the user's real-time audio stream from the FreeSWITCH and sends it to the DUI platform for voice recognition; The DUI platform calls the ASR engine to perform speech recognition and generate a response text; The DUI platform calls the TTS engine to convert the answer text into an audio stream and returns it to the module mod_ai; The module mod_ai writes the audio stream into the FreeSWITCH and sends it to the user; During the data interaction process, the module mod_ai parses the control command received from the DUI platform and adjusts the state of the audio playback according to the control command.
2. According to claim 1, the voice call method based on FreeSWITCH with flexible interruption of broadcasting is characterized in that: The FreeSWITCH-based development module mod_ai, wherein the FreeSWITCH is connected to the DUI platform through the module mod_ai, comprises: The module mod_ai is connected to the read interface and the write interface of the FreeSWITCH, the read interface is used to obtain the audio data sent by the calling user from the FreeSWITCH, and the write interface is used to receive the audio data processed by the DUI platform and return the audio data to the calling user through FreeSWITCH; The module mod_ai establishes a connection with the DUI platform via the WebSocket protocol.
3. According to claim 1, the voice call method based on FreeSWITCH with flexible interruption of broadcasting is characterized in that: The module mod_ai reads the user's real-time audio stream from the FreeSWITCH and sends it to the DUI platform for voice recognition, including: The module mod_ai reads the real-time audio stream of the calling user through the read interface of FreeSWITCH; The module mod_ai transmits the real-time audio stream in the form of a stream to the full-link service interface of the DU I platform.
4. The voice calling method based on FreeSWITCH with flexible interruption of broadcast according to claim 1, characterized in that: The DU I platform calls the ASR engine to perform speech recognition and generate a response text including: The DUI platform transmits the real-time audio stream to the ASR speech recognition engine, extracts the speech content and converts it into text; The DUI platform processes the recognition results according to pre-configured semantics and speech logic to generate corresponding response text.
5. The voice calling method based on FreeSWITCH with flexible interruption of broadcast according to claim 1, characterized in that: In the data interaction process, the module mod_ai parses the control command received from the DU I platform, and adjusts the state of the audio playback according to the control command, including: The module mod_ai divides the data stream read from the DU I platform into two categories according to the content: Audio data: audio stream, used for voice broadcast; Control commands: control the audio playback process, including pausing, resuming playback, and clearing the buffer; For audio data, the module mod_a i stores it in an internal audio buffer.
6. The voice calling method based on FreeSWI TCH with flexible interruption of broadcast according to claim 5, characterized in that: In the data interaction process, the module mod_ai parses the control command received from the DU I platform, and adjusts the state of the audio playback according to the control command, and further includes: If a clear command is received, the module mod_a i will clear the data in the current audio buffer to prepare space for new audio content; If a playback-related command is received, the module mod_a i adjusts the state of the audio playback according to the command instruction.
7. The voice calling method based on FreeSWITCH with flexible interruption of broadcast according to claim 6 is characterized in that: In the data interaction process, the module mod_ai parses the control command received from the DU I platform, and adjusts the state of the audio playback according to the control command, and further includes: The specific playback process is based on the audio data in the buffer and the instructions of the control command, including: If there is audio data in the buffer and the play condition is play, the module mod_a i starts playing the audio until the audio in the buffer is finished playing; During the playback process, if a pause command is received, the module mod_ai will pause the current audio playback; If the audio playback is paused and a resume command is received, the module mod_a i will continue to play the remaining audio data in the buffer and resume the voice broadcast; If a clear command is received, the module mod_a i will clear the data in the current audio buffer to prepare for the playback of new audio data, thus interrupting the current audio playback; After the audio resumes playing, if a pause command or a clear command is received, the module mod_a i can perform pause, resume, and interrupt operations multiple times before the audio playback ends until the audio playback ends.
8. A voice call system based on FreeSWI TCH that can flexibly interrupt broadcasting, characterized in that: include: A two-way communication module is used for developing a module mod_ai based on FreeSWITCH, wherein the FreeSWITCH is connected to the DU I platform through the module mod_ai communication; A data reading module, used for the module mod_ai to read the user's real-time audio stream from the FreeSWITCH and send it to the DU I platform for voice recognition; A speech recognition module, used for the DU I platform to call the ASR engine for speech recognition and generate a response text; An audio conversion module, used for the DU I platform to call the TTS engine to convert the answer text into an audio stream and return it to the module mod_ai; An audio feedback module, used for the module mod_a i to write the audio stream into the FreeSWI TCH and send it to the user; The audio control module is used to parse the control command received from the DU I platform during the data interaction process and adjust the state of the audio playback according to the control command.
9. An electronic device, characterized in that: The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a voice call method based on FreeSWI TCH that can flexibly interrupt broadcasting as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and when the program is executed by the processor, it is used to implement a voice call method based on FreeSWITCH that can flexibly interrupt broadcasting as described in any one of claims 1 to 7.
Citation Information
Cited By
ASR real-time character transferring method for voice communication in customer service calling platform
CN121148393A