FreeSwitch real-time voice monitoring and call control system and method based on WebSocket

By loading dynamic modules and a voice gateway into FreeSwitch, real-time voice monitoring and call control based on WebSocket were achieved, solving the problems of latency and high cost of real-time voice stream monitoring in FreeSwitch call centers, and realizing millisecond-level real-time control and visualized operation and maintenance.

CN121000710APending Publication Date: 2025-11-21TIANJIN AUTOHOME DATA INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511254019.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies in FreeSwitch call centers suffer from high latency, high cost, and high technical difficulty in real-time voice stream monitoring, making it impossible to achieve real-time voice stream monitoring and call control.

Method used

The FreeSwitch real-time voice monitoring and call control system, based on WebSocket, is adopted. By loading dynamic modules into FreeSwitch, a voice gateway is used to maintain a long connection with FreeSwitch, and voice streams are pushed to the ASR service for parsing in real time. Based on quality inspection rules or agent instructions, executable commands for FreeSwitch are generated to achieve real-time closed-loop control.

Benefits of technology

It achieves millisecond-level end-to-end latency from voice stream to ASR to quality inspection to FreeSwitch commands, supports 500 concurrent connections, reduces hardware and maintenance costs, simplifies the architecture, and supports rapid expansion and visualized maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121000710A_ABST
    Figure CN121000710A_ABST
Patent Text Reader

Abstract

The invention discloses a FreeSwitch real-time voice monitoring and call control method and a FreeSwitch real-time voice monitoring and call control system based on WebSocket. After the voice monitoring module is dynamically loaded in the FreeSwitch, the real-time voice flow of the call channel is pushed to the voice gateway through the WebSocket according to a uidaudio command; and the gateway completes ASR identification and quality inspection rule matching, and generates a FreeSwitch executable command in real time according to a violation event or an agent manual instruction, thereby realizing closed-loop control such as millisecond-level hang-up, switching and the like. The system does not need an MRCP server, supports multi-channel concurrent monitoring, hot plugging and log tracing, is simple in deployment, low in delay and high in expansibility, and can be widely applied to call center real-time quality inspection and seat auxiliary scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of call center softswitch technology, and in particular to a FreeSwitch real-time voice monitoring and call control system and method based on WebSocket. Background Technology

[0002] Call centers are frequently established for business purposes, with agents making calls to promote services or customers calling customer service numbers for agents to answer inquiries. During these calls, various scenarios arise, such as agents being unfamiliar with the company's business and providing superficial answers, or users verbally abusing agents. To review the call content, it's necessary to capture the audio, convert it to text using ASR (Automatic Speech Retrieval System), and then perform text analysis. Currently, when building call centers using FreeSwitch (FS), a common method is to record calls using FS's built-in recording function and then upload the recordings to a backend for centralized processing. This method has significant latency and can only be used for post-event review. To achieve real-time audio streaming, the industry standard is to build a UniMRCP server and perform the listening within that server. This approach involves setting up a new service, increasing time and costs, and requires ensuring the stable operation of multiple servers simultaneously, making it technically challenging. Summary of the Invention

[0003] The purpose of this invention is to provide a FreeSwitch real-time voice monitoring and call control system and method based on WebSocket, thereby solving the aforementioned problems existing in the prior art.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0005] A WebSocket-based FreeSwitch real-time voice monitoring and call control system includes:

[0006] The FreeSwitch voice switching platform is used to establish call channels and generate channel identifiers.

[0007] The middle layer of the voice gateway maintains a long connection with FreeSwitch through an event channel. It is used to receive channel identifiers and send dynamic listening commands carrying WebSocket addresses, receive real-time voice streams pushed by FreeSwitch via WebSocket and forward them to external ASR services, receive text parsed by ASR and quality inspection results, and generate executable commands for FreeSwitch according to predefined rules.

[0008] The agent interface receives text messages from the voice gateway via a long connection and sends manual control commands to the voice gateway. The voice gateway translates quality inspection rule-triggered events or manual commands into FreeSwitch commands to achieve real-time closed-loop control.

[0009] Furthermore, the parameters for the dynamic monitoring command include:

[0010] Channel identifier, WebSocket target address, channel type, sampling rate, supports multi-channel concurrent monitoring without modifying the FreeSwitch source code.

[0011] Furthermore, the voice gateway interacts with FreeSwitch through an event channel, which is built on FreeSwitch's native event mechanism and requires no additional protocol conversion.

[0012] Furthermore, the events that trigger quality inspection rules include:

[0013] The system detects abusive keywords and business-sensitive terms, triggering actions such as automatically disconnecting the call or transferring it to a quality control agent.

[0014] Furthermore, the agent's interactive terminal receives text and sends control commands through the same WebSocket long connection, enabling human-machine collaborative control;

[0015] The agent interaction terminal also includes: a visual page for displaying the text content parsed by ASR in real time, and providing buttons for hanging up, transferring, and mute; after clicking the button, control commands are sent to the voice gateway through the same WebSocket connection.

[0016] Furthermore, the middleware layer of the voice gateway includes:

[0017] The buffer management module is used to receive voice data frames copied by the FreeSwitch module, and resample and convert them according to the configured sampling rate and channel type to ensure that the WebSocket transmitted data format is consistent.

[0018] The rules engine module is used to parse the text returned by ASR and call the company's quality inspection service API to determine whether there is any illegal content; if a rule is triggered, the corresponding FreeSwitch command identifier and parameters are generated.

[0019] The logging and tracing module records the channel identifier, timestamp, ASR parsing result, and triggered FreeSwitch command for each monitoring command to support post-event auditing and rule optimization.

[0020] Furthermore, FreeSwitch modules are dynamically loaded, enabling hot-swapping through FreeSwitch's module loading interface. This allows you to enable or disable the voice monitoring function without restarting the FreeSwitch service.

[0021] A real-time voice monitoring and call control method based on FreeSwitch, based on the same concept, includes the following steps:

[0022] S1. The agent triggers a dialing request through a webpage. The voice gateway converts the request into the FreeSwitch originate command and establishes a call channel. At the same time, it records the mapping relationship between the channel identifier and the agent session identifier.

[0023] S2. After the call channel is established, the voice gateway sends a uuid_audio_fork command to FreeSwitch. The command carries the channel identifier, target WebSocket address, channel type, sampling rate and custom user data to start real-time audio stream copying for the channel.

[0024] S3 and FreeSwitch copy the real-time audio stream frame by frame to the buffer through a dynamically loaded audio monitoring module, and then resample it according to the channel type and sampling rate before pushing it to the audio gateway via the WebSocket address.

[0025] S4. The voice gateway receives the voice stream, forwards it to the ASR service for real-time speech recognition, obtains the text result, and then calls the quality inspection service for rule matching.

[0026] S5. If a violation rule is matched, the voice gateway queries the preset command template according to the rule ID, converts the violation event into a FreeSwitch executable command, and immediately issues it for execution.

[0027] S6. Regardless of whether the rule is triggered, the voice gateway pushes the text result to the agent's webpage in real time through the same WebSocket connection. The agent issues manual control instructions based on the text result. The voice gateway converts the manual control instructions into FreeSwitch commands and issues them for execution immediately.

[0028] S7. The voice gateway continuously records the channel identifier, timestamp, ASR text, quality inspection results, and the final FreeSwitch command issued for each uuid_audio_fork command, forming a traceable log.

[0029] Furthermore, in step S2, the parameters of the uuid_audio_fork command are dynamically generated by the voice gateway through the HTTP / JSON interface to enable concurrent monitoring of different channels, different sampling rates, and different channel types without restarting the FreeSwitch service;

[0030] In step S3, the voice monitoring module is mounted on the target channel inside FreeSwitch using MediaBug and is automatically unloaded when the channel is destroyed to avoid memory leaks.

[0031] The command template in step S5 includes: when the rule type is "abuse", the generated FreeSwitch command is uuid_kill. <channel-id> <channel-id> <target-number>;

[0032] Furthermore, in step S6, the agent's web page simultaneously receives text streams and sends control commands through the same WebSocket connection, realizing bidirectional communication on a single link, and the end-to-end delay of the control commands is less than 100ms.

[0033] In step S7, the logs are persisted to the local time-series database of the voice gateway with the channel identifier as the primary key, for post-event rule effect evaluation and channel-level fault troubleshooting.

[0034] The beneficial effects of this invention are as follows: This invention discloses a method and system for real-time voice monitoring and call control in FreeSwitch based on WebSocket; after dynamically loading the voice monitoring module in FreeSwitch, the real-time voice stream of the call channel is pushed to the voice gateway via WebSocket through the uuid_audio_fork command; the gateway completes ASR recognition and quality inspection rule matching, and generates FreeSwitch executable commands in real time according to violation events or manual instructions from agents, realizing closed-loop control such as millisecond-level hang-up and call transfer. This invention has the following beneficial effects:

[0035] Extremely simple architecture: Only one dynamic module needs to be loaded in FreeSwitch, and together with a lightweight voice gateway, it can replace the traditional "FreeSwitch+MRCP+ASR" triple server architecture, reducing deployment time from several days to minutes, and reducing hardware and maintenance costs by more than 50%.

[0036] Millisecond-level closed loop: The end-to-end latency of voice stream → ASR → quality inspection → FreeSwitch command is stabilized within 100ms, enabling "second-level" automatic disconnection, transfer, or manual intervention by agents for prohibited content such as abusive language and sensitive words.

[0037] Elastic scaling: Listening channels can be dynamically added or removed via parameterized commands of uuid_audio_fork. A single gateway instance has been verified to support 500 concurrent connections. Expansion only requires horizontally adding gateway nodes without restarting FreeSwitch.

[0038] Operation and maintenance visibility: Channel-level logs (channel ID, timestamp, text, control commands) are centrally stored, reducing fault location time from hours to minutes, and supporting hot rule updates and canary releases. Attached Figure Description

[0039] Figure 1 This is a diagram of the overall system architecture of the present invention;

[0040] Figure 2 This is a flowchart of the uuid_audio_fork command in FreeSwitch of this invention;

[0041] Figure 3 This is a diagram of the real-time call monitoring and voice parsing architecture based on UniMRCP of the present invention;

[0042] Figure 4 This is a diagram illustrating the call flow and quality inspection intervention timing of the present invention;

[0043] Figure 5 This is a schematic diagram of the real-time text display interface on the agent terminal of the present invention;

[0044] Figure 6 This is a flowchart of the method of the present invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0046] Reference Figures 1 to 5 The FreeSwitch real-time voice monitoring and call control system based on WebSocket, shown below, includes:

[0047] The FreeSwitch voice switching platform is used to establish call channels and generate channel identifiers.

[0048] The middle layer of the voice gateway maintains a long connection with FreeSwitch through an event channel. It is used to receive channel identifiers and send dynamic listening commands (uuid_audio_fork) carrying WebSocket addresses, receive real-time voice streams pushed by FreeSwitch via WebSocket and forward them to external ASR services, receive ASR parsed text and quality inspection results, and generate FreeSwitch executable commands according to predefined rules.

[0049] The agent's interactive terminal receives text returned by the voice gateway via a long connection and issues manual control commands to the voice gateway. The voice gateway translates the quality inspection rule trigger events or manual commands into FreeSwitch commands (such as hang-up or transfer) to achieve real-time closed-loop control.

[0050] Furthermore, the parameters for the dynamic monitoring command include:

[0051] The system includes channel identifier (channel-id), WebSocket target address (ws-url), channel type (mono / stereo), and sampling rate (8k / 16k), supporting multi-channel concurrent monitoring without modifying the FreeSwitch source code.

[0052] Furthermore, the voice gateway interacts with FreeSwitch through an event channel, which is established based on FreeSwitch's native event mechanism (such as CHANNEL_ANSWER) without requiring additional protocol conversion.

[0053] Furthermore, the events that trigger quality inspection rules include:

[0054] The system detects abusive keywords and business-sensitive terms, triggering actions such as automatically disconnecting the call or transferring it to a quality control agent.

[0055] Furthermore, the agent's interactive terminal receives text and sends control commands through the same WebSocket long connection, enabling human-machine collaborative control;

[0056] The agent interaction terminal also includes: a visual page for displaying the text content parsed by ASR in real time, and providing buttons for hanging up, transferring, and mute; after clicking the button, control commands are sent to the voice gateway through the same WebSocket connection.

[0057] Furthermore, the middleware layer of the voice gateway includes:

[0058] The buffer management module is used to receive voice data frames copied by the FreeSwitch module, and resample and convert them according to the configured sampling rate and channel type to ensure that the WebSocket transmitted data format is consistent.

[0059] The rules engine module is used to parse the text returned by ASR and call the company's quality inspection service API to determine whether there is any violation content; if a rule is triggered, the corresponding FreeSwitch command identifier (such as "uuid_kill" or "uuid_transfer") and parameters are generated.

[0060] The logging and tracing module records the channel identifier, timestamp, ASR parsing result, and triggered FreeSwitch command for each monitoring command to support post-event auditing and rule optimization.

[0061] Furthermore, FreeSwitch modules are dynamically loaded, enabling hot-swapping through FreeSwitch's module loading interface (mod_interface). This allows you to enable or disable the voice monitoring function without restarting the FreeSwitch service.

[0062] Reference Figure 6 The method for real-time voice monitoring and call control based on FreeSwitch, which is based on the same concept, includes the following steps:

[0063] S1. The agent triggers a dialing request through a webpage. The voice gateway converts the request into the FreeSwitch originate command and establishes a call channel. At the same time, it records the mapping relationship between the channel identifier and the agent session identifier.

[0064] S2. After the call channel is established, the voice gateway sends a uuid_audio_fork command to FreeSwitch. The command carries the channel identifier, target WebSocket address, channel type, sampling rate and custom user data to start real-time audio stream copying for the channel.

[0065] S3 and FreeSwitch copy the real-time audio stream frame by frame to the buffer through a dynamically loaded audio monitoring module, and then resample it according to the channel type and sampling rate before pushing it to the audio gateway via the WebSocket address.

[0066] S4. The voice gateway receives the voice stream, forwards it to the ASR service for real-time speech recognition, obtains the text result, and then calls the quality inspection service for rule matching.

[0067] S5. If a violation rule is matched, the voice gateway queries the preset command template according to the rule ID, converts the violation event into a FreeSwitch executable command, and immediately issues it for execution.

[0068] S6. Regardless of whether the rule is triggered, the voice gateway pushes the text result to the agent's webpage in real time through the same WebSocket connection. The agent issues manual control instructions based on the text result. The voice gateway converts the manual control instructions into FreeSwitch commands and issues them for execution immediately.

[0069] S7. The voice gateway continuously records the channel identifier, timestamp, ASR text, quality inspection results, and the final FreeSwitch command issued for each uuid_audio_fork command, forming a traceable log.

[0070] Furthermore, in step S2, the parameters of the uuid_audio_fork command are dynamically generated by the voice gateway through the HTTP / JSON interface to enable concurrent monitoring of different channels, different sampling rates, and different channel types without restarting the FreeSwitch service;

[0071] In step S3, the voice monitoring module is mounted on the target channel inside FreeSwitch using MediaBug and is automatically unloaded when the channel is destroyed to avoid memory leaks.

[0072] The command template in step S5 includes: when the rule type is "insult", the generated FreeSwitch command is uuid_kill. <channel-id> <channel-id> <target-number>;

[0073] Furthermore, in step S6, the agent's web page simultaneously receives text streams and sends control commands through the same WebSocket connection, realizing bidirectional communication on a single link, and the end-to-end delay of the control commands is less than 100ms.

[0074] In step S7, the logs are persisted to the local time-series database of the voice gateway with the channel identifier as the primary key, for post-event rule effect evaluation and channel-level fault troubleshooting.

[0075] Reference Figure 1 As shown in [Process 1], the agent dials the customer's number.

[0076] Step 1: Start the logic.

[0077] Step 2: The agent uses the webpage button to make a call. The agent number is 1001, and the customer number is 1002.

[0078] Step 3: The web page backend establishes a link with the voice gateway. After the link is successfully established, the voice gateway will simultaneously record the correspondence between the link and the agent's number. This link will be used continuously during subsequent data transmission with the agent until the agent or the called party hangs up the phone.

[0079] Step 4: The voice gateway establishes a communication channel with the FS. After the channel is established, commands can be sent to the FS through this channel, and related FS behavior events can also be transmitted to the voice gateway through this channel.

[0080] Step 5: The voice gateway converts the dialing command into an FS command. The FS naming format is:

[0081] originate user / 1001&bridge(user / 1002)

[0082] This command tells FS to call both user 1001 and user 1002 simultaneously. After a successful call, the channel identifiers created by the two calls are bridged, allowing the two parties to communicate.

[0083] Step 6: After receiving the command, the FS executes the call command. Once the call begins, the FS triggers the "CHANNEL_CREATE" event. This event contains a lot of information, including the event name, channel identifier, calling party, and called party. The channel identifier will be used in subsequent steps [Process 2], [Process 3], [Process 4], and [Process 5].

[0084] Step 7: After receiving the "Channel Establishment" event, the voice gateway records the channel identifier, the calling party, and the called party. Step 8: After the FS executes the command, if both the calling and calling parties answer, the "Channel Response (CHANNEL_ANSWER)" event is triggered. The event is also triggered if the call fails. This article primarily deals with voice monitoring during successful calls; therefore, the situation when the call fails is not detailed here.

[0085] Step 9: After receiving the "Channel Response" event, the voice gateway locates the channel identifier and issues a command to the FS to listen for and receive the voice stream. The command format is:

[0086] uuid_audio_fork<start|stop> <channel-id> <ws-url> <type> <sampling-rate> <user-data> < / user-data> < / sampling-rate> < / type> < / ws-url> < / channel-id>

[0087] The command parameters are: `start` to start monitoring; `stop` to stop monitoring; `channel-id` to identify the channel to monitor; `ws-url` to send the audio stream; `type` to indicate the audio stream channel type (mono or stereo); `sampling-rate` to indicate the audio sampling rate (currently 8kHz or 16kHz); and `user-data` including the monitoring event name and other parameters (this value can be customized). The parameters passed in the command will be used in [Process 2], and its internal execution flow can be found in [Process 2].

[0088] Step 10: The voice gateway prepares to receive the voice stream from the call.

[0089] Step 11: End logic.

[0090] Reference Figure 2 The following is the voice acquisition and transmission process, shown in [Process 2]:

[0091] Step 1 Start Logic

[0092] Step 2: After the FS receives the `uuid_audio_fork` command, it begins parsing the command. After parsing, the execution process is divided into two parts: Process 1 and Process 2, which occur simultaneously. (See reference...) Figure 2 .

[0093] Step 3 begins by introducing Process 1. Process 1 first performs initial configuration operations on the channel to be copied, including the channel identifier, sampling rate, number of channels, etc. These parameters are passed through Process 1.

[0094] Step 4: Register your custom listener function with the FS channel listening interface. FS internally reserves a channel listening interface. When you need to listen for channel information, simply add the function pointer of your custom function to the listening interface. When there is data on the channel, the custom function will be called.

[0095] After registering the listening function in step 5, when data is available on the channel, the listening function will copy the channel's audio data into a buffer and then wait for other functions to read it. During copying, data resampling and other operations will be performed based on the parameters read in step 3 to ensure that the data copied to the buffer is in the target format.

[0096] Step 6, process 1, will repeatedly execute step 5 until the channel is closed.

[0097] Step 7: This step describes process 2.

[0098] In step 8, process 2, the module first establishes a connection with the WebSocket server read from the command.

[0099] Step 9 calls a function to read the buffer. If the buffer contains data, it is read directly, and then the data is sent to the server via the link obtained in Step 8. The server-side processing flow can be found in [Flow 3].

[0100] Step 10, process 2, will repeatedly execute step 9 until the channel is closed.

[0101] Step 11 ends the logic.

[0102]

Process 3

[0103] Step 1 begins the logic.

[0104] Step 2: Start the WebSocket service in the voice gateway.

[0105] Step 3 involves receiving the audio stream from [Process 2] in the `onMessage` method of the WebSocket server. After receiving the audio stream, it is processed according to the channel configuration. If it is stereo, the audio stream is separated into channels. If it is mono, no processing is performed. The configuration file is predefined in the project.

[0106] Step 4: Launch the ASR client and transmit the audio stream to the ASR server for parsing. The ASR service is one of the company's basic services, accessible via a WebSocket interface; simply follow the interface documentation for instructions. This document does not explain the internal mechanisms of the ASR service.

[0107] Step 5: The ASR service completes speech parsing and sends the text back to the calling client.

[0108] Step 6 involves text processing. This processing includes two parts. The first part involves text quality control. If the text triggers predefined rules, corresponding event handling is performed. These predefined rules are based on business needs, such as prohibiting abusive language and profanity, and avoiding inquiries about prices. Corresponding event handling is also predefined, such as hanging up the call when abusive language is detected. The quality control rules used in this document are part of the company's basic services and are accessed via API. This document does not explain the internal mechanisms of the quality control service. The second part involves transmitting the text content to the call agent, who can view it in real-time on a webpage or software, facilitating script review. The detailed process for Part 1 can be found in [Process 4], and the detailed process for Part 2 can be found in [Process 5].

[0109] Step 7 concludes the logic.

[0110] [Process 4] Real-time call control based on quality inspection results:

[0111] Step 1 begins the logic.

[0112] Step 2: When the quality inspection service detects that the call content contains content that violates the rules, an alarm event will be triggered.

[0113] Step 3: The language gateway finds the corresponding processing flow based on the alarm event.

[0114] In steps 3 and 4, different alarm events correspond to different processing procedures. The voice gateway translates these procedures into corresponding FS commands based on the specific procedures.

[0115] Step 5: The voice gateway sends the corresponding command to the FS for execution.

[0116] Step 6: FS executes the corresponding command.

[0117] Step 7 concludes the logic.

[0118] [Step 5] The parsing results are sent to the calling end, where the agent can view and control the call:

[0119] Step 1 begins the logic.

[0120] Step 2: At this point, the speech has been converted to text. In step 3 of [Process 1], the voice gateway recorded the agent's information and the link between the agent and the voice gateway. The text content can now be transmitted through this link to the webpage the agent used at the start of the call.

[0121] Step 3: After receiving the text content, the webpage displays the text content on the webpage.

[0122] Step 4: Based on the text content, the agent can choose to take further action.

[0123] After the agent selects the next operation in step 5, the command can be transmitted to the voice gateway using the same link from step 2.

[0124] Step 6: After receiving the agent's operation, the voice gateway converts the operation into a corresponding FS command and sends it for execution.

[0125] Step 7: FS executes the corresponding command.

[0126] Step 8 ends the logic.

[0127] Furthermore, firstly, compare Figure 1 and Figure 3 ,exist Figure 1 In this context, the audio stream is transmitted via the WebSocket protocol.

[0128] WebSocket is supported by many language frameworks, and its interface is simple to use. It typically involves creating a WebSocket connection instance and then implementing message handling logic for four events: onConnection, onMessage, onClose, and onError. Data transmission is mainly handled through onMessage; you only need to process the received data in this method. If UniMRCP is used for language stream acquisition... Figure 3 As shown, firstly, you need to set up a UniMRCP service; secondly, you need to understand the framework's development interface; and thirdly, you need to understand the RCP protocol. UniMRCP is a professional server for the audio and video field, while RCP is a real-time transmission protocol with complex parsing. These two technical modules add extra development and learning costs.

[0129] The general procedure for an agent to make a phone call is as follows: Figure 4 As shown in the diagram, the solution described in this paper involves monitoring in step 4, with the voice stream being transmitted to the receiving end for processing in real time, ensuring timely monitoring. In contrast, the monitoring solution in Scheme 1 only begins recording in step 4, and after the call ends, the recording file needs to be reported before monitoring can be completed, making its timeliness far from sufficient.

[0130] The monitoring module is loaded using FS's built-in module loading method. The module is relatively independent and not tightly coupled with the FS project. The channel identifier to be monitored can be passed as a command parameter, internally implementing monitoring for specific channels. To monitor multiple channels, simply re-execute the command, passing different channel identifiers.

[0131] Within the monitoring module, the voice stream is copied to the buffer via data frame replication and then sent to the voice gateway. This real-time replication method enables real-time monitoring of the call content, eliminating the need to wait for the FS (Fixed Streaming Service) to complete call recording and then perform quality checks on the recording file, ensuring real-time monitoring. Furthermore, the call can be controlled in real-time. The following example illustrates the process of an agent dialing customer number 1002 from number 1001. In step 4 of [Process 1], a communication channel between the voice gateway and the FS has been established. Through [Process 3], voice stream monitoring and content quality checks can be implemented. If any inappropriate content is detected, an alarm event is triggered. Taking the abusive language alarm rule as an example, the current measure for "abuseous language" is to hang up the phone. For this rule, the hang-up command is translated into an FS command.

[0132] uuid_kill <channel-id>

[0133] The channel-id is the call channel identifier. Then, through the information channel established in step 4 of [Process 1], commands are sent directly, and FS executes them. This achieves real-time control of the call.

[0134] In [Process 5], the text content parsed by ASR can be sent to the agent, and then displayed on the agent's end, as shown in the display page. Figure 5 As shown:

[0135] After viewing the content on the webpage, the agent can control the call through the webpage, such as transferring the call to another agent, hanging up the call, and searching and organizing the script based on the call content. Taking call transfer as an example, after the agent clicks the "Transfer" button, the webpage will send the user's operation to the voice gateway through the link established in [Process 1]

[003] . The voice gateway will convert the operation into an FS transfer command.

[0136] uuid_kill <channel-id><cal lee-number>

[0137] Here, `channel-id` is the call channel identifier, and `call-number` is the target number to be transferred. The voice gateway still sends commands through the information channel established in step 4 of [Process 1], which the FS then executes. In this way, real-time control of the call by the agent is achieved.

[0138] The beneficial effects of this invention are as follows: This invention discloses a method and system for real-time voice monitoring and call control in FreeSwitch based on WebSocket; after dynamically loading the voice monitoring module in FreeSwitch, the real-time voice stream of the call channel is pushed to the voice gateway via WebSocket through the uuid_audio_fork command; the gateway completes ASR recognition and quality inspection rule matching, and generates FreeSwitch executable commands in real time according to violation events or manual instructions from agents, realizing closed-loop control such as millisecond-level hang-up and call transfer. This invention has the following beneficial effects:

[0139] Extremely simple architecture: Only one dynamic module needs to be loaded in FreeSwitch, and together with a lightweight voice gateway, it can replace the traditional "FreeSwitch+MRCP+ASR" triple server architecture, reducing deployment time from several days to minutes, and reducing hardware and maintenance costs by more than 50%.

[0140] Millisecond-level closed loop: The end-to-end latency of voice stream → ASR → quality inspection → FreeSwitch command is stabilized within 100ms, enabling "second-level" automatic disconnection, transfer, or manual intervention by agents for prohibited content such as abusive language and sensitive words.

[0141] Elastic scaling: Listening channels can be dynamically added or removed via parameterized commands of uuid_audio_fork. A single gateway instance has been verified to support 500 concurrent connections. Expansion only requires horizontally adding gateway nodes without restarting FreeSwitch.

[0142] Operation and maintenance visibility: Channel-level logs (channel ID, timestamp, text, control commands) are centrally stored, reducing fault location time from hours to minutes, and supporting hot rule updates and canary releases.

[0143] By adopting the above-disclosed technical solution of this invention, the following beneficial effects are obtained:

[0144] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. < / channel-id> < / channel-id> < / channel-id> < / channel-id>

Claims

1. A FreeSwitch real-time voice monitoring and call control system based on WebSocket, characterized in that, include: The FreeSwitch voice switching platform is used to establish call channels and generate channel identifiers. The middle layer of the voice gateway maintains a long connection with FreeSwitch through an event channel. It is used to receive the channel identifier and send dynamic listening commands carrying the WebSocket address, receive real-time voice streams pushed by FreeSwitch via WebSocket and forward them to external ASR services, receive text parsed by ASR and quality inspection results, and generate executable commands for FreeSwitch according to predefined rules. The agent interaction terminal receives text returned by the voice gateway via a long connection and issues manual control commands to the voice gateway; wherein, the voice gateway uniformly translates the quality inspection rule trigger events or manual commands into FreeSwitch commands to achieve real-time closed-loop control.

2. The system according to claim 1, characterized in that, The parameters of the dynamic monitoring command include: Channel identifier, WebSocket target address, channel type, sampling rate, supports multi-channel concurrent monitoring without modifying the FreeSwitch source code.

3. The system according to claim 1, characterized in that, The voice gateway interacts with FreeSwitch through an event channel, which is established based on FreeSwitch's native event mechanism and requires no additional protocol conversion.

4. The system according to claim 1, characterized in that, The quality inspection rule triggering events include: The system detects abusive keywords and business-sensitive terms, triggering actions such as automatically disconnecting the call or transferring it to a quality control agent.

5. The system according to claim 1, characterized in that, The agent interaction terminal receives text and sends control commands through the same WebSocket long connection to achieve human-machine collaborative control; The agent interaction terminal also includes: a visual page for displaying the text content parsed by ASR in real time, and providing buttons for hanging up, transferring, and mute operations; When the button is clicked, a control command is sent to the voice gateway via the same WebSocket connection.

6. The system according to claim 1, characterized in that, The middle layer of the voice gateway includes: The buffer management module is used to receive voice data frames copied by the FreeSwitch module, and resample and convert them according to the configured sampling rate and channel type to ensure that the WebSocket transmitted data format is consistent. The rules engine module is used to parse the text returned by ASR and call the company's quality inspection service API to determine whether there is any illegal content; if a rule is triggered, the corresponding FreeSwitch command identifier and parameters are generated. The logging and tracing module records the channel identifier, timestamp, ASR parsing result, and triggered FreeSwitch command for each monitoring command to support post-event auditing and rule optimization.

7. The system according to claim 1, characterized in that, The FreeSwitch module is a dynamically loaded module that is hot-swappable through the FreeSwitch module loading interface, allowing the voice monitoring function to be enabled or disabled without restarting the FreeSwitch service.

8. A method for real-time voice monitoring and call control based on FreeSwitch, characterized in that, Includes the following steps: S1. The agent triggers a dialing request through a webpage. The voice gateway converts the request into the FreeSwitch originate command and establishes a call channel. At the same time, it records the mapping relationship between the channel identifier and the agent session identifier. S2. After the call channel is established, the voice gateway sends a uuid_audio_fork command to FreeSwitch. The command carries the channel identifier, target WebSocket address, channel type, sampling rate and custom user data to start real-time audio stream copying for the channel. S3 and FreeSwitch copy the real-time audio stream frame by frame to the buffer through the dynamically loaded audio monitoring module, and after resampling according to the channel type and sampling rate, push it to the audio gateway through the WebSocket address. S4. The voice gateway receives the voice stream, forwards it to the ASR service for real-time speech recognition, obtains the text result, and then calls the quality inspection service for rule matching. S5. If a violation rule is matched, the voice gateway queries the preset command template based on the rule ID, converts the violation event into a FreeSwitch executable command, and immediately issues it for execution. S6. Regardless of whether the rule is triggered, the voice gateway pushes the text result to the agent's webpage in real time through the same WebSocket connection. The agent issues a manual control command based on the text result. The voice gateway converts the manual control command into a FreeSwitch command and issues it for execution immediately. S7. The voice gateway continuously records the channel identifier, timestamp, ASR text, quality inspection results, and the final FreeSwitch command issued for each uuid_audio_fork command, forming a traceable log.

9. The method according to claim 8, characterized in that, The parameters of the uuid_audio_fork command in step S2 are dynamically generated by the voice gateway through the HTTP / JSON interface to enable concurrent monitoring of different channels, different sampling rates and different channel types without restarting the FreeSwitch service; In step S3, the voice monitoring module is mounted on the target channel in FreeSwitch using MediaBug and is automatically unloaded when the channel is destroyed to avoid memory leaks. The command template mentioned in step S5 includes: when the rule type is "insults", the generated FreeSwitch command is uuid_kill. <channel-id> <channel-id> <target-number> 。< / target-number> < / channel-id> < / channel-id> 10. The method according to claim 8, characterized in that, In step S6, the agent webpage simultaneously receives text streams and sends control commands through the same WebSocket connection, realizing bidirectional communication on a single link, and the end-to-end delay of the control commands is less than 100ms. In step S7, the logs are persisted to the local time-series database of the voice gateway with the channel identifier as the primary key, for post-event rule effect evaluation and channel-level fault troubleshooting.

Citation Information

Cited By

  • AI call supervision method, system and device, and storage medium

    CN121887920A