A system and method for enhancing a call characterization of voice calls
A platform computes audio metrics in real-time for wholesale international calls, addressing scalability and latency issues, enhancing call characterization and operational insights through eCDRs, ensuring efficient and reliable communication.
Patent Information
- Application Number
- PCT/EP2025/069103
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-09
- Filing Date
- 2025-07-04
- Publication Date
- 2026-01-15
AI Technical Summary
Existing wholesale international voice call management systems fail to utilize the vast amount of audio information within calls for enhancing operational capabilities and insights, despite handling large volumes of calls, and face challenges in real-time processing and scalability.
A platform that computes metrics from two-channel audio streams, including speech transcription, speaker ID, language, and paralinguistics, within a switch engine, using DSP, AS, and MC engines, generating enhanced Call Detail Records (eCDRs) to store and process massive call traffic efficiently, ensuring low and constant latency.
Enables real-time processing of tens of thousands of calls per hour with continuous availability and horizontal scalability, providing enhanced call characterization metrics for seamless, reliable, and cost-effective international communication.
Smart Images

Figure EP2025069103_15012026_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] A SYSTEM AND METHOD FOR ENHANCING A CALL CHARACTERIZATION OF VOICE CALLS
[0003] OBJECT OF THE INVENTION
[0004] The present invention introduces a unique and comprehensive method for enhancing the telephone voice records (i.e., voice calls) with new call characterization information within the wholesale telecommunication business in real-time. This invention holds particular significance for service providers that offer customer applications that employ telephonic traffic for business operations involving audio processing. Examples include insurance agencies and call centers utilizing speech analytics and transcription tools, as well as online banking employing biometric authentication.
[0005] TECHNICAL FIELD
[0006] The invention belongs to the telecommunications technology sector, and more specifically to those technologies that monitor and evaluate the quality of phone calls.
[0007] BACKGROUND OF THE INVENTION
[0008] The wholesale international call traffic consists of a wide array of communication exchanges between different countries involving both, traditional circuit-switched networks and modern IP-based networks. This traffic includes two types of voice call exchanges routed through telecommunications networks, the traditional telephone calls made between individuals or businesses located in different countries and VoIP (Voice over Internet Protocol) calls often used by businesses for cost-effective international calling. Wholesale international call traffic is typically managed by carriers and service providers who lease capacity from each other to facilitate seamless communication between different regions of the world. These calls can range from personal conversations to business dealings, and they play a crucial role in global connectivity and commerce.
[0009] The objective of the wholesale international traffic of voice calls (or wholesale voice management) is to efficiently handle the routing, termination, and billing of large volumes of voice calls between telecommunications carriers and service providers. This involves cost optimization through negotiating favorable rates and optimizing routing paths, ensuring high-quality voice calls by maintaining low latency and minimal jitter, and capacity planning to meet peak call volumes. Interconnection agreements are established and managed to facilitate the exchange of voice traffic across borders, while measures are implemented to prevent fraud and ensure accurate billing and settlement processes. Compliance with regulatory requirements, including licensing, taxation, and local telecommunications laws, is also a crucial aspect of wholesale voice management. Overall, the goal is to facilitate seamless, reliable, and cost-effective international voice communication while upholding standards of service quality and regulatory compliance.
[0010] Despite the wholesale international call traffic handling large volumes of voice calls, the information into the audio itself is rarely employed for any issue. For instance, the quality requirements are mostly addressed through the Average Length of Conversation (ALOC) and the Answer Seizures Ratio (ASR), when even with good levels of ALOC and ASR, the voice signal inside the call could be noisy or contain acoustic artifacts that harm the quality of the conversation. Furthermore, the audio contains a significant amount of information about the conversation, including technical quality cues, the spoken text, speaker id, language, conversation flow, nonverbal aspects of the conversation such as emotions, speakers’ health state, engagement, etc. All this information is implicitly transported inside the call traffic without any knowledge or use of it.
[0011] The US patent document with a publication number US10375234B1 , published on August 6, 2019 discloses a PBX (Private Branch Exchange) platform that analyzes audio for volume level changes as enhanced information within a voice call for controlled-environment facilities, like correctional facilities. Therefore, this document presents limitations regarding the number of calls the platform can handle simultaneously (local environment means a few calls at the same time) and also regarding the type of information that can be added to the CDR. Regarding how an audio call is classified, the patent application US 20180144040 A1 , published on May 24. 2018, discloses a classification method that uses a combination of Conventional AMD (Answering Machine Detection), Acoustic fingerprinting system and Real-time audio stream processing. Specifically, this document discloses a sample rate of 8 kHz, a frame duration of 30 milliseconds, a step size of 20 milliseconds, an overlap of 1 / 3 and a FFT size of 256. BRIEF DESCRIPTION OF THE INVENTION
[0012] The present invention elevates the value of telephone voice records, addressing the challenge of processing vast data volumes inherent in wholesale telecommunications environments in real-time. The present invention equips customers with additional information computed over the audio, thereby enhancing their operational capabilities and insights.
[0013] In order to do the above, the present invention describes the two-channel streams of audio within a telephone call by computing a set of metrics from different information dimensions. These information dimensions encompass the technical quality in speech and non-speech segments, the flow of the conversation within the call, the speech transcription, the speaker id, the language, the paralinguistics aspects of speech, among others.
[0014] In addition, the platform of the present invention is designed to handle massive call traffic — on the order of tens or even hundreds of thousands per hour — which implies much more demanding technical requirements in terms of resources. Unlike a known PBX platform for controlled-environment facilities, which is generally dedicated to corporate environments or specific applications with limited traffic volume, the platform of the present invention ensures continuous availability, horizontal scalability, and efficient processing of simultaneous calls on multiple protocols (such as SIP, H.323, etc.). No matter how large the PBX would be configured, it cannot meet the operational complexity or real-time requirements demanded by the platform of the present invention, whose architecture must be optimized to handle traffic spikes, provide geographic redundancy, load balancing, and fault tolerance at the infrastructure level.
[0015] Furthermore, the present invention must operate in strict real-time conditions, which means that the latency must be kept not only low but also constant, to avoid fluctuations that generate synchronization problems perceptible to users. In this regard, the present invention uses the “Real Time Factor” (RTF) defined as
[0016] RTF = latency / frame size (seconds) wherein RTF must be less or equals to “1”. In the first aspect of the invention, a method for enhancing a call characterization of voice calls is disclosed. Voice calls (tens or even hundreds of thousands per hour) are transmitted from speakers over carrier networks to different speakers over different carrier networks, the carrier networks being intercommunicated by means of the platform. The platform comprises a switch engine. The switch engine comprises a media server engine. The method comprises the steps of:
[0017] • receiving, at the platform, each voice call segmented in audio packets, each audio packet comprising an audio segment;
[0018] • calculating a frame size for each audio packet, wherein each frame size is calculated as the time of each audio segment, i.e.: o frame size (seconds) = audio segment (seconds)
[0019] • generating a Call Detail Record “CDR” comprising call parameters which are calculated based on a set of rules related to business criteria and technical criteria;
[0020] • generating an URL to a database for each voice call.
[0021] Wherein the media server engine comprises a call characterization engine that in turn comprises a Digital Signal Processing “DSP” engine, an Audio Segmentation “AS” engine and a Metrics computation engine. The method further comprises the following steps processed in the frame size (to ensure that RTF <= 1) by the call characterization engine:
[0022] • processing, by the Digital Signal Processing “DSP” engine, the audio packets of the voice call, wherein the processing step in turn comprises: o computing audio spectral representations; o accumulating a predetermined number of audio packets of the voice call until a sufficient number is gathered to compute the metrics of the call characterization;
[0023] • evaluating, by the Audio Segmentation “AS” engine, an audio type present in the accumulated audio packets;
[0024] • generating temporal marks, by the Audio Segmentation “AS” engine, indicating which audio segment correspond to each audio type;
[0025] • computing metrics, by the Metrics Computation “MC” engine, using the accumulated audio packets and the temporal marks;
[0026] • storing the metrics in the database;
[0027] • generating an enhanced Call Detail Record “eCDR” by adding the URL to the Call Detail Record “CDR”.
[0028] In one embodiment of the present invention, the metrics comprises at least one of the groups consisting of: speech transcription; audio types comprising at least: music, nonspeech, speech of each speaker, overlap between speakers and waits; several parameters of the technical quality evaluation for speech and non-speech segments; a description of the conversation flow within the call; speaker and language ids of the speakers in the conversation; and, non-verbal aspects of speech.
[0029] The call parameters comprise at least call timestamps, voice call duration, speaker direction, audio type, speaker and called number, billing, termination, and routing information.
[0030] In another embodiment of the invention, the step of evaluating an audio type present in the accumulated audio packets consists of comparing the audio (or accumulated audio packets) with a trained model with different audio types (such as music, non-speech, speech from each speaker, overlaps between speakers and pauses).
[0031] In another embodiment of the invention, the step of computing audio spectral representations comprises to calculate Energy and / or a Fast Fourier Transform (FFT).
[0032] In another embodiment of the invention, the set of rules consist of: commercial agreement, including data protection and privacy; revenues and rates; currency exchange rates; volume traffic compromises; and, quality indicators.
[0033] In another embodiment of the invention, the method further comprises monitoring the metrics with applications. The applications show the performance of the new metrics as part of the call registers. These applications generate reports and corresponding alerts based on the information in the enhanced Call Detail Record “eCDR” to follow the behaviour of the traffic exchange.
[0034] In a second aspect of the invention, a system for enhancing a call characterization of voice calls is disclosed. The system at least comprises a platform, which in turn comprises at least a call characterization engine. The system is configured to carry out the method of the first aspect of the invention.
[0035] The platform further comprises a backend for a call management and a frontend for a call traffic monitoring. The backend comprises a switch engine, a routing engine, a rule engine and a billing engine. The switch engine comprises a media server engine and the call characterization engine. The frontend comprises applications for the call traffic monitoring.
[0036] BRIEF DESCRIPTION OF THE FIGURES
[0037] In order to help with a better understanding of the features of the invention and to complement this description, the following figures are attached as an integral part of the same, by way of illustration and not limitation:
[0038] FIG. 1 is a schematic diagram that illustrates the platform for managing the wholesale international traffic of voice calls according to the state of the art.
[0039] FIG. 2 is a schematic diagram that illustrates the required devices for establishing a communication between two users from different countries. It also illustrates that the call characterization engine of the present invention is deployed in the platform.
[0040] FIG. 3 is a schematic diagram that illustrates the platform for managing the wholesale international traffic of voice calls according to the present invention.
[0041] FIG. 4 is a schematic diagram that illustrates the call characterization engine of the present invention in detail.
[0042] FIG. 5 shows the routing rules.
[0043] DETAILED DESCRIPTION OF THE INVENTION
[0044] A prior art platform (i.e. , a system that serves as a basis for operating hardware or software modules with which it is compatible) for the management of the wholesale international traffic of voice calls is illustrated in FIG. 1. The platform 1 is deployed for carrier’s interconnect traffic across a telecommunication network. The platform 1 has the backend 10 for the call management, and the frontend 11 for the call traffic monitoring. The backend 10 encompasses the switch engine 101, the routing engine 102, the rule engine 103 and the billing engine 104. The switch engine 101 takes care of network protocols and signaling. The routing engine 102 decides how to handle each call based on parameters of the call and meta-information retrieved from external systems, as well as incorporates the information of the quality, prices, commercial agreements and business decisions to the routing path optimization. The rule engine 103 or profiler checks technical criteria mainly focusing on reviewing call parameters to detect fraud. The billing engine 104 computes the call rates and generates invoices. The backend 10 produces the Call Detail Record (CDR) 12 that contains the information associated with the call transportation, including call timestamps, voice call duration, direction, type, caller and called number, billing, termination, routing information, and additional metadata. The frontend 11 is a web-based system 105 with a variety of applications that allow online monitoring the voice call 30 transportation process and all the metadata associated.
[0045] As depicted in FIGs. 2 and 3, the call characterization engine 3 is strategically positioned at the heart of the carrier's interconnection, i.e., at the heart of the platform 1 . The context of the present invention is shown in FIG. 2, where a user employs the carrier network A 201 in a country to communicate with other user in a different country through the carrier network B 202 of the latter, being the carrier network A 201 and the carrier network B 202 intercommunicated by means of the platform 1. Unlike the known systems, the carrier network B not only receives the voice call 30 and the “standard” CDR 12 from the platform 1 , but also the metrics 32 including the URL 32’ to the database 106 within the enhanced Call Detail Record (eCDR) 33 provided by the call characterization engine 3. As shown in FIG. 3, the platform 1 of the present invention comprises the database 106 configured to store the metrics 32. Thanks to the metrics 32 provided in the enhanced Call Detail Record (eCDR) 33, the user of the carrier network B 202 can monitor the voice call 30 features using customer apps 20 that manage the metrics 32.
[0046] Fig. 4 shows part of the same known platform 1 of FIG. 1 , in which the database 106 and the call characterization engine 3 of the present invention are deployed. The call characterization engine 3 comprises the Digital Signal Processing “DSP” engine 300, the Audio Segmentation “AS” engine 301 and a Metrics Computation “MC” engine 302. The call characterization engine 3 communicates with the routing engine 102, the rule engine 103 and the database 106 to carry out the call characterization method for enhancing voice calls of the present invention.
[0047] The call characterization method for enhancing voice calls of the present invention is initiated by the routing engine 102 of the platform 1 for managing the wholesale international traffic of voice calls 30. The routing engine 102 holds the information of the commercial agreement wherein the service provider of each telephone call to be transported has requested the call characterization method. Therefore, upon the arrival of a voice call 30 at platform 1 seeking routing to its destination, the routing engine 102 checks whether the call characterization method is deployed and configured. If so, the routing engine 102 sends to the switch engine 101 the entire routing data adjusted to facilitate the customized routing according to the method of the present invention. In the absence of such confirmation, the voice call follows a traditional routing.
[0048] The customized routing directs the voice call 30 to the media server engine 1101 responsible for managing the voice communication. The media server engine 1101 handles the audio within the voice call 30 encapsulated in the corresponding voice managing protocol, e.g. RTP, webRTC, websocket. Alongside typical tasks, such as transcoding between G.711 and the required codec for the destination server and the injection of Dual Tone Multi-Frequency (DTMF) events into ongoing two-channel streams of audio within a telephone call. The media server engine 1101 , thanks to the call characterization engine 3, is enhanced to include the functionality of call characterization as an additional feature for audio traffic processing. This way, the call characterization engine 3, illustrated in FIGs. 2 to 4, is implemented and compiled as part of the media server engine 1101.
[0049] As shown in FIG. 4, the media server engine 1011 operates by iteratively processing audio packets of a frame size duration, for instance of 20 milliseconds, from the voice call 30. These packets go to the call characterization engine 3 that pre-process the packets with a Digital Signal Processing “DSP” engine 300. The DSP engine 300 computes various audio spectral representations such as Energy, Fast Fourier Transform (FFT), among others, and accumulates these packets (voice call accumulated 30.1) until a sufficient number is gathered to compute the metrics 32 of the call characterization. Once the accumulated packets 30.1 reach a predefined threshold, typically 120 seconds (configurable value), the Audio Segmentation “AS” engine 301 evaluates the type of audio present in the entire set of packets. This evaluation consists of comparing the audio within the voice call 30, with a trained model with different types of audios such as music, non-speech, speech from speaker in channel 1 , speech from speaker in channel 2, overlaps between speakers, pauses, etc., that recognize the type of audio present in the voice call under analysis. This evaluation results in the establishment of temporal marks 31 , indicating which audio segment, namely a temporal segment of the audio under analysis, corresponds to each audio type mentioned. Subsequently, the Metrics Computation “MC” engine 302 utilizes both, the voice call accumulated 30.1 and the corresponding temporal marks 31 , to compute the metrics 32 and to generate the URL 32’ to the database 106. The temporal marks 31 are instrumental in defining the audio segments where it is meaningful to compute the metrics. For example, speaker and language identity metrics are computed over speech segments, while quality metrics of the transmission channel are computed over non-speech segments.
[0050] This cyclical approach accounts for the constantly changing traffic of audio packets. Overall, the media server engine 1101 processes the audio stream as usual, along with the call characterization method of the present invention. The resulting metrics 32 from the call characterization engine 3 are then stored in the data database 106 and the URL 32’ to the database is appended to the standard CDR 12, enhancing its content with the extracted information to produce the enhanced Call Detail Record “eCDR” 33. This URL 32’ can be stored within the CDR either as vector of values or via a link directing to the online location of a monitoring dashboard within the frontend 11 , where they are displayed for further analysis.
[0051] The main challenge of this process lies in meeting the traffic time requirement for the media server engine’s package processing. In order to do this, the present invention uses the “Real Time Factor” (RTF) defined as:
[0052] RTF = latency I frame size (seconds) wherein RTF <= 1. An acceptable delay (= latency), the media server engine 1011 can hold a package of 20 milliseconds (=frame size), would be in the range of 4 to 8 milliseconds. Thus, the cumulative delay introduced by the three engines (300, 301 , 302) comprised in call characterization engine 3 must remain within 8 milliseconds, fulfilling the RTF condition as RTF would be 0.4 (=8 milliseconds / 20 milliseconds). Therefore, optimizing the implementation of the call characterization method is paramount to achieve these traffic time requirements. On the other side, the call characterization engine 3 must have the capability to process multiple calls simultaneously in parallel. This is essential as the media server engine 1011 typically handles several calls concurrently. Consequently, packets from multiple concurrent calls are sent to the call characterization engine 3 in a matrix format, such that call characterization engine 3 manages the DSP processing and accumulation, AS and MC tasks for each call. To ensure the optimal routing path of a voice call incorporating call characterization, the switch engine 101 must factor in the routing rules in FIG. 5 associated with the call characterization service. Initially, it is essential to verify the commercial agreements 401 to facilitate the computation of accurate characterization metrics based on the agreed-upon information dimensions stipulated by the service provider. Also, checks the revenues and rates 402 necessary to accommodate the additional costs incurred by the call characterization service. As well as, the updated currency exchange rates 403 and previous compromises related to the assignment of certain traffic volume 404 that may change the rates. Moreover, it may be prudent to elevate the service quality indicators of available routes 405 to reflect a premium service level. In essence, all business decisions must consider the integration of the call characterization engine and method into the routing path optimization process.
[0053] The platform 1 maintains the database 106 that retains CDRs for transported calls over a certain period of time. This database serves the purpose of monitoring traffic performance and resolving any subsequent disputes arising from payment discrepancies. Enhanced eCDRs 33, are generated upon completion of the call characterization method and are stored within this database. Consequently, the database requires reconfiguration to accommodate the new fields introduced by the metrics.
[0054] Then, the billing engine 104 charges the service according to the type and number of metrics included. Namely, checking the commercial agreement with the service provider for getting the agreed price at including information about, for instance, the speaker id and the speech transcription, or for instance the speech analytics and language id for other service providers. The billing engine usually handles a distributed server with the updated information about prices and rates.
[0055] The monitoring applications 20 are also modified to be able to show the performance of the new metrics 32 as part of the call registers. These applications generate reports and corresponding alerts based on the information in the eCDRs 33 to follow the behaviour of the traffic exchange.
Claims
CLAIMS1.- A method for enhancing a call characterization of voice calls, wherein voice calls (30) are transmitted from speakers over a carrier networks (201) to different speakers over different carrier networks (202), the carrier networks intercommunicated by means of a platform (1), which comprises a switch engine (101), which in turn comprises a media server engine (1011); wherein the method comprises:• receiving, at the platform (1), each voice call (30) segmented in audio packets, each audio packet comprising an audio segment;• calculating a frame size for each audio packet, wherein each frame size is calculated as the time of each audio segment;• generating a Call Detail Record “CDR” (12) comprising call parameters which are calculated based on a set of rules (401-405) related to business criteria and technical criteria;• generating an URL (32’) to a database (106) for each voice call; wherein, the media server engine (1011) comprises a call characterization engine (3) that in turn comprises a Digital Signal Processing “DSP” engine (300), an Audio Segmentation “AS” engine (301) and a Metrics computation engine (302); wherein the method further comprises the following steps processed in the frame size by the call characterization engine (3):• processing, by the Digital Signal Processing “DSP” engine (300), the audio packets of the voice call (30), wherein the processing step in turn comprises: o computing audio spectral representations; o accumulating a predetermined number of audio packets (30.1) of the voice call;• evaluating, by the Audio Segmentation “AS” engine (301), an audio type present in the accumulated audio packets;• generating temporal marks (31), by the Audio Segmentation “AS” engine (301), indicating which audio segment correspond to each audio type;• computing metrics (32), by the Metrics Computation “MC” engine (302), using the accumulated audio packets (30.1) and the temporal marks (31);• storing the metrics (32) in the database (106);• generating an enhanced Call Detail Record “eCDR” (33) by adding the URL to the Call Detail Record “CDR” (12).2.- The method of claim 1, wherein the metrics (32) comprises at least one of the groups consisting of:• speech transcription;• audio types comprising at least: music, non-speech, speech of each speaker, overlap between speakers and waits;• several parameters of the technical quality evaluation for speech and non-speech segments;• a description of the conversation flow within the call;• speaker and language ids of the speakers in the conversation;• non-verbal aspects of speech.3.- The method of claim 2, wherein the step of evaluating an audio type present in the accumulated audio packets consists of comparing the audio with a trained model with different audio types.4.- The method of claim 1, wherein the step of computing audio spectral representations comprises to calculate Energy and / or a Fast Fourier Transform (FFT).5.- The method of claim 1, wherein the call parameters comprise at least call timestamps, voice call duration, speaker direction, audio type, speaker and called number, billing, termination, and routing information.6.- The method of claim 1, wherein the set of rules consisting of:• commercial agreement (401), including data protection and privacy;• revenues and rates (402);• currency exchange rates (403);• volume traffic compromises (404);• quality indicators (405).7.- The method of claim 1, wherein the method further comprises monitoring the metrics (32) with applications (20).8.- A system for enhancing a call characterization of voice calls, characterized in that the system at least comprises a platform (1), which in turn comprises at least a call characterization engine (3); wherein the system is configured to carry out the method ofclaims 1 to 7.9.- The system of claim 8, wherein the platform (1) further comprises a backend (10) for a call management and a frontend (11) for a call traffic monitoring.10.- The system of claim 9, wherein the backend (10) comprises a switch engine (101), a routing engine (102), a rule engine (103) and a billing engine (104).11.- The system of claim 10, wherein the switch engine (101) comprises a media server engine (1011) and the call characterization engine (3).12.- The system of claim 9, wherein the frontend (11) comprises applications (105) for the call traffic monitoring.