System and method for voice signal analytic on telephone calls
The method and system analyze voice signals to evaluate the entire telephone call, incorporating non-speech segments, addressing the limitations of existing methods by providing a comprehensive assessment of call quality and identifying key factors affecting communication.
Patent Information
- Application Number
- PCT/EP2025/061205
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-04-24
- Publication Date
- 2025-10-30
AI Technical Summary
Existing methods for assessing telephone call quality focus solely on speech segments, neglecting non-speech segments and providing an incomplete evaluation of the overall call quality, which hinders the identification of specific issues affecting communication.
A method and system that analyze voice signals using acoustic metrics, dialogue-related metrics, and paralinguistic aspects to evaluate the entire telephone call, including non-speech segments, with real-time computation capabilities.
Provides a comprehensive assessment of telephone call quality by identifying specific factors influencing communication, enabling proactive measures to enhance call performance and user experience.
Smart Images

Figure EP2025061205_30102025_PF_FP_ABST
Abstract
Description
[0001]DESCRIPTION SYSTEM AND METHOD FOR VOICE SIGNAL ANALYTIC ON TELEPHONE CALLS OBJECT OF THE INVENTIONThe present system and method for analysing the voice signal on telephone calls allows tomonitor and evaluate the overall quality as part of a telecommunication service. Additionally, the present invention helps to identify specific issues causing poor callperformance using objective metrics to quantify voice quality problems. It observes andassesses comprehensive quality attributes of telephonic conversations, incorporatingseveral metrics of the conversational dynamics among participants, the technical acousticalfidelity of the transmitted signals, and paralinguistic aspects of the conversation within thevoice signal. These metrics allow for the identification and characterization of objective situations that frequently occur in telephone calls and that have a direct influence on thequality of the call. To this end, in the context of the present invention, it is defined the qualityof the telephone call as those acoustic artifacts that hinder communication between the interlocutors and motivate them to end the call. First, these are acoustic artifacts that reduce intelligibility and therefore may render the call impossible. However, beyond that, even if some intelligibility remains, these are acoustic artifacts that annoy the interlocutors, reducing the comfort in the conversation and also leading to the termination of the call.In order to provide an exhaustive evaluation, the framework uses statistical methodologiesand machine learning models to analyse acoustic features, e.g., signal clarity, andparalinguistic cues, e.g., non-verbal elements within speech. The method is languageindependent, and its design accommodates multi-channel and multi-speaker telephonicinteractions, and considers both, speech and non-speech segments. This evaluationencompasses factors related to the callers and the acoustic environment in which the calltakes place, including the participants' environment and the transmission channel. Thisvoice signal analytic method facilitates the online auditory monitoring through a low latencycomputation of the metrics allowing both online and offline analysis. This feature is relatedto the fact that the information is mostly obtained from the audio data of the voice signals, instead of the transcription that is the usual source of information in methods for speechanalytics in the state of the art. Nevertheless, the voice signal analytic method allows alsooffline assessment of the telecommunication service by area of interest (e.g. geographical regions, call centers, commuters). TECHNICAL FIELD The invention belongs to the telecommunications technology sector, and more specifically to those technologies that monitor and evaluate the quality of phone calls. BACKGROUND OF THE INVENTION Methods for analysing the telephone calls in order to assess the quality of a telecommunication service are known in the prior art. The assessment carried out in priorart documents is typically based on two main ideas. Firstly, evaluating text files generatedfrom audio transcriptions of telephone calls. Secondly, assessing speech segments only.Speech segments are those parts of the telephone calls in which the speaker speaks, andnon-speech segments are those parts of the telephone calls in which no one speaks, butsome audio may still persist on the line, such as music, background noise, acoustic artifacts, and other sounds. To generate the text files, prior art methods apply a transcription process to the voice signal of the telephone call. This audio transcription captures only the conversation itself—what was said—while disregarding other valuable information present in the audio file, such as acoustic artifacts, the speaking style of the interlocutors, background noise, and more. On the other side, to be able of analysing the speech segments of the telephone call the methods of the prior art apply a diarization process to each single channel stream from the voice signal. The diarization process generates a temporal map in seconds that delineates the audio segments attributed to each speaker. The temporal map only contains the speech segments of each speaker, losing the information that can be obtained from the non-speech segments. An example of a prior art document disclosing the analysis based only on speech segmentsis the patent application with publication number US 2007 / 071206 A1 (Gainsboro Jay L, etal.) published on March 29, 2007. Another example of a prior art document disclosing the above transcription and diarization process including only speech segments, is the patent application with publication numberUS9672825B2 (Arslan et al.) published on June 6, 2017.Therefore, there is a need to provide a method and system capable of providing moreexhaustive evaluation of the telephone calls, i.e., an assessment of the entire telephone callthat includes the non-speech segments and a more in-depth evaluation of the speechsegments, such that it will be able of characterise specific situations in telephone calls thatinfluence the technical audio quality perceived by the interlocutors. The following described invention proposes the use of acoustic metrics derived from the voice signal to identify specific situations during phone calls that affect the perceived technical audio quality for the participants. Unlike the two previously cited documents, which assess the overall quality of audio and conversation, this approach focuses on characterizing distinct influencing factors within the call. DESCRIPTION OF THE INVENTION In order to overcome the above cited issues of the prior art documents, the present invention discloses a method and system for analysing voice signals on telephone calls to assess the quality of telecommunication calls not only for the entire telephone call but also in a comprehensive way for a human by providing the causes of the quality loss. The evaluationof the entire telephone call carried out by the present invention, includes the evaluation ofthe non-speech segments and a more in-depth evaluation of the speech segments.The architecture of the voice signal analytic method comprises three distinct computationaldimensions for metrics computation: 1. Analysis of conversation dynamics through the calculation of dialogue-relatedmetrics, with an emphasis on the fluidity and coherence of the dialogue at all ends of the call. 2. Evaluation of call fidelity through the calculation of acoustic quality metrics,dedicated to the objective analysis of clarity, intelligibility, and signal quality. 3. Interpretation of paralinguistic aspects of the call, by calculating metrics that identifyelements and nuances not present in the semantic content of the call but in the voices of the participants and in the broader context of the conversation. Subsequent to individualized assessments, an interpretation layer is employed to analysethe gathered metrics and affix score labels, thereby synthesizing evaluations ofperformance within the defined dimensions. Furthermore, a logical schema delineates inter- channel relationships within concurrent conversations, extending the interpretation to the full telephone call. The tallying of alarms, categorized according to established criteria reflecting application-specific performance indicators, adds quantitative insight. Optionally, a metadata-centric approach enriches the evaluation process by leveraging source- destination identifiers, thereby refining performance assessments and furnishing detailed insights into predefined areas of interest.In the first aspect of the invention, a method of voice signal analytic on telephone callsbased on acoustic quality, paralinguistic, and dialogue-related metrics is disclosed. Themethod of the present invention is applied to a multichannel voice signal corresponding toa telephone call between speakers and is language independent.Thus, the method for voice signal analytic on telephone calls of the present inventioncomprises: ^a pre-processing stage, which comprises the following pre-processing steps:o converting the voice signal to an uncompressed format signal;o dividing the uncompressed format signal into single channel streams;o generating an audio segmentation (segmentation process), for each singlechannel stream and from the uncompressed format signal, a temporal map in seconds that delineates audio segments attributed to each speaker, synthetic speech, music and waits, wherein the temporal map comprisestimestamps for non-speech segments and speech segments, bothcomprised in the audio segments; oapplying a Voice Activity Detection “VAD” to each single channel stream forgenerating speech and non-speech segments (these segments are different to those generated in the previous step by the segmentation process because they are computed with an alternative method for detecting speech / non-speech segments), having initial and end timestamps attributedto each speech segment; ^a metrics computation stage, which comprises the following computational steps:o computing dialogue-related metrics over the temporal map, wherein the dialogue-related metrics at least comprises the overall duration of the voice signal “dur”, the length of time the speaker spends talking “durspk”, theperiods of silence throughout the conversation “sil”, duration of eachspeaker’s turn “spkturn", the response times “restime” and the wait times“waits”; ocomputing acoustic quality metrics over the single channel streams, whereinthe acoustic quality metrics at least comprises:^ an average noise power “noise power” for the non-speech segments;and, for the speech segments: ^Signal to Noise Ratio “SNR”;^ Signal to Reverberation Modulation Ratio “SRMR”;^ speech level “speech level”;^ saturation “saturation”;^ tone presence “tones”,^ microcuts “microcuts”;o computing paralinguistic metrics over the single channel streams, whereinthe paralinguistic metrics at least comprises: ^intonation patterns “intonation”;^ speaking rate “speaking rate”;o computing combined metrics by combining the dialogue-related metrics, theacoustic quality metrics, the paralinguistic metrics and the single channel streams; wherein the step of computing combined metrics comprises calculating at least one of the following combined metrics: -“qindex”: ^^ ∗ ^^(^^^^^ ^^^^^ ),N, R, I, Ne, and A represent the predefined weights of the sum, indicating theirrelative importance in determining the overall acoustic quality; the function ft(x)is a non-linear transformation that normalizes the metrics “SNR”, “SRMR”, “speech level”, “saturation”, “tones” and “microcuts” to a Gaussian distribution; -“spkoverlap”: wherein “confidence”: where TVAD,iare the timestamps from the VAD output, i.e. the initial and end timestamps of the speech / non-speech segments, TSeg,iare the corresponding timestamps from the segmentation process output (temporal map), duriis the duration of the segment in seconds, N is the total number of segments, |TVAD,i -TSegD,I | is the absolute difference between the timestamps of the VAD andsegmentation for each segment. The result is subtracted from “1” to give the final“confidence”, which represents the percentage of correctly matched segments, with a higher value indicating better alignment between the VAD and segmentation process; -“artifacts”:^^^^^^^^^ = ^^^^^^^^^^ + ^^^^^ + ^^^^^^^^^;“Spkoverlap” identifies instances of overlap, indicating interruptions between speakers. To obtain this metric the minimum value of the metric “restime” is weighted with the confidenceas presented in the following. The “artifacts” metric provides an overall insight about thepresence of signal distortions.“Restime” are the minimum and maximum values of the set of segments in seconds wherea speech segment of one single channel stream is below or overlaps a speech segment ofthe other single channel stream. “Waits” are the minimum and maximum values of the setof segments in seconds where the start of the following speech segment is over certain amount of time. The present invention measures the “saturation” in a different way as it is known in the prior art. Specifically, the present invention applies a “saturation detection method” that aims to identify instances where the single channel stream reaches its maximum level, potentially causing distortion. It starts by initializing buffers for the detection process and then preprocesses the segment by subtracting the mean of the single channel stream. Subsequently, it applies various saturation criteria, including detecting local maxima, counting the number of maxima within each frame, and clustering to remove false positives. Additionally, it computes ratios between samples at high and low zones of the histogram and ratios between energy at high and low frequencies. Based on certain conditions and thresholds empirically set according to the data, the function decides whether a frame is saturated or not. The present invention measures the “microcuts” in a different way as it is known in the prior art. Specifically, the present invention computes the microcuts based on the energy histogram of the speech segments. The energy of each frame is calculated, and a mean average filter is applied to smooth the energy values. Subsequently, the method computes the difference in energy between adjacent frames and extracts significant differences based on a threshold (th) empirically set according to the data. A histogram of these significant differences is constructed with bins ranging from [0,20], or alternative values according tothe data. The metric “microcuts” is then calculated as the sum of the normalized histogramvalues within a specific range of bins following the rule: ^^^^^^^^^ = ^1 ^^ ^^^^^^^^^ > ^ℎ0 ^^ℎ^^^^^^In the context of the present invention, “intonation” is computed in a different way as it is known in the prior art. The goal of this metric is to quantify intonation by analysing the variation in the fundamental frequency, or its inverse, the fundamental period, in a speech signal. This is achieved through Pitch contour extraction. Afterwards, an interpolation for the unvoiced segments is performed to smooth the sequence. Calculation of pitchdifferences is made (how much they rise or fall over time) using its derivative. The intonationmetric is based on the histogram of the derivative of the fundamental period to quantify its stability or variation. The value of the center element of this histogram determines whether the pitch remains stable in a region or whether it has consistent increases or decreases. This metric is based on an analysis of segments of variable and selectable duration. If an aggregate metric is desired for several segments, the median of the values obtained for each segment is taken. In the context of the present, the “speaking rate” can be computed from the audio data in the voice signal. The speaking rate is computed by first filtering the speech through a series of seven sub-band specific filters. Then, band-specific envelopes are extracted using Hilbert transformation, followed by downsampling and low-pass filtering. Temporal and sub-band correlations are calculated from a subset of the filtered envelopes. The resulting correlation is normalized and peaks with a predominance greater than a threshold are identified as syllables. Finally, the syllables rate is calculated by dividing the number of detected syllables by the total speech duration. Furthermore, the present invention provides a real-time assessment of the telephone callssince it carries out on-the-fly computation of the metrics from the audio data in the voicesignal. It is “real-time” method since the provided metrics require low-latency computation requirements. The method may further comprise a scoring stage, which comprises: oscoring each single channel stream based on the combined metrics, thedialogue-related metrics, the acoustic quality metrics and the paralinguistic metrics to obtain at least two partial scores, one score per channel; oscoring the entire telephone call with a global score.In another embodiment, the scoring stage comprises to apply at least one of the followingtwo different approaches: 1. Rule-based: the scores (partials and global) are assigned based on pre-definedrules; these rules use specific thresholds (or limits) that are applied to the numerical measurements of each metric. For example, if a metric exceeds a certain threshold, aspecific score is automatically applied;2. Machine Learning (Data-Driven) techniques: the scores (partials and / or global) canalso be assigned using a machine learning model. This model learns patterns and relationships from the metric data provided to it. Based on what it learns, it predicts the appropriate score for each instance.In preferred embodiment, the pre-processing stage may further comprise the step of generating, for each single channel stream and from uncompressed format signal, a textcontaining spoken words of each speaker. For this particular case, the “speaking rate” of the speaker quantifies the number of syllables uttered per second as the following: ^^^^^^^^ ^^^^ =^^^^^^_^^^^^^^^ ^^^^^^ where the number of syllables is computed by using the text transcription aligned to thesingle channel stream and “durspk” is the length of time the speaker spends talking, i.e.,the sum of speech segments determined by the segmentation process.In preferred embodiment, the step of computing the acoustic quality metrics is repeated forall segments of the audio segments.In another embodiment, the method further comprises establishing alarms for atypical voicesignals by leveraging insights gleaned from the global score, which are the interpretation ofthe metrics. Optionally, the method further comprises generating comprehensive reports onbehaviours observed within the telephone calls.In another embodiment, the step of converting the voice signal to an uncompressed formatsignal comprises decoding the encoded telephonic line and generating a Linear Pulse Code Modulation “LPCM” uncompressed format signal. In another embodiment, the step of generating the temporal map is performed by aSegmentation engine “SE”. The Segmentation engine “SE” merges the timestamps (outputof the Diarization process) with the music, waits and the synthetic audio. In another embodiment, the step of generating the text is performed by a Speech-to-Text “STT” engine.In a second aspect of the invention, a system for voice signal analytic on telephone calls isdisclosed. The telephone call comprises a voice signal between speakers. The systemcomprises a first pre-processing module, a second metrics computation module, a third tagging module, a fourth alarming module, and a fifth reporting module that provides a finalreport of the telephone call. The system, and thus, all its modules, is configured to carry outthe method of the first aspect of the invention. BRIEF DESCRIPTION OF THE FIGURES In order to help with a better understanding of the features of the invention and to complement this description, the following figures are attached as an integral part of the same, by way of illustration and not limitation:FIG. 1 shows a block diagram of the pre-processing module which encompasses fourfunctionalities for taking the voice signal from the telephonic line, if the processing is online,or a local folder, if the processing is offline, to prepare the conditions for the next computation of metrics.FIG.1.1 shows a block diagram of the segmentation process that merges the output of thediarization process with the music, waits and the synthetic audio.FIG. 2 shows a block diagram of the metrics computation module wherein the metricsrelated to the dialogue, acoustic quality and paralinguistic aspects are computed over thevoice signal.FIG. 3 shows a block diagram of the computation of the confidence metric.FIG. 4A shows a block diagram of the rule-based scoring process where the metricscorresponding to a voice signal are interpreted according to their values and various scoresare associated to the voice signal for each channel and the entire call.FIG.4B shows a block diagram of the Machine Learning scoring process where the metricscorresponding to a voice signal are interpreted according to their values and various scores are associated to the voice signal for each channel and the entire call.FIG. 5 shows a block diagram of the process for categorizing alarms and counting thenumber of voice signals that encompass each alarm category from the whole group of calls under analysis.FIG. 6 shows the block diagram of the full system for voice signal analytic on telephonecalls. DESCRIPTION OF A PREFERRED EMBODIMENT OF THE INVENTION A. PreprocessingAn embodiment of a preprocessing module 100 of the present invention will now bedescribed with reference to FIG. 1. The pre-processing module 100 facilitates the auxiliarytools required for the subsequent computation of voice signal metrics. It encompassesconverting the voice signal to a convenient uncompressed format signal 101. Subsequently,the voice signal, that can be presented in a two-channel format, is divided into twoindependent streams for subsequent analysis 102. The following steps of this module focuson voice signal segmentation 103, generating a temporal map in seconds that delineatesthe voice signal segments attributed to each speaker along with the music, waits andsynthetic segments detection (synthetic audio is also the audio of a speaker, which insteadof being a human speaker is a synthetic speaker, but it is still a segment attributed to aspeaker). The temporal map comprises timestamps for non-speech segments and speechsegments, both comprised in the audio segments. Optionally, the preprocessing module100 may be configured to carried out the transcription 104 that furnishes the text associatedwith the voice signal, along with corresponding timestamps. Finally, the preprocessingmodule 100 may be configured to carried out the Voice Activity Detection “VAD” 105 thatoutputs a set of discrete segments 21a, 21b, with initial and end timestamps, correspondingto the speech and non-speech areas within the single channel stream 13a, 13b.A1. Signal reformat and channel separationIn this phase of the process, the encoded voice signal sourced directly from the telephonicline is the initial input 11, typically featuring a sampling rate of 8000 Hz and encoded usingformats such as G729, GSM, u-law, a-law, or similar. The encoded voice signal undergoesa decoding process specific to its encoding format, transforming it into a LPCMuncompressed form 12. Given that the voice signal may be in a two-channel format, asubsequent step involves channel separation, dividing the voice signal into two distinctsingle channel streams 13a, 13b. These streams retain the original sampling rate of 8000Hz and are in LPCM. The outcome of this stage is the decoded telephonic voice signalcomprising two separate single channel streams, each single channel stream is ready forfurther processing or analysis, while maintaining the integrity of the original voice signalcontent. Intermediate data obtained in the preprocessing stage is computed on-the-fly and there is no need for storage. This also allows to perform a real-time processing of the voice signal through a low latency computation. A2. SegmentationIn the following, the method undertakes the segmentation process 103 (FIG.1.1) carried outby the segmentation engine “SE”. The segmentation process 103 starts with the diarizationprocess 103.1 of each single channel stream 13a, 13b, that often encompass conversationsinvolving multiple speakers. The diarization process generates a temporal map in seconds that delineates the audio segments attributed to each speaker. This process is oftenperformed by a Speaker Diarization (SD) engine. The outputs of this phase encompass thetemporal map (diarization), delineated into discrete audio segments corresponding toindividual speakers. In these audio segments there is a set of timestamp pairs thatencapsulates the audio segments where a solitary speaker is active. Additionally, the audiosegments can be accompanied by speaker identification labels denoting the speaker’s identity or a distinctive identifier, which is known as Speaker Identity Attribution, and isoptionally included in the diarization process. After the diarization process 103.1, thesynthetic speech detection process 103.2, performed by a Synthetic Speech Detector(SSD) engine, determines the audio segments corresponding to answering machines or IVR (Interactive Voice Response) systems. Additionally, a music and wait detection process 103.3, performed by a Music Detector (MD) engine, is employed to define the audio segments that correspond to music and waits. These timestamps, including the syntheticspeech, the music and waits, are merged with the ones obtained from the initial diarizationprocess giving place to the final audio segmentation timestamps by overwriting the temporalmap 14a, 14b.A3. TranscriptionSubsequently, the method carries out the transcription process of each single channelstream, being this transcription process optionally. The transcription process is oftenperformed by a Speech to Text (STT) engine. The transcription procedure applies automaticspeech recognition techniques to convert the speech into text. Optionally, speaker adaptation techniques can be employed to further refine the transcription based onidentified speaker characteristics. The output of the transcription process yields texttranscriptions 15a and 15b of the two distinct single channel streams 13a and 13b, wherein each segment is transcribed into text and attributed to the respective speaker. A4. VADIn the following, the method undertakes the VAD process 105 carried out by the VAD engine“VAD”. In this invention the VAD engine might employed the LTSD (Long-Term SpectralDivergence) method, described in the following, or other related VAD that provides thetimestamps of the speech segments in the audio file and implements an alternativeapproach to the used in the SE.For implementing the VAD-based on LTSD method, the audio signal is called x(n) and it issegmented into overlapping frames, and X(k,l) is its N-th bin of the Fast Fourier Transform(NFFT) spectral amplitude for the k-th band in frame l, defining the LTSE (Long-Term Spectral Envelope) of order N: The VAD decision rule is formulated using the LTSD of order N, defined as the deviation ofthe LTSE with respect to the average spectral magnitude of the noise B(k) for the k-th band,with k = 0, 1, …, NFFT - 1. A frame l is considered to contain speech if ^^^^^ (^) is greater than a predefined threshold^ ^^^^^ (^) > ^Finally, the method provides the timestamps 21a, 21b with the start and end time in secondsof the speech segments in the audio signal. B. Metrics computationAn embodiment of a metrics computation module 200 of the present invention will now bedescribed with reference to FIG.2. The process of computing metrics encompasses variousstages related to type of information: dialogue-related metrics 201, metrics of the acousticquality 202, and paralinguistic metrics 203. The inputs of this embodiment for metricscomputation are the single channel streams 13a and 13b, the corresponding temporal map(segmentation) 14a and 14b with the timestamps of the speech segments of individualspeakers, speech, music and waits; and, optionally the corresponding text transcription 15aand 15b. Initially, metrics are computed over the single channel stream capturing variousmeasurable attributes. The resulting dialogue-related metrics 16a-16b, quality metrics 17a-17b, and paralinguistic metrics 18a-18b, may include amplitude, frequency, and durationparameters, and other characteristics that provide insights into the properties of thewaveform and are objectively defined and computed. Subsequently, these metrics (16a,17a, 18a and 16b, 17b, 18b) are post-processed by aggregation, manipulation, or analysisof the individual metrics to generate new combined metrics 19a and 19b that offer a morecomprehensive understanding of the overall quality of the telephone call 2. By leveragingthe information gleaned from the dialogue-related metrics 16a-16b, quality metrics 17a-17b,and paralinguistic metrics 18a-18b, the computation of combined metrics 19a and 19bfacilitates deeper analysis and interpretation, enabling more informed decision-making invarious applications. The quality metrics 17a, 17b are obtained every one or more seconds,along with a general record of the entire single channel stream, typically averaging orsumming the records obtained in these intervals. A single value for the entire single channelstream is obtained for dialogue-related metrics 16a-16b and paralinguistic metrics 18a-18b.Then, a selection of the metrics 16a, 17a, 18a and 16b, 17b, 18b are used for computingthe combined metrics 19a, 19b, which are also a single value by metric corresponding tothe entire single channel streams 13a, 13b. The computation of the objective metrics 16a,17a, 18a, 19a and 16b, 17b, 18b 19b are detailed described in the following paragraphs.Dialogue-related metrics computation 201 provide a comprehensive understanding of theflow and interaction within the telephone call 2. They shed light on how the telephone call 2progresses and how speakers engage with each other and enable to identify patterns and areas for improvement in communication dynamics. The dialogue-related metrics arecomputed over the two distinct single channel streams 13a and 13b and use thecorresponding temporal map (segmentation) 14a, 14b as guide of the segments withspeech and non-speech. The dialogue-related metrics comprise, but are not tied to: -The overall duration of the single channel stream “dur”,- The length of time the speaker spends talking “durspk”,- The periods of silence throughout the conversation “sil”,- Response times “restime”,- Wait times “wait”,- Duration of each speaker's turn “spkturn”.The resulting output of this computation is a set of dialogue-related metrics 16a, 16b.Note that, this set of metrics 16a, 16b are computed over the output of the audiosegmentation, i.e., the timestamps in seconds, allowing to obtain the dialogue-related metrics with a very low latency respect to the result that would be obtain if the transcription was the source. To compute the mentioned metrics, the timestamps in seconds with the discrete segmentswith speech and non-speech are employed. First, “dur” is obtained as the full number ofseconds of the single channel stream (13a or 13b). Then “durspk” is the sum of speechsegments determined by the segmentation process, while “sil” is the sum of non-speechsegments determined by the corresponding segmentation 14a or 14b.“Restime” are the minimum and maximum values of the set of segments in seconds wherea speech segment of one single channel stream (e.g. 13a) is below or overlaps a speechsegment of the other single channel stream (e.g. 13b). “Waits” are the minimum andmaximum values of the set of segments in seconds where the start of the following speechsegment is over certain amount of time (e.g. more than 5 seconds). “Spkturn” are theminimum and maximum values of the set of speech segments of the single channel stream(13a or 13b) in seconds.Acoustic quality metrics computation 202 assess the technical aspects of a telephone call2, offering insights into the fidelity and clarity. They are computed over the single channelstreams 13a and 13b. The acoustic quality metrics comprise, but are not tied to:- Signal to Noise Ratio “SNR”,- Signal to Reverberation Modulation Ratio “SRMR”,- Speech level “speech level”,- Saturation “saturation”,- Tone presence “tones”,- Microcuts “microcuts”, all previous quality metrics are computed in speech segments. -Average noise power “noise power”,computed in non-speech segments.This set of metrics allows to characterize the typical acoustic artifacts in a telephoneconversation beyond the standard environmental noise. It includes artifacts found in both speech and non-speech segments. Additionally, since the metrics are based on the voicesignal, they can be computed in real-time without having to wait for the entire conversationto end.The definition of speech and non-speech segments are taken from the temporal map(segmentation process). The resulting output of this computation is a set of quality metrics17a, 17b.The step of computing the acoustic quality metrics 17a, 17b is repeated 205 for all segmentsof the audio segments.First, SNR and SRMR collectively indicate the level of distortion and acoustic quality in thesingle channel stream (13a or 13b). As previously mentioned, the temporal map(segmentation process) 14a and 14b are employed for establishing the audio segmentswhere compute the power of speech (i.e. those corresponding to speech in thesegmentation), and the power of noise (i.e. those corresponding to non-speech in thesegmentation). Then, the value of SNR is computed as follows: where ^^^^^^^^is the power of speech segments and ^^^^^^^ the power of non-speechsegments computed as ^ =^ ^∑^^^^^^ |^(^)|^ with N the number of samples of thecorresponding single channel stream (13a or 13b) segment, and X(n) the samples.The SRMR measures the quality of a single channel stream by quantifying the balancebetween the speech segments and its reverberant components. First, the speech segmentsof the single channel stream (13a or 13b) are divided into short-time frames and thereverberant components are estimated for each frame using a model-based approach, such as a room impulse response estimation or blind source separation. The modulation spectraof both, the original speech segments and the reverberant components, are then extracted using a filterbank or a similar method. Finally, the metric is calculated as the ratio of the modulation energy of the speech segments to the modulation energy of the reverberant components: where ^ is the average per modulation band energy. A higher value indicates a strongerpresence of speech relative to reverberation, thus reflecting better quality with reduced reverberation effects.Speech level represents the volume of the speech segments referred to full scale known asdBFS (dB Full-Scale), where 0 dBFS represents the root mean square (RMS) level of a fullscale sinusoidal. Working with 16 bits quantification, the full-scale value is 2^^. This iscomputed as follows: where ^^^^^^^^is the power of speech segments previously defined in SNR metric, and 2^^ is the power of a full-scale sinusoid. Acoustic artifacts may impact the quality of the single channel stream, including saturation, tones, and microcuts. These artifacts have the potential to affect the overall listening experience by introducing distortion, interference, or interruptions. Detection of these artifacts enables proactive measures to mitigate their effects and enhance the quality of telephone calls.Saturation detection method aims to identify instances where the single channel streamreaches its maximum level, potentially causing distortion. It starts by initializing buffers forthe detection process and then preprocesses the segment by subtracting the mean of thesingle channel stream. Subsequently, it applies various saturation criteria, including detecting local maxima, counting the number of maxima within each frame, and clustering to remove false positives. Additionally, it computes ratios between samples at high and low zones of the histogram and ratios between energy at high and low frequencies. Based oncertain conditions and thresholds empirically set according to the data, the function decideswhether a frame is saturated or not.Tones are obtained through a Linear Predictive Coding (LPC) analysis. It begins by emphasizing high-frequency components to distinguish tones from other typical acousticfeatures in speech, such as the fundamental frequency F0. Thanks to this analysis, the LPCcoefficients are obtained and they allow the representation of the spectral envelope of thesingle channel stream. The LPC coefficients are used to derive the prediction filter in the Z-domain and this filter helps in predicting future samples of the signal based on past samples. The denominator of this prediction filter in the Z-domain is what we call LPC polynomial, derived from the LPC coefficients, and its roots provide valuable insights into the signal's spectral properties. The roots of the LPC polynomial allow to detect the presence of narrow-band components (tones) based on their module and bandwidths. If a root meets certaincriteria empirically set according to the data, it is potentially a tone. Then, the method checksif these tones correspond to standard Dual Tone Multi-Frequency (DTMF) values (697, 770,852, 941, 1209, 1336, 1477, 1633 Hz) within a small tolerance. Detected DTMF tones aremarked, and their conjugate roots are also labelled accordingly. Finally, the method aggregates the detection results to generate binary flags indicating the presence of DTMF tones and other tones.Microcuts are extremely short segments, in the order of milliseconds, where the voice signalis interrupted or lost. This issue may correspond to particular acoustic distortions observedin recorded telephone data, similar to packet loss or telephone service coverage loss, thatmight occur during the telephone call 2 is in progress or when it is recorded. Microcutsmetric is computed based on the energy histogram of the speech segments. The energy of each frame is calculated, and a mean average filter is applied to smooth the energy values. Subsequently, the method computes the difference in energy between adjacent frames andextracts significant differences based on a threshold (th) empirically set according to thedata. A histogram of these significant differences is constructed with bins ranging from[0,20], or alternative values according to the data. The metric microcuts is then calculatedas the sum of the normalized histogram values within a specific range of bins following the rule: ^^^^^^^^^ = ^1 ^^ ^^^^^^^^^ > ^ℎ0 ^^ℎ^^^^^^Furthermore, the analysis of the average power of non-speech segments, called averagenoise power “noise power”, helps identify the persistent noise floor. This metric providesvaluable information about the ambient noise throughout the telephone call 2, and iscomputed as follows:where ^^^^^^^(^) is the power of the m-th non-speech segment and M the number of non-speech segments detected.Paralinguistic metrics computation 203 offer valuable insights into the non-verbal aspects.These metrics capture subtle nuances of the prosody at speaking, that contribute to theoverall expression and communication style. They are computed over the single channelstream 13a and 13b. Paralinguistic metrics include, but are not tied to:- Intonation patterns “intonation”,- Speaking rate “speaking rate”.The resulting output of this metrics computation is a set of paralinguistic metrics 18a, 18b.As with the previously discussed metrics (16a, 16b, 17a, 17b) the paralinguistic metrics18a, 18b are derived from the audio data within the voice signal, which renders themappropriate for real-time application due to their minimal latency in computation. In the present invention, this specific set of paralinguistic metrics was selected due to their abilityto be consistently extracted from the audio, even under conditions of low quality.Furthermore, these metrics are proficient in identifying alterations in speech patterns, particularly those indicative of discomfort or irritation resulting from a decline in call quality. In the context of the present invention, “intonation” is computed in a different way as it is known in the prior art. It contains an evolution of the fundamental frequency, by providing an overview of vocal inflections and variations of intonation throughout the whole conversation. Initially, it calculates the fundamental period (T0) of the single channel streams, which is the inverse of the fundamental frequency (F0). Then, it identifies unvoicedsegments —intervals without F0— and uses them as anchors to interpolate the pitchcontour, thereby smoothing the overall trajectory. After extracting this pitch contour, the method computes pitch differences over time by calculating the derivative of T0, denoted as ΔT0. A histogram of these differences is then created using a specified number of bins and range, forming the basis of a metric that quantifies the stability or variation in the pitch— what could be described as pitch flatness. Optionally, the signal can be segmented into tokens of variable and selectable duration. For each segment, the metric evaluates the central value of the histogram: values around zero indicate stable pitch; consistently positive values indicate rising pitch; and negative values reflect falling pitch. When an aggregated intonation value is needed across segments, the median of the individual segment metrics is computed. This final value serves as a descriptor of pitch variation: higher values suggest lower variation and more monotone speech, whereas lower values imply a wider and more dynamic range of intonation.“Speaking rate” of the speaker quantifies the number of syllables uttered per second as thefollowing: ^^^^^^^^ ^^^^ =^^^^^^_^^^^^^^^ , ^^^^^^ where the number of syllables is computed by using the text transcription aligned to thesingle channel stream or by using the audio in the voice signal itself. In the latter case,syllables are detected by counting the peaks in the normalized cross-band correlation obtained from temporal band-specific envelope correlations that exceed a defined threshold. To compute the speaking rate from the audio, the voice signal is first passed through a series of seven specific filters. Then, band-specific envelopes are extracted using Hilbert transformation, followed by downsampling and low-pass filtering. Temporal and sub- band correlations are calculated from a subset of the filtered envelopes. The resulting correlation is normalized and peaks with a predominance greater than a threshold are identified as syllables. Finally, the syllables rate is calculated by dividing the number of detected syllables by the total speech duration.In this exemplary embodiment, the “speaking rate” is computed from the audio data in thevoice signal. The speaking rate is computed by first filtering the speech through a series of seven sub-band specific filters, obtaining seven different signals ^^(^). Then the Hilbert transform for each one of them is obtained: ^^(^) = ℋ(^^(^))where i=1,2,…,7 is the sub-band and n is the temporal index. After that, band-specific envelopes are estimated by means of their analytic signals obtained using the aforementioned Hilbert transforms. These band-specific envelopes are then downsampled after low-pass filtering. Temporal and cross-band correlations are calculated from a subset of the filtered band-specific envelopes.where ^(^) is a finite-length window function such as Hamming, Hann, or rectangular. Theresulting cross-band correlation is normalized and peaks with a predominance greater than a threshold are counted as syllables. The syllable rate is then estimated by dividing the number of detected peaks by the total speech duration. This metric provides insights into the pace and fluency of speech. In this case, lower values than 3-4 syllables per second indicate slow speech, while higher values indicate fast speech. By analysing these paralinguistic metrics, analysts can gain a deeper understanding of the emotional tone, emphasis, and engagement levels in spoken communication.Combined metrics computation 204 provide additional insights at generating more metrics19a, 19b from the previously computed metrics 16a,16b,17a,17b,18a,18b, and the sideinformation provided by the segmentation 14a, 14b. Essentially, they are analysing theentire single channel stream. These combined metrics are typically represented as binary(indicating presence or absence) or as probabilities (ranging from 0 to 1). Combined metrics comprise, but are not tied to: -Quality index “qindex”,- Speaker overlap index “spkoverlap”,- Acoustic artifacts “artifacts”.The resulting output of this computation is a set of combined metrics 19a, 19b. Thecombined metrics 19a, 19b encompass the acoustic features of the telephone call, optimized by the non-acoustic metric, Speech Mask Confidence, referred to as "confidence". Confidence, reflects the reliability of the segmentation established by the accuracy of the segmentation process. This is crucial for assessing the reliability of obtained metrics(combined metrics, dialogue-related metrics, quality metrics, paralinguistic metrics),especially those that depend on the segmentation in speech and non-speech segmentssuch as the acoustic quality metrics. To compute this metric (see FIG.3) the single channelstreams 13a, 13b are processed with a Voice Activity Detection (VAD) method 105, usuallyperformed by a VAD engine, that outputs a set of discrete segments 21a, 21b, with initialand end timestamps, corresponding to the speech and non-speech areas within the singlechannel stream. The set of timestamps from the VAD (21a, 21b) and from the temporal map(14a, 14b) from the segmentation process is considered as two different experts. Thus, theyare compared 206 among them and the Speech Mask Confidence “confidence” 19.1a,19.1b is determined by the percentage of time in seconds they differ with respect of “dur”,considering a permissible window of error in seconds that can be adjusted according to theapplication allowing for flexibility in determining how much difference is acceptable: where TVAD,i are the timestamps from the VAD output, i.e. the start and end of the speech / non-speech segments, TSeg,i are the corresponding timestamps from thesegmentation process output, duri is the duration of the segment in seconds, N is the totalnumber of segments, |TVAD,i - TSegD,I | is the absolute difference between the timestamps ofthe VAD and segmentation for each segment. The result is subtracted from 1 to give the final “confidence”, which represents the percentage of correctly matched segments, with a higher value indicating better alignment between the VAD and segmentation processes.“Qindex” is a composite metric used to assess the overall technical quality of the singlechannel stream by drawing upon several quality factors expressed by the following metricsin 17a, 17b: SNR, SRMR, speech level, saturation, tones, microcuts, and noise power. Tocompute the qindex metric, first the quality metrics 17a, 17b are transformed into theirprobabilistic representations and then they are combined by means of a weighted sum: ^^ ∗ ^^(^^^^^ ^^^^^ ),where N, R, I, Ne, and A represent the predefined weights of the sum, indicating theirrelative importance in determining the overall quality. The function ft(x) is a non-lineartransformation that normalizes the metrics 17a, 17b to a Gaussian distribution. It builds atransformation model using a set of single channel streams as training data and applies it to both training and testing data. The Gaussianization is obtained by mapping data quantiles to a standard normal distribution. The function outputs the transformed testing data and their cumulative distribution function (CDF), streamlining the process of preparing the 17a,17b metrics for mixing among them. Thus, “qindex” provides a holistic evaluation of thequality by considering various aspects that can affect the listener's experience, including clarity, presence of distortions or artifacts, and overall fidelity. It offers a convenient way toquantify and compare the quality of different single channel streams.“Spkoverlap” identifies instances of overlap, indicating interruptions between speakers. Toobtain this metric the minimum value of the metric “restime” is weighted with the confidenceas presented in the following: Finally, the “artifacts” metric provides an overall insight about the presence of signaldistortions. This is computed as:^^^^^^^^^ = ^^^^^^^^^^ + ^^^^^ + ^^^^^^^^^C. ScoringDuring the scoring stage 300, single channel streams 13a, 13b undergo scoring, initially bychannel 301 and call 302 as presented in FIGs.4A, 4B. The scoring stage aims to providean objective interpretation of the voice signal quality from the information on the acousticmetrics. Scores can be computed using two different approaches:Rule-Based 31: scores can be assigned based on pre-defined rules initially bychannel 31.1 and subsequently by call 31.2, as presented in FIG.4A. These rules usespecific thresholds (or limits) that are applied to the numerical measurements ("metrics"). For example, if a metric (like silence duration) exceeds a certain threshold, aspecific score is automatically applied.Machine Learning (Data-Driven) 32: scores can also be assigned using a machinelearning model, either by channel 32.1 than by call 32.2 according to the inputs, aspresented in FIG.4B. This model learns patterns and relationships from the metric data provided to it. Based on what it learns, it predicts the appropriate scoring label for each instance.Both methods (31,32) might be used together, namely they are not mutually exclusive, tointerpret the information provided by the metrics. However, it is important to note that whilerule-based scoring is easily understood and adjusted by experts, global scores calculatedby machine learning involve a level of complexity that cannot be easily deciphered by humans without the model's support. The development of a robust machine learning modulecapable of accurately inferring scores from the underlying metrics requires significantexpertise and careful consideration of model architecture, feature engineering, and validation strategies. C1. Scoring based on rulesScoring rules are established based on reasonable thresholds of metric values, guided bythe metric definition and its manifestation in the data. Initially, this process occurs at thechannel level, namely a partial score 20a from the metrics 16a, 17a, 18a, 19a correspondsto the single channel stream 13a and another partial score 20b from the metrics 16b, 17b,18b, 19b corresponds to the single channel stream 13b. Then, both partial scores 20a, 20bare merged to create a score at the call level, namely a global score 20 for the wholetelephone call 2.At the channel level (31.1), scoring labels are:- "General" is activated when a metric value exceeds or falls below the mean by morethan two standard deviations. The mean and standard deviation values are computed for each metric 16a, 17a, 18a, 19a, 16b, 17b, 18b, 19b considering the full group ofsingle channel streams under analysis.- "Empty" denotes instances where the content is negligible or “durspk” is less than the5% of dur.- "Noisy" indicates poor SNR (< 5dB).- "Distortion" indicates significant number of acoustic artifacts, namely when “artifacts”exceeds the 10% of time of “dur”.- "Quality" indicates low “qindex” < 50.- "Background noise" indicates high noise pattern in the non-speech segmentsdetermined by the “noise power” > -10dB.- “Interruptions” when the “spkoverlap” > 4 seconds.- "Emotions" when “intonation” < 40.- “Slow” when “speaking rate” < 2 syllables per seconds.- “Fast” when “speaking rate” > 8 syllables per seconds.At the call level (31.2), further scoring occurs to encompass broader call-relatedphenomena, these comprise, but are not tied to:- “General call” when one or both single channel streams have the scoring label “General”activated,- "Unanswered call" is applied when one single channel stream remains empty, i.e., hasthe “Empty” scoring label, while the other does not.- "Empty call" when both single channel streams lack substantial content, i.e., both havethe “Empty” scoring label.- "Noisy call" scoring label is indicative of overall poor quality throughout the call duration,i.e., when both single channel streams have activated one or more of the followingscore labels: “Noisy”, “Distortion”, “Quality”, or “Background noise”.- "Interrupted conversation" scoring label when one or both single channel streams haveactivated the scoring label “Interruptions”- “Monotone conversation” when “Slow” scoring label is activated in one or both singlechannel streams, and there is not activated the “Emotions” scoring label. “Emotional conversation” when “Emotions” scoring label is activated in one or bothsingle channel streams.C2. Scoring based on modelIn addition to the previously described scoring rules, the method calculates partial scoresfor each channel 20.1a and 20.1b (Fig.4B), using the corresponding metrics 16a, 16b, 17a,17b, 18a, 18b, 19a, and 19b. Additionally, it generates a global score 20.1 that reflects theoverall quality of the ongoing call based on the full set of acoustic metrics 16a, 16b, 17a, 17b, 18a, 18b, 19a, and 19b. These metrics are combined to generate an indicator that compares the ongoing call with a set of pre-selected calls, which are stored in a database and are labelled according to the quality of the telecommunication service. Furthermore, other relevant data—such as sales records, product characteristics, and client profiles— can be incorporated to complement the acoustic metrics derived from the voice signals.This data is processed by a task-oriented machine learning “ML” model 32 to assess thequality of the telecommunication service. The model establishes connections to KPIs that are closely aligned with operations and the expected return from the task. The partial scores20.1a and 20.1b, and the global score 20.1 are estimated in a supervised (with scoringlabels) manner, with a strong emphasis on operational performance. The scoring labels are objectively defined based on the quality of calls delivered by the telecommunications service, as evaluated by industry experts. For instance, in a call center environment, calls may be categorized as low, medium, or high quality. The ML model consists of a Deep Neural Network (DNN). The architecture of the model may be based on a Convolutional Neural Networks (CNN), Residual Neural Networks(ResNets), or other proper models like transformers. The neural networks take multipleinputs, consisting of the acoustic metrics 16a, 16b, 17a, 17b, 18a, 18b, 19a, 19b and otherrelevant business data, like sales records, product characteristics, client profiles, amongothers. The output of each neural network is the partials scores 20.1a, 20.1b or the globalscore 20.1, since each neural network is specifically trained for the partial scores or globalscore, respectively. The DNN of the present invention is trained using a large dataset of voice signals (hundreds of hours of speech) and their corresponding scoring labels, which reflect the quality of the calls according to the telecommunication service. The model istrained with a cross-entropy loss function during the training phase, defined as: where ^^represents the predicted probability distribution of the model for the class i, ^^is the true distribution, and the sum is over all possible classes or categories.In this context, the partial scores 20.1a, 20.1b, and the global score 20.1 can be either areal scalar or vector. In the case of scalar, the score allows for ranking calls based on their quality. In the case of a vector, it enables grouping calls based on their global quality, with both approaches used to evaluate the quality of the call.The fact of obtaining the partial scores 20.1a, 20.1b, and the global score 20.1 from acousticmetrics such as the description of quality, the call flow, the paralinguistic description of the call, etc., not only allows for obtaining a better global metric but also enables understanding the interpretability of that value. This provides the ability to explain or comprehend why a call was considered that impact more or less in the telecommunication service.C2. Rule-based Scoring vs. Machine Learning-Driven Global ScoringRule-based scoring labels 20 are accessible to experts because they are derived frompredefined thresholds and criteria applied to specific metrics, allowing for easy understanding and direct adjustment. Experts can interpret and modify the rules easily, as they operate on a fixed structure of data with limited variability. In contrast, the global score20.1 calculated through a machine learning model, such as a Deep Neural Network,combines multiple metrics from dialogue, quality, and paralinguistic dimensions in an abstract and nonlinear way, making the result of the interactions between these dimensions difficult for a human to obtain. The model processes and adapts the data dynamically, identifying complex patterns that an expert could not detect using only rules, as this type of analysis requires a level of abstraction and calculation beyond what a human can perform directly. D. Set alarmsFollowing the scoring process, and optionally, the system proceeds to establish alarms foratypical voice signals (or telephone calls), see FIG. 5. The alarm stage 400 is dedicated todelineating the KPIs by leveraging insights gleaned from the scoring labels 20 and 20.1,which are the interpretation of the metrics. Moreover, metadata furnished by the clientpertaining to each telephone call 2 is incorporated. These alarms are classified into fourfundamental categories: general, agent protocol, technical anomalies, and campaign objectives. So, the output of this process is a list of the alarm categories with the corresponding number of voice signals that have those alarms 22.Noteworthy examples of them comprise but are not tied to:^ Global metric alarms according to the label “global score” corresponding to the qualityof the telecommunication service in general; ●General metric alarms according to the score label “General”;● Agent protocol:○ Interruptions;○ Speech mannerisms (monotone or fast speech);● Technical issues:○ Empty outgoing / incoming calls;○ Technical call quality;○ Distortions and noises;● Campaign objectives:○ Satisfaction;○ Emotional activity;E. Reporting Finally, and optionally, the reports are aggregated and presented in an interactivedashboard. The reporting phase 500 involves generating comprehensive reports on thebehaviours observed within client-specific areas of interest, including call centers, switches,recorders, campaigns, and agents. Inputs for this process include the metadata related tothe telephone calls, along with the metrics 16a, 17a, 18a, 19a, 16b, 17b, 18b, 19b, scoringlabels 20a, 20b, 20, 20.1a, 20.1b and 20.1 and alarms 22 derived from the analysis of thevoice signals. The reports are meticulously crafted to outline behaviours and facilitatecomparisons among different elements, enabling a clear depiction of service functionality and the identification of any potential issues. The reports can be structured into distinct categories, for instance: ●Global service metrics: This category provides a comprehensive summary of themetrics from both agent and client channels. The summary is based on the quantity of alarms detected for each type, including dialogue-related metrics, acoustic quality metrics, and paralinguistic metrics.● Comparison between centers and switches: Reports in this category offer acomparative analysis of alarms quantified across different call centers and switches.Detailed breakdowns of alarm types are provided to facilitate a deeper understanding of performance variations. ●Recorder-based quality analysis: Focusing specifically on acoustic quality-relatedalarms, this category evaluates the technical quality of recorded voice signals. Given the recorder's influence in this aspect, the analysis aims to identify and address any issues impacting voice signal (or telephone call) fidelity.● Agent-specific comparison: Metrics related to agent protocol, encompassingparalinguistic and dialogue-related aspects, are examined in this category. By analysing agent-specific data, insights into individual performance and adherence to protocol standards are gained. ●Temporal evolution analysis: This category involves a longitudinal study of alarmquantities over varying time periods, including days, weeks, and months. The analysis spans across centers, switches, agents, and other relevant factors, providing insights into trends and patterns over time. F. System for voice signal analytic on telephone callsThe whole system 1 for implementing the analysis of a given telephone call 2 or a collectionof telephone calls, can be configured according to the use-case requirements and goalsand is delineated in FIG. 6. The system 1 of the present invention comprises a first pre-processing module 100, a second metrics computation module 200, a third scoring module300, a fourth alarming module 400 (optionally), and finally a fifth reporting module 500(optionally) that provides a final report of the telecommunication service.For a group of telephone calls, or a single telephone call 2, the first pre-processing module100, described in FIG. 1, proceeds to reformat 101 and obtains the single channel stream102. If the system operates offline the voice signals are in a local folder, but if the systemoperates online the voice signals comes directly from the telephone line. Additionally, thesegmentation 103, VAD engine 105 and transcription 104 provide useful information of thesingle channel streams under analysis. The second metrics computation module 200,described in FIG.2, computes the metrics of the single channel streams in the dimensionsof the dialogue-related aspects 201, the acoustic quality 202, and paralinguistic aspects203 of the telephone conversation, and finally a combined metrics dimension 204 computesa set of metrics based on the previous. The third scoring module 300, described in FIG. 4,interprets the metrics performance for each single channel stream included in theconversation 301, and then links up the results of the entire telephone call 302. The fourthalarming module 400, described in FIG. 5, count and categorizes the number of voicesignals with alarms 401. Finally, the fifth reporting module 500 is optional and organizes theanalysis by region / area according to a metadata provided by the telecommunication serviceprovider in charge of telephone calls. It creates reports of results that include the results ofthe quality performance by region, call center, recorder, commutator, etc.
Claims
CLAIMS1.- A method for voice signal analytic on telephone calls, wherein a telephone call comprisesa voice signal between speakers; the method comprises:^ a pre-processing stage (100), which comprises the following pre-processing steps:o converting (101) the voice signal (11) to an uncompressed format signal (12);o dividing (102) the uncompressed format signal (12) into single channelstreams (13a,13b);o generating an audio segmentation (103), for each single channel stream(13a, 13b) and from the uncompressed format signal, a temporal map (14a,14b) in seconds that delineates audio segments attributed to each speaker,synthetic speech, music and waits, wherein the temporal map comprisestimestamps for non-speech segments and speech segments, both comprised in the audio segments; oapplying a Voice Activity Detection “VAD” (105) to each single channelstream (13a, 13b) for generating speech and non-speech segments (21a,21b), having initial and end timestamps attributed to each speech segment;^ a metrics computation stage (200), which comprises the following computationalsteps: ocomputing (201) dialogue-related metrics (16a, 16b) over the temporal map(14a, 14b), wherein the dialogue-related metrics at least comprises anoverall duration of the voice signal “dur”, a length of time the speaker spendstalking “durspk”, periods of silence throughout the conversation “sil”, durationof each speaker’s turn “spkturn", response times “restime” and wait times “waits”; ocomputing (202) acoustic quality metrics (17a, 17b) over the single channelstreams (13a, 13b), wherein the acoustic quality metrics at least comprises:^ an average noise power “noise power” for the non-speech segments;and, for the speech segments: ^Signal to Noise Ratio “SNR”;^ Signal to Reverberation Modulation Ratio “SRMR”;^ speech level “speech level”;^ saturation “saturation”;^ tone presence “tones”,^ microcuts “microcuts”;o computing (203) paralinguistic metrics (18a, 18b) over the single channelstreams (13a, 13b), wherein the paralinguistic metrics at least comprises: ^intonation patterns “intonation”;^ speaking rate “speaking rate”;o computing (204) combined metrics (19a, 19b) by combining the dialogue-related metrics (16a, 16b), the acoustic quality metrics (17a, 17b), theparalinguistic metrics (18a, 18b) and the single channel streams (13a, 13b); wherein the step of computing combined metrics comprises calculating at least one of thefollowing combined metrics (19a, 19b):- “qindex”:^^ ∗ ^^(^^^^^ ^^^^^ ),N, R, I, Ne, and A represent the predefined weights of the sum, indicating theirrelative importance in determining the overall quality. The function ft(x) is a non- linear transformation that normalizes the metrics “SNR”, “SRMR”, “speech level”, “saturation”, “tones” and “microcuts” to a Gaussian distribution; -“spkoverlap”:whereinwhere TVAD,iare the initial and end timestamps from the VAD, TSeg,iare the initial and end timestamps from the temporal map, duriis the duration of the segment in seconds, N is the total number of segments;- “artifacts”:^^^^^^^^^ = ^^^^^^^^^^ + ^^^^^ + ^^^^^^^^^.2.- The method for voice signal analytic on telephone calls of claim 1, wherein the methodfurther comprises a scoring stage (300), which comprises:o scoring (301) each single channel stream (13a, 13b) based on the combinedmetrics (19a, 19b), the dialogue-related metrics (16a, 16b), the acousticquality metrics (17a, 17b) and the paralinguistic metrics (18a, 18b) to obtainat least two partial scores (20a, 20b, 20.1a, 20.1b), one score per channel;o scoring (302) the telephone call (2) with a global score (20, 20.1).3.- The method for voice signal analytic on telephone calls of claim 2, the scoring stage(300) comprises to apply at least one of the following two different approaches:^ Rule-based (31): the scores, partials (20a,20b) and global (20), are assigned basedon pre-defined rules that use specific thresholds applied to each metric; ^Machine Learning techniques (32): the scores, partials (20.1a, 20.1b) and global(20.1), are assigned using a machine learning model.4.- The method for voice signal analytic on telephone calls of claim 2, wherein the pre-processing stage (100) further comprises generating (104), for each single channel streamand from uncompressed format signal, a text (15a,15b) containing spoken words of each speaker.5.- The method for voice signal analytic on telephone calls of claim 1, the step of computingthe acoustic quality metrics (17a, 17b) is repeated (205) for all segments of the audiosegments.6.- The method for voice signal analytic on telephone calls of claim 2 or 3, wherein themethod further comprises establishing (400) alarms for atypical voice signals by leveraginginsights gleaned from the global score (20, 20.1).7.- The method for voice signal analytic on telephone calls of claim6, wherein the methodfurther comprises generating (500) comprehensive reports on behaviours observed withinthe telephone calls (2).8.- The method for voice signal analytic on telephone calls of claim 1, wherein the step ofconverting (101) the voice signal to an uncompressed format signal comprises decodingthe encoded telephonic line (11) and generating a Linear Pulse Code Modulation “LPCM”uncompressed format signal (12).9.- The method for voice signal analytic on telephone calls of claim 1, wherein the step ofgenerating (103) the temporal map is performed by a Segmentation Engine “SE” .10.- The method for voice signal analytic on telephone calls of claim 4, wherein the step ofgenerating (104) the text is performed by a Speech-to-Text “STT” engine.11.- A system (1) for voice signal analytic on telephone calls, wherein a telephone call (2)comprises a voice signal between speakers; the system comprises a first pre-processingmodule (100), a second metrics computation module (200), a third scoring module (300), afourth alarming module (400), and a fifth reporting module (500) that provides a final reportof the telephone call; the system is configured to carry out the method of claims 1 to 10.
Citation Information
Patent Citations
Multi-party conversation analyzer & logger
US20070071206A1
Real-time speaker state analytics platform
US20170084295A1
Method and system for conversation transcription with metadata
US20220115019A1
Speech analytics system and methodology with accurate statistics
US9672825B2