Procedure for the recognition of answering machines in call center activities'
The pattern matching method for voicemail detection in call centers enhances recognition accuracy and efficiency by comparing audio features with pre-recorded messages, addressing latency and compliance issues in call centers.
Patent Information
- Application Number
- PCT/IB2025/056956
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-09
- Filing Date
- 2025-07-09
- Publication Date
- 2026-01-15
AI Technical Summary
Current voicemail detection systems in call centers suffer from high latency, low recognition rates, and false positives/negatives, leading to inefficient call handling and regulatory compliance issues, especially in multilingual contexts.
A pattern matching method that transforms audio signals into features, compares them with a database of pre-recorded messages, and identifies answering machines by assigning scores to correspondences, allowing real-time recognition.
Achieves a recognition rate of over 99.9% with near-instant response times, eliminating false positives and optimizing call management by minimizing unproductive calls and adhering to regulatory standards.
Smart Images

Figure IB2025056956_15012026_PF_FP_ABST
Abstract
Description
[0001] TITLE:
[0002] "PROCEDURE FOR THE RECOGNITION OF ANSWERING MACHINES IN CALL
[0003] CENTER ACTIVITIES''
[0004] * * * * *
[0005] TECHNICAL FIELD
[0006] The present invention relates to a method for the recognition of answering machines in call center activities.
[0007] PRIOR ART
[0008] Voicemail detection technology, commonly known as Answering Machine Detection (AMD), tries to identify when a phone call is answered from a voicemail. Recognition allows you to avoid passing the call unnecessarily to an operator and to quickly clear the busy phone line. This technology is used in call centers to optimise performance in telemarketing and outbound campaigns in general.
[0009] Recently, there has been a transformation in the way answering machines are used, especially in the personal sphere. Currently, in fact, most of the systems with dedicated devices, connected to the telephone line, have been replaced by centralised secretarial services that can be used by all users. In Italy, telemarketers have found that, for calls to mobile phones, more than 30% of calls are forwarded to the user's centralised answering machine. With these numbers, it is clear that identifying these calls brings significant productivity benefits.
[0010] The voicemail recognition systems currently on the market use a combination of heuristic techniques to distinguish a person's response from that of a voicemail.
[0011] The main techniques used include the following.
[0012] Speech analysis
[0013] This technique analyzes the audio of the early stages of the call trying to identify sound patterns associated with the human voice. This includes detecting specific strings of words and analyzing the intonation and rhythm patterns of the voice. It is clear that this technique remains linked to the language on which it was developed, usually English, so it is quite expensive to adapt to multilingual contexts. This technique requires specific calibrations to obtain the best compromise between recognition rate and latency of the algorithm.
[0014] Background noise
[0015] When voice is not present, answering machines often have constant background noise or silence. Algorithms implementing this technique can detect this feature and use it to signal a possible voicemail.
[0016] Pause detection
[0017] During audio analysis, algorithms monitor pauses between words: answering machines, in fact, tend to have longer and more regular breaks than a conversation with a real person.
[0018] Response pattern
[0019] Another technique concerns the analysis of typical answering patterns of answering machines: a typical answering pattern involves 1 -2 seconds of audio followed by 1 .2 - 2.4 seconds of silence before further significant energy audio and the characteristic acoustic signal.
[0020] The analysis of this type of pattern involves rather long algorithm response times, which could make the caller seem to receive a silent call. This technique also requires calibrations that can be very costly in terms of time for the necessary and appropriate configuration.
[0021] Beep detection
[0022] Typically, voicemails emit a short beep, commonly called a beep, at the beginning or end of the voicemail. Algorithms can detect this signal and use it as an indicator of the presence of a voicemail. However, if the acoustic signal is present at the end of the message, the use of this algorithm does not make any useful contribution to the detection.
[0023] By combining the above techniques, current voicemail recognition algorithms can determine whether a call was answered by a person or a voicemail. As highlighted, these techniques require calibrations that affect the progress of the phone call and can also cause it to fail.
[0024] The fundamental parameter that influences the quality of all these algorithms is the recognition latency: typically, in fact, there is a trade-off between recognition rate and response speed in the sense that with high response speed, there is a low recognition rate and vice versa, with low response speed, there is a high recognition rate.
[0025] Trying to maintain acceptable latencies, the recognition rates indicated by the manufacturers are quite far from 100%. A low response speed, in fact, creates a long pause of silence after the call connection. This can lead to a call drop by the called user hearing the mute call. In this case, the algorithm generates a failure without bringing any benefit to the operation. It is important to note that although almost all answering machines are centralized and the response message is standard, this does not make the recognition more reliable. To benefit from the simplification due to the standardization of messages, with these heuristic techniques it would be necessary to carry out a specific calibration for each message. However, this operation is hardly feasible in systems that use traditional recognition techniques based on generic patterns.
[0026] When the algorithm misses a recognition, two distinct situations occur.
[0027] False positive
[0028] This error occurs when the call is mistakenly identified as answering by answering machine, when in fact it was answered by a person. This situation involves the call being dropped by the algorithm and the consequent loss of contact. In addition, it is very annoying for the called user and negatively affects the image of the company engaged in the telephone campaign.
[0029] False negative
[0030] It occurs when the answering machine is not detected by the algorithm which is then transferred to a human operator of the call center. The false negative means that the algorithm does not make any improvement to the overall performance of the call center. Mute Calls and Regulatory Aspects
[0031] Telemarketing activities in Italy, as in other countries of the world, are generally regulated. One aspect of applying the algorithms described above, which has regulatory impacts, is the latency of the algorithm, with the consequence of generating so-called silent calls, or false positives generated by the algorithm that produce dropped calls without conversation.
[0032] In Italy, the provision of the Privacy Guarantor no. 83 of 20 February 2014 regulated silent calls. According to Italian legislation, companies that use automated systems to make advertising or promotional calls must ensure that these are not mute at the time of response. If an automatically generated call is not connected to an operator within three seconds after the recipient has answered, it must be terminated.
[0033] The number of silent calls that an automatic system can make with respect to correctly established calls is strictly defined by law. It is clear, therefore, that by reaching this limit, the algorithm can no longer be used until it falls within the legal parameters. This, in practice, in addition to decreasing the efficiency of the call center, lowers the real recognition rate of the algorithms mentioned above.
[0034] US 2004 / 037397 concerns the classification of outgoing telephone calls, with particular attention to the detection of automatic voice messages. A server analyzes received voice messages, distinguishing between human responses and automated messages. If it detects a non- “hello” message, it extracts the voice patterns, comparing them with a database of known patterns. If the model is known, it performs a predefined action; otherwise, it stores the voice pattern in a candidate database for future analysis. Speech patterns can be, for example, sequences of relevant sounds or phonemes, simplified textual representations (without specific numbers or details), data extracted from speech recognition algorithms (ASR), for example the only key terms or phrase structures typical of automatic messages, for example "The number you have called is no longer in service".
[0035] US 10,014,006 describes a system and method for the automatic identification of audio messages transmitted by network operators, based on acoustic recognition techniques (pattern matching). The system is based on a pre-configured database containing audio tracks representative of standardised messages, used by telephone operators. Each sample message is associated with metadata indicating the operator of origin and the purpose (e.g. voicemail, unavailability).
[0036] The analysis is carried out in real time, through a comparison between the audio received at the start of the call and the tracks stored. However, this approach is burdensome, especially in terms of response times that are not acceptable in the context of call center calls.
[0037] An object of the invention is therefore to solve the aforementioned problems of the prior art by drastically increasing the recognition performance of answering machines in call center activities.
[0038] A further object of the invention is to solve the above mentioned problems in a rational and economical way.
[0039] SUMMARY OF THE INVENTION
[0040] The above purposes are achieved thanks to a procedure for the recognition of answering machines in call center activities, where the aforementioned procedure includes the following phases: a) receiving the audio signal of a call; b) transforming the audio signal into a sequence of features representative of said audio signal; c) comparing the features of the audio signal with a set of pre-recorded and preset messages in a database using a pattern matching method; d) assigning scores to the correspondences between the audio signal and the prerecorded messages; e) identification of the presence of a centralised answering machine when the score of a pre-recorded message exceeds a predefined threshold.
[0041] Among the important advantages of the invention is the fact that the invention presented addresses the problem of the recognition of voicemails using a completely innovative method: pattern matching. This approach is based on the assumption that, currently, almost all answering machines are provided as a service by telephone operators, certainly covering 100% of the answering machines used in mobile telephony.
[0042] Further important advantages of the invention will be explained in the course of the present description.
[0043] Further features of the invention may be deduced from the dependent claims.
[0044] BRIEF DESCRIPTION OF THE FIGURES
[0045] Further characteristics and advantages of the invention will be evident from reading the following description provided by way of example and not limitation, with the help of the attached figures, where: figure 1 illustrates an example of call distribution in real call centers; figure 2 shows a graph relating to the analysis of an audio file used for the tests; figure 3 was made by concatenating in a single test file different audio segments to be recognized; figure 4 indicates with black triangles in the upper part the moment in time at which the algorithm issues the decision relating to the recognized segment; figure 5 illustrates a first operating scenario in which the invention is implemented; figure 6 illustrates a second scenario in which the invention is implemented; and figure 7 illustrates a third scenario in which the invention is implemented.
[0046] DETAILED DESCRIPTION OF THE FIGURES
[0047] The invention will now be described with reference to the accompanying drawings on the understanding that centralized answering machines have a pre-recorded message which, as a rule, is unique to each telephone operator. The present invention exploits this feature extremely efficiently.
[0048] The system identifies in almost real time the correspondence of the analyzed phone with respect to a set of pre-recorded and preset messages in the system; in this way it is able to determine the status of the call and, in particular, if it has been forwarded to a centralized answering machine.
[0049] The use of this method allows to obtain consistent advantages especially in predictive systems. In predictive systems, in fact, calls are generated automatically by the system and, therefore, transferred to the operator as soon as answers are received, avoiding occupying the operator with calls that will not be answered.
[0050] In figure 1 it is possible to see that, at least as far as the Italian panorama is concerned, most of the secretariats are managed by a small number of operators, and this means that a limited number of pre-recorded messages is sufficient to recognize most of the secretariats used. This is a factor that justifies the approach used in the invention presented here and allows its efficient implementation.
[0051] The system is based on acoustic analysis algorithms that allow you to identify which message, among a set of known messages, is present at a given time on the telephone line. In this way, the system is able to identify the status of the call with high precision. PM (Pattern Matching)
[0052] The system provides a database of audio tracks that represent the possible messages used by the managers. Each message is associated with the information of the manager and the purpose for which the message is reproduced. Mainly the messages indicate the voicemail service or the unreachability of a telephone device. The system recognizes in real time which message is sent by the manager and behaves accordingly.
[0053] When a message is recognized, the information of the event and the possible call drop is written in a log and transferred to the predictive call system, thus making it possible to optimize contact management strategies.
[0054] Implementation used
[0055] The approach used for the recognition of answering machines involves the transformation of the audio signal into a sequence of features that represent it. These features are chosen in such a way that they are robust and stable as the transformations that can take place on the telephone network change.
[0056] According to an embodiment of the invention, the system processes PCM (Pulse Code Modulation) audio streams at 8 kHz, segmenting the signal into 40 ms frames with a step of 20 ms. Spectral analysis is applied by FFT (Fast Fourier Transform), filtering out irrelevant components and identifying up to three significant spectral peaks for each frame. The peaks are combined into ennuples characterized by frequencies and time distances, encoded in a compact form for quick comparison. During the training phase, the most representative ennuples are stored in an audio pattern database. In the recognition phase, the ennuples extracted from the input signals are compared with those stored, calculating a score that determines the probability of recognition of a known message.
[0057] The system uses a movable time window of 100 frames (2 seconds) and a scoring function to correctly classify messages, even in the presence of time variations. This method allows for high computational efficiency and accuracy in the recognition of answering machines.
[0058] In today's fully digitized telephone network, the major sources of variability depend primarily on the chain of audio codecs used during media transport. The call can be processed by multiple telephone operators using different codecs in their transmission networks.
[0059] Recognition method
[0060] During recognition, for each prememorized segment and for each time window, the correspondences of the features with the analyzed audio are measured. These correspondences, if inserted in a temporal graph having as abscissa the time of the analyzed audio and as ordinate the time of the segment to be recognized, assume a diagonal arrangement. This arrangement is used by the algorithm to assign a score, variable in time, associated with each of the audio segments that you want to recognize.
[0061] Figure 2 shows a graph relating to the analysis of an audio file used for testing. This file is deliberately long (about 28 seconds) and contains only one recognizable segment at the beginning of the file. The diagonal arrangement of the matches displayed with purple crosses is easily visible.
[0062] Score associated with each segment
[0063] During the processing of the audio, scores, time variables, associated with each segment are calculated. The algorithm decides to recognize a certain voicemail / audio segment when the associated score unequivocally exceeds any other score associated with all other segments.
[0064] The graph in Figure 3 was made by concatenating different audio segments to be recognized into a single test file. These segments consist of portions of audio used by operators for voicemail alert messages or for mobile device unreachability alert messages.
[0065] For each time frame, the highest score is displayed graphically, while for each segment a different color was used.
[0066] You can see how easily the maximum scores are recognizable and it is easy to understand how selective the algorithm is. This is one of the most important features because it is the one that allows to achieve, even in real installations in call centers, the total elimination of false positives which, as seen above, constitute an important limit for the use of AMD systems.
[0067] In the graph in Figure 4, the black triangles at the top indicate the moment in time at which the algorithm issues the decision relating to the recognized segment.
[0068] From the graph in figure 4 it is therefore possible to see that, given the high selectivity and the high slope of the scores in correspondence with a recognizable message, the decision can be taken in a decidedly short time, less than one second, from when the message begins.
[0069] As already seen, this is also a very important feature for the use of an AMD system, as it allows you to avoid silent calls and the negative consequences already analyzed above.
[0070] The system has been tested on a confidential basis in the call center and its performance has been verified.
[0071] The following table lists the main features.
[0072] Table 1
[0073] Turning to the application of the invention, it is noted that the invention presented here is particularly useful in the call canter field during outbound campaigns that use automatic and predictive systems to make telephone calls to users.
[0074] As seen above, the recognition of answering machines and their management (destruction or transfer to the operator) is of fundamental importance to increase the productivity of the call center and to make the best use of available telephone resources.
[0075] The following sections describe two typical operational scenarios in which it is shown how the system works and actively intervenes on the call flow.
[0076] Vehicle Operating Scenarios
[0077] It is possible to distinguish two main operational scenarios in which the system intervenes effectively. These scenarios depend on the type of message that is transmitted and the telephone operator that is in charge of the called user. The latter can send the transfer notice to the voicemail in the pre-connection phase or in the postconnection phase.
[0078] The first scenario indicates the voicemail warning message during pre-connection (setup).
[0079] With the telephone operators shown in Figure 5, when a voicemail is recognised, no human operator is engaged.
[0080] In addition, since the answering machine is identified in the pre-connection phase, the commitment of telephone resources is minimized and the possible costs related to connection fee are avoided.
[0081] The second scenario indicates the voicemail warning message in the post-connection phase.
[0082] In the case shown in figure 6, the human operator can be engaged by the call, however, the system knocks down the call much faster than the operator would, eliminating the unproductive times in which he should listen, even partially, to the message, understand what it is about and, therefore, manually knock down the call, declaring the outcome of the same to the management program.
[0083] Additional options
[0084] In the telephone environment, voice messages are used to signal the status of a connection before the recipient is actually contacted. These messages are directed to the human caller.
[0085] Although there are some abatement codes in telecommunications standards that cover virtually all cases, there is no obligation for operators to use them correctly. In the mobile phone scenario, for example, messages such as the temporary unreachability of a mobile phone are never uniquely translated into an international standard code (e.g. Q.850).
[0086] The method presented here is able to recognize these phonies and provide the automatic calling system with a correct disconnection cause that can be used by specific algorithms to decide when to try again to contact the out-of-coverage user. The data collected, in fact, can be used to optimise contact management and thus increase overall efficiency. For example, it would be unhelpful to call a number that is incorrect, or immediately call a number that is unreachable; in these cases it is better to exclude contact or postpone the call by reusing telephone resources for new potentially productive calls.
[0087] Message of unreachability or incorrect numbering.
[0088] In this case, visible in figure 7, the unreachability message can last for several seconds, even more than 15 seconds if sent in multiple languages. So, even in this context, the system is useful because it breaks down these types of unproductive calls and quickly frees up telephone lines, with the possibility of reusing them immediately.
[0089] The non-recognition of the message in this scenario is however less relevant, rather than in the case of the voicemail, because it does not imply the occupation of the human operator (the call never goes into the connect state), although it can occupy the telephone line for several seconds.
[0090] It is desired here to enumerate the main advantages of the invention.
[0091] In particular, the system presented manages to overcome all the main problems of the currently used AMD systems, allowing the following advantages to be obtained.
[0092] High recognition rate
[0093] The system presented is characterized by a recognition rate higher than 99.9% under real operating conditions. The recognition rate close to 100% significantly exceeds that obtained with the traditional analysis methods used by the main VoIP operators worldwide.
[0094] Response times
[0095] Response times are typically less than 1 second from the beginning of the audio message to be recognized. This result is unfeasible in systems that use the heuristics or response patterns described above, because, by their nature, you need to listen to much longer sections of audio to have enough data to make a thoughtful decision.
[0096] Operation during pre-connection
[0097] As mentioned above, the system is primarily designed to handle centralized answering machines. A very common feature among these services is that many operators send a message to inform the caller that the person sought is currently unavailable and that the call will be redirected to a centralised answering machine. This message is sent before the call is connected, in practice it is transmitted exactly as a tone indicating the status of the call (e.g. free tone, busy tone, etc.). Typically, voicemail recognition systems do not analyze this message, but only intervene after the call is connected.
[0098] The system presented here, on the other hand, also analyzes the warning message and, under typical conditions, intervenes before the call is connected. By dropping the call even before it is connected, the efficiency of the operators is further increased because they are not in any way impacted by these calls. In addition, freeing the line occupied by the call that is about to be transferred to the answering machine in such a timely manner also increases the number of call attempts that can be made by the call center on the available lines, positively affecting productivity.
[0099] Absence of false positives
[0100] The system used in real conditions does not present false positives. This is a very important feature because it avoids missing potentially productive calls. In addition, it avoids abandoning a call with a user mistakenly mistaken for a voicemail, which is quite negative for the quality of the service and the reputation of the company that is carrying out the outbound campaign, as well as for the aforementioned laws on silent calls.
[0101] Ease of set up
[0102] The system does not need complex configuration tuning steps. For use, a simple database of audio tracks containing the possible messages used by telephone operators is sufficient. This allows for cost savings while still providing high performance.
[0103] Regardless of the language used The system takes advantage of the audio characteristics of the messages to be recognized, without relying on the language used or the recognition of certain words. It is therefore completely language independent and does not require changes if applied in countries with different languages, other than updating the database with the selected messages related to the region of interest.
[0104] No impact on silent calls
[0105] The method presented never generates silent calls. In fact, the method can:
[0106] Act before connecting the call, on the free transfer message.
[0107] In this mode, the call is never answered, so it cannot fall within the category of silent calls.
[0108] Act immediately after connecting the call, with the operator already on the line. In this mode, the algorithm intervenes in parallel with the operator's response. Therefore:
[0109] If the call is a voicemail, it is taken down, but there is no answer from an operator.
[0110] If the call is not a voicemail, the algorithm does not perform any action and the call center operator is already in conversation, as required by law.
[0111] It is wished to specify that the invention as described can be modified or improved for contingent or particular reasons, without departing from the scope of the invention.
Claims
CLAIMS1. Procedure for the recognition of answering machines in call center activities, where the aforementioned procedure is characterized by the following phases: a) receiving the audio signal of a call in PCM stream format; b) transforming the audio signal into a sequence of features representative of said audio signal by frame segmentation and subsequent spectral analysis; c) comparing the features of the audio signal with a set of pre-recorded and preset messages in a database using a pattern matching method; d) assigning scores to the correspondences between the audio signal and the prerecorded messages; e) identification of the presence of a centralised answering machine when the score of a pre-recorded message exceeds a predefined threshold.
2. Method according to claim 1 , wherein the comparison of the audio signal with the pre-recorded messages takes place in real time, allowing the recognition of the voicemail with a latency of less than one second from the beginning of the voice message.
3. Method according to any one of the preceding claims, wherein the prerecorded messages relate to centralized answering machines provided by mobile and fixed telephone operators.
4. Process according to any one of the preceding claims, wherein said pattern matching method comprises the steps of: a) transformation of the audio signal into a sequence of representative features by extracting ennuples of spectral peaks based on frequencies and temporal distances; b) comparison of the ennuples with the audio segments stored in the database; c) assigning scores to the matches between the audio signal in question and the prerecorded segments; d) recognition decision based on exceeding a score threshold.
5. Method according to any one of the preceding claims, wherein the scores are displayable in a time graph where the abscissa represents the time of the audio signal under examination and the ordinate represents the time of the prerecorded segments.
6. Method according to any one of the preceding claims, wherein the recognition decision takes place when the score of one segment unequivocally exceeds the scores of the other segments.
7. A method according to any one of the preceding claims, wherein, once the voicemailmessage is acknowledged, information relating to that acknowledgement is recorded in a log and transferred to a predictive call system to optimise contact management.
8. Method according to any one of the preceding claims, wherein the in PCM stream format is at 8 KHz and the segmentation of the audio signal is operated in 40 ms frames with 20 ms pitch.
9. Method according to any one of the preceding claims, wherein the audio signal transformation step involves the use of a Fourier transform (FFT) and the filtering of non-significant spectral components.
10. Method according to any one of the preceding claims, wherein the management of the ennuples takes place by means of a moving time window of 100 frames, equivalent to about 2 seconds, and the application of a scoring function that considers time tolerances to increase the reliability of recognition.