Cumulative average spectral entropy analysis for tone and voice classification.

Cumulative average spectral entropy analysis enhances call progress analysis by accurately distinguishing tones from voice in degraded audio signals, improving the efficiency of call handling in contact centers.

JP7771097B2Active Publication Date: 2025-11-17GENESIS CLOUD SERVICES CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022574354
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-06-30
Filing Date
2021-06-30
Publication Date
2025-11-17
Estimated Expiration
2041-06-30

AI Technical Summary

Technical Problem

Current signal processing algorithms for call progress analysis struggle to accurately distinguish between tones and voice in audio signals, especially when the audio is degraded by poor transmission networks, due to the lack of a consistent global standard for tone frequencies and patterns.

Method used

Implementing cumulative average spectral entropy analysis to calculate a difference measure between the cumulative average of entropy and cumulative average spectral entropy of audio signals, allowing for robust tone and voice classification even in degraded conditions.

Benefits of technology

The method significantly improves the accuracy of call progress analysis by reducing classification errors, particularly in noisy environments, enabling efficient handling of outbound calls through automated systems or agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007771097000037
    Figure 0007771097000037
  • Figure 0007771097000038
    Figure 0007771097000038
  • Figure 0007771097000039
    Figure 0007771097000039
Patent Text Reader

Abstract

According to one embodiment, a contact center system for performing call progress analysis including tone and voice classification includes at least one processor and at least one memory containing a plurality of instructions stored thereon, the instructions, responsive to execution by the at least one processor, causing the contact center system to determine a cumulative average of entropy of audio signals received by the contact center system, determine a cumulative average power spectral amplitude and a cumulative average spectral entropy of the audio signals, calculate a difference measure of the audio signals as the difference between the cumulative average of entropy and the cumulative average spectral entropy, distinguish between tones and voice in the audio signals based on the difference measure, and process one or more tones in the audio signals in response to identifying one or more tones in the audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 045,908, filed June 30, 2020, entitled "Cumulative Average Spectral Entropy Analysis for Tone and Speech Classification," the contents of which are incorporated herein by reference in their entirety.

[0002] Call analysis or call progress analysis (CPA) is a term for a set of signal processing algorithms that operate on audio signals during call setup, consisting of both tones and voice, to determine the outcome of the call. Humans can easily hear and detect various tones (e.g., dial tone before dialing, ring back, busy, answer, etc.). However, for a machine to be able to do the same with the same accuracy requires considerable care in its implementation, especially when it must distinguish human voice from various tones in network carrier messages.

[0003] Telephony applications involving outbound calling capabilities require the ability to accurately and quickly interpret call progress tones (e.g., ringback and busy) delivered by the network to the calling entity. While the International Telecommunication Union has published recommended tone definitions for each country, and these definitions are generally adhered to, there remains no consistent standard set of tone frequencies and patterns used worldwide by all telephone providers to signal specific events, which complicates call progress analysis. Providers use a variety of approaches in attempting to detect and identify the different tones involved in the process of analyzing call progress. However, most currently employed signal processing algorithms tend to perform poorly when the audio signal under analysis is degraded by a poor transmission network or otherwise. Summary of the Invention

[0004] One embodiment relates to unique systems, components, and methods for cumulative average spectral entropy analysis for tone and audio classification. Other embodiments are directed to apparatus, systems, devices, hardware, methods, and combinations thereof for cumulative average spectral entropy analysis for tone and audio classification.

[0005] According to one embodiment, a contact center system for performing call progress analysis using tone and voice classification may include at least one processor and at least one memory containing a plurality of instructions stored thereon, the instructions, responsive to execution by the at least one processor, causing the contact center system to determine a cumulative average of entropy of audio signals received by the contact center system; determine a cumulative average power spectral amplitude of the audio signals and a cumulative average spectral entropy of the audio signals based on the cumulative average power spectral amplitude of the audio signals; calculate a difference measure for the audio signals as the difference between the cumulative average of the entropy of the audio signals and the cumulative average spectral entropy of the audio signals; distinguish between tones and voice in the audio signals based on the difference measure for the audio signals; and process one or more tones of the audio signals in response to identifying one or more tones in the audio signals.

[0006] In some embodiments, processing one or more tones of the audio signal may include identifying a call progression tone pattern within the one or more tones of the audio signal, and transferring the telephone call from a first one of the contact center systems to a second one of the contact center systems in response to identifying the call progression tone pattern within the one or more tones of the audio signal.

[0007] In some embodiments, processing one or more tones of the audio signal may include connecting the outbound call to an automated interactive voice response (IVR) system of the contact center system.

[0008] In some embodiments, processing one or more tones of the audio signal may include connecting the outbound call to an agent in a contact center system.

[0009] In some embodiments, one or more tones of the audio signal may include a call progress tone pattern.

[0010] In some embodiments, the call progress tone pattern may be a busy signal pattern, a ringback pattern, or a special information tone pattern.

[0011] In some embodiments, processing one or more tones of the audio signal may include determining a corresponding frequency of each of the one or more tones of the audio signal.

[0012] In some embodiments, determining a cumulative average of the entropy of the audio signal may include calculating the entropy of the audio signal.

[0013] According to another embodiment, one or more non-transitory machine-readable storage media including a plurality of instructions stored thereon, the instructions, when executed by at least one processor, may cause a contact center system to: calculate an entropy of audio signals received by the contact center system; calculate a cumulative average of the entropies of the audio signals; calculate a cumulative average power spectral amplitude of the audio signals; calculate a cumulative average spectral entropy of the audio signals based on the cumulative average power spectral amplitude of the audio signals; calculate a difference measure for the audio signals as the difference between the cumulative average of the entropies of the audio signals and the cumulative average spectral entropy of the audio signals; classify tones and voices of the audio signals based on the difference measure for the audio signals; and process one or more tones of the audio signals in response to identifying one or more tones in the audio signals.

[0014] In some embodiments, processing one or more tones of the audio signal may include transferring a telephone call from a first one of the contact center systems to a second one of the contact center systems in response to identifying a call progress tone pattern within the one or more tones of the audio signal.

[0015] In some embodiments, processing one or more tones of the audio signal may include connecting the outbound call to an automated interactive voice response (IVR) system of the contact center system.

[0016] In some embodiments, processing one or more tones of the audio signal may include connecting the outbound call to an agent in a contact center system.

[0017] In some embodiments, one or more tones of the audio signal may include a call progress tone pattern.

[0018] In some embodiments, the call progress tone pattern may be a busy signal pattern, a ringback pattern, or a special information tone pattern.

[0019] In some embodiments, processing one or more tones of the audio signal may include determining a corresponding frequency of each of the one or more tones of the audio signal.

[0020] According to yet another embodiment, a method for performing call progress analysis using tone and voice classification within a contact center system may include receiving, by the contact center system, an audio signal; determining, by the contact center system, an entropy of the audio signal received by the contact center system; determining, by the contact center system, a cumulative average of the entropy of the audio signal; determining, by the contact center system, a cumulative average power spectral amplitude of the audio signal; determining, by the contact center system, a cumulative average spectral entropy of the audio signal based on the cumulative average power spectral amplitude of the audio signal; determining, by the contact center system, a difference measure of the audio signal as the difference between the cumulative average of the entropy of the audio signal and the cumulative average spectral entropy of the audio signal; classifying, by the contact center system, the tone and voice of the audio signal based on the difference measure of the audio signal; and processing, by the contact center system, one or more tones of the audio signal in response to identifying one or more tones in the audio signal.

[0021] In some embodiments, processing one or more tones of the audio signal may include identifying a call progression tone pattern within the one or more tones of the audio signal, and transferring the telephone call from a first one of the contact center systems to a second one of the contact center systems in response to identifying the call progression tone pattern within the one or more tones of the audio signal.

[0022] In some embodiments, processing one or more tones of the audio signal may include connecting the outbound call to an agent of a contact center system or to an automated interactive voice response (IVR) system.

[0023] In some embodiments, one or more tones of the audio signal may include a call progress tone pattern.

[0024] In some embodiments, processing one or more tones of the audio signal may include determining a corresponding frequency of each of the one or more tones of the audio signal.

[0025] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used as an aid in limiting the scope of the claimed subject matter. Further embodiments, forms, features, and aspects of the present application will become apparent from the description and figures provided herewith. [Brief explanation of the drawings]

[0026] The concepts described herein are illustrated by way of example, and not by way of limitation, in the accompanying drawings. For simplicity and clarity of illustration, elements shown in the figures have not necessarily been drawn to scale. Where considered appropriate, reference labels have been repeated among the figures to indicate corresponding or analogous elements. [Figure 1A] FIG. 1 is a simplified flow diagram of at least one embodiment of a method for performing call progress analysis using tone and voice classification. [Figure 1B] FIG. 1 is a simplified flow diagram of at least one embodiment of a method for performing call progress analysis using tone and voice classification. [Figure 2] FIG. 1 is a simplified block diagram of at least one embodiment of a call center system. [Figure 3] FIG. 1 is a simplified block diagram of at least one embodiment of a computing system. [Figure 4] 1 is a spectrogram of an exemplary audio signal that includes both a tone signal and a voice signal. [Figure 5] 1 is a graph of the spectral flatness of an audio signal during tones and speech. [Figure 6]1 is a graph of entropy measurements of an audio signal during tones and speech; [Figure 7] 1 is a spectrogram of an exemplary poor-quality tone audio signal observed in a call center. [Figure 8] 1 is a graph of entropy measurements of an audio tone signal. [Figure 9] 1 is a graph of three entropy measures of an audio signal having a degraded tone. [Figure 10] 1 is a graph of three entropy measures of an audio signal during tones and speech. DETAILED DESCRIPTION OF THE INVENTION

[0027] While the concepts of the present disclosure are susceptible to various modifications and alternative forms, specific embodiments have been shown by way of example in the drawings and are herein described in detail. It should be understood, however, that there is no intention to limit the concepts of the present disclosure to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives consistent with the scope of the present disclosure and the appended claims.

[0028] References such as "one embodiment," "an embodiment," "an exemplary embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but that all embodiments may or may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. It should be further understood that references to "preferred" components or features may indicate the desirability of a particular component or feature with respect to one embodiment, but that the present disclosure is not so limited with respect to other embodiments in which such component or feature may be omitted. Furthermore, when a particular feature, structure, or characteristic is described in connection with an embodiment, the implementation of such feature, structure, or characteristic with other embodiments, whether or not explicitly described, is within the knowledge of one of ordinary skill in the art. Furthermore, particular features, structures, or characteristics may be combined in any suitable combinations and / or subcombinations in various embodiments.

[0029] Additionally, it should be understood that items included in a list in the format "at least one of A, B, and C" can mean (A), (B), (C), (A and B), (B and C), (A and C), or (A, B, and C). Similarly, it should be understood that items listed in the format "at least one of A, B, or C" can mean (A), (B), (C), (A and B), (B and C), (A and C), or (A, B, and C). Further, with respect to the claims, the use of words and phrases such as "a," "an," "at least one," and / or "at least one portion" should not be construed as limiting to only one of such elements unless specifically stated to the contrary, and the use of phrases such as "at least a portion" and / or "a portion" should be construed to encompass both embodiments including only a portion of such an element and embodiments including the whole of such an element unless specifically stated to the contrary.

[0030] The disclosed embodiments may, in some cases, be implemented in hardware, firmware, software, or a combination thereof. The disclosed embodiments may also be implemented as instructions stored on or executed by one or more transitory or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. A machine-readable storage medium may be embodied as any storage device, mechanism, or other physical structure for storing or transmitting information in a form readable by a machine (e.g., volatile or non-volatile memory, a media disk, or other media device).

[0031] In the figures, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, such features may be arranged in a different manner and / or order than that shown in the illustrative figures, unless indicated to the contrary. Furthermore, the inclusion of a structural or method feature in a particular figure does not imply that such feature is required in all embodiments, and in some embodiments, it may not be included or may be combined with other features.

[0032] Providers use various approaches in attempting to detect and identify various tones involved in the process of analyzing call progress. For example, in some embodiments, the Goertzel algorithm or Fast Fourier Transform (FFT) may be used, each of which exhibits different advantages and disadvantages. Tone detection and identification are performed simultaneously using the Goertzel algorithm. That is, the algorithm not only detects the presence of a tone but also identifies which tone is being detected. However, when an FFT algorithm is used, the process is generally separated into two steps. In the first step, generic tone versus speech classification is performed. In the second step, once a tone is known to be present, its frequency is identified in order to interpret the call outcome. The technology described herein improves the FFT algorithm by incorporating "cumulative average spectral entropy" into the analysis, thereby making the process more robust when the audio signal is degraded by poor transmission networks.

[0033] Telephone networks are not standardized around the world. For example, in Europe, a single tone is primarily used to report events, and the ringback tone in most European countries is 425 Hz. However, in North America, dual tones are preferred, using dual tones between 350 and 440 Hz to indicate ringback. Therefore, relatively complex setups are often required by any telephone provider wishing to serve different countries. The Goertzel algorithm and the FFT algorithm represent two approaches to tone detection and identification.

[0034] The Goertzel algorithm is a technique used in digital signal processing (DSP) for the efficient evaluation of individual terms in the Discrete Fourier Transform (DFT), and is used to calculate the kth DFT component of a signal {x(n), n = [0, N]}. The Goertzel algorithm analyzes one selectable frequency component from a discrete signal.

[0035]

number

[0036]

number

[0037] It should be appreciated that the Goertzel algorithm is relatively simple and performs well in most situations. However, one important limitation of the Goertzel algorithm is that the specific frequencies of interest must be known a priori. The frequencies of interest are defined by an index k, chosen from index numbers k∈{0, 1, 2, ..., N-1}. Assuming a sampling frequency of 8000 Hz, which is common in telephones, the frequency resolution Δf is given by

[0038]

number

[0039] In telephony, the Goertzel algorithm can be used to detect dual-tone multi-frequency (DTMF) signals, where the meaning of the signaling is determined by the simultaneous presence of two of a total of eight frequencies. As shown in Table 1 below, because eight different frequencies are evaluated simultaneously, the Goertzel algorithm is evaluated eight times for each value of n with different values ​​of k defining the frequency ω. The Goertzel algorithm has a higher order of complexity than the FFT, but remains efficient for computing a small number of selected frequency components.

[0040] [Table 1]

[0041] Other audible signals used in telephones include call-progression tone patterns that indicate the progress or nature of a telephone call. Busy signals, ringback, and special information tones (SITs) are all examples of such call-progression tone patterns. CPA requires distinguishing between a large set of tones to be able to classify the tone patterns. For example, in North America, up to 16 tone frequencies, as shown in Table 2 below, are used to create different tone patterns, each of which consists of one or more frequencies played at specific intervals. As mentioned above, when FFT is used for CPA, the process is separated into two steps: detecting the presence of a tone and identifying its frequency.

[0042] [Table 2]

[0043] Spectral flatness and entropy measures may be calculated in the frequency domain of the audio signal. As described below, spectral flatness and entropy measures work relatively well to distinguish between tones and speech. Specifically, the audio signal may first be segmented into overlapping frames. For example, in some embodiments, each frame may have a length of 0.03 seconds and be weighted by a Hamming window of the same length. The overlap used may be 2 / 3, such that the window is advanced by 0.01 seconds at each time step. A 256-point FFT may be performed over each frame. In some embodiments, because for real signals the power spectral amplitude is symmetric, the power spectral amplitude coefficient X k Only half of k is retained at index k ∈ [1,128]. Figure 4 shows a spectrogram of an exemplary audio signal that includes both a tone signal and a speech signal.

[0044] Spectral flatness or tonality coefficient is a measurement used in DSPs to characterize the audio spectrum. Spectral flatness or tonality coefficient provides a way to quantify how tone-like a sound is, as opposed to being noise-like. Spectral flatness measurements are

[0045]

number

[0046] In some embodiments, entropy may be used as a measure of randomness to distinguish between tones and voices. The definition of entropy is, in terms of a discrete set:

[0047]

number

[0048] Both flatness and entropy measures often vary significantly during speech (see Figures 5 and 6), so cumulative averages are generally preferred over instantaneous values. For example, for entropy measures, the cumulative average of entropy is

[0049]

number

[0050]

number

[0051] However, audio quality is often low, for example, due to a poor transmission network, which typically causes CPA systems to perform substandardly. FIG. 7 shows an example of a low-quality tone audio signal observed in a call center, and FIG. 8 shows the corresponding entropy measurement of the audio signal in FIG. 7. While it would be expected that each block of the tone signal would have the same approximate value, it is clear from FIG. 8 that a low-quality tone signal results in each block of the tone signal having a different entropy value, which can make it difficult to find an appropriate threshold for distinguishing between tone signals and speech signals. Thus, when presented with such a poor audio signal, CPA algorithms often struggle to properly classify the tones.

[0052] As noted above, the cumulative average of entropy measures that can be used to distinguish between tones and speech

[0053]

number

[0054]

number

[0055]

number

[0056] The improved techniques described herein overcome this problem caused by degraded audio signals by utilizing cumulative average spectral analysis to generate a cumulative average spectral entropy measure for use in distinguishing between voice and tone. In an exemplary embodiment, the cumulative average power spectral amplitude is:

[0057]

number

[0058]

number

[0059] Then, the cumulative average power spectrum amplitude value

[0060]

number

[0061]

number

[0062]

number

[0063] Essentially, the cumulative average power spectral amplitude

[0064]

number

[0065]

number

[0066]

number

[0067]

number

[0068]

number

[0069] FIG. 9 shows the instantaneous entropy measurements, the average entropy measurements, and the average entropy measurements for the same exemplary degraded tone signal under analysis.

[0070]

number

[0071]

number

[0072]

number

[0073]

number

[0074]

number

[0075] To verify the performance and robustness of the improved techniques and algorithms (i.e., the new method), a set of audio signals carried over a very poor transmission network was analyzed. More specifically, the audio files used to evaluate the robustness of the system were all audio signals from contact centers that were reported as problematic. In a first step, 525 files containing primarily tone signals with occasional speech were analyzed to measure the difference measure.

[0076]

number

[0077]

number

[0078] [Table 3]

[0079] As shown, the new method produces fewer classification errors than alternative methods, including outperforming the NN model by producing approximately 20% fewer errors for voices marked as tones and 65% fewer errors for tones marked as voices. It should therefore be appreciated that such results provide substantial confidence in the feasibility of the new method / technology.

[0080] In a second step, 225 files containing primarily speech audio signals with occasional tones were analyzed using the same method, with the results provided in Table 4 below.

[0081] [Table 4]

[0082] Thus, the analysis of an audio signal containing primarily speech with occasional tones yields similar comparative results, with the new method / technique producing fewer classification errors than the other two methods analyzed.

[0083] Being able to accurately and efficiently detect when tones are present in an audio stream is a critical step in call progress analysis (CPA) systems. Assuming that a signal represents a tone, identifying a particular tone (e.g., as a 400 Hz tone or a 679 Hz tone) is typically relatively straightforward. The methods and techniques described herein, involving cumulative average spectral analysis, are robust in their ability to distinguish between tones and speech, even when encountering degraded audio signals, and thus will help significantly improve the performance of CPA systems worldwide.

[0084] It should be appreciated that the audio signals may be received and / or analyzed by one or more devices of a contact center system (e.g., the contact center system of FIG. 2 ), e.g., using the techniques described herein to distinguish between tones and voices in the audio signals in order to automatically interpret / process the call. For example, in some embodiments, the call initiator (e.g., an automated outbound dialer system) is interested in knowing whether the line is busy, whether someone answered, etc., in order to take the next appropriate action, such as connecting the outbound call to an agent or an automated interactive voice response (IVR) system.

[0085] 1A and 1B, in use, a system may execute method 100 for performing call progress analysis using tone and voice classification. It should be understood that in some embodiments, the system may be embodied as a computing device (e.g., computing device 300 of FIG. 3) and / or a contact center system (e.g., contact center system 200 of FIG. 2) or systems / devices thereof. It should be understood that certain blocks of method 100 are shown as examples, and that such blocks may be combined, divided, added, removed, and / or reordered in whole or in part, depending on the particular embodiment, unless stated to the contrary.

[0086] The exemplary method 100 begins at block 102 of FIG. 1A where a system (e.g., computing device 300 or contact center system 200) receives an audio signal. At block 104, the system determines the entropy of the received audio signal. In doing so, at block 106, the system:

[0087]

number

[0088] At block 108, the system determines a cumulative average of the entropy of the audio signal. In doing so, at block 110, the system:

[0089]

number

[0090] At block 112, the system determines the cumulative average power spectral amplitude of the audio signal. In doing so, at block 114, the system:

[0091]

number

[0092] 1B, the system determines a cumulative average spectral entropy of the audio signal based on the cumulative average power spectral amplitude of the audio signal. In doing so, at block 118, the system:

[0093]

number

[0094] At block 120, the system determines a difference measure for the audio signal as the difference between the cumulative average of the entropy of the audio signal and the cumulative average spectral entropy of the audio signal. In doing so, at block 122, the system:

[0095]

number

[0096] At block 124, the system classifies tones and voices based on the difference measure of the audio signal. For example, as described above, the difference measure may be close to or near zero during tones, indicating that the portion of the audio signal (e.g., frequency range) being analyzed likely corresponds to a tone. Accordingly, in some embodiments, the system may utilize one or more thresholds to distinguish between tone and voice portions of the audio signal. Specifically, the system may identify portions of the audio signal that are below a predetermined threshold as being the tone portion of the audio signal and portions of the audio signal that are above (or at least at) the predetermined threshold as being the voice portion of the audio signal. It should be understood that in other embodiments, the system may distinguish / classify tones and voices differently based on the difference measure of the audio signal.

[0097] At block 126, the system determines whether one or more tones are identified within the audio signal. If so, method 100 proceeds to block 128, where the system processes (or attempts to process) the one or more tones. Otherwise, method 100 may end. In situations where tones are identified within the audio signal, it should be understood that the identified tones may be or include one or more call-progression tone patterns. For example, in some embodiments, the tones may include or represent a busy signal pattern, a ringback pattern, or a special information tone (SIT) pattern. It should be understood that the system may process the one or more tones using any suitable technique and / or algorithm, for example, by determining the corresponding frequency of each of the tones identified within the audio signal. For example, in some embodiments, the system may identify a call-progression tone pattern within one or more tones of the audio signal to transfer a telephone call to another entity. Specifically, in the context of a contact center system, a telephone call may be transferred from a first system of the contact center system to a second system of the contact center system. In another embodiment, processing tones of audio signals (e.g., call progress tone patterns) may include connecting the outbound call to an agent or an automated interactive voice response (IVR) system of a contact center system.

[0098] Although blocks 102-128 are described relatively serially, it should be understood that various blocks of method 100 may be performed in parallel in some embodiments.

[0099] 2, there is shown a simplified block diagram of at least one embodiment of a communications infrastructure and / or content center system 200 that may be used in conjunction with one or more of the embodiments described herein. The contact center system 200 may be embodied as any system capable of providing contact center services (e.g., call center services, chat center services, SMS center services, etc.) to end users and otherwise performing the functions described herein. The exemplary contact center system 200 includes a customer device 205, a network 210, a switch / media gateway 212, a call controller 214, an interactive media response (IMR) server 216, a routing server 218, a storage device 220, a statistics server 226, agent devices 230A, 230B, 230C, a media server 234, a knowledge management server 236, a knowledge system 238, a chat server 240, a web server 242, an interaction (iXn) server 244, a universal contact server 246, a reporting server 248, a media services server 249, and an analytics module 250.The exemplary embodiment of FIG. 2 includes one customer device 205, one network 210, one switch / media gateway 212, one call controller 214, one IMR server 216, one routing server 218, one storage device 220, one statistics server 226, one media server 234, one knowledge management server 236, one knowledge system 238, one chat server 240, one iXn server 244, one universal contact server 246, one reporting server 248, one media services server 249, and one analytics module. Although only customer device 205, network 210, switch / media gateway 212, call controller 214, IMR server 216, routing server 218, storage device 220, statistics server 226, media server 234, knowledge management server 236, knowledge system 238, chat server 240, iXn server 244, universal contact server 246, reporting server 248, media services server 249, and / or analytics module 250 are shown. Additionally, in some embodiments, one or more of the components described herein may be excluded from system 200, and one or more of the components described as being independent may form part of another component, and / or one or more of the components described as forming part of another component may be independent.

[0100] It should be understood that, as used herein, the term "contact center system" is used to refer to the system and / or components thereof illustrated in Figure 2, while the term "contact center" is used more generally to refer to contact center systems, the customer service providers that operate these systems, and / or the organizations or companies associated therewith. Thus, unless otherwise limited, the term "contact center" generally refers to a contact center system (e.g., contact center system 200), associated customer service providers (e.g., a particular customer service provider that provides customer service through contact center system 200), and the organization or company on behalf of which customer service is provided.

[0101] By way of background, customer service providers may offer many types of services through contact centers. Such contact centers may be staffed with employees or customer service agents (or simply “agents”) who serve as an interface between a company, enterprise, government agency, or organization (hereinafter interchangeably referred to as an “organization” or “enterprise”) and people, such as users, individuals, or customers (hereinafter interchangeably referred to as “individuals” or “customers”). For example, contact center agents may assist customers in making purchasing decisions, placing orders, or resolving issues with products or services they have already received. Within a contact center, such interactions between contact center agents and external entities or customers may occur via various communication channels, such as, for example, voice (e.g., telephone calls or voice over IP, i.e., VoIP calls), video (e.g., video conferencing), text (e.g., email and text chat), screen sharing, co-browsing, and / or other communication channels.

[0102] Operationally, contact centers generally strive to provide quality service to customers while minimizing costs. For example, one way contact centers operate is to handle all customer interactions with live agents. While this approach may be entirely successful from a service quality perspective, it would likely be prohibitively expensive due to the high cost of agent labor. For this reason, most contact centers utilize some level of automated processes, such as interactive voice response (IVR) systems, interactive media response (IMR) systems, Internet robots (i.e., "bots"), automated chat modules (i.e., "chatbots"), and / or other automated processes, in place of live agents. In many cases, this has proven to be a successful strategy, as automated processes can be highly efficient at handling certain types of interactions and effective in reducing the need for live agents. Such automation allows contact centers to target the use of human agents to more difficult customer interactions, while the automated processes handle more repetitive or routine tasks. Furthermore, automated processes can be structured in a way that optimizes efficiency and promotes repeatability. Although human agents, i.e., live agents, may forget to answer certain questions or follow up on certain details, such errors are typically avoided through the use of automated processes. While customer service providers increasingly rely on automated processes to interact with customers, the use of such technologies by customers remains far less developed. Thus, on the contact center side of the interaction, IVR systems, IMR systems, and / or bots are used to automate parts of the interaction, while actions on the customer side remain manually performed by the customer.

[0103] It should be understood that the contact center system 200 can be used by customer service providers to provide various types of services to customers. For example, the contact center system 200 can be used to participate in and manage interactions in which automated processes (or bots) or human agents communicate with customers. As will be appreciated, the contact center system 200 can be an in-house facility of a business or enterprise for performing sales and customer service functions related to products and services available through the enterprise. In another embodiment, the contact center system 200 can be operated by a third-party service provider contracted to provide services on behalf of another organization. Furthermore, the contact center system 200 can be deployed on equipment dedicated to the enterprise or third-party service provider and / or in a remote computing environment, such as, for example, a private or public cloud environment with infrastructure to support multiple contact centers for multiple enterprises. The contact center system 200 can include software applications or programs that can be executed on-premise, remotely, or some combination thereof. Furthermore, it should be understood that various components of the contact center system 200 can be distributed across various geographic locations and are not necessarily contained in a single location or computing environment.

[0104] Furthermore, unless specifically limited otherwise, it should be understood that any of the computing elements of the present invention may be implemented within a cloud-based or cloud computing environment. As used herein and further described below with reference to computing device 300, "cloud computing" or simply "cloud" is defined as a model for enabling ubiquitous, convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) that can be rapidly provisioned via virtualization, released with minimal management effort or service provider interaction, and then scaled accordingly. Cloud computing can consist of a variety of characteristics (e.g., on-demand self-service, wide area network access, resource pooling, rapid elasticity, scalable service, etc.), service models (e.g., Software as a Service ("SaaS"), Platform as a Service ("PaaS"), Infrastructure as a Service ("IaaS")), and deployment models (e.g., private cloud, community cloud, public cloud, etc.). Cloud execution models, often referred to as "serverless architectures," generally involve a service provider dynamically managing the allocation and provisioning of remote servers to achieve desired functions.

[0105] It should be understood that any of the computer-implemented components, modules, or servers described in connection with Figure 2 may be implemented via one or more types of computing devices, such as, for example, computing device 300 of Figure 3. As will be appreciated, contact center system 200 generally manages resources (e.g., personnel, computers, telecommunications equipment, etc.) to enable delivery of services via telephone, email, chat, or other communication mechanisms. Such services may vary depending on the type of contact center and may include, for example, customer service, help desk functions, emergency response, telemarketing, order taking, and / or other features.

[0106] A customer desiring to receive service from the contact center system 200 may initiate inbound communications (e.g., phone calls, emails, chats, etc.) to the contact center system 200 via a customer device 205. While FIG. 2 shows one such customer device, customer device 205, it should be understood that any number of customer devices 205 may be present. The customer device 205 may be, for example, a communication device such as a telephone, smartphone, computer, tablet, or laptop. According to the functionality described herein, a customer may generally use the customer device 205 to initiate, manage, and conduct communications with the contact center system 200, such as telephone calls, emails, chats, text messages, web browsing sessions, and other multimedia transactions.

[0107] Inbound and outbound communications from and to customer device 205 may typically traverse network 210, with the nature of the network depending on the type of customer device and mode of communication being used. By way of example, network 210 may include telephone, cellular, and / or data service communication networks. Network 210 may be a private or public switched telephone network (PSTN), a local area network (LAN), a private wide area network (WAN), and / or a public WAN such as the Internet. Additionally, network 210 may include a wireless carrier network, including a code division multiple access (CDMA) network, a global system for mobile communications (GSM) network, or any wireless network / technology conventional in the art, including, but not limited to, 3G, 4G, LTE, 5G, etc.

[0108] The switch / media gateway 212 may be coupled to the network 210 to receive and transmit telephone calls between customers and the contact center system 200. The switch / media gateway 212 may include a telephone switch or a communication switch configured to function as a central switch for agent-level routing within the center. The switch may be a hardware switching system or may be implemented via software. For example, the switch 212 may include an automatic call distributor, a private branch exchange (PBX), an IP-based software switch, and / or any other switch having dedicated hardware and software configured to receive Internet-sourced and / or telephone network-sourced interactions from customers and route these interactions to, for example, one of the agent devices 230. Thus, in general, the switch / media gateway 212 establishes a connection between the customer device 205 and the agent device 230, thereby establishing a voice connection between the customer and the agent.

[0109] As further shown, the switch / media gateway 212 may be coupled to a call controller 214, which functions, for example, as an adapter or interface between the switch and other routing, monitoring, and communication processing components of the contact center system 200. The call controller 214 may be configured to process PSTN calls, VoIP calls, and / or other types of calls. For example, the call controller 214 may include computer-telephone integration (CTI) software for interfacing with the switch / media gateway and other components. The call controller 214 may include a session initiation protocol (SIP) server for processing SIP calls. The call controller 214 may also extract data about incoming interactions, such as the customer's phone number, IP address, or email address, and then communicate this with other contact center components when processing the interaction.

[0110] The interactive media response (IMR) server 216 can be configured to enable self-help or virtual assistant functions. Specifically, the IMR server 216 can be similar to an interactive voice response (IVR) server, except that the IMR server 216 is not limited to voice and can cover a variety of media channels. In an example illustrating voice, the IMR server 216 can be configured with IMR scripts to query customers about their needs. For example, a bank contact center may instruct customers via an IMR script to "press 1" if they want to look up their account balance. Through ongoing interaction with the IMR server 216, customers can receive service without having to speak with an agent. The IMR server 216 can also be configured to verify the reason the customer is contacting the contact center so that the communication can be routed to the appropriate resource. IMR configuration can be implemented through the use of self-service and / or assisted-service tools, including web-based tools for developing IVR and routing applications that run within the contact center environment (e.g., Genesys® Designer).

[0111] The routing server 218 may function to route incoming interactions. For example, once it is determined that an inbound communication should be handled by a human agent, functionality within the routing server 218 may select the most appropriate agent and route the communication to that agent. This agent selection may be based on which available agent is best suited to handle the communication. More specifically, the selection of the appropriate agent may be based on a routing strategy or algorithm implemented by the routing server 218. In doing so, the routing server 218 may query data related to the incoming interaction, such as data related to the particular customer, available agents, and type of interaction, which may be stored in a particular database, as described herein. Once an agent is selected, the routing server 218 may interact with the call controller 214 to route (i.e., connect) the incoming interaction to a corresponding agent device 230. As part of this connection, information about the customer may be provided to the selected agent via their agent device 230. This information is intended to enhance the service the agent can provide to the customer.

[0112] It should be appreciated that the contact center system 200 may include one or more mass storage devices (generally represented by the storage device 220) for storing data in one or more databases related to the contact center's functions. For example, the storage device 220 may store customer data maintained in a customer database. Such customer data may include, for example, customer profiles, contact information, service level agreements (SLAs), and interaction history (e.g., details of previous interactions with particular customers, including the nature of the previous interactions, disposition data, wait times, handling times, and actions taken by the contact center to resolve the customer's issues). As another example, the storage device 220 may store agent data in an agent database. The agent data maintained by the contact center system 200 may include, for example, agent availability and agent profiles, schedules, skills, handling times, and / or other related data. As another example, the storage device 220 may store interaction data in an interaction database. The interaction data may include, for example, data related to numerous past interactions between customers and the contact center. More generally, unless otherwise specified, it should be understood that storage device 220 may include databases and / or be configured to store data related to any of the types of information described herein, and that these databases and / or data may be accessible to other modules or servers of contact center system 200 in a manner that facilitates the functions described herein. For example, a server or module of contact center system 200 may query such databases to retrieve data stored therein or transmit data to a database for storage. Storage device 220 may take the form of, for example, any conventional storage medium and may be housed locally or operated from a remote location.By way of example, the database may be a Cassandra database, a NoSQL database, or a SQL database, and may be managed by a database management system such as Oracle, IBM DB2, Microsoft SQL Server, Microsoft Access, PostgreSQL, or the like.

[0113] Statistics server 226 may be configured to record and aggregate data related to the performance and operational aspects of contact center system 200. Such information may be compiled by statistics server 226 and made available to other servers and modules, such as reporting server 248, which may then use the data to generate reports used to manage operational aspects of the contact center and perform automated actions in accordance with the functionality described herein. Such data may relate to the status of contact center resources, such as average wait times, abandon rates, agent occupancy rates, and others as may be required by the functionality described herein.

[0114] The agent devices 230 of the contact center system 200 may be communication devices configured to interact with the various components and modules of the contact center system 200 in a manner that facilitates the functionality described herein. For example, the agent devices 230 may include telephones adapted for regular telephone calls or VoIP calls. The agent devices 230 may further include computing devices configured to communicate with the servers of the contact center system 200, perform data processing associated with operations, and interface with customers via voice, chat, email, and other multimedia communication mechanisms in accordance with the functionality described herein. While FIG. 2 shows three such agent devices 230, namely, agent devices 230A, 230B, and 230C, it should be understood that any number of agent devices 230 may be present in certain embodiments.

[0115] The multimedia / social media server 234 may be configured to facilitate media interactions (other than voice) with the customer device 205 and / or server 242. Such media interactions may relate to, for example, email, voicemail, chat, video, text messaging, web, social media, co-browsing, etc. The multimedia / social media server 234 may take the form of any IP router conventional in the art having dedicated hardware and software for receiving, processing, and forwarding multimedia events and communications.

[0116] The knowledge management server 236 may be configured to facilitate interactions between customers and the knowledge system 238. Generally, the knowledge system 238 may be a computer system capable of receiving questions, i.e., queries, and providing answers in response. The knowledge system 238 may be included as part of the contact center system 200 or may be operated remotely by a third party. The knowledge system 238 may include an artificial intelligence computer system capable of answering questions posed in natural language by retrieving information from sources such as encyclopedias, dictionaries, newswire articles, literary works, or other documents submitted to the knowledge system 238 as reference material. As an example, the knowledge system 238 may be embodied as an IBM Watson or similar system.

[0117] Chat server 240 may be configured to conduct, orchestrate, and manage electronic chat communications with customers. Generally, chat server 240 is configured to conduct and maintain chat conversations and generate chat transcripts. Such chat communications may be conducted by chat server 240 in a manner such that customers communicate with automated chatbots, human agents, or both. In an exemplary embodiment, chat server 240 may function as a chat orchestration server that dispatches chat conversations between chatbots and available human agents. In such cases, the processing logic of chat server 240 may be rule-driven to leverage intelligent workload distribution among available chat resources. Chat server 240 may also implement, manage, and facilitate user interfaces (UIs) associated with the chat feature, including user interfaces (UIs) generated on either customer device 205 or agent device 230. Chat server 240 may be configured to transfer chats between automated and human sources within a single chat session with a particular customer, for example, so that the chat session moves from a chatbot to a human agent or from a human agent to a chatbot. The chat server 240 may also be coupled to the knowledge management server 236 and the knowledge system 238 to receive suggestions and answers to queries posed by guests during chat, for example, so that links to related articles can be provided.

[0118] Web server 242 may be included to provide site hosts for various social interaction sites to which customers subscribe, such as Facebook, Twitter, and Instagram. While shown as part of contact center system 200, it should be understood that web server 242 may be provided by a third party and / or maintained remotely. Web server 242 may also host web pages for businesses or organizations supported by contact center system 200. For example, customers may browse web pages to receive information about a particular business's products and services. Within such business web pages, mechanisms may be provided for initiating interactions with contact center system 200, for example, via web chat, voice, or email. One example of such a mechanism is a widget that may be deployed on a web page or website hosted on web server 242. As used herein, a widget refers to a user interface component that performs a specific function. In some implementations, a widget may include a graphical user interface control that may be overlaid on a web page displayed to customers over the Internet. A widget may show information, such as in a window or text box, or include buttons or other controls that allow a user to access a particular function, such as sharing or opening a file or initiating a communication. In some implementations, a widget includes a user interface component with a portable portion of code that can be installed and executed within a separate web page without compilation. Some widgets may include corresponding or additional user interfaces and may be configured to access various local resources (e.g., calendar or contact information on the customer device) or remote resources over a network (e.g., instant messaging, email, or social networking updates).

[0119] The interaction (iXn) server 244 may be configured to manage contact center deferrable activities and their routing to human agents for completion. As used herein, deferrable activities include back-office work that can be performed offline, such as responding to email, attending training, and other activities that do not require real-time communication with customers. As an example, the interaction (iXn) server 244 may be configured to interact with the routing server 218 to select an appropriate agent to handle each deferrable activity. Once assigned to a particular agent, the deferrable activity is pushed to that agent, and as a result, the deferrable activity is displayed on the selected agent's agent device 230. The deferrable activity may be displayed in a work bin as a task for the selected agent to complete. The work bin functionality may be implemented via any conventional data structure, such as a linked list, an array, and / or other suitable data structure. Each of the agent devices 230 may include a work bin. As an example, the work bin may be maintained in a buffer memory of the corresponding agent device 230.

[0120] A universal contact server (UCS) 246 may be configured to retrieve information stored in a customer database and / or send information to a customer database for storage in the customer database. For example, UCS 246 may be utilized as part of a chat function to facilitate maintaining a history of how chats with particular customers were handled, which may then be used as a reference for how future chat communications should be handled. More generally, UCS 246 may be configured to facilitate maintaining a history of customer preferences, such as preferred media channels and best times to contact. To do this, UCS 246 may be configured to identify data related to each customer's interaction history, such as data regarding comments from agents, customer communication history, etc. Each of these data types may then be stored in customer database 222 or other modules and retrieved as needed by the functionality described herein.

[0121] Reporting server 248 may be configured to generate reports from data compiled and aggregated by statistics server 226 or other sources. Such reports may include near real-time or historical reports and may relate to the status of contact center resources and performance characteristics, such as average wait times, abandon rates, and / or agent occupancy. Reports may be generated automatically or in response to specific requests from requestors (e.g., agents / managers, contact center applications, etc.). The reports may then be used to manage the operation of the contact center in accordance with the functionality described herein.

[0122] Media services server 249 may be configured to provide audio and / or video services to support contact center functions, such as IVR or IMR system prompts (e.g., playing audio files), music on hold, voicemail / single-party recording, multi-party recording (e.g., of audio and / or video calls), speech recognition, dual tone multi-frequency (DTMF) recognition, fax, audio and video transcoding, secure real-time transport protocol (SRTP), audio conferencing, video conferencing, coaching (e.g., support for a coach to eavesdrop on an interaction between a customer and an agent and for the coach to provide comments to an agent without the customer hearing the comments), call analysis, keyword spotting, and / or other related functions, according to the functionality described herein.

[0123] The analytics module 250 may be configured to provide a system and method for performing analytics on data received from multiple different data sources, as may be required by the functionality described herein. According to an exemplary embodiment, the analytics module 250 may also generate, update, train, and modify predictors or models based on collected data, such as, for example, customer data, agent data, and interaction data. The models may include customer or agent behavior models. The behavior models may be used to predict, for example, customer or agent behavior in various situations, thereby enabling embodiments of the present invention to adjust interactions based on such predictions or allocate resources in preparation for predicted characteristics of future interactions, thereby improving overall contact center performance and customer experience. While the analytics module is described as being part of the contact center, it will be understood that such behavior models may also be implemented in customer systems (or, as used herein, the “customer side” of an interaction) and used to the benefit of the customer.

[0124] According to an example embodiment, analytics module 250 may have access to data stored in storage device 220, including a customer database and an agent database. Analytics module 250 may also have access to an interaction database that stores data related to interactions and interaction content (e.g., transcripts of interactions and events detected therein), interaction metadata (e.g., customer identifier, agent identifier, interaction medium, interaction length, interaction start and end times, department, tagged categories), and application settings (e.g., interaction path through the contact center). Additionally, analytics module 250 may be configured to search the data stored in storage device 220 for use in developing and training algorithms and models, for example, by applying machine learning techniques.

[0125] One or more of the included models may be configured to predict customer or agent behavior and / or aspects related to contact center operation and performance. Additionally, one or more of the models may be used for natural language processing, including, for example, intent recognition. The models may be developed based on known first-principles equations describing the system, data resulting in empirical models, or a combination of known first-principles equations and data. When developing models for use in the present embodiments, first-principles equations are often not available or easily derived, so building empirical models based on collected and stored data may generally be preferred. To adequately capture the relationships between manipulated / disturbance variables and controlled variables of a complex system, in some embodiments, it may be preferable for the model to be nonlinear. This is because nonlinear models may exhibit curvilinear relationships between manipulated / disturbance variables and controlled variables rather than the linear relationships common in complex systems such as those discussed herein. Given the aforementioned requirements, machine learning or neural network-based approaches may be preferred for implementing the models. For example, neural networks may be developed based on empirical data using advanced regression algorithms.

[0126] The analysis module 250 may further include an optimizer. As will be appreciated, an optimizer may be used to minimize a "cost function" to which a set of constraints are applied, where the cost function is a mathematical expression of a desired objective or system behavior. Because the model may be nonlinear, the optimizer may be a nonlinear programming optimizer. However, it is contemplated that the techniques described herein may be implemented using a variety of different types of optimization approaches, individually or in combination, including, but not limited to, linear programming, quadratic programming, mixed-integer nonlinear programming, stochastic programming, global nonlinear programming, genetic algorithms, particle / swarm techniques, etc.

[0127] According to some embodiments, the model and optimizer may be used together in an optimization system. For example, analytics module 250 may utilize the optimization system as part of an optimization process in which aspects of contact center performance and operation are optimized or at least enhanced. This may include, for example, features related to customer experience, agent experience, interaction routing, natural language processing, intent recognition, or other functionality related to automated processes.

[0128] The various components, modules, and / or servers in FIG. 2 (as well as other figures contained herein) may each include one or more processors that execute computer program instructions and interact with other system components to perform the various functions described herein. Such computer program instructions may be stored in memory implemented using standard memory devices such as, for example, random-access memory (RAM), or may be stored on other non-transitory computer-readable media such as, for example, a CD-ROM, a flash drive, or the like. While each of the server functions is described as being provided by a particular server, those skilled in the art should understand that the functions of various servers may be combined or integrated into a single server, or that the functions of a particular server may be distributed across one or more other servers without departing from the scope of the present invention. Furthermore, the terms “interaction” and “communication” are used interchangeably and generally refer to any real-time and non-real-time interaction using any communication channel, including, but not limited to, phone calls (PSTN or VoIP calls), email, Vmail, video, chat, screen sharing, text messages, social media messages, WebRTC calls, etc. Access to and control of components of contact center system 200 may be affected through user interfaces (UIs), which may be generated on customer devices 205 and / or agent devices 230. As previously mentioned, contact center system 200 may be operated as a hybrid system in which some or all components are hosted remotely, such as in a cloud-based or cloud computing environment. It should be understood that each of the devices of contact center system 200 may be embodied as part of, include, or form part of one or more computing devices similar to computing device 300 described below with reference to FIG. 3.

[0129] 3, a simplified block diagram of at least one embodiment of a computing device 300 is shown. The exemplary computing device 300 illustrates at least one embodiment of each of the computing devices, systems, servicers, controllers, switches, gateways, engines, modules, and / or computing components (e.g., which for brevity herein may be interchangeably referred to as computing devices, servers, or modules) described herein. For example, various computing devices may be processes or threads running on one or more processors of one or more computing devices 300, which may execute computer program instructions and interact with other system modules to perform various functions described herein. Unless otherwise limited, functionality described in the context of multiple computing devices may be integrated into a single computing device, or various functionality described in the context of a single computing device may be distributed across several computing devices. Additionally, with respect to a computing system described herein, such as the contact center system 200 of FIG. 2, the various servers and computer devices of that system may be located on a local computing device 300 (e.g., on-site in the same physical location as the contact center agents), a remote computing device 300 (e.g., off-site, i.e., in a cloud-based environment, or in a cloud computing environment, e.g., in a remote data center connected via a network), or some combination thereof.In some embodiments, functionality provided by servers located on off-site computing devices may be accessed and provided via a virtual private network (VPN) as if such servers were on-site, or functionality may be provided using Software as a Service (SaaS) accessed over the Internet using various protocols, e.g., by exchanging data via extensible markup language (XML), JSON, and / or functionality may be accessed / utilized in other ways.

[0130] In some embodiments, computing device 300 may be embodied as a server, a desktop computer, a laptop computer, a tablet computer, a notebook, a netbook, an Ultrabook™, a mobile phone, a mobile computing device, a smartphone, a wearable computing device, a personal digital assistant, an Internet of Things (IoT) device, a processing system, a wireless access point, a router, a gateway, and / or any other computing, processing, and / or communication device capable of performing the functions described herein.

[0131] The computing device 300 includes a processing device 302 that executes algorithms and / or processes data according to operational logic 308, an input / output device 304 that enables communication between the computing device 300 and one or more external devices 310, and a memory 306 that stores data received from the external device 310 via, for example, the input / output device 304.

[0132] The input / output devices 304 enable the computing device 300 to communicate with external devices 310. For example, the input / output devices 304 may include a transceiver, a network adapter, a network card, an interface, one or more communication ports (e.g., a USB port, a serial port, a parallel port, an analog port, a digital port, VGA, DVI, HDMI, FireWire, CAT5, or any other type of communication port or interface), and / or other communication circuitry. The communication circuitry of the computing device 300 may be configured to perform such communication using any one or more communication technologies (e.g., wireless or wired communication) and associated protocols (e.g., Ethernet, Bluetooth, Wi-Fi, WiMAX, etc.), depending on the particular computing device 300. The input / output devices 304 may include hardware, software, and / or firmware suitable for performing the techniques described herein.

[0133] External device 310 may be any type of device that allows data to be input or output from computing device 300. For example, in various embodiments, external device 310 may be embodied as one or more of the devices / systems described herein and / or portions thereof. Further, in some embodiments, external device 310 may be embodied as another computing device, a switch, a diagnostic tool, a controller, a printer, a display, an alarm, a peripheral device (e.g., a keyboard, a mouse, a touchscreen display, etc.), and / or any other computing, processing, and / or communication device capable of performing the functionality described herein. Furthermore, it should be understood that in some embodiments, external device 310 may be integrated into computing device 300.

[0134] The processing device 302 may be embodied as any type of processor capable of performing the functions described herein. In particular, the processing device 302 may be embodied as one or more single-core or multi-core processors, microcontrollers, or other processors or processing / control circuitry. For example, in some embodiments, the processing device 302 may include or be embodied as an arithmetic logic unit (ALU), a central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or another suitable processor. The processing device 302 may be of a programmable type, a dedicated hardwired state machine, or a combination thereof. A processing device 302 having multiple processing units may utilize distributed processing, pipelined processing, and / or parallel processing in various embodiments. Furthermore, processing device 302 may be dedicated to performing only the operations described herein, or may be utilized in one or more additional applications. In an exemplary embodiment, processing device 302 is programmable, executing algorithms and / or processing data according to operating logic 308 defined by programming instructions (e.g., software or firmware) stored in memory 306. Additionally or alternatively, operating logic 308 of processing device 302 may be defined at least in part by hardwired logic or other hardware. Furthermore, processing device 302 may include one or more components of any type suitable for processing signals received from input / output device 304 or from other components or devices and providing a desired output signal. Such components may include digital circuits, analog circuits, or a combination thereof.

[0135] Memory 306 may be one or more types of non-transitory computer-readable media, such as solid-state memory, electromagnetic memory, optical memory, or a combination thereof. Furthermore, memory 306 may be volatile and / or non-volatile, and in some embodiments, some or all of memory 306 may be portable, such as a disk, tape, memory stick, cartridge, and / or other suitable portable memory. During operation, memory 306 may store various data and software used during operation of computing device 300, such as an operating system, applications, programs, libraries, and drivers. It should be understood that memory 306 may store data representing signals received from and / or sent to input / output devices 304 in addition to, or instead of, storing data manipulated by operating logic 308 of processing device 302, e.g., programming instructions defining operating logic 308. As shown in FIG. 3 , memory 306 may be included in and / or coupled to processing device 302, depending on the particular embodiment. For example, in some embodiments, the processing device 302, memory 306, and / or other components of the computing device 300 may form part of a system on a chip (SoC) and be integrated into a single integrated circuit chip.

[0136] In some embodiments, various components of computing device 300 (e.g., processing device 302 and memory 306) may be communicatively coupled via an input / output subsystem, which may be embodied as circuits and / or components for facilitating input / output operations with processing device 302, memory 306, and other components of computing device 300. For example, the input / output subsystem may be embodied as or may include a memory controller hub, an input / output control hub, firmware devices, communication links (i.e., point-to-point links, bus links, wires, cables, light guides, printed circuit board traces, etc.), and / or other components and subsystems for facilitating input / output operations.

[0137] Computing device 300, in other embodiments, may include other or additional components, such as those commonly found in a typical computing device (e.g., various input / output devices and / or other components). It should further be understood that one or more of the components of computing device 300 described herein may be distributed across multiple computing devices. In other words, the techniques described herein may be employed by a computing system including one or more computing devices. Furthermore, while only a single processing device 302, I / O device 304, and memory 306 are illustratively shown in FIG. 3 , it should be understood that in other embodiments, a particular computing device 300 may include multiple processing devices 302, I / O devices 304, and / or memory 306. Furthermore, in some embodiments, multiple external devices 310 may communicate with computing device 300.

[0138] The computing device 300 may be one of multiple devices connected by a network or to other systems / resources via a network. The network may be embodied as any one or more types of communications network capable of facilitating communication between various devices communicatively connected via the network. Thus, the network may include one or more networks, routers, switches, access points, hubs, computers, client devices, endpoints, and / or other intervening network devices. For example, the network may be embodied as or otherwise include one or more cellular networks, telephone networks, local or wide area networks, publicly available global networks (e.g., the Internet), ad hoc networks, short range communications links, or combinations thereof. In some embodiments, the network may include circuit-switched voice or data networks, packet-switched voice or data networks, and / or any other network capable of carrying voice and / or data. In particular, in some embodiments, the network may include an Internet Protocol (IP)-based network and / or an asynchronous transfer mode (ATM)-based network. In some embodiments, the network may handle voice traffic (e.g., via a Voice over IP (VOIP) network), web traffic, and / or other network traffic, depending on the particular embodiment and / or the devices in the system that communicate with each other.In various embodiments, the networks may include analog or digital wired and wireless networks (e.g., IEEE 802.11 networks, Public Switched Telephone Networks (PSTN), Integrated Services Digital Networks (ISDN), and Digital Subscriber Lines (xDSL)), Third Generation (3G) mobile networks, Fourth Generation (4G) mobile networks, Fifth Generation (5G) mobile networks, wired Ethernet networks, private networks (e.g., intranets, etc.), radio, television, cable, satellite, and / or any other distribution or tunneling mechanism for carrying data, or any suitable combination of such networks. It should be understood that various devices / systems may communicate with each other over different networks depending on the source and / or destination device / system.

[0139] It should be understood that computing device 300 may communicate with other computing devices 300 via any type of gateway or tunneling protocol, such as Secure Sockets Layer or Transport Layer Security. Network interfaces may include built-in network adapters, such as network interface cards, suitable for interfacing a computing device to any type of network capable of performing the operations described herein. Furthermore, the network environment may be a virtual network environment in which various network components are virtualized. For example, various machines may be virtual machines implemented as software-based computers running on a physical machine. The virtual machines may share the same operating system, or in other embodiments, different operating systems may run on each virtual machine instance. For example, a "hypervisor" type of virtualization is used in which multiple virtual machines run on the same host physical machine, each functioning as if it had its own dedicated box. Other types of virtualization may be employed in other embodiments, such as for networks (e.g., via software-defined networking) or functions (e.g., via network function virtualization).

[0140] Accordingly, one or more of the computing devices 300 described herein may be embodied as or form a part of one or more cloud-based systems. In cloud-based embodiments, the cloud-based system may be embodied, for example, as a server-ambiguous computing solution that executes instructions on demand, executes instructions only when prompted by specific activities / triggers, and does not consume computing resources when not in use. That is, the system may be embodied as a virtual computing environment residing “on” a computing system (e.g., a distributed network of devices) in which various virtual functions (e.g., Lambda functions, Azure functions, Google Cloud Functions, and / or other suitable virtual functions) may be executed corresponding to the functionality of the system described herein. For example, when an event occurs (e.g., data is transferred to the system for processing), the virtual computing environment may be communicated (e.g., via a request to the virtual computing environment's API), which may then route the request to the correct virtual function (e.g., a particular server-ambiguous computing resource) based on a set of rules. Thus, when a request to transmit data is made by a user (e.g., via an appropriate user interface to the system), an appropriate virtual function may be executed to perform the action before deleting the instance of the virtual function.

Claims

1. 1. A contact center system for performing call progress analysis using tone and voice classification, comprising: at least one processor; at least one memory including a plurality of instructions stored therein, the instructions, in response to execution by the at least one processor, for causing the contact center system to: Calculating instantaneous entropy from the power spectrum amplitude of the audio signal received by the contact center system at one timing; determining an average entropy by calculating a cumulative average of a plurality of said instantaneous entropies determined over a predetermined period of time; determining a cumulative average power spectral amplitude of the audio signal, which is an average of the power spectral amplitudes of the audio signal over a predetermined period of time; and determining a cumulative average spectral entropy of the audio signal, which is an entropy based on the cumulative average power spectral amplitude of the audio signal; calculating a difference measure of the audio signal as the difference between the average entropy of the audio signal and the cumulative average spectral entropy of the audio signal; distinguishing between tones and voices in the audio signal by determining portions of the difference measure of the audio signal that are below a predetermined threshold as tones in the audio signal and portions that are equal to or above the predetermined threshold as voices in the audio signal; at least one memory that, in response to identifying one or more tones in the audio signal, causes the one or more tones of the audio signal to be processed; A contact center system comprising:

2. processing the one or more tones of the audio signal, identifying a call progress tone pattern within the one or more tones of the audio signal; and transferring a telephone call from a first one of the contact center systems to a second one of the contact center systems in response to identifying the call progress tone pattern within the one or more tones of the audio signal; 2. The contact center system according to claim 1, wherein the first system is a system that performs processing to identify the call progress tone pattern, and the second system is a system different from the first system that determines the call progress tone pattern according to the identification result.

3. 10. The contact center system of claim 1, wherein processing the one or more tones of the audio signal includes connecting an outbound call to an automated interactive voice response (IVR) system of the contact center system.

4. The contact center system of claim 1 , wherein processing the one or more tones of the audio signal includes connecting an outbound call to an agent of the contact center system.

5. The contact center system of claim 1 , wherein the one or more tones of the audio signal include a call progress tone pattern.

6. The contact center system of claim 5 , wherein the call progress tone pattern includes one of a busy signal pattern, a ringback pattern, or a special information tone pattern.

7. The contact center system of claim 1 , wherein processing the one or more tones of the audio signal comprises determining a corresponding frequency of each of the one or more tones of the audio signal.

8. One or more non-transitory machine-readable storage media having a plurality of instructions stored thereon, the instructions, in response to execution by at least one processor, causing a contact center system to: Calculating instantaneous entropy from the power spectrum amplitude of the audio signal received by the contact center system at one timing; determining an average entropy by calculating a cumulative average of a plurality of said instantaneous entropies determined over a predetermined period of time; calculating a cumulative average power spectral amplitude by averaging the power spectral amplitude of the audio signal over a predetermined period of time; calculating a cumulative average spectral entropy of the audio signal, the entropy being based on the cumulative average power spectral amplitude of the audio signal; calculating a difference measure for the audio signal as the difference between the average entropy of the audio signal and the cumulative average spectral entropy of the audio signal; classifying the tones and voices of the audio signal by determining portions of the difference measure of the audio signal that are below a predetermined threshold as tones of the audio signal and portions that are above the predetermined threshold as voices of the audio signal; one or more non-transitory machine-readable storage media that, in response to identifying one or more tones in the audio signal, cause the one or more tones of the audio signal to be processed.

9. processing the one or more tones of the audio signal includes transferring a telephone call from a first one of the contact center systems to a second one of the contact center systems in response to identifying a call progress tone pattern within the one or more tones of the audio signal; 9. The one or more non-transitory machine-readable storage media of claim 8, wherein the first system is a system that performs processing to identify the call progress tone pattern, and the second system is a system different from the first system that determines the call progress tone pattern according to the identification result.

10. 10. The one or more non-transitory machine-readable storage media of claim 8, wherein processing the one or more tones of the audio signal includes connecting an outbound call to an automated interactive voice response (IVR) system of the contact center system.

11. 10. The one or more non-transitory machine-readable storage media of claim 8, wherein processing the one or more tones of the audio signal comprises connecting an outbound call to an agent of the contact center system.

12. 10. The one or more non-transitory machine-readable storage media of claim 8, wherein the one or more tones of the audio signal include a call progress tone pattern.

13. 13. The one or more non-transitory machine-readable storage media of claim 12, wherein the call progress tone pattern comprises one of a busy signal pattern, a ringback pattern, or a special information tone pattern.

14. 10. The one or more non-transitory machine-readable storage media of claim 8, wherein processing the one or more tones of the audio signal comprises determining a corresponding frequency of each of the one or more tones of the audio signal.

15. 1. A method for performing call progress analysis using tone and voice classification in a contact center system, comprising: receiving an audio signal by the contact center system; determining, by the contact center system, instantaneous entropy from the power spectrum amplitude of the audio signal at one time point received by the contact center system; determining an average entropy by calculating a cumulative average of a plurality of the instantaneous entropies determined by the contact center system over a predetermined period of time; determining, by the contact center system, a cumulative average power spectral amplitude by averaging the power spectral amplitude of the audio signal over a predetermined period of time; determining, by the contact center system, a cumulative average spectral entropy of the audio signal, the entropy being based on the cumulative average power spectral amplitude of the audio signal; determining, by the contact center system, a difference measure of the audio signal as the difference between the average entropy of the audio signal and the cumulative average spectral entropy of the audio signal; classifying, by the contact center system, the tones and voices of the audio signal by determining portions of the difference measure of the audio signal that are below a predetermined threshold as tones of the audio signal and portions that are equal to or above the predetermined threshold as voices of the audio signal; and processing, by the contact center system, one or more tones of the audio signal in response to identifying the one or more tones in the audio signal.

16. processing the one or more tones of the audio signal, identifying a call progress tone pattern within the one or more tones of the audio signal; and transferring a telephone call from a first one of the contact center systems to a second one of the contact center systems in response to identifying a call progress tone pattern within the one or more tones of the audio signal; 16. The method of claim 15, wherein the first system is a system that performs processing to identify the call progress tone pattern, and the second system is a system different from the first system that determines the call progress tone pattern according to the identification result.

17. 16. The method of claim 15, wherein processing the one or more tones of the audio signal includes connecting an outbound call to one of an agent or an automated interactive voice response (IVR) system of the contact center system.

18. The method of claim 15 , wherein the one or more tones of the audio signal include a call progress tone pattern.

19. 16. The method of claim 15, wherein processing the one or more tones of the audio signal comprises determining a corresponding frequency of each of the one or more tones of the audio signal.

Citation Information

Patent Citations

  • Apparatus and method for analyzing progress of call

    JP1987166640A

  • Voice section determination device, voice section determination method, and program

    JP2012215600A

  • Voice detection method, device and storage medium

    JP2018532155A

  • Call progress analysis on the edge of a VOIP network

    US20100189249A1

  • Detection of signal tone in audio signal

    US20200184995A1