Optimization for Calls Waiting in the Queue
A software module adjusts device volume and alerts users when staff is available, enhancing the call center experience by enabling multitasking and timely engagement.
Patent Information
- Application Number
- JP2024033199
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-13
- Filing Date
- 2024-03-05
- Publication Date
- 2025-07-09
- Estimated Expiration
- 2040-05-26
AI Technical Summary
Users waiting for staff service in call centers often experience long wait times and repetitive prompts, which can be boring and may cause them to miss important announcements if they are not paying attention.
A software module that adjusts the device volume when repetitive prompts are detected and alerts the user when a staff member becomes available, allowing the user to engage promptly.
Improves user experience by allowing users to multitask during wait times and ensuring they are alerted when staff is ready to assist, reducing the likelihood of missed announcements.
Smart Images

Figure 0007705215000001 
Figure 0007705215000002 
Figure 0007705215000003
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer software. More particularly, embodiments relate to a method, system, and computer program product for managing calls with respect to calls waiting in a queue.
Background Art
[0002] Today, call centers are widely used in many industries, such as the financial industry and other service industries.
Summary of the Invention
[0003] In one aspect, a method for managing calls to a call center waiting in a queue when attempting to receive service from a member of the call center staff is disclosed. According to the method, a first voice segment received in a call made by a device is first recorded. Next, it is determined whether a portion of the first voice segment is related to a first predefined voice segment. Finally, in response to the portion of the first voice segment being related to the first predefined voice segment, the volume of the device is adjusted, while in response to the portion of the first voice segment not being related to the first predefined voice segment, an alert is issued to the user of the device.
[0004] In another aspect, a system implemented by a computer is disclosed. The system may include a computer processor coupled to a computer-readable memory unit, the memory unit comprising instructions for implementing the aforementioned method when executed by the computer processor.
[0005] In yet another aspect, a computer program product is disclosed. The computer program product comprises a computer-readable storage medium having program instructions embodied thereon. When executed on one or more processors, the instructions may cause the one or more processors to perform the methods described above.
[0006] Through the more detailed description of some embodiments of the present disclosure in the accompanying drawings, the above as well as other objects, features, and advantages of the present disclosure will become clearer. In the drawings, the same reference numerals generally refer to the same components in the embodiments of the present disclosure.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0008] Some embodiments are described in more detail with reference to the accompanying drawings in which embodiments of the present disclosure are illustrated. However, the present disclosure can be implemented in various ways and should not be construed as being limited to the embodiments disclosed herein.
[0009] Although this disclosure includes a detailed description regarding cloud computing, it should be understood that the implementation of the teachings described herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment, now known or later developed.
[0010] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a service provider. This cloud model can include at least five characteristics, at least three service models, and at least four deployment models.
[0011] The characteristics are as follows.
[0012] On-demand self-service Cloud consumers can unilaterally provision computing capabilities such as server time and network storage as needed, automatically, without the need for human interaction with the service provider.
[0013] Broad network access Capabilities are available over the network and accessed through a standard mechanism that promotes use by heterogeneous thin-client platforms or thick-client platforms (e.g., mobile phones, laptops, and PDAs).
[0014] The computing resources of a resource pooling provider are pooled so that they can be utilized by multiple consumers using a multi-tenant model, and various physical and virtual resources are dynamically allocated and reallocated according to demand. Consumers generally do not control or know the exact location of the provided resources, but there is a sense of location independence in that they can specify the location at a higher level of abstraction (e.g., country, state, or data center).
[0015] Rapid elasticity The ability can be provisioned quickly, elastically, and in some cases automatically so that it can scale out rapidly and be released quickly to scale in rapidly. To the consumer, the available capacity for provisioning often appears to be infinite and can be purchased in any quantity at any time.
[0016] Measured service The cloud system automatically controls and optimizes resource usage by leveraging metering capabilities at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, reported, and made transparent to both the provider and consumer of the utilized service.
[0017] The service model is as follows.
[0018] The capabilities provided to Software as a Service (SaaS) consumers are to use the provider's applications that run on cloud infrastructure. Those applications are accessible from various client devices via a thin client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even the individual application capabilities, with the possible exception of limited user-specific application configurations.
[0019] The capabilities provided to Platform as a Service (PaaS) consumers are to deploy on cloud infrastructure the applications created by, acquired by, or created by the consumers using the programming languages and programming tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including the network, servers, operating system, or storage, but govern the deployed applications and, optionally, the application hosting environment configuration.
[0020] The capabilities provided to Infrastructure as a Service (IaaS) consumers are to provision processing, storage, networks, and other basic computing resources on which consumers can deploy and run any software that may include an operating system and applications. Consumers do not manage or control the underlying cloud infrastructure, but govern the operating system, storage, deployed applications, and, optionally, have limited control over selected networking components (e.g., host firewalls).
[0021] The deployment model is as follows.
[0022] Private Cloud The cloud infrastructure is operated exclusively for an organization. The cloud infrastructure may be managed by the organization or by a third party and may exist on-premises or off-premises.
[0023] Community Cloud The cloud infrastructure is shared by several organizations and supports a specific community with shared interests (e.g., missions, security requirements, policies, and compliance considerations). The cloud infrastructure may be managed by the organization or by a third party and may exist on-premises or off-premises.
[0024] Public Cloud The cloud infrastructure is made available for use by the general public or a large industry group and is owned by an organization that sells cloud services.
[0025] Hybrid Cloud The cloud infrastructure remains a distinct entity but is a composition of two or more clouds (private, community, or public) tied together by standardized or proprietary technologies (e.g., cloud bursting for load balancing between clouds) that enable data portability and application portability.
[0026] The cloud computing environment focuses on the stateless nature, loose coupling, modularity, and semantic interoperability and is service-oriented. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0027] Next, referring to FIG. 1, a schematic diagram of an embodiment of a cloud computing node is shown. Cloud computing node 10 is merely one example of a suitable cloud computing node and is not intended to suggest any limitation as to the scope of use or functionality of the embodiments described herein. In any event, cloud computing node 10 can be implemented as or perform any of the functions shown above, or any combination thereof.
[0028] In cloud computing node 10, there is a computer system / server 12, such as a communication device, or a portable electronic device, that is operable in many other general purpose or special purpose computing system environments or computing system configurations. Examples of well-known computing systems, computing environments, or computing system configurations, or combinations thereof, that may be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld devices or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the foregoing systems or devices, and the like.
[0029] Computer system / server 12 may be described in the general context of computer system executable instructions, such as program modules, being executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer system / server 12 may be implemented in a distributed cloud computing environment where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules may be located in both a local computer system storage medium including a memory storage device and a remote computer system storage medium.
[0030] As shown in FIG. 1, computer system / server 12 in cloud computing node 10 is shown in the form of a general purpose computing device. The components of computer system / server 12 may include, but are not limited to, one or more processors or processing devices 16, a system memory 28, and a bus 18 that couples various system components including system memory 28 to processor 16.
[0031] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor bus or local bus, using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0032] Computer system / server 12 typically includes various computer system-readable media. Such media can be any available media accessible by computer system / server 12, and such media includes both volatile and non-volatile media, removable and non-removable media.
[0033] System memory 28 can include computer system-readable media in the form of volatile memory such as random access memory (RAM) 30 or cache memory 32, or both. Computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 can be provided to read from and write to a non-removable, non-volatile magnetic medium (not shown, typically referred to as a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM, or other optical media can be provided. In such instances, each media can be connected to bus 18 by one or more data media interfaces. As further shown and described hereinafter, memory 28 can include at least one program product having a set (e.g., at least one) of program modules configured to execute the functions of embodiments of the present invention.
[0034] By way of example, and not limitation, a program / utility 40 having a set (at least one) of program modules 42, as well as an operating system, one or more application programs, other program modules, and program data may be stored in the memory 28. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation example of a networking environment. The program modules 42 generally execute the functions or methods of the embodiments of the present invention described herein, or a combination thereof.
[0035] In addition, computer system / server 12 may communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc., one or more devices that enable a user to interact with computer system / server 12, or any device that enables computer system / server 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.), or a combination thereof. Such communication can be performed via input / output (I / O) interface 22. Further, computer system / server 12 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof. As shown, network adapter 20 communicates with other components of computer system / server 12 via bus 18. Although not shown, it should be understood that other hardware components or software components, or a combination thereof, may also be used in conjunction with computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems, etc.
[0036] Next, referring to FIG. 2, an exemplary cloud computing environment 50 is shown. As illustrated, the cloud computing environment 50 includes one or more cloud computing nodes 10 that may communicate with a local computing device used by a cloud consumer, such as, for example, a personal digital assistant (PDA) or cellular phone 54A, a desktop computer 54B, a laptop computer 54C, or an automotive computer system 54N, or a combination thereof. The nodes 10 may communicate with each other. The nodes 10 may be physically or virtually grouped (not shown) in one or more networks such as a private cloud, community cloud, public cloud, or hybrid cloud, or a combination thereof, described above. This enables the cloud computing environment 50 to provide, as a service, an infrastructure, platform, software, or a combination thereof for which a cloud consumer need not maintain resources on a local computing device. The types of computing devices 54A-N shown in FIG. 2 are merely intended to be exemplary, and it is understood that the computing nodes 10 and the cloud computing environment 50 may communicate with any type of computerized device via any type of network or network addressable connection or both (e.g., using a web browser).
[0037] Next, referring to FIG. 3, a set of functional abstractions provided by the cloud computing environment (FIG. 2) is shown. It should be understood in advance that the components, layers, and functions shown in FIG. 3 are merely intended to be exemplary and that embodiments of the present invention are not limited thereto. As illustrated, the following layers and corresponding functions are provided.
[0038] The hardware and software layer 60 includes hardware components and software components. Examples of hardware components include mainframe 61, RISC (Reduced Instruction Set Computer) architecture-based server 62, server 63, blade server 64, storage device 65, and network and networking components 66. In some embodiments, the software components include network application server software 67 and database software 68.
[0039] The virtualization layer 70 provides an abstraction layer that may provide the following examples of virtual entities, namely, virtual server 71, virtual storage 72, virtual network 73 including a virtual private network, virtual applications and virtual operating systems 74, and virtual client 75.
[0040] In one example, the management layer 80 is capable of providing the functions described later. Resource provisioning 81 provides for the dynamic procurement of computing resources and other resources used to execute tasks within a cloud computing environment. Metering and pricing 82 provides for cost tracking as resources are utilized within a cloud computing environment, as well as billing or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides for the verification of identification information regarding cloud consumers and tasks, as well as protection regarding data and other resources. The user portal 83 provides access to the cloud computing environment to consumers and system administrators. Service level management 84 provides cloud computing resource allocation and cloud computing resource management such that the required service level is met. Service level agreement (SLA) planning and fulfillment 85 provides for the prearrangement regarding cloud computing resources for which future requirements are anticipated by the SLA, and the procurement of such resources.
[0041] The workload layer 90 provides examples of functions for which a cloud computing environment may be utilized. Examples of workloads and functions that may be provided from this layer include mapping and navigation 91, software development and lifecycle management 92, virtual classroom education delivery 93, data analysis processing 94, transaction processing 95, and call management 96.
[0042] With the increase in labor costs, many companies currently prefer to provide fewer service staff for call centers. This means that many users' calls will subsequently have to wait in the queue for a long time before they can talk to a call center staff member to receive the support or service called staff service. These waiting users have to listen to repeated prompt tones from the call center, such as "There is no available service staff, please wait...", until a service staff member becomes available. This waiting process is very boring for users, but if users do not pay attention to those prompt tones, users may miss the staff service they are waiting for. For example, those prompt tones may include voices from available staff, such as "Hello, please let us know your matter."
[0043] Under the above circumstances, if the user can immediately put down the user's device and perform other tasks during the waiting process, and an alarm can be immediately issued to the user so that the user can talk to the service staff in a timely manner after a service staff member becomes available, the user experience can be significantly improved.
[0044] Existing technologies for call management in call centers include robots that attempt to solve users' problems using artificial intelligence (AI) technology. However, there are still some problems that require help from a service staff member. For example, the AI robot may not be able to understand the user's problem due to the user's accent, or the AI robot may not provide a solution to a new problem from the user, and so on.
[0045] Therefore, there is a need to provide an approach for managing calls related to requests for staff service to improve the user experience when users have to wait for staff service.
[0046] Embodiments of the present invention provide a method for managing calls requesting staff services to improve the user experience. According to the method, when a user requests staff services and calls a call center, and there is no available service staff member, the user can activate a software module that implements the method of the present invention. Next, the user can place the device and perform other tasks. The software module can monitor the response of the call center while adjusting the volume of the device so that the user is not disturbed. When the software module finds that there is an available service staff member, the software module can immediately alert the user so that the user can talk to the service staff in a timely manner.
[0047] FIG. 4 shows a schematic flowchart of a method 400 for managing calls requesting staff services according to an embodiment of the present invention. In some embodiments, method 400 can be implemented in a module with two threads, one being a recording thread and the other being a call thread. Method 400 can be initiated by a user of the device. For example, the operating system of the device can provide a button in the call interface that the user presses to activate the module of the present invention. When the user sets up a call to the call center and the user's service request enters the waiting process, the user can press the button to activate the module of the present invention so that the module can implement the method of the present invention on the device. In some embodiments, the method can be automatically initiated by the configuration of the module by the user.
[0048] Referring to FIG. 4, in step 410, the first voice segment received from the call center during a call made by the device is recorded by the call thread. In some embodiments, the first voice segment can be continuously stored in one voice file. In some embodiments, the first voice segment can be continuously stored in memory. The recording thread may be executed in parallel with the call thread.
[0049] In step 420, it is determined whether a portion of the first voice segment is related to the first predefined voice segment in the call thread. The first predefined voice segment can be a repeated voice segment received from the call center, such as the voice segment "There is no available service staff, please wait...". If the current voice sub-segment (the voice sub-segment selected from the first voice segment) is related to the first predefined voice segment, it can be concluded that there is still no available service staff. However, if the current voice sub-segment is completely different from the first predefined voice segment, for example, if the current voice segment is "Service staff member number 123 is available, good morning, please let me know your request", it can be concluded that there is an available service staff.
[0050] In some embodiments, the first predefined voice segment may comprise several predefined voice sub-segments. For example, one predefined voice sub-segment may be "There is no available service staff, please wait...", and another predefined voice sub-segment may be the notification voice segment from the call center. For simplicity of explanation, the first predefined voice segment is hereinafter assumed to comprise only one predefined voice segment.
[0051] In some embodiments, the first predefined audio segment can be determined by the user by recording a repeated audio segment, hereinafter referred to as the repeated audio segment, from the call center. For example, the module of the present invention can provide an option for the user of the device to record the repeated audio segment as the first predefined audio segment. The user may press the "start recording" button before the repeated audio segment corresponding to the text generated by the speech-to-text feature, "There is no available service staff, please wait...", is received, and after the repeated audio segment is recorded, the user may press the "end recording" button. Next, the repeated audio segment can be stored as the first predefined audio segment. In some embodiments, the first predefined audio segment can be obtained from a third party. For example, the first predefined audio segment can be downloaded from the call center's website. In some embodiments, the first predefined audio segment may be selected by the user from existing audio segments.
[0052] In some embodiments, the first predefined audio segment can be determined automatically. For example, the module may record a second audio segment received over the call via the device over a predefined time such as 20 seconds. Next, the module may identify a repeated audio segment from the second audio segment.
[0053] In some embodiments, the pitch of the second voice segment may be used to identify repeating voice segments, and then the identified repeating voice segments may be used as the first predefined voice segments. Specifically, the second voice segment may be divided into a plurality of voice sub-segments using a sliding window (there may be an overlap between two voice sub-segments). For example, assuming that the sliding window corresponds to a voice segment with a width of 5 seconds and the sliding length corresponds to a voice segment with a length of 1 second (the parameters can be defined as other values as needed, and the width of the window and the sliding length can also be defined by the user as needed), the first voice sub-segment corresponds to the second voice segment up to the point (start) corresponding to 5 seconds from the start, the second voice sub-segment corresponds to the second voice segment from the point (start) corresponding to 1 second to the point (start) corresponding to 6 seconds, and the third voice sub-segment corresponds to the second voice segment from the point (start) corresponding to 2 seconds to the point (start) corresponding to 7 seconds, and so on. Next, a plurality of sets of the pitches of the aforementioned plurality of voice sub-segments can be determined. Next, a repeating set of pitches can be identified from the plurality of sets of pitches, and one of the plurality of voice sub-segments corresponding to the repeating set of pitches can be identified as the first predefined voice segment (for example, the difference between two sets of pitches is within a predefined threshold range). For example, considering that there are four voice sub-segments in the second voice segment and the four sets of pitches are {A, A, B, C}, {A + 0.05A, B + 0.06B, C + 0.08C, D}, {A + 0.01A, A + 0.02A, B + 0.04B, C + 0.03C}, {B + 0.09B, C + 0.09C, D + 0.02D, E} respectively, {A, A, B, C} can be identified as a repeating set of pitches.An audio sub-segment corresponding to a repeated set of pitches {A, A, B, C} can be identified as a first predefined audio segment. Those skilled in the art may understand that the above four sets of pitches are for illustrative purposes only, and the values of the sets of pitches can be determined by those skilled in the art using existing techniques.
[0054] In some embodiments, Mel Frequency Cepstral Coefficients (MFCCs) can replace the above-mentioned pitches to identify the repeated audio segments in the second audio segment. Specifically, for each set of MFCCs of each of the above-mentioned plurality of audio sub-segments, which may be known to those skilled in the art, can be determined first. Next, a repeated set of MFCCs can be identified from the plurality of sets of MFCCs (e.g., the difference between two sets of MFCCs is within a predefined threshold range). Next, one of the plurality of audio sub-segments corresponding to the repeated set of MFCCs can be identified as the first predefined audio segment. The determination of MFCCs for audio segments is well known to those skilled in the art and is omitted here.
[0055] In some embodiments, the second audio segment may be converted into a first text, and then a second text (hereinafter referred to as "repeated text") repeated within the first text can be identified using text recognition techniques. For example, two identical words within the first text are first searched for, and then their respective next words for those identical words within the first text are compared, and the process is repeated until the repeated text is found. Next, one of the plurality of audio sub-segments corresponding to the repeated text in the second audio segment can be obtained as the first predefined audio segment.
[0056] In some embodiments, the plurality of audio sub-segments described above can be converted into a plurality of texts. Next, the repeated text can be identified from the plurality of texts. For example, if two texts are substantially related (e.g., 80% of the words are the same), one of the two texts is identified as the repeated text. Thereafter, it is possible to determine that a certain audio sub-segment among the plurality of audio sub-segments corresponding to the repeated text is the first pre-defined audio segment.
[0057] One of ordinary skill in the art may understand that other audio features can also be used to identify the repeated audio segment in the second audio segment. Also, some common audio processing steps such as filtering, etc. are omitted here because those steps are well known to one of ordinary skill in the art. Generally, the manner of using audio-text conversion is more excellent than other manners in terms of accuracy, availability, and resource utilization efficiency.
[0058] In some embodiments, when a first audio segment is stored in one audio file, a sliding window similar to the sliding window used to identify repeated audio segments can be used to obtain a portion of the first audio segment. FIG. 5 shows an exemplary diagram for continuously comparing a portion of a first audio segment with a first predefined audio segment according to an embodiment of the present invention. The first predefined audio segment 501 has a width of 5 seconds, which is the same width as the sliding window. Each time, the sliding window is slid along the first audio segment by a length of the audio segment corresponding to, for example, 1 second to obtain the next audio sub-segment within the sliding window as a portion of the first audio segment. The smaller the sliding length, the better the comparison result. As shown in FIG. 5, the first audio segment 502 is divided into a plurality of audio sub-segments (having an overlap between the current audio sub-segment and the next audio sub-segment), where the first audio sub-segment 503 corresponds to the first audio segment up to the point (start) corresponding to 5 seconds from the start, and the second audio sub-segment 504 corresponds to the first audio segment from the point (start) corresponding to 1 second to the point (start) corresponding to 6 seconds, and so on. In this case, each audio sub-segment (e.g., a portion of the first audio segment) of the first audio segment, such as 503, 504, etc., can be compared with the first predefined audio segment 501 to determine whether both audio segments are related. If both audio segments are related, the next audio sub-segment obtained by using the sliding window within the first audio segment can be compared with the first predefined audio segment 501, and the comparison process is repeated until an audio sub-segment 505 that is not related to the first predefined audio segment 501 is found.
[0059] Whether a portion of the first audio segment is related to the first predefined audio segment can be determined based on each set of pitches, a set of MFCCs, or the text of two audio sub-segments. In some embodiments, that portion of the first audio segment is converted to a first text, the first predefined audio segment is converted to a second text, and whether the first text is related to the second text can be determined by comparing those texts. Next, whether that portion of the first audio segment is related to the first predefined audio segment can be determined based on a comparison between the first text and the second text. In an example, if the first text is related to the second text, it is possible to determine that that portion of the first audio segment and the first predefined audio segment are related. In another example, if the difference between the first text and the second text is less than a predetermined threshold, it is possible to determine that that portion of the first audio segment and the first predefined audio segment are related.
[0060] In some embodiments, a set of pitches of a portion of the first audio segment is first determined, and then a set of pitches of the first pre-defined audio segment is determined. Next, the set of pitches of the portion of the first audio segment and the set of pitches of the first pre-defined audio segment are compared to determine whether the set of pitches of the portion of the first audio segment is related to the set of pitches of the first pre-defined audio segment. Thus, it is possible to determine whether the portion of the first audio segment is related to the first pre-defined audio segment based on the comparison of the set of pitches of the portion of the first audio segment and the set of pitches of the first pre-defined audio segment. In an example, if the set of pitches of the portion of the first audio segment is related to the set of pitches of the first pre-defined audio segment, it is possible to determine that the portion of the first audio segment and the first pre-defined audio segment are related. In another example, if the difference between the set of pitches of the portion of the first audio segment and the set of pitches of the first pre-defined audio segment is less than a predetermined threshold, it is possible to determine that the portion of the first audio segment and the first pre-defined audio segment are related.
[0061] In some embodiments, a set of MFCCs of a portion of the first audio segment is first determined, and then a set of MFCCs of the first pre-defined audio segment is determined. Next, the set of MFCCs of the portion of the first audio segment and the set of MFCCs of the first pre-defined audio segment are compared to determine whether the set of MFCCs of the portion of the received audio is related to the set of MFCCs of the first pre-defined audio segment. Thus, it is possible to determine whether the portion of the first audio segment is related to the first pre-defined audio segment based on the comparison between the set of MFCCs of the portion of the first audio segment and the set of MFCCs of the first pre-defined audio segment. In an example, if the set of MFCCs of the portion of the first audio segment is related to the set of MFCCs of the first pre-defined audio segment, it is possible to determine that the portion of the first audio segment and the first pre-defined audio segment are related. In another example, if the difference between the set of MFCCs of the portion of the first audio segment and the set of MFCCs of the first pre-defined audio segment is less than a predetermined threshold, it is possible to determine that the portion of the first audio segment and the first pre-defined audio segment are related.
[0062] In some embodiments, it is possible to determine whether a portion of the first audio segment is related to a first predefined audio segment based on the correlation between that portion of the first audio segment and the first predefined audio segment. The correlation between that portion of the first audio segment and the first predefined audio segment can be represented as the correlation between the text of that portion of the first audio segment obtained via audio-text and the text of the first predefined audio segment obtained via audio-text, or the correlation between the set of pitches of that portion of the first audio segment and the set of pitches of the first predefined audio segment, or the correlation between the set of MFCCs of that portion of the first audio segment and the set of MFCCs of the first predefined audio segment, among others. Furthermore, various correlations can be defined. In some embodiments, if the correlation between that portion of the first audio segment and the first predefined audio segment exceeds a predetermined threshold, it is possible to determine that that portion of the first audio segment and the first predefined audio segment are related. For example, the correlation can be defined as the number of identical words contained within the two texts corresponding to these two audio sub-segments. In some embodiments, if the correlation between that portion of the first audio segment and the first predefined audio segment is less than a predetermined threshold, it is possible to determine that that portion of the first audio segment and the first predefined audio segment are related. For example, the correlation can be defined as the accumulated difference between the two sets of pitches of the two audio sub-segments.
[0063] Referring again to FIG. 4, at step 430, in response to that portion of the first voice segment being related to the first predefined voice segment, the volume of the device is adjusted by, for example, reducing the volume of the call in the call thread and even muting the call. As shown in FIG. 5, each of the two voice sub-segments 503 and 504 is related to the first predefined voice segment 501, and it is possible to determine that the call center is transmitting a repeated voice segment during the corresponding period, in other words, during that period, there is no available service staff member. Therefore, the volume of the device is adjusted so as not to disturb the user.
[0064] At step 440, in response to that portion of the first voice segment not being related to the first predefined voice segment, an alert is issued to the user of the device in the call thread. As shown in FIG. 5, when it is determined that the voice sub-segment 505 is not related to the first predefined voice segment 501, it is possible to determine that the call center is transmitting a different voice sub-segment during that period, in other words, a member of the available service staff for the service call now exists. The user can be alerted using existing methods such as voice alerts, device vibrations, optical signals from the device, information displayed on the device's screen, and the device's call sound, among others. Next, method 400 ends. Since a member of the available service staff exists at this point, the user can speak directly to that service staff. As shown in FIG. 5, an alert is issued to the user at the end of that portion of the first voice segment 505. Prior to that point, the volume of the device is adjusted so as not to disturb the user.
[0065] As can be seen from FIG. 5, it may be found that the user may miss the voice sub-segment 505 that is expected to be announced by a service staff member who becomes available when a staff member becomes available. For that purpose, embodiments of the present invention may include repeating the voice sub-segment 505 for the user.
[0066] Figure 6 shows a schematic flowchart of a method 600 for improving the user experience during a call waiting in a queue according to an embodiment of the present invention. Similar to the method 400 of Figure 4, the method of Figure 6 also includes recording a first voice segment received from a call center during a call made by a device in a call thread 410, determining whether a portion of the first voice segment is related to a first predefined voice segment in the call thread 420, adjusting the volume of the device 430, and issuing an alert to the user in the call thread 440 in response to the portion of the first voice segment not being related to the first predefined voice segment. In the method of Figure 6, in response to the portion of the first voice segment not being related to the first predefined voice segment, in step 610, that portion of the first voice segment (such as the voice sub-segment 505 in Figure 5) is played in the call thread at a faster speed compared to the normal speaking speed. For example, the speed can be twice the normal speaking speed or any other faster speaking speed compared to the normal speaking speed. In step 620, the next voice segment received in the call (such as the voice segment 506 in Figure 5) is recorded in the recording thread. Both step 610 and step 620 can be executed in parallel in different threads. After that portion of the first voice segment (505) is played, in step 630, the next voice segment (506) is played in the call thread at a faster speed compared to the normal speaking speed until the end of this next voice segment. Here, "end" means that the playback process for this next voice segment and the recording process for this next voice segment reach the same point in time of this next voice segment. Next, the user can talk directly to the service staff. It can be found that step 610 and step 620 can be executed at substantially the same time in different threads, and step 610 and step 440 can be executed in any order. For example, step 440 can follow step 610, or step 610 can follow step 440.
[0067] In some embodiments, prior to the end of method 400, method 400 further includes a step of using a second predefined voice segment to call the other side of the call (e.g., a call center) in response to a portion of the first voice segment not being related to a first predefined voice segment. In an example, the second predefined voice segment can be, for example, a voice segment such as "The caller is in a waiting process and will answer the call as soon as possible. Please wait a moment." so that a corresponding service staff member can know the current situation and can wait a little while to continue the call. After alerting the caller, the call thread that implements most of method 400 can send the aforementioned second predefined voice segment to alert the corresponding service staff so that when the user picks up the device and speaks, the corresponding service staff can repeat a repeated voice segment such as 505 to improve the user experience. Those skilled in the art may understand that this step can be combined with the method of FIG. 6. The playback speed is faster than the recording speed so that both the playback process and the recording process can ultimately end simultaneously. The user is listening to the missed voice sub-segment and the next voice segment, during which the corresponding service staff is listening to the second predefined voice segment, and then the service staff can wait for the user. Next, the user and the service staff can speak directly at the end of method 400 of FIG. 4 or method 600 of FIG. 6.
[0068] Note that the process of managing calls for staff services according to embodiments of the present disclosure can be implemented by the computer system / server 12 of FIG. 1.
[0069] The present invention may be a system, method, or computer program product, or a combination thereof, at any possible level of integration of technical details. The computer program product may include a computer-readable storage medium (or media) having thereon computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0070] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes, hereinafter, namely, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed to be a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.
[0071] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage in a computer-readable storage medium within each respective computing / processing device.
[0072] The computer-readable program instructions for carrying out the operations of the present invention may be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, configuration data for integrated circuits, or object-oriented programming languages such as Smalltalk(R), C++, or the like, and procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on a remote computer or server. In a scenario where execution is entirely on a remote computer or server, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or wide area network (WAN), or the connection may be made to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit for carrying out aspects of the present invention.
[0073] Aspects of the present invention will be described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.
[0074] These computer readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, or other devices to function in a particular manner, such that the storage medium storing the instructions comprises an article of manufacture including instructions which implement the function / act specified in one or more blocks of the flowchart illustrations and / or block diagrams.
[0075] Alternatively, the computer readable program instructions may be loaded onto a computer, other programmable apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams.
[0076] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of instructions that includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions described in the blocks may occur out of the order described in the figures. For example, two blocks shown in succession may in fact be implemented as one step, may be executed simultaneously, may be executed substantially simultaneously partially or fully overlapping in time, or the blocks may sometimes be executed in the reverse order depending on the functions involved. It should also be noted that each block of the block diagram or flowchart or both, and combinations of blocks in the block diagram or flowchart or both, can be implemented by a dedicated hardware-based system that performs the specified function or operation or a combination of dedicated hardware instructions and computer instructions.
[0077] The descriptions of the various embodiments of the present invention have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the disclosed embodiments. Many variations and modifications will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles of the embodiments, the practical application, or a technical improvement found in the marketplace over the technology, or to enable other skilled artisans to understand the embodiments disclosed herein.
[0078] [Item 1] A method implemented by a computer, comprising: recording, by one or more processors, a first audio segment received in a call made by a device; An action by one or more processors to determine whether a portion of the first voice segment is related to a first predefined voice segment, An action by one or more processors to adjust the volume of the device in response to the portion of the first voice segment being related to the first predefined voice segment, An action by one or more processors to alert the user of the device in response to the portion of the first voice segment not being related to the first predefined voice segment A method comprising. [Item 2] In response to the portion of the first voice segment not being related to the first predefined voice segment, An action by one or more processors to play back the portion of the first voice segment at a faster speed compared to the normal speaking speed, An action by one or more processors to record the next voice segment received in the call, An action by one or more processors to play back the next voice segment at a faster speed compared to the normal speaking speed until the end of the next voice segment in response to the end of the action of playing back the portion of the first voice segment The method according to item 1, further comprising. [Item 3] The method according to item 1, further comprising an action by one or more processors to transmit a second predefined voice segment to the other side of the call in response to the portion of the first voice segment not being related to the first predefined voice segment. [Item 4] The action of determining whether the portion of the first voice segment is related to the first predefined voice segment is An action by one or more processors to convert the portion of the first voice segment into a first text, An action of converting the first pre - defined voice segment into a second text by one or more processors, An action of determining, by one or more processors, based on a comparison between the first text and the second text, whether the portion of the first voice segment is related to the first pre - defined voice segment, the method according to claim 1. [Claim 5] The action of determining whether the portion of the first voice segment is related to the first pre - defined voice segment, An action of determining, by one or more processors, a set of pitches of the portion of the first voice segment, An action of determining, by one or more processors, a set of pitches of the first pre - defined voice segment, An action of determining, by one or more processors, based on a comparison between the set of pitches of the portion of the first voice segment and the set of pitches of the first pre - defined voice segment, whether the portion of the first voice segment is related to the first pre - defined voice segment, the method according to claim 1. [Claim 6] The action of determining whether the portion of the first voice segment is related to the first pre - defined voice segment, An action of determining, by one or more processors, a set of Mel - Frequency Cepstral Coefficients (MFCCs) of the portion of the first voice segment, An action of determining, by one or more processors, a set of MFCCs of the first pre - defined voice segment, An action, performed by one or more processors, of determining whether a portion of the first audio segment is related to the first pre - defined audio segment based on a comparison between the set of MFCCs of the portion of the first audio segment and the set of MFCCs of the first pre - defined audio segment, the method according to claim 1. [Claim 7] The first pre - defined audio segment is An action, performed by one or more processors, of recording a second audio segment received in the call over a pre - defined time; An action, performed by one or more processors, of converting the second audio segment into a third text; An action, performed by one or more processors, of identifying text repeated within the third text; The method according to claim 1, obtained by an action, performed by one or more processors, of obtaining a portion of the second audio segment corresponding to the repeated text as the first pre - defined audio segment. [Claim 8] The first pre - defined audio segment is An action, performed by one or more processors, of recording a second audio segment received in the call over a pre - defined time; An action, performed by one or more processors, of splitting the second audio segment into a plurality of audio sub - segments using a sliding window; An action, performed by one or more processors, of converting the plurality of audio sub - segments into a plurality of texts; An action, performed by one or more processors, of identifying text repeated from the plurality of texts; The method according to claim 1, obtained by an action, performed by one or more processors, of identifying an audio sub - segment of the plurality of audio sub - segments corresponding to the repeated text as the first pre - defined audio segment. [Item 9] wherein the first predefined voice segment is obtained by: an action by one or more processors to record a second voice segment received by the call over a predefined time; an action by one or more processors to divide the second voice segment into a plurality of voice sub-segments using a sliding window; an action by one or more processors to determine a plurality of sets of pitches of the plurality of voice sub-segments; an action by one or more processors to identify a set of repeating pitches from the plurality of pitches; and an action by one or more processors to identify a certain voice sub-segment among the plurality of voice sub-segments corresponding to the set of repeating pitches as the first predefined voice segment, the method according to claim 1. [Item 10] wherein the first predefined voice segment is obtained by: an action by one or more processors to record a second voice segment received during the call over a predefined time; an action by one or more processors to divide the second voice segment into a plurality of voice sub-segments using a sliding window; an action by one or more processors to determine a plurality of sets of Mel Frequency Cepstral Coefficients (MFCCs) of the plurality of voice sub-segments; an action by one or more processors to identify a set of repeating MFCCs from the plurality of sets of MFCCs; and an action by one or more processors to identify a certain voice sub-segment among the plurality of voice sub-segments corresponding to the set of repeating MFCCs as the first predefined voice segment, the method according to claim 1. [Item 11] The method according to claim 1, wherein the action of alerting the user of the device uses at least one of an audio alert, device vibration, an optical signal from the device, information displayed on a screen of the device, and a ringing sound of the device to alert the user of the device. [Claim 12] A system comprising: one or more processors; a memory coupled to at least one of the processors; a set of computer program instructions stored in the memory, the set of computer program instructions being: an action of recording a first audio segment received in a call made by a device; an action of determining whether a portion of the first audio segment is related to a first predefined audio segment; in response to the portion of the first audio segment being related to the first predefined audio segment, an action of adjusting a volume of the device; and in response to the portion of the first audio segment not being related to the first predefined audio segment, an action of alerting a user of the device, the set of computer program instructions being executed by at least one of the processors. A system comprising the above. [Claim 13] The action is: in response to the portion of the first audio segment not being related to the first predefined audio segment, an action of playing back the portion of the first audio segment at a faster speed compared to a normal speaking speed; an action of recording a next audio segment received in the call; In response to the end of the action of playing the portion of the first voice segment, the system according to claim 12 further includes an action of playing the next voice segment at a speed faster than the normal speaking speed until the end of the next voice segment. [Claim 14] The action is In response to the portion of the first voice segment not being related to the first predefined voice segment, the system according to claim 12 further includes an action of transmitting a second predefined voice segment to the other side of the call. [Claim 15] The action of determining whether the portion of the first voice segment is related to the first predefined voice segment is an action of converting the portion of the first voice segment into a first text; an action of converting the first predefined voice segment into a second text; and an action of determining whether the portion of the first voice segment is related to the first predefined voice segment based on a comparison between the first text and the second text, the system according to claim 12. [Claim 16] The first predefined voice segment is obtained by an action of recording a second voice segment received in the call over a predefined time, an action of converting the second voice segment into a third text, an action of identifying text repeated within the third text, and an action of obtaining a portion of the second voice segment corresponding to the repeated text as the first predefined voice segment, the system according to claim 12. [Claim 17] The first predefined voice segment is An action of recording a second voice segment received over the call for a pre-defined period of time, An action of splitting the second voice segment into a plurality of voice sub-segments using a sliding window, An action of converting the plurality of voice sub-segments into a plurality of texts, An action of identifying repetitive text from the plurality of texts, The system according to claim 12, obtained by an action of identifying a certain voice sub-segment among the plurality of voice sub-segments corresponding to the repetitive text as the first pre-defined voice segment. [Claim 18] The system according to claim 12, wherein the action of alerting the user of the device includes an action of alerting the user of the device using at least one of an audio alert, device vibration, an optical signal from the device, information displayed on the screen of the device, and a ringing tone of the device. [Claim 19] A computer program product, By a processor, to the processor, An action of recording a first voice segment received in a call made by a device, An action of determining whether a portion of the first voice segment is related to a first pre-defined voice segment, In response to the portion of the first voice segment being related to the first pre-defined voice segment, an action of adjusting the volume of the device, In response to the portion of the first voice segment not being related to the first pre-defined voice segment, an action of alerting the user of the device, and a computer-readable storage medium storing program instructions executable to cause the processor to perform the actions. [Claim 20] The program instructions, by a processor, to the processor, In response to the portion of the first voice segment not being related to the first pre-defined voice segment, an action of playing back the portion of the first voice segment at a speed faster than the normal speaking speed, and an action of recording the next voice segment received in the call, and The computer program product according to item 19, which is further executable to perform an action of playing back the next voice segment at a speed faster than the normal speaking speed until the end of the next voice segment in response to the end of the action of playing back the portion of the first voice segment.
Claims
【Claim 1】 A method implemented by a computer, comprising: an action of recording, by one or more processors, a first audio segment received in a call made by a device; an action of determining, by one or more processors, whether a portion of the first audio segment is related to a first predefined audio segment; an action of adjusting the volume of the device, by one or more processors, in response to the portion of the first audio segment being related to the first predefined audio segment; an action of alerting a user of the device, by one or more processors, in response to the portion of the first audio segment not being related to the first predefined audio segment and wherein the action of determining whether the portion of the first audio segment is related to the first predefined audio segment includes any one of the following (a), (b), or (c): a) the action includes an action of converting, by one or more processors, the portion of the first audio segment into a first text; an action of converting, by one or more processors, the first predefined audio segment into a second text; an action of determining, by one or more processors, whether the portion of the first audio segment is related to the first predefined audio segment based on a comparison between the first text and the second text or b) the action includes an action of determining, by one or more processors, a set of pitches of the portion of the first audio segment; an action of determining, by one or more processors, a set of pitches of the first predefined audio segment; an action of determining, by one or more processors, whether the portion of the first audio segment is related to the first predefined audio segment based on a comparison between the set of pitches of the portion of the first audio segment and the set of pitches of the first predefined audio segment or c) the action includes An action of determining a set of Mel Frequency Cepstral Coefficients (MFCCs) of the portion of the first voice segment by one or more processors, An action of determining a set of MFCCs of the first predefined voice segment by one or more processors, An action of determining whether the portion of the first voice segment is related to the first predefined voice segment based on a comparison between the set of MFCCs of the portion of the first voice segment and the set of MFCCs of the first predefined voice segment by one or more processors comprising, the first predefined voice segment is obtained by any of the following (a'), (b'), (c') or (d'): (a') the first predefined voice segment is An action of recording, by one or more processors, a second voice segment received in the call over a predefined time, An action of converting, by one or more processors, the second voice segment into a third text, An action of identifying, by one or more processors, text repeated within the third text, An action of obtaining, by one or more processors, a portion of the second voice segment corresponding to the repeated text as the first predefined voice segment obtained by, or (b') the first predefined voice segment is An action of recording, by one or more processors, a second voice segment received in the call over a predefined time, An action of splitting, by one or more processors, the second voice segment into a plurality of voice sub-segments using a sliding window, An action of converting, by one or more processors, the plurality of voice sub-segments into a plurality of texts, An action of identifying, by one or more processors, text repeated from the plurality of texts, An action of identifying, by one or more processors, a certain voice sub-segment among the plurality of voice sub-segments corresponding to the repeated text as the first predefined voice segment obtained by, or (c') the first predefined voice segment is an action of recording, by one or more processors, a second voice segment received by the call over a predefined time; an action of splitting, by one or more processors, the second voice segment into a plurality of voice sub-segments using a sliding window; an action of determining, by one or more processors, a plurality of sets of pitches of the plurality of voice sub-segments; an action of identifying, by one or more processors, a set of repeating pitches from the plurality of pitches; an action of identifying, by one or more processors, a certain voice sub-segment among the plurality of voice sub-segments corresponding to the set of repeating pitches as the first predefined voice segment obtained by, or (d') the first predefined voice segment is an action of recording, by one or more processors, a second voice segment received over the call over a predefined time; an action of splitting, by one or more processors, the second voice segment into a plurality of voice sub-segments using a sliding window; an action of determining, by one or more processors, a plurality of sets of mel-frequency cepstral coefficients (MFCCs) of the plurality of voice sub-segments; an action of identifying, by one or more processors, a set of repeating MFCCs from the plurality of sets of MFCCs; an action of identifying, by one or more processors, a certain voice sub-segment among the plurality of voice sub-segments corresponding to the set of repeating MFCCs as the first predefined voice segment obtained by the method. **Claim 2** The method according to claim 1, further comprising an action of transmitting, by one or more processors, a second predefined voice segment to the other side of the call in response to a portion of the first voice segment not being related to the first predefined voice segment. **Claim 3** In response to a portion of the first voice segment not being related to the first predefined voice segment, An action of playing back, by one or more processors, the portion of the first audio segment at a speed faster than the normal speaking speed, An action of recording, by one or more processors, the next audio segment received in the call, An action of playing back, by one or more processors, the next audio segment at a speed faster than the normal speaking speed until the end of the next audio segment in response to the end of the action of playing back the portion of the first audio segment The method according to claim 1 or 2, further comprising: **Claim 4** The action of alerting the user of the device includes an action of alerting the user of the device using at least one of an audio alert, device vibration, an optical signal from the device, information displayed on the screen of the device, and a ringing tone of the device. The method according to any one of claims 1 to 3. **Claim 5** A method implemented by a computer, comprising: An action of recording, by one or more processors, a first audio segment received in a call made by a device, An action of determining, by one or more processors, whether a portion of the first audio segment is related to a first predefined audio segment, An action of adjusting the volume of the device, by one or more processors, in response to the portion of the first audio segment being related to the first predefined audio segment, An action of alerting the user of the device, by one or more processors, in response to the portion of the first audio segment not being related to the first predefined audio segment including The first predefined audio segment is The first predefined audio segment is obtained by any one of the following (a'), (b'), (c') or (d'): (a') The first predefined audio segment is An action of recording, by one or more processors, a second audio segment received in the call over a predefined time, An action of converting, by one or more processors, the second audio segment into a third text, an action of identifying text repeated within the third text by one or more processors, an action of obtaining, by one or more processors, a portion of the second voice segment corresponding to the repeated text as the first predefined voice segment obtained by, or (b') the first predefined voice segment is an action of recording, by one or more processors, the second voice segment received in the call over a predefined time, an action of splitting, by one or more processors, the second voice segment into a plurality of voice sub-segments using a sliding window, an action of converting, by one or more processors, the plurality of voice sub-segments into a plurality of texts, an action of identifying, by one or more processors, text repeated from the plurality of texts, an action of identifying, by one or more processors, a certain voice sub-segment among the plurality of voice sub-segments corresponding to the repeated text as the first predefined voice segment obtained by, or (c') the first predefined voice segment is an action of recording, by one or more processors, the second voice segment received by the call over a predefined time, an action of splitting, by one or more processors, the second voice segment into a plurality of voice sub-segments using a sliding window, an action of determining, by one or more processors, a plurality of sets of pitches of the plurality of voice sub-segments, an action of identifying, by one or more processors, a set of pitches repeated from the plurality of pitches, an action of identifying, by one or more processors, a certain voice sub-segment among the plurality of voice sub-segments corresponding to the repeated set of pitches as the first predefined voice segment obtained by, or (d') the first predefined voice segment is an action of recording, by one or more processors, the second voice segment received over the call over a predefined time, An action of dividing the second audio segment into a plurality of audio sub - segments using a sliding window by one or more processors; An action of determining a plurality of sets of Mel - Frequency Cepstral Coefficients (MFCCs) of the plurality of audio sub - segments by one or more processors; An action of identifying a repeating set of MFCCs from the plurality of sets of MFCCs by one or more processors; An action of identifying a certain audio sub - segment among the plurality of audio sub - segments corresponding to the repeating set of MFCCs as the first pre - defined audio segment by one or more processors obtained by; the method. **Claim 6** A computer program for causing a computer to execute each step of the method according to any one of Claims 1 to 5. **Claim 7** A computer - readable storage medium storing the computer program according to Claim 6. **Claim 8** A system, one or more processors, and a memory coupled to at least one of the processors, the memory storing the computer program according to Claim 6 A system comprising the above.
Citation Information
Patent Citations
Method and apparatus for replacing telephone on-hold music at caller's side
JP2015220755A
Systems and methods for fingerprinting datasets
JP2015515646A
Method and system for informing a user that a call is no longer on hold
US20150189089A1
On-Hold Processing for Telephonic Systems
US20150326716A1
On hold detection
US20150341763A1