System and method for providing context-matched desired vocal properties in the audio portion of a communication
By using AI to adjust the audio and video images of customer service agents in real time, the problem of mismatch between the agent's facial expressions and voice attributes during video calls was solved, improving the interaction effect and customer satisfaction.
Patent Information
- Application Number
- CN202211087152.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-09-07
- Filing Date
- 2022-09-07
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-09-07
AI Technical Summary
Customer service agents struggle to naturally display appropriate facial expressions and vocal attributes during video calls, resulting in poor interaction, especially in purely audio calls where emotional delivery is insufficient.
By using artificial intelligence to monitor and modify the agent's audio and video images in real time, adjusting voice attributes and facial expressions to match the emotional needs of the interaction, and using AI to determine and apply appropriate vocalizations and expressions to optimize the interaction effect.
This improved the effectiveness of customer service interactions, enhanced customer trust and satisfaction with the agent, and ensured the accuracy and consistency of emotional communication.
Smart Images

Figure CN115776552B_ABST
Abstract
Description
[0001] Priority requirements
[0002] This application is a continuation-in-part claim to the benefit of U.S. Patent Application No. 16 / 535,169, filed August 8, 2019, entitled “Optimizing Interaction Results Using AI-Guided Manipulated Video,” which is incorporated herein by reference.
[0003] Copyright Notice
[0004] This patent document contains a portion of copyrighted material. The copyright holder does not object to any fax copying of the patent document or patent disclosure appearing in the Patent and Trademark Office's patent documents or records, but otherwise reserves all copyright. Technical Field
[0005] This invention generally relates to systems and methods for video image processing, and more particularly to altering the facial expressions of people in video images. Background Technology
[0006] Video images of customer service agents allow customers to see the agent and can facilitate better interaction between them to resolve specific issues. It is known that human behavior, including facial expressions, can positively or negatively influence the outcome of an interaction. Currently, agents are instructed and trained to provide appropriate facial expressions to drive the best possible outcome of the interaction—regardless of whether they know what to do or whether they are in a good mood. Pre-call training, instruction, and real-time prompts / guidance "during the call" are helpful, but they still rely on the agent's ability to translate such instructions into appropriate facial expressions. While it is possible to use avatars that can be programmed to provide the exact facial expressions expected, avatars do not convey the personal aspects of a real person. When a customer interacts with an agent's avatar, the interaction may not provide the expected benefits of seeing the agent.
[0007] Customer service is one of the key aspects defining a business's success. For many businesses, the outcome of customer service interactions largely depends on the agents interacting with customers. It is known that the mannerisms of individuals, including their different voice attributes, especially during vocal interactions, can have a positive or negative impact on the outcome of the interaction.
[0008] This becomes especially important in purely voice-based calls. However, even when communication includes audio and video, the agent's correct vocal attributes are crucial for accurately conveying emotions or feelings. Some studies suggest that speech has a stronger influence than visual cues (e.g., facial expressions) in detecting the other party's emotions. When the visual and auditory content of the participants in a communication carries attributes that do not convey the same emotional content, the communication is perceived as negative, such as affected or insincere. Summary of the Invention
[0009] In many contact centers, agents rely on “doing the right thing,” using the right facial expressions and / or vocal attributes to drive the best possible outcome of their interactions with customers—regardless of whether they consciously know what to do or not to do, and without considering their own current emotional state. Pre-call guidance and real-time prompts / guidance during the call can be helpful, but delivering appropriate emotional content still depends on the agent’s ability to deliver communication with the appropriate vocal attributes suited to the specific context of the conversation.
[0010] In one embodiment, artificial intelligence provides altered audio, such as deepfake audio or voice cloning, to the agent's voice attributes. The agent is monitored in real time to see if the modification is appropriate. If appropriate, then when the determined desired emotion or sentiment audio content does not match the currently observed emotion or sentiment audio content, an appropriate emotion or sentiment is applied to the call. If a mismatch exists, then systems and methods are provided to modify the audio so that the audio delivered to the communication channel includes emotions or sentiments determined to be suitable for communication but not present in the agent's unaltered speech.
[0011] In another embodiment, real-time audio processing allows audio to be modified (e.g., auditory depth forgery, alteration of tone, speech rate, accent, inflection, lilt, etc.) to have different speech attributes, thereby having speech attributes that are determined to be more suitable for a given communication content or client, such as those that convey emotion or mood.
[0012] In another embodiment, AI (such as a trained neural network) acquires the linguistic content (e.g., words, phrases, utterances) of the agent's speech and modifies the audio to include altered audio content, such as deepfakes, altered speech attributes, or vocal qualities, such as specific emotional content determined to be absent (e.g., empathy), or present but determined to require removal from the utterances provided by the agent's unaltered speech (e.g., stimulation). As used herein, "words" and "phrases" have their common and conventional meanings. As used herein, "utterances" include sounds uttered by the agent, quasi-words, (one or more) pseudo-words, etc. For example, utterances such as "uh huh" are generally understood as "yes," confirmation of understanding, confirmation of hearing, etc.; the utterance of "hmm" is generally understood as expressing confusion, curiosity, uncertainty, etc.; the utterance of "oh" is generally understood as expressing surprise, confusion, disappointment, etc.; and so on. Other examples of utterances may include specific sounds provided, such as sounds used by non-English speakers.
[0013] In one embodiment, the agent's spoken words are maintained, while the agent's voice is modified. The modified voice can vary in attributes or qualities such as pitch, intonation, rhythm (e.g., too halting or too fluent), speed (for all or part of the speech), resonance, beat, and texture. Human communication is complex and often subtle, especially when it is performed verbally. This is further complicated when the parties cannot see each other, such as during a purely audio call over a network. Words can convey a meaning that can be amplified or offset based on specific vocal qualities. For example, during a technical support call, an agent might say "I see your problem." However, emphasizing the word "your" might convey additional or alternative meanings, such as dishonesty or an attempt to blame the customer. Conversely, emphasizing "I see" could convey surprise, realization, or one of the following, which is more conducive to conveying support, understanding, and care to the customer. Similar variations in meaning can be conveyed based on speech rate, intonation, and so on. For example, a slow, laborious, laborious, and monotonous "Oh, I am sorry to hear about the problem" can convey aggression, indifference, or lack of attention, while concise and clear wording with appropriate pace and / or tone (e.g., conveying gentleness or sympathy) can convey support, attention, concern, or confidence in resolving a particular problem.
[0014] For example, after saying "Oh, I am sorry to hear about the problem," the system can determine that it needs to convey confidence that the problem will be solved. If the agent then says, "Nevertheless, we are here and will get your problem sorted out today," the system can change the pitch to higher, the tone to more cheerful and less soft, to convey assurance to the customer that their problem is being properly considered and will be resolved.
[0015] In another example, when a customer calls for an emergency, such as requesting an ambulance or roadside assistance after an accident, the customer service agent's voice should be presented to the customer in a softer tone and higher pitch to make their voice sound sweeter and reassuring. Here, the resonance of the voice can also be adjusted to make it sound more like it's coming from deep in the throat, giving the customer an impression of care, rather than a voice that sounds like it's coming from the nose, which might convey to the customer that the agent is annoying.
[0016] In another embodiment, human speech, such as speech delivered by an agent to a customer during a communication, is observed and quantified. Speech is quantified to record when the agent uses specific vocal attributes, along with the state and circumstances of the call (e.g., duration of interaction, mood via text transcribed from speech, emotional trajectory, information just / about to be exchanged), and the outcome of one or more interaction sub-phases, to create a training database. Quantization can be performed automatically, and additionally or alternatively, humans can provide input, including training input, to recognize the voices of untrained or mistrained neural networks or other artificial intelligence agents. Human input can be provided to the system to give humans a perception of all humans or a specific group of humans. For example, one person, such as someone with specific demographic attributes, might find specific qualities of the agent's speech to be harsh, cold, indifferent, and generally negative, while another person, such as someone with different demographic attributes, might find specific speech to be efficient and effective, and generally positive. Thus, neural networks can be trained to apply vocal attribute modifications to communications presented to customers, these modifications being more specifically tailored to a particular customer or customer category.
[0017] In another embodiment, to optimize interaction, machine learning is used to determine best practices for appropriate vocal attributes based on observations of specific vocalizations and situations (e.g., the emotion, mood, customer demographics, etc. of the call). The success of an interaction can be determined based on the duration of the interaction, customer feedback, and / or automated analysis of the transcription of the interaction.
[0018] In another embodiment, intentional, manual or automatic changes can be applied to determined appropriate vocal attributes, such as presenting them to numerous customers as control and experimental groups or A / B tests, to determine whether the changes improve the success rate of interactions. Feedback from the tests is provided to the neural network or other AI agent to provide training, thereby further strengthening the vocal attributes of the control group if the success of the experimental group is the same or worse, or changing the vocal attributes used when the success of the experimental group improves.
[0019] In another embodiment, the audio recording may be made from original and / or modified speech. Markings such as color coding may be applied to speech portions where the quality of the speech is modified and / or to specific modified markings used.
[0020] In another embodiment, audio quality manipulation features are enabled for a pre-configured duration, such as through contact center system management, to gain insights into how the agent performs, and the collected data can be further used for quality control, training, and reporting. Reports may include metrics such as the number of agents whose voice attributes are manipulated by the system, the number of times manipulation occurs on a particular agent within a given time interval when the agent's emotion / feeling deviates from standards or best practices (and the resulting success / failure indicators), the emotions, feelings, and / or circumstances that trigger manipulation for a particular agent, the degree of impact of the manipulation on the success of the interaction, and / or other metrics, feedback to the system, and / or training or evaluation of agent performance and the need or absence of manipulation of the agent's speech attributes or the degree or type thereof.
[0021] Humans use vision to receive nonverbal cues about others they interact with. If a person's words don't match their facial expressions, that person may be perceived as untrustworthy or insincere. For example, a traveler who misses a flight due to an unfortunate event might contact an agent to rebook their trip. If the agent is smiling, while words of comfort and understanding are offered, the customer might conclude that these are just words, devoid of any genuine sincerity. Conversely, if the agent's expression conveys surprise or concern, the spoken words are imbued with extra sincerity. However, conflicting verbal content and facial expressions, when appropriate, may not always lead to negative perceptions. For instance, an agent smiling and saying, "I'm sorry, but don't worry, I'll get you on the next available flight," can be perceived as friendly and supportive, easing a potentially stressful situation for the traveler. However, facial expressions can be a matter of degree. A wide grin might be perceived as amusement at the traveler's predicament or distraction from something amusing that the camera didn't capture, but a slight smile is often better perceived as reassuring and friendly.
[0022] These and other needs are addressed through various embodiments and configurations of the present invention. The present invention can provide numerous advantages depending on the specific configuration. These and other advantages will be apparent from the disclosure of one or more of the inventions contained herein.
[0023] In one embodiment, face transformation technology (FTT) is provided to manipulate video images of a proxy face presented to a customer or other party watching a video, in order to improve customer perception and enhance the outcome of the interaction.
[0024] In another embodiment, human labeling and / or machine observation are provided to observe during the interaction to record when the agent uses specific facial expressions (e.g., smiling, raised eyebrows, concerned expressions, etc.), along with the state and circumstances of the call (e.g., topic, duration of the interaction, emotion conveyed via voice, emotional trajectory, information exchanged or to be exchanged), and the results of one or more sub-stages of the interaction or interaction that create / modify a training database. Machine learning is then used to identify the best outcomes and associated facial expressions in the state and / or circumstances of the call to determine best practices for facial expressions in order to optimize future interactions. In another embodiment, a real-time system is provided that uses best practices determined by machine learning and / or other inputs to manipulate agent faces in a video stream during the interaction to alter or further secure the outcome of the interaction.
[0025] In another embodiment, the interaction can be paired with and provide different video modifications and evaluation results to further optimize the interaction and identify which modifications are successful and / or when certain modifications are successful.
[0026] In another embodiment, actual proxies and human overlays can be recorded, such as for quality management and review processes. In a further embodiment, codes such as color codes can be applied to easily categorize quality management records, such as records indicating when an operation was used or not.
[0027] In one embodiment, a system for providing context-matched facial expressions in video images is disclosed, comprising: a communication interface configured to receive video images of a human agent interacting with a client utilizing a client communication device via a network; a processor having accessible memory; a data storage device configured to maintain data records accessible to the processor; and the processor being configured to: receive the video images of the human agent; determine a desired facial expression of the human agent; modify the video images of the human agent to include the desired facial expression; and present the modified video images of the human agent to the client communication device.
[0028] In another embodiment, a method is disclosed, comprising: receiving a video image of a human agent interacting with a customer via a network through an associated customer communication device; determining a desired facial expression of the human agent; modifying the video image of the human agent to include the desired facial expression; and presenting the modified video image of the human agent to the customer communication device.
[0029] In another embodiment, a system is disclosed, comprising: components for receiving video images of a human agent interacting with a customer via a network through an associated customer communication device; components for determining a desired facial expression of the human agent, wherein the desired facial expression is selected based on facial expressions associated with attributes of the interaction and successful outcomes of past interactions having those attributes; components for modifying the video images of the human agent to include the desired facial expression; and components for presenting the modified video images of the human agent to the customer communication device.
[0030] In another embodiment, a system for providing context-matched vocal attributes in the audio portion of a communication is disclosed. The system includes: a communication interface configured to receive audio including speech from a human agent interacting with a client using a client communication device via a network; a processor having accessible memory; a data storage device configured to maintain data records accessible to the processor; and the processor configured to: receive the audio of the human agent's speech; determine desired vocal attributes of the human agent's speech; modify the audio of the human agent's speech to include the desired vocal attributes; and present the modified audio of the human agent's speech to the client communication device.
[0031] In another embodiment, a method is disclosed, comprising: receiving audio of speech by a human agent during an interaction with a customer via an associated customer communication device over a network; determining desired vocal attributes of the human agent's speech; modifying the audio of the human agent's speech to include the desired vocal attributes; and presenting the modified audio of the human agent's speech to the customer communication device.
[0032] In another embodiment, a system is disclosed, comprising: means of receiving audio of speech by a human agent during an interaction with a customer via an associated customer communication device over a network; means of determining a desired vocal attribute of the human agent's speech, wherein the desired vocal attribute is selected based on a vocal attribute associated with an attribute of the interaction and a successful outcome of a past interaction having the vocal attribute; means of modifying the audio of the human agent's speech to include the desired vocal attribute; and means of presenting the modified audio of the human agent's speech to the customer communication device.
[0033] Any one or more aspects of the above embodiments include one or more of the following: Determining the desired vocal attributes of a human agent includes accessing data records containing records that match the topic of the interaction, and wherein the records identify the desired vocal attributes. The determination of the desired vocal attributes of the human agent includes accessing the current customer attributes of the customer and a record in the data record that has stored customer attributes that match the current customer attributes, wherein the record identifies the desired vocal attributes. The determination of the expected vocal attributes of the human agent includes accessing records of expected customer impressions that have human agent attributes that match the topics of interaction, and wherein the record identifies the expected vocal attributes. It also includes the processor storing at least one of the speech of a human agent or modified audio of the speech of a human agent in a data storage device; The processor modifies the speech of the human agent to include desired vocal attributes, including applying changes to at least one of the following: speech rate, intonation, vibrato, rhythm, accent, pitch variation, or brisk tone. Specifically: once the current vocal attributes are determined, the processor modifies the human agent's speech to include the desired vocal attributes; and the processor determines that the current vocal attributes do not match the desired vocal attributes.
[0034] Specifically, once it is determined that the current vocal attribute and the expected vocal attribute provide the same emotional expression to a degree of mismatch in the same emotional expression, the processor determines that the current vocal attribute does not match the expected vocal attribute. It also includes the processor storing a successful interaction flag in the data storage device and at least one of the associated human agent’s expected vocalization attribute or current vocalization attribute; The processor determines the desired vocal attributes of the human agent by determining at least one of the desired vocal attributes or the current vocal attributes of the human agent having a success tag stored in a data storage device. Determining the expected vocal attributes of the speech of a human agent further includes accessing a data storage device containing a record of a topic that matches the topic of the interaction, wherein the record identifies the expected vocal attributes. Determining the expected vocal attributes of a human agent's speech further includes accessing the current customer attributes of a customer in a record of stored customer attributes that match the current customer attributes in a data storage device, wherein the record identifies the expected vocal attributes. The determination of the expected vocal attributes of the human agent's speech also includes records of expected customer impressions in access data records that have human agent attributes that match the topic of the interaction, and wherein the records identify the expected vocal attributes. It also includes storing at least one of audio or modified audio in a data storage device; The modification includes altering the audio of the human agent's speech to include desired vocal attributes, and also includes applying changes to at least one of the following: speech rate, intonation, vibrato, rhythm, accent, pitch variation, or upbeat tone of the human agent's speech. The modification of the human agent's audio to include the desired vocal attributes also occurs when the current vocal attributes are first determined and it is determined that the current vocal attributes do not match the desired vocal attributes. Determining a mismatch between the current vocal attribute and the expected vocal attribute also includes determining that the current vocal attribute and the expected vocal attribute provide the same expression to the same degree of mismatch; and It also includes storing in the data storage device a successful interaction marker and at least one of the associated human agent’s expected vocalization attribute or current vocalization attribute.
[0035] The phrases “at least one,” “one or more,” “or,” and “and / or” are open-ended expressions that are both connected and separated in operation. For example, each of the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” “A, B, and / or C,” and “A, B, or C” means only A, only B, only C, A and B together, A and C together, B and C together, or A, B, and C together.
[0036] The term "a" refers to one or more of the same entity. Therefore, the terms "a," "one or more," and "at least one" are used interchangeably herein. It should also be noted that the terms "comprising," "including," and "having" are used interchangeably.
[0037] As used herein, the term "automatic" and its variations refer to any process or operation that is typically continuous or semi-continuous and can be performed without material human input when executed. However, a process or operation can be automatic even if the execution of the process or operation uses material or non-material human input, if that input is received prior to the execution of the process or operation. Human input is considered material if it influences how the process or operation will be performed. Human input consenting to the execution of a process or operation is not considered "material."
[0038] Various aspects of this disclosure may take the form of a completely hardware embodiment, a completely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which are generally referred to herein as a "circuit," "module," or "system." Any combination of one or more computer-readable media may be used. The computer-readable media may be a computer-readable signal medium or a computer-readable storage medium.
[0039] Computer-readable storage media can be, for example, but not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination thereof. More specific examples (not an exhaustive list) of computer-readable storage media will include the following: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable optical disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the context of this document, a computer-readable storage medium can be any tangible, non-transitory medium that can contain or store programs used by or in connection with an instruction execution system, apparatus, or device.
[0040] Computer-readable signal media may include propagated data signals (e.g., in baseband or as part of a carrier wave) on which computer-readable program code is implemented. Such propagated signals may take any of a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer-readable storage medium and can transmit, propagate, or transfer a program used by or in connection with an instruction execution system, apparatus, or device. Program code implemented on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, or any suitable combination thereof.
[0041] As used herein, the terms “determine,” “calculate,” and their variations are used interchangeably and include any type of method, process, mathematical operation, or technique.
[0042] As used herein, the term "component" shall be given its broadest possible interpretation in accordance with Section 112(f) and / or paragraph 6 of Section 112 of 35 USC. Therefore, claims containing the term "component" shall cover all structures, materials, or actions set forth herein and all their equivalents. Furthermore, structures, materials, or actions and their equivalents shall include all that is described in the summary of the invention, the description of the drawings, the detailed description, the abstract of the specification, and the claims themselves.
[0043] The foregoing is a simplified summary of the invention to provide an understanding of some aspects thereof. This summary is neither a broad overview nor an exhaustive summary of the invention and its various embodiments. It is not intended to identify key or defining elements of the invention, nor to define the scope of the invention, but rather to present selected concepts of the invention in a simplified form as an introduction to the more detailed description that follows. As will be appreciated, other embodiments of the invention may use one or more features set forth above or described in detail below, alone or in combination. Moreover, although this disclosure is presented in the form of exemplary embodiments, it should be understood that various aspects of this disclosure may be claimed separately. Attached Figure Description
[0044] This disclosure is described in conjunction with the accompanying drawings: Figure 1 A first system according to an embodiment of the present disclosure is described; Figure 2 Video image manipulation according to embodiments of the present disclosure is depicted; Figure 3 A second system according to an embodiment of the present disclosure is described; Figure 4 A first data structure according to an embodiment of the present disclosure is described; Figure 5 A second data structure according to an embodiment of the present disclosure is described; Figure 6 A first process according to an embodiment of the present disclosure is described; Figure 7 A fourth system according to an embodiment of the present disclosure is described.
[0045] Figure 8 Audio manipulation according to embodiments of the present disclosure is described; Figure 9 A fifth system according to an embodiment of the present disclosure is described; Figure 10 A third data structure according to an embodiment of the present disclosure is described; Figure 11 A fourth data structure according to embodiments of the present disclosure is described; and Figure 12 A second process according to an embodiment of the present disclosure is described. Detailed Implementation
[0046] The following description provides only examples and is not intended to limit the scope, applicability, or configuration of the claims. Rather, the following description is intended to provide those skilled in the art with an enabling description for implementing the embodiments. It will be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of the appended claims.
[0047] When a sub-element identifier is present in the accompanying drawings, any reference numeral in the description that includes an element number but does not have a sub-element identifier, when used in the plural form, is intended to refer to any two or more elements with similar element numbers. When such reference numeral is used in the singular form, it is intended to refer to one of the elements with similar element numbers, but not limited to a specific element. Any explicit use to the contrary or any further qualification or identification provided herein shall prevail.
[0048] Exemplary systems and methods of this disclosure will also be described with respect to analysis software, modules, and associated analysis hardware. However, to avoid unnecessarily obscuring this disclosure, well-known structures, components, and devices are omitted in the following description, or may be shown in a simplified form or otherwise summarized in the accompanying drawings.
[0049] For purposes of explanation, numerous details have been set forth in order to provide a thorough understanding of this disclosure. However, it should be understood that this disclosure may be practiced in various ways beyond the specific details set forth herein.
[0050] Now for reference Figure 1Based on at least some embodiments of this disclosure, a communication system 100 is discussed. The communication system 100 may be a distributed system and, in some embodiments, includes a communication network 104 connecting one or more communication devices 108 to a work assignment agency 116, which may be owned and operated by an enterprise managing a contact center 102, in which multiple resources 112 are distributed to handle incoming work items (in the form of contacts) from client communication devices 108.
[0051] Contact center 102 is implemented in various ways to receive and / or send messages as work items and the processing and management of work items (e.g., scheduling, allocation, routing, generation, billing, receiving, monitoring, viewing, etc.) or associated with one or more resources 112. A work item is generally a message or a component thereof generated electronically and / or electromagnetically transmitted in response to a request generated and / or received by processing resource 112. Contact center 102 may include more or fewer components and / or provide more or fewer services than shown. The boundaries indicating contact center 102 may be physical boundaries (e.g., buildings, campuses, etc.), legal boundaries (e.g., companies, enterprises, etc.), and / or logical boundaries (e.g., resources 112 providing services to customers of contact center 102).
[0052] Furthermore, the boundaries of contact center 102 may be as shown in the figures, or in other embodiments may include changes and / or more and / or fewer components than shown. For example, in other embodiments, one or more of resources 112, customer database 118, and / or other components may be connected to routing engine 132 via communication network 104, such as when these components are connected via a public network (e.g., the Internet). In another embodiment, communication network 104 may be a dedicated use of at least partially public networks (e.g., VPNs) that can be used to provide electronic communications for the components described herein; a private network located at least partially within contact center 102; or a hybrid of private and public networks. Furthermore, it should be recognized that components shown as external (such as social media server 130 and / or other external data sources 134) may be physically and / or logically within contact center 102, but are still considered external for other purposes. For example, contact center 102 may operate social media server 130 (e.g., a website operable to receive user messages from customers and / or resources 112) as a means of interacting with customers via its customer communication device 108.
[0053] Customer communication devices 108 are implemented outside contact center 102 because they are under the more direct control of their respective users or customers. However, embodiments may be provided whereby one or more customer communication devices 108 are physically and / or logically located within contact center 102 and are still considered outside contact center 102, such as when a customer uses customer communication device 108 located at a kiosk and attached to a private network of contact center 102 that is within or controlled by contact center 102 (e.g., a WiFi connection to the kiosk, etc.).
[0054] It should be recognized that the description of contact center 102 provides at least one embodiment, thereby making the following embodiments easier to understand without limiting these embodiments. Further modifications, additions, and / or reductions to contact center 102 may be made without departing from the scope of any embodiments described herein, and unless expressly provided, this does not limit the scope of the embodiments or claims.
[0055] In addition, contact center 102 may combine and / or utilize social media server 130 and / or other external data sources 134, which can be used as a means for resource 112 to receive and / or retrieve contacts and connect to contact center 102. Other external data sources 134 may include data sources such as service bureaus or third-party data providers (e.g., credit agents, public and / or private records, etc.). Customers can use their respective customer communication devices 108 to send / receive communications using social media server 130.
[0056] According to at least some embodiments of this disclosure, communication network 104 may include any type of known communication medium or set of communication media, and may use any type of protocol to transmit electronic messages between endpoints. Communication network 104 may include wired and / or wireless communication technologies. The Internet is an example of communication network 104, constituting an Internet Protocol (IP) network composed of numerous computers, computing networks, and other communication devices located around the world connected by numerous telephone systems and other means. Other examples of communication network 104 include, but are not limited to, standard ordinary legacy telephone systems (POTS), Integrated Services Digital Network (ISDN), Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Session Initiation Protocol (SIP) network, Voice over IP (VoIP) network, cellular network, and any other type of packet-switched or circuit-switched network known in the art. Furthermore, it is appreciated that communication network 104 is not required to be limited to any one network type, but may include many different networks and / or network types. As an example, embodiments of this disclosure may be used to improve the efficiency of grid-based contact center 102. An example of a grid-based contact center 102 is described more fully in U.S. Patent Publication No. 2010 / 0296417 granted to Steiner, the entire contents of which are incorporated herein by reference. Furthermore, the communication network 104 may include a variety of different communication media, such as coaxial cables, copper / electrical cables, fiber optic cables, antennas for transmitting / receiving wireless messages, and combinations thereof.
[0057] Communication device 108 may correspond to a customer's communication device. According to at least some embodiments of this disclosure, a customer may use their communication device 108 to initiate work items. Exemplary work items include, but are not limited to, communications received at and directed to contact center 102, web page requests received at and directed to a server group (e.g., a set of servers), media requests, application requests (e.g., requests for application resource locations on remote application servers, such as SIP application servers), etc. Work items may take the form of messages or sets of messages sent over communication network 104. For example, work items may be sent as telephone calls, packets or sets of packets (e.g., IP packets sent over an IP network), email messages, instant messages, SMS messages, faxes, and combinations thereof. In some embodiments, communication may not be directed to work assignment agency 116, but may be received by work assignment agency 116 on another server (such as social media server 130) within communication network 104, which generates work items for the received communication. Examples of such received communication include social media communications received by work assignment agency 116 from a social media network or server 130. Exemplary architectures for harvesting social media communications and generating work items based thereon are described in U.S. Patent Applications Nos. 12 / 784,369, 12 / 706,942, and 12 / 707,277, filed March 20, 2010, February 17, 2010, and February 17, 2010, respectively; the entire contents of each application are incorporated herein by reference.
[0058] The format of a work item can depend on the capabilities of communication device 108 and the format of the communication. Specifically, a work item is executed within contact center 102 in association with providing services for communications received at contact center 102 (and more specifically, work assignment agency 116). Communication can be received and maintained at work assignment agency 116, switches connected to work assignment agency 116, or servers, until resource 112 is assigned to the work item representing that communication. At this point, work assignment agency 116 passes the work item to routing engine 132 to connect the initiating communication device 108 to the assigned resource 112.
[0059] Although routing engine 132 is depicted as separate from work assignment mechanism 116, routing engine 132 may be incorporated into work assignment mechanism 116, or its functions may be performed by work assignment engine 120.
[0060] According to at least some embodiments of this disclosure, communication device 108 may include any type of known communication equipment or a collection of communication equipment. Examples of suitable communication devices 108 include, but are not limited to, personal computers, laptop computers, personal digital assistants (PDAs), cellular phones, smartphones, telephones, or combinations thereof. Generally, each communication device 108 may be adapted to support video, audio, text, and / or data communication with other communication devices 108 and processing resources 112. The type of medium used by communication device 108 for communicating with other communication devices 108 or processing resources 112 may depend on the communication applications available on communication device 108.
[0061] According to at least some embodiments of this disclosure, work items are sent to a set of processing resources 112 via a combined effort of work assignment mechanism 116 and routing engine 132. Resources 112 may be fully automated resources (e.g., interactive voice response (IVR) units, microprocessors, servers, etc.), human resources using communication devices (e.g., human agents using computers, telephones, laptops, etc.), or any other resources known to be used in contact center 102.
[0062] As discussed above, the work assignment agency 116 and resources 112 can be owned and operated by a public entity in the form of a contact center 102. In some embodiments, the work assignment agency 116 can be managed by multiple enterprises, each with its own dedicated resources 112 connected to the work assignment agency 116.
[0063] In some embodiments, the work assignment agency 116 includes a work assignment engine 120 that enables the work assignment agency 116 to make intelligent routing decisions for work items. In some embodiments, the work assignment engine 120 is configured to manage and make work assignment decisions in a queueless contact center 102, as described in U.S. Patent Application Serial No. 12 / 882,950, the entire contents of which are incorporated herein by reference. In other embodiments, the work assignment engine 120 may be configured to perform work assignment decisions in a conventional queue-based (or skill-based) contact center 102.
[0064] The job assignment engine 120 and its various components may reside within the job assignment agency 116 or on multiple different servers or processing devices. In some embodiments, a cloud-based computing architecture may be employed, thereby making one or more components of the job assignment agency 116 available in the cloud or network, allowing them to be shared resources among multiple different users. The job assignment agency 116 may access the customer database 118, such as to retrieve records, profiles, purchase histories, previous work items, and / or other aspects of customers known to the contact center 102. The customer database 118 may be updated in response to work items and / or input from the resource 112 that processes work items.
[0065] It should be recognized that, in addition to embodiments that are entirely on-premises, one or more components of contact center 102 may be implemented, in whole or in part (e.g., in a hybrid) in a cloud-based architecture. In one embodiment, customer communication equipment 108 is connected to one of resources 112 via components that are entirely hosted by a cloud-based service provider, wherein processing and data storage elements may be dedicated to the operator of contact center 102 or shared or distributed among multiple service provider customers of contact center 102.
[0066] In one embodiment, a message is generated by a client communication device 108 and received at a work assignment agency 116 via a communication network 104. Messages received by a contact center 102, such as at a work assignment agency 116, are generally and herein referred to as “liaisons”. A routing engine 132 routes liaisons to at least one of resources 112 for processing.
[0067] Figure 2 Video image manipulation 200 according to embodiments of the present disclosure is depicted. In one embodiment, the original video image 202 includes an image of a human agent captured by a camera during interaction with a client (such as a client communication device 108 implemented to present video images). As will be discussed more fully with respect to the following embodiments, the processor receives the original video image 202 and, once it determines a mismatch between the expression of the human agent presented in the original video image 202 and a desired expression, applies a modification to the original video image 202 to become a modified video image 204, which presents the image of the human agent to the client via the client communication device 108, while the original video image 202 is not provided to the client communication device 108.
[0068] Due to distraction (e.g., thinking about lunch), physical limitations, inadequate training, misunderstanding of the client's task, or other reasons, a person (such as an agent) may not provide the expected facial expressions. When using video, this can be off-putting and reduce the chances of a successful interaction. Presenting the client with an image of an agent that has been proven to increase the success of the interaction can improve the outcome, which could be a reason for the work items associated with the interaction.
[0069] Systems and methods for manipulating real-time video images (such as original video image 202) to transform into modified video image 204 are now more widely available. Older techniques for manipulating still images, when used on computing systems with processors possessing sufficient processing power, memory, and bandwidth, can manipulate individual frames of video images to create a desired manipulated image. In one embodiment, such as mapping a human agent's face by electronically applying markers 210 (e.g., dots) to an image, when manipulation is desired, the geometry of a polygon with swirling patterns identified by markers 210 can be altered, and portions of the original video image 202 are adjusted (e.g., stretched, shrunk, etc.) to fill the modified polygon image and become at least a portion of the modified video image 204. For example, polygon 206A is reshaped to become polygon 206B, and portions of the original image 202 captured within polygon 206A are modified to fill polygon 206B.
[0070] The image of the human agent's face can be manipulated to interpret some of the desired modifications to the image. Changes to the agent's face can also be applied. For example, a smile line 208 can be added. Additionally or alternatively, graphic elements, such as those present when the agent is smiling (e.g., with a smile line), can be removed, and the modified image will make the agent appear more serious, and the smile line is removed among other things. Such graphic elements can be stored as images, such as still or video images of the agent, providing multiple facial expressions to be mapped to polygons (e.g., polygon 206A) and / or algorithmically determined manipulations (e.g., polygon 206A needs to be modified to polygon 206B, selectively applying shading to create smile line 208, etc.). The selection of specific manipulation techniques can be made at least in part by the available bandwidth and / or attributes of the client communication device 108 for the interaction between the human agent and the client communication device 108. For example, if a customer is watching a video on a low-resolution screen (e.g., a cellular phone), more subtle changes in facial expressions can be omitted because the resolution presented to the customer may not be able to render such subtle image components due to the small screen size or the data contained in the low-bandwidth video. Instead, a modified image 204, which includes a greater number of operations from the original video image 202, can be presented to customers using a high-resolution / high-bandwidth connection.
[0071] Figure 3 A system 300 according to an embodiment of the present disclosure is depicted. In one embodiment, a client 302 and a human agent 310 engage in an interaction that includes at least a real-time video image of the human agent 310. The human agent 310 may be implemented as a resource 112 when further implemented as a human utilizing an agent communication device 306. The interaction may also include audio (e.g., speech), text messages, emails, co-browsing, etc., from the human agent 310, the client 302, or both. The human agent 310 has a current facial expression. A server 304 includes at least one processor having memory and also includes, or is such as, a data storage library 314 accessible via a communication interface. The server 304 may receive real-time video images via a camera 308 of the agent communication device 306 and monitor the interaction between the client 302 and the human agent 310. The server 304 may determine a desired facial expression for the human agent 310. The desired facial expression may include a facial expression determined to indicate a desired emotional state that has previously been identified as leading to a greater likelihood of successfully completing the interaction. For example, the server 304 may determine that the desired facial expression includes a smile. If server 304 determines that there is no mismatch (e.g., human agent 310 is smiling and the expected facial expression is also smiling), then server 304 can provide the unmodified original image of human agent 310, as presented in video image 320. However, if server 304 determines that there is a mismatch between the facial expression of human agent 310 and the expected facial expression, then server 304 can access data structure 316 and select a replacement image 318.
[0072] Those skilled in the art will recognize that the data structure 316, including the replacement image 318, is shown with graphically different facial expressions to facilitate understanding of the embodiments and avoid unnecessarily complicating the diagrams and descriptions. The data structure 316 may include multiple records, such as the replacement image 318, for each desired facial expression. It may also be implemented as a computer-readable data structure and / or algorithmic modifications to an image or portions thereof, including but not limited to polygon identifiers mapped to a portion of the image of the human agent 310 and / or manipulations applied to one or more polygons mapped to the face of the human agent 310, graphical elements to be added and / or removed (e.g., smile lines), vectors of markers associated with polygon vertices for repositioning, and / or other graphical image manipulation data / instructions. Therefore, the server 304 can select desired facial expressions and apply the replacement image 318 to a real-time image of the human agent 310 so that the presented video image 320 is a manipulated facial expression of the human agent 310, at least partially determined by the replacement image 318.
[0073] Figure 4 Data structure 400 according to embodiments of the present disclosure is depicted. In one embodiment, data structure 400 is used by at least one processor of server 304 to determine the desired facial expression of human agent 310, such as when the overall facial expression is known and a particular degree or level is known. Server 304 may determine a particular level 402 suitable for a previously selected facial expression 404 and one of the records 406 selected from it. Data structure 400 includes records 406 that identify and / or include specific image manipulations for the desired facial expression. For example, a secondary frown (FR2) may subsequently be identified, such as within data structure 316, to access the specific image manipulations required to make such a facial expression be presented as a video image 320.
[0074] Even within the same type of facial expression, not every facial expression is equivalent. For example, a smile can be a friendly expression, while another smile might be appropriate when something funny happens, and the same can be true for other emotions. As another example, customer 302 might interact with human agent 310, and server 304 might determine that the human agent 310's desired facial expression is a frown, such as showing sadness to sympathize with customer 302 after learning that customer 302 lost their bag on an airline flight. However, a level one frown might be appropriate if customer 302 indicates that the bag contains only a few old clothes, while a frown might be appropriate if customer 302 indicates that the bag contains a very expensive camera. Therefore, data structure 316 could include data structures of various degrees or levels for a particular desired facial expression.
[0075] Figure 5 Data structure 500 according to embodiments of this disclosure is depicted. Humans learn what facial expressions are appropriate and what are inappropriate. This determination is often very intuitive and difficult to quantify. For example, a smile can be perceived as friendly or derogatory (e.g., being mocked). Humans may wish to sympathize with another person and therefore exhibit facial expressions associated with that person's emotions. However, opposite emotions and associated facial expressions can provide assurance, authority, or other states of being associated with a particular interaction. For example, a smiling agent might be presented to a traveler who has lost their bag, saying, "That's easy, we'll get that taken care of," to give the traveler the impression that the problem will be successfully resolved and that the agent is able to facilitate the resolution.
[0076] Therefore, and in one embodiment, data structure 500 includes data records 506 associated with topic 504 and expected emotional response 502. For example, the processor of server 304 may determine that the interaction between customer 302 and human agent 310 includes “topic attribute 2” (e.g., in-flight food). Server 304 may also determine that the expected emotional response includes “understanding”. Thus, a surprise level 1 (SP1) may be selected from data structure 316 and applied to the image of human agent 310. Additionally or alternatively, attributes of customer 302 may be used to determine a particular expected facial expression and / or its degree. For example, customer 302 may be highly expressive and well-associated with a human agent who is also highly expressive. Thus, a particular image manipulation or image manipulation level may be selected. Conversely, customer 302 may be uncomfortable with a highly expressive agent, thus providing a different image manipulation or image manipulation level. Such differences may be based, individually or in part, on customer 302’s culture, gender, geography, age, education, occupation, and / or other attributes, specifically or as part of a particular demographic.
[0077] In another embodiment, machine learning can be provided to determine specific desired facial expressions. For example, server 304 can select alternative desired facial expressions that are not the desired facial expression for human agent 310. If the interaction between client 302 and human agent 310 results in success, then weights are applied to the alternative desired facial expressions, causing them to be selected more frequently or become the desired facial expression.
[0078] Figure 6 A process 600 according to an embodiment of the present disclosure is depicted. In one embodiment, at least one processor (such as the processor of client 302) is configured to execute process 600 when it is implemented as machine-readable instructions for execution. Process 600 begins, and optionally step 602 accesses client attributes. For example, a particular client 302 may prefer highly expressive agents or belong to a demographic that prefers highly expressive agents. Step 604 accesses real-time video of the agent, such as accessing a real-time image of the human agent 310 provided by camera 308 during interaction with client 302. Next, step 608 analyzes the topic of the interaction. Step 608 may be performed by the agent alone (such as by indicating the topic the client wishes to address), by the client alone (such as via input to an interactive voice response (IVR) or other input prior to initiating an interaction with the agent), and / or by monitoring keywords or phrases provided in the interaction.
[0079] Next, step 610 determines the desired impression the agent should convey. For example, it may have been previously determined that, for a particular client and / or topic, the agent should leave a specific impression, such as authority, empathy, respect, friendliness, etc., to increase the prospect of successfully resolving the interaction. Step 612 then selects the desired facial expression based on the desired impression, and in step 614, the agent's current expression is observed. Test 618 determines whether the agent's current expression matches the desired facial expression. If test 618 is determined to be affirmative, then step 620 provides the client with the unmodified image of the agent. However, if test 618 is determined to be negative, then step 622 modifies the facial expression applied to the agent and provides the client with the modified image of the agent. Process 600 can then proceed back to step 608 to analyze subsequent topics, or if the interaction is completed, then process 600 can end.
[0080] Figure 7 A system 700 according to an embodiment of the present disclosure is depicted. In one embodiment, the proxy communication device 306 and / or server 304 may be implemented, in whole or in part, as a device 702 including various components and connections to other components and / or systems. The components are implemented differently and may include a processor 704. The processor 704 may be implemented as a single electronic microprocessor or a multiprocessor device (e.g., multi-core) having components such as control units(s), input / output units(s), arithmetic logic units(s), registers(s), main memory, and / or other components that access (e.g., received via bus 714) information (e.g., data, instructions, etc.), execute instructions, and (again, via bus 714) output data.
[0081] In addition to the components of processor 704, device 702 may also utilize memory 706 and / or data storage device 708 to store accessible data (such as instructions, values, etc.). Communication interface 710 facilitates communication between components such as processor 704 and components not accessible via bus 714. Communication interface 710 may be implemented as a network port, card, cable, or other configured hardware device. Additionally or alternatively, input / output interface 712 is connected to one or more interface components to receive and / or present information (e.g., instructions, data, values, etc.) to human and / or electronic devices. Examples of input / output devices 730 that may be connected to input / output interface 712 include, but are not limited to, keyboards, mice, trackballs, printers, displays, sensors, switches, relays, etc. In another embodiment, communication interface 710 may include or be included with input / output interface 712. Communication interface 710 may be configured to communicate directly with networked components or utilize one or more networks (such as network 720 and / or network 724).
[0082] The communication network 104 may be implemented in whole or in part as network 720. Network 720 may be a wired network (e.g., Ethernet), a wireless network (e.g., WiFi, Bluetooth, cellular, etc.), or a combination thereof and enables device 702 to communicate with network component(s) 722(s).
[0083] Additionally or alternatively, one or more other networks may be utilized. For example, network 724 may represent a second network that facilitates communication with the components used by device 702. For example, network 724 may be the internal network of contact center 102, whereby the components are trusted (or at least more trusted) than the networking component 722, which may be connected to network 720, including potentially untrusted public networks (e.g., the Internet). Components attached to network 724 may include memory 726, data storage device 728, one or more input / output devices 730, and / or other components accessible to processor 704. For example, memory 726 and / or data storage device 728 may completely supplement or replace memory 706 and / or data storage device 708, either entirely or for a specific task or purpose. For example, memory 726 and / or data storage device 728 may be an external data repository (e.g., a server farm, array, "cloud," etc.) and allow device 702 and / or other devices to access the data thereon. Similarly, processor 704 may access one or more input / output devices 730 via input / output interface 712 and / or via communication interface 710, or directly, via network 724, via network 720 (not shown), or via networks 724 and 720.
[0084] It should be recognized that computer-readable data can be sent, received, stored, processed, and presented by various components. It should also be recognized that the components shown can control other components, whether illustrated herein or otherwise. For example, an input / output device 730 may be a router, switch, port, or other communication component, such that a specific output of processor 704 enables (or disables) the input / output device 730 that may be associated with network 720 and / or network 724 to allow (or disallow) communication between two or more nodes on network 720 and / or network 724. For example, a connection between a specific client using specific client communication device 108 and a specific networking component 722 and / or a specific resource 112 may be enabled (or disabled). Similarly, a specific networking component 722 and / or resource 112 may be enabled (or disabled) to communicate with specific other networking components 722 and / or resources 112 (in some embodiments, including device 702), and vice versa. Those skilled in the art will recognize that other communication equipment may be used in addition to or in lieu of those described herein without departing from the scope of the embodiments.
[0085] Figure 8 Audio manipulation 800 according to embodiments of the present disclosure is depicted. In one embodiment, raw audio 802 includes speech, utterances, and / or other vocal attributes of a human agent captured by a microphone during interaction with a client (such as a client using a client communication device 108 implemented to present audio). As will be discussed more fully with respect to the following embodiments, the processor receives raw audio 802 and, once it determines a mismatch between the vocal attributes of the human agent presented in the raw audio 802 and the desired vocal attributes, applies a modification to the raw audio 802 to become modified audio 804, thereby presenting the human agent's audio to the client via the client communication device 108, and the raw audio 802 is not provided to the client communication device 108.
[0086] Due to distraction (e.g., thinking about lunch), physical limitations, inadequate training, misunderstanding of the customer's task, or other reasons, a person (such as an agent) may not provide the desired audio attributes. When using audio, especially when audio is provided without video, this can be off-putting and reduce the chances of a successful interaction. However, if the agent presents the customer with speech that has audio attributes identified as increasing the success of the interaction, the outcome of the interaction can be improved, which can be the reason for the work items associated with the interaction.
[0087] Systems and methods are provided to manipulate real-time audio, including speech (such as raw audio 802), into modified audio 804. In one embodiment, the audio of a human agent is analyzed to determine attributes other than explicitly spoken words, such as ratings and / or categories of various attributes. Vocal attributes may include tone, speech rate, flutter, overall pitch, breathing, and / or other vocal attributes and / or changes, rates of change, degrees of change, or differences between two or more parts of speech. For example, speech part 806A may be modified to become speech part 808A (e.g., with greater amplitude / louder), speech part 806B may be modified to become speech part 808B (e.g., slower), and speech part 806C may be modified to become speech part 808C (e.g., lower bass).
[0088] Manipulating the speech of a human agent can interpret some or all of the desired modifications to the audio. In other embodiments, manipulation of the agent's vocal content can also be applied, such as adding or removing pauses, adding or removing sighs, or adding or removing nonverbal utterances.
[0089] Figure 9A system 900 according to an embodiment of the present disclosure is depicted. In one embodiment, a client 302 and a human agent 310 engage in an interaction that includes at least audio containing speech provided by the human agent 310. The human agent 310 may be implemented as a resource 112 when further implemented as a person using an agent communication device 306. The interaction may also include video, text messages, emails, co-browsing, etc., from the human agent 310, the client 302, or both. The current vocal attributes of the human agent 310 convey expressions (e.g., words) embedded in the speech utterances 904 (e.g., words) provided by the human agent 310. A server 304 includes at least one processor with memory and also includes, or is such as, a data storage library 314 accessible via a communication interface. The server 304 may receive audio from the microphone 904 of the agent communication device 306 and monitor the interaction between the client 302 and the human agent 310. The server 304 may determine desired vocal attributes for the human agent 310. The expected vocal attributes may include vocal attributes determined to indicate a desired emotional state that has previously been identified as having a greater likelihood of leading to a successful interaction. For example, server 304 may determine that the expected vocal attributes include a cheerful tone (e.g., rising and falling tones typically associated with friendliness). If server 304 determines that there is no mismatch (e.g., human agent 310 is speaking in a manner that provides a cheerful tone), then server 304 may provide the unmodified raw audio of human agent 310, as presented in presented audio 910. However, if server 304 determines that there is a mismatch between the vocal attributes of human agent 310 and the expected vocal attributes, then server 304 may access data structure 906 and choose to modify 908.
[0090] Those skilled in the art will recognize that the data structure 906, including modification 908, is shown as having graphically different sound waves in order to facilitate understanding of the embodiments and avoid unnecessarily complicating the figures and descriptions. Data structure 906 may include multiple records for each desired sound attribute, such as modification 908, and may also be further implemented as a computer-readable data structure and / or algorithmic modifications (one or more) to the sound waves or portions thereof (including, but not limited to, portions mapped to manipulated sound attributes), such as which syllables should be stressed, which parts should be spoken faster or slower, which parts should be spoken at higher or lower pitches, which parts should be spoken evenly or varied in different parts, etc.
[0091] Therefore, server 304 can select desired vocal attributes and apply modification 908 to the real-time speech of human agent 310 so that the presented audio 910 is speech of human agent 310 as if manipulated, having at least in part the vocal attributes determined by modification 908.
[0092] As is known in the art and in one embodiment, the neural network self-configures layers with logical nodes having inputs and outputs. If the output is below a self-determined threshold level, then the output is ignored (i.e., the input is within the inactive response portion of the scale and no output is provided); if the self-determined threshold level is above the threshold, then an output is provided (i.e., the input is within the active response portion of the scale and an output is provided). Specific placement of active and inactive delineation is provided as one or more training steps. A multidimensional plane (e.g., a hyperplane) is generated for multiple inputs to the nodes to delineate combinations of active or inactive inputs.
[0093] In one embodiment, a computer-implemented method for training a neural network to address affective intonation mismatch in spoken communication on the internet includes: collecting a set of previous audio recordings, the audio recordings including speech from agents participating in communication with clients on the internet; applying one or more transformations to at least a portion of each audio recording, including pitch increase, pitch decrease, speed increase, speed decrease, increase vibrato, decrease vibrato, increase brisk tone, decrease brisk tone, increase smoothness, decrease smoothness, increase volume, decrease volume, increase monotone, decrease monotone, increase any one or more of the above changes, and decrease any one or more of the following changes to create a set of modified audio recordings; creating a first training set comprising the collected set of audio recordings, the set of modified audio recordings, and a set of affective neutral audio recordings (e.g., flat, monotone, etc.); training a neural network using the first training set in a first phase; creating a second training set for a second phase of training, comprising the first training set and affective neutral audio recordings incorrectly detected as having affective content after the first phase of training; and training the neural network using the second training set in a second phase.
[0094] Figure 10A data structure 1000 according to embodiments of the present disclosure is depicted. In one embodiment, the data structure 1000 is used by at least one processor of server 304 to determine desired vocal attributes for human agent 310, such as when the overall vocal attributes are known and a particular degree or level is known. Server 304 may determine a particular level 1002 suitable for a previously selected vocal attribute 1004 and one of the selected records 1006. Data structure 1000 includes records 1006 that identify and / or include specific audio manipulations for the desired vocal attributes. For example, a secondary flutter (FL2), such as within data structure 906, may subsequently be identified to access specific audio manipulations required to make such vocal attributes available as presented audio 910.
[0095] Even within the same type of vocal attributes, not every vocal attribute is equivalent. For example, an increase in tone or vibrato can be an expression of friendliness, while another increase can be an indication of happiness, and so can other emotions. As another example, customer 302 may interact with human agent 310, and server 304 may determine that the human agent 310's desired vocal attribute is a frown, such as expressing sadness to sympathize with customer 302 after learning that customer 302 lost a bag on an airline flight. However, if customer 302 indicates that the bag contains only a few old clothes, then lowering the pitch by one level might be appropriate, while if customer 302 indicates that the bag contains a very expensive camera, then a higher level of pitch reduction might be appropriate. Thus, data structure 316 can include data structures for various degrees or levels of a particular desired vocal attribute.
[0096] Figure 11 Data structure 1100 according to embodiments of this disclosure is depicted. Humans learn which vocal attributes are appropriate and which are inappropriate. This determination is often very intuitive and difficult to quantify. For example, a giggle can be interpreted as friendly or derogatory (e.g., being mocked). Humans may wish to sympathize with another person and therefore exhibit vocal attributes associated with that person's emotions. However, the opposite emotion and associated vocal attributes can provide assurance, authority, or other status associated with a particular interaction. For example, a cheerful agent might be introduced to a traveler who has lost their bag, saying, "That's easy, we'll get that taken care of," to give the traveler the impression that the problem will be successfully resolved and that the agent can facilitate the resolution.
[0097] Therefore, and in one embodiment, data structure 1100 includes data records 1106 associated with topic 1104 and expected emotional response 1102. For example, the processor of server 304 may determine that the interaction between customer 302 and human agent 310 includes “topic attribute 2” (e.g., in-flight food). Server 304 may also determine that the expected emotional response includes “understanding”. Thus, a surprise level 1 (with a speech rate level 3 (PA3)) may be selected from data structure 906 and applied to the speech of human agent 310. Additionally or alternatively, attributes of customer 302 may be used to determine specific expected facial expressions and / or their degree. For example, customer 302 may be highly expressive and well-associated with a human agent who is also highly expressive. Thus, a specific audio manipulation or audio manipulation level may be selected. Conversely, customer 302 may be uncomfortable with a highly expressive agent, thus providing a different audio manipulation or audio manipulation level. Such differences may be based, individually or in part, on customer 302’s culture, gender, geography, age, education, occupation, and / or other attributes, specifically or as part of a particular demographic.
[0098] In another embodiment, machine learning can be provided to determine specific desired vocal attributes. For example, server 304 can select alternative desired vocal attributes that are not the desired vocal attribute for human agent 310. If the interaction between client 302 and human agent 310 is successful, then weights are applied to the alternative desired vocal attributes, causing them to be selected more frequently or become the desired vocal attribute.
[0099] Figure 12 A process 1200 according to an embodiment of the present disclosure is depicted. In one embodiment, at least one processor (such as the processor of client 302) is configured to execute process 1200 when it is implemented as machine-readable instructions for execution. Process 1200 begins, and optionally step 1202 accesses client attributes. For example, a particular client 302 may prefer a highly expressive agent or belong to a demographic that prefers highly expressive agents. Step 1204 accesses real-time audio of the agent, such as via microphone 904, capturing real-time speech provided by human agent 310 during interaction with client 302. Next, step 1208 analyzes the topic of the interaction. Step 1208 may be performed by the agent alone (such as by indicating the topic the client wishes to address), by the client alone (such as via input to an interactive voice response (IVR) or other input prior to initiating an interaction with the agent), and / or by monitoring keywords or phrases provided in the interaction.
[0100] Next, step 1210 determines the desired impression the agent should provide. For example, it may have been previously determined that, for a specific client and / or topic, the agent should leave a specific impression, such as authority, empathy, respect, friendliness, etc., to increase the prospect of successfully resolving the interaction. Step 1212 then selects the desired vocal attributes based on the desired impression, and in step 1214, the agent's current vocal attributes are determined. Test 618 determines whether the agent's current vocal attributes match the desired vocal attributes. If test 1218 is determined to be affirmative, then step 1220 provides the client with the agent's unmodified audio. However, if test 1218 is determined to be negative, then step 1222 modifies the agent's vocal expression and provides the client with the agent's modified audio. Process 1200 can then proceed back to step 1208 to analyze subsequent topics, or if the interaction is completed, then process 600 can end.
[0101] In the foregoing description, the methods have been described in a particular order for illustrative purposes. It should be understood that, in alternative embodiments, the methods can be performed in a different order than described without departing from the scope of the embodiments. It should also be understood that the methods described above can be executed as an algorithm by a hardware component (e.g., a circuit system) designed to perform one or more algorithms or portions thereof described herein. In another embodiment, the hardware component may include a general-purpose microprocessor (e.g., a CPU, a GPU) that is first converted into a dedicated microprocessor. The dedicated microprocessor, having then loaded with encoded signals, enables the dedicated microprocessor to maintain machine-readable instructions that allow it to read and execute a set of machine-readable instructions derived from the algorithms and / or other instructions described herein. The machine-readable instructions for performing one or more algorithms or portions thereof are not unlimited but utilize a finite set of instructions known to the microprocessor. The machine-readable instructions may be encoded in the microprocessor as signals or values in signal generation components and, in one or more embodiments, include voltages in memory circuitry, configurations of switching circuitry, and / or selective use of specific logic gates. Additionally or alternatively, machine-readable instructions may be microprocessor-accessible and encoded in a medium or device as magnetic fields, voltage values, charge values, reflective / non-reflective portions, and / or physical markings.
[0102] In another embodiment, the microprocessor also includes one or more of the following: a single microprocessor, a multi-core processor, multiple microprocessors, a distributed processing system (e.g., one or more arrays, one or more blades, one or more server farms, a "cloud," one or more multi-purpose processor arrays, one or more clusters, etc.) and / or may be located in the same location as the microprocessor performing other processing operations. Any one or more microprocessors may be integrated into a single processing device (e.g., a computer, server, blade, etc.) or may be wholly or partially located in discrete components connected via communication links (e.g., buses, networks, backplanes, etc., or multiple thereof).
[0103] Examples of general-purpose microprocessors may include a central processing unit (CPU) having data values encoded in an instruction register (other circuitry that maintains instructions) or data values including memory locations, which in turn include values used as instructions. Memory locations may also include memory locations external to the CPU. Such external CPU components may be implemented as one or more of field-programmable gate arrays (FPGAs), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), random access memory (RAM), bus-accessible storage devices, network-accessible storage devices, etc.
[0104] These machine-executable instructions can be stored on one or more machine-readable media, such as CD-ROMs or other types of optical discs, floppy disks, ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory, or other types of machine-readable media suitable for storing electronic instructions. Alternatively, these methods can be executed by a combination of hardware and software.
[0105] In another embodiment, a microprocessor may be a system or collection of processing hardware components, such as microprocessors on client devices and microprocessors on servers, a collection of devices having their respective microprocessors, or a shared or remote processing service (e.g., a "cloud-based" microprocessor). A system of microprocessors may include task-specific allocation of processing tasks and / or shared or distributed processing of tasks. In yet another embodiment, a microprocessor may execute software to provide services to emulate one or more different microprocessors. Thus, a first microprocessor, comprising a first set of hardware components, may virtually provide services to a second microprocessor, wherein the hardware associated with the first microprocessor can operate using the instruction set associated with the second microprocessor.
[0106] While machine-executable instructions may be stored locally and executed on a particular machine (e.g., a personal computer, mobile computing device, laptop computer, etc.), it should be recognized that the storage of data and / or instructions and / or the execution of at least a portion of the instructions may be provided via connectivity to remote data storage devices and / or processing devices or collections of devices (generally referred to as the “cloud,” but may include public, private, dedicated, shared, and / or other service providers, computing services, and / or “server farms”).
[0107] Examples of microprocessors described herein may include, but are not limited to, at least one of the following: Qualcomm® Snapdragon® 800 and 801, Qualcomm® Snapdragon® 610 and 615 with 4G LTE integration and 64-bit computing, Apple® A7 microprocessor with 64-bit architecture, Apple® M7 motion coprocessor, Samsung® Exynos® series, Intel® Core TM Series of microprocessors, Intel® Xeon® series microprocessors, Intel® Atom TM Intel® Itanium® series microprocessors, Intel® Core® i5-4670K and i7-4770K 22nm Haswell, Intel® Core® i5-3570K 22nm Ivy Bridge, AMD® FX TM Series microprocessors, AMD® FX-4300, FX-6300 and FX-8350 32nm Vishera, AMD® Kaveri microprocessors, Texas Instruments® Jacinto C6000 TM Automotive infotainment microprocessor, Texas Instrument® OMAP TM Automotive-grade mobile microprocessors, ARM® Cortex TM -M microprocessor, ARM® Cortex-A and ARM926EJ-S TM Microprocessors, other industrial equivalent microprocessors, and can perform computational functions using any known or future-developed standards, instruction sets, libraries, and / or architectures.
[0108] Any steps, functions, and operations discussed in this article can be performed continuously and automatically.
[0109] Exemplary systems and methods of the present invention have been described in conjunction with communication systems and components and methods for monitoring, enhancing, and modifying communications and messages. However, to avoid unnecessarily obscuring the invention, many known structures and devices have been omitted from the foregoing description. This omission should not be construed as a limitation on the scope of the claimed invention. Specific details have been set forth to provide an understanding of the invention. However, it should be recognized that the invention can be practiced in various ways beyond the specific details set forth herein.
[0110] Furthermore, while the exemplary embodiments illustrated herein depict various components of a co-located system, some components of the system may be located remotely in distant portions of a distributed network (such as a LAN and / or the Internet), or within a dedicated system. Therefore, it should be appreciated that components of the system, or portions thereof (e.g., microprocessors, memory / storage devices, interfaces, etc.), may be combined into one or more devices (such as servers, computers, computing devices, terminals, “clouds”, or other distributed processing), or co-located on specific nodes of a distributed network (such as analog and / or digital telecommunications networks, packet-switched networks, or circuit-switched networks). In another embodiment, components may be physically or logically distributed across multiple components (e.g., microprocessors may include a first microprocessor on one component and a second microprocessor on another component, each microprocessor performing a portion of a shared task and / or an assigned task). As will be appreciated from the foregoing description, and for computational efficiency reasons, components of the system can be arranged anywhere within the distributed network of components without affecting the operation of the system. For example, various components may be located in switches (such as PBXs) and media servers, gateways, in one or more communication devices, in the homes of one or more users, or some combination thereof. Similarly, one or more functional parts of the system may be distributed among one or more telecommunications devices and associated computing devices.
[0111] Furthermore, it should be recognized that the various links connecting the elements can be wired or wireless links or any combination thereof, or any other known or later developed element(s) capable of supplying and / or transmitting data to or from the connected element. These wired or wireless links can also be secure links and can be capable of transmitting encrypted information. For example, the transmission medium used as the link can be any suitable carrier for electrical signals, including coaxial cables, copper wires, and optical fibers, and can take the form of sound waves or light waves, such as those generated in radio waves and infrared data communications.
[0112] Furthermore, although a flowchart has been discussed and illustrated with respect to a specific sequence of events, it should be recognized that this sequence can be altered, added to, or omitted without substantially affecting the operation of the invention.
[0113] Many variations and modifications of the present invention may be used. Some features of the present invention may be provided without providing others.
[0114] In yet another embodiment, the systems and methods of the present invention may be implemented in combination with a dedicated computer, a programmable microprocessor or microcontroller and one or more peripheral integrated circuit elements, an ASIC or other integrated circuit, a digital signal microprocessor, hardwired electronic or logic circuitry (such as discrete component circuitry), a programmable logic device or gate array (such as a PLD, PLA, FPGA, PAL), a dedicated computer, any similar device, etc. In general, any device or apparatus capable of implementing the methods illustrated herein can be used to implement various aspects of the present invention. Exemplary hardware that may be used in the present invention includes computers, handheld devices, telephones (e.g., cellular, internet-enabled, digital, analog, hybrid, and others), and other hardware known in the art. Some of these devices include microprocessors (e.g., single or multiple microprocessors), memory, non-volatile storage devices, input devices, and output devices. Furthermore, alternative software implementations, including but not limited to distributed processing or component / object distributed processing, parallel processing, or virtual machine processing, may be constructed to implement the methods described herein.
[0115] In yet another embodiment, the disclosed method can be readily implemented using software from an object-oriented or object-based software development environment that provides portable source code usable on various computer or workstation platforms. Alternatively, the disclosed system can be implemented partially or entirely in hardware using standard logic circuitry or VLSI design. Whether to implement the system according to the invention using software or hardware depends on the system's speed and / or efficiency requirements, specific functions, and the particular software or hardware system or microprocessor or microcomputer system used.
[0116] In yet another embodiment, the disclosed method can be implemented in part with software that can be stored on a storage medium and executed on a general-purpose computer, special-purpose computer, microprocessor, or the like, which is programmed in cooperation with a controller and memory. In these cases, the system and method of the present invention can be implemented as a program (such as an applet, JAVA®, or CGI script) embedded in a personal computer, as a resource residing on a server or computer workstation, as a routine embedded in a dedicated measurement system, a system component, and so on. The system can also be implemented by physically integrating the system and / or method into a software and / or hardware system.
[0117] The software embodiments described herein are executed or stored by one or more microprocessors for subsequent execution and are executed as executable code. The executable code is selected to execute instructions including those specific to the embodiment. The instructions to be executed are a constrained set of instructions selected from a discrete native instruction set understood by the microprocessor and are submitted to microprocessor-accessible memory prior to execution. In another embodiment, human-readable "source code" software is first converted into system software to include a platform-specific (e.g., computer, microprocessor, database, etc.) set of instructions selected from the platform's native instruction set.
[0118] While this invention describes components and functions implemented in embodiments with reference to specific standards and protocols, the invention is not limited to these standards and protocols. Other similar standards and protocols not mentioned herein exist and are considered to be included in this invention. Moreover, the standards and protocols mentioned herein, as well as other similar standards and protocols not mentioned herein, are periodically replaced by faster or more efficient equivalents with substantially the same functionality. Such replacement standards and protocols with the same functionality are considered to be equivalents included in this invention.
[0119] This invention includes, in various embodiments, configurations, and aspects, components, methods, processes, systems, and / or apparatuses substantially as depicted and described herein, including various embodiments, sub-combinations, and subsets thereof. Upon understanding this disclosure, those skilled in the art will understand how to make and use this invention. This invention includes, in various embodiments, configurations, and aspects, providing devices and processes without items not depicted and / or described herein, or in various embodiments, configurations, or aspects thereof, providing devices and processes without such items as those already used in prior devices or processes, for example, to improve performance, achieve ease of use, and / or reduce implementation costs.
[0120] The foregoing discussion of the invention has been presented for purposes of illustration and description. The foregoing is not intended to limit the invention to the one or more forms disclosed herein. For example, in the foregoing detailed description, various features of the invention are combined in one or more embodiments, configurations, or aspects for the purpose of fluent disclosure. Features of embodiments, configurations, or aspects of the invention may be combined in alternative embodiments, configurations, or aspects other than those discussed above. The approach of this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than those expressly set forth in each claim. Rather, as reflected in the following claims, the inventive aspect lies in fewer than all features of a single foregoing disclosed embodiment, configuration, or aspect. Therefore, the appended claims are hereby incorporated into the detailed description, wherein each claim itself is a separate preferred embodiment of the invention.
[0121] Furthermore, although the description of the invention has included descriptions of one or more embodiments, configurations, or aspects, as well as certain variations and modifications, other variations, combinations, and modifications are also within the scope of the invention, for example, as can be understood by those skilled in the art within the skill and knowledge following this disclosure. It is intended to obtain rights to alternative embodiments, configurations, or aspects to a permissible extent, including alternatives, interchangeable, and / or equivalent structures, functions, scopes, or steps for those claimed, whether or not such alternatives, interchangeable, and / or equivalent structures, functions, scopes, or steps are disclosed herein, and is not intended to publicly contribute any patentable subject matter.
Claims
1. A system for providing context-matched desired vocal attributes in the audio portion of a communication, comprising: The communication interface is configured to receive audio, including speech, from a human agent who interacts with a client using a client communication device via a network. A processor with accessible memory; The data storage device is configured to maintain data records accessible to the processor; as well as The processor is configured as follows: Receive audio of speech spoken by a human agent; Determine the expected vocal properties of the speech of a human agent, including: Access the customer's current customer attributes and the records in the data records that contain stored customer attributes that match the current customer attributes. This record identifies the desired vocal attributes; Modify the audio of the human agent's speech to include the desired vocalization attributes; and Presents modified audio of human agent speech to customer communication devices.
2. The system of claim 1, wherein determining the desired vocal attributes of the human agent includes accessing records in the data records that have a topic matching the topic of the interaction, and wherein the record identifies the desired vocal attributes.
3. The system of claim 1, wherein determining the desired vocal attributes of a human agent includes accessing records of desired customer impressions that have human agent attributes matching the topic of the interaction, and wherein the records identify the desired vocal attributes.
4. The system of claim 1, further comprising the processor storing in a data storage device at least one of human agent speech or modified audio of human agent speech.
5. The system of claim 1, wherein the processor modifies the speech of the human agent to include desired vocal attributes by: The application of at least one of the following alterations to the speech rate, intonation, vibrato, rhythm, accent, intonation variation, or brisk tone of a human agent's speech.
6. The system of claim 1, wherein: Once the current vocal attributes are determined, the processor modifies the human agent's speech to include the desired vocal attributes; and The processor determines that the current vocal attributes do not match the expected vocal attributes.
7. The system of claim 6, wherein, Once the processor determines the degree to which the current vocal attributes and the desired vocal attributes provide the same emotional expression in terms of mismatch, it determines that the current vocal attributes do not match the desired vocal attributes.
8. The system of claim 1, further comprising: The processor stores in the data storage device a marker indicating successful interaction and at least one of the associated human agent's expected vocal attributes or current vocal attributes; as well as The processor determines the expected vocal attributes of the human agent by determining at least one of the expected vocal attributes or the current vocal attributes of the human agent that have a success tag stored in the data storage device.
9. A method for providing context-matched desired vocal attributes in the audio portion of a communication, comprising: Receive audio of speech spoken by a human agent when interacting with a customer via a network through the customer's communication device; Determine the expected vocal properties of the speech of a human agent, including: Access data records containing desired customer impressions that match the topics of interaction with human agent attributes. This record identifies the desired vocal attributes; Modify the audio of the human agent's speech to include the desired vocalization attributes; and Presents modified audio of human agent speech to customer communication devices.
Citation Information
Patent Citations
Pre-qualified or history-based customer service
US20100235218A1
Grid-based contact center
US20100296417A1
Method for determining response channel for a contact center from historic social media postings
US20110125793A1
Stalking social media users to maximize the likelihood of immediate engagement
US20110125826A1
Optimizing interaction results using ai-guided manipulated video
US20210042508A1