System and method for image and sound communication
Patent Information
- Application Number
- PCT/EP2026/054786
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-24
- Filing Date
- 2026-02-23
- Publication Date
- 2026-08-27
Smart Images

Figure EP2026054786_27082026_PF_FP_ABST
Abstract
Description
[0001] tellma GmbH
[0002] Hiebelerstr. 40
[0003] 87629 Füssen
[0004] System and procedure for image and sound communication
[0005] The invention relates to a system for audio-visual communication between a user and a contact person of the user. The system comprises a communication unit for the user and a communication station for the contact person. The communication unit and the communication station are each equipped with a computer, a camera, a microphone, a screen, and a loudspeaker, and are interconnected via data transmission.
[0006] Such an arrangement is also known as a "video call center." In this setup, the user (for example, a bank customer) and the contact person (for example, a bank advisor), who may be located far apart, communicate not only through voice but can also see each other, which significantly enhances the personal, confidential, and intimate nature of the conversation. The basic principle of a call center is that calls are automatically distributed to the contact person. While this leads to high productivity among the contact persons, it inevitably results in a queue if the number of callers exceeds the number of available contact persons. Many users then abandon their calls while in the queue, and the contact person is left waiting.The employer of the contact person loses a potentially lucrative customer or customer inquiry.
[0007] The invention therefore aims to reduce user dropout rates in known video call center applications.
[0008] To solve this problem, the invention first proposes a system for image and sound communication between a user and a contact person of the user, wherein the system comprises a communication unit, a communication station, a control unit and a digital communicator.
[0009] The communication unit, the communication workstation, and the communicator are all connected to the control unit via data transmission. The communication unit includes a computer, camera, microphone, screen, and speaker for the user, while the communication workstation includes a computer, camera, microphone, screen, and speaker for the person communicating. The communicator includes a speech analysis unit, a communication module, and an image and sound generation unit.
[0010] Based on the user's interaction with the system, the communicator produces image information in the form of images of an artificial human (an avatar) and transmits this to the screen of the communication unit.
[0011] The digital communicator (a virtual video agent) is an AI-supported system component that acts as a visually present avatar and comprises several different modules.
[0012] The digital communicator is able to understand the user's spoken language and independently respond to the user's questions and information. Based on the user's interaction with the system, the control unit establishes a connection either between the communication unit and the communicator or between the communication unit and the communication workstation. The control unit orchestrates the connections between all components.
[0013] It establishes connections between
[0014] Communication unit communicator or
[0015] Communication unit Communication station
[0016] and switches dynamically between these modes. It monitors the availability of contacts at the respective communication location. And it reacts to triggers, i.e., active actions by the user, communicator, or contact person.
[0017] The clever aspect of the invention lies in the fact that the user is kept in the queue as long as possible by an interactive AI (artificial intelligence) technology until a transfer to the appropriate contact person in the video call center can take place. The technology used is designed, for example, to either conduct trivial communication or to pre-qualify the user / customer for subsequent transfer to the contact person.
[0018] The system provides the user (the bank customer) with a communication unit. This communication unit is equipped with a camera, a microphone, a screen, and a speaker. The communication unit also includes a computer (a device that processes data using programmable instructions). The camera and microphone are input devices for the computer, while the screen and speaker are output devices. Such equipment is available, for example, on a standard laptop. Therefore, in the context of the invention, a laptop is a variant of this communication unit.
[0019] Advantageously, any internet-enabled device (smartphone, tablet, PC, notebook) is a communication unit within the meaning of the invention. All these devices have a computer, a speaker, a screen, a camera, and a microphone. These devices are then integrated into the system via a browser application. User identification then takes place, for example, when the password-protected application is started.
[0020] A communication unit can also consist of a computer housed in its own enclosure, equipped with a camera, microphone, screen, and speaker located outside the enclosure. These elements are nevertheless connected to the computer via data transmission. This can be, for example, a desktop computer setup or a kiosk-style configuration. In a kiosk-style configuration, the screen is mounted on a stand or column so that the user can view it while standing or sitting in front of it. The screen is sized so that the upper body or head of the displayed contact person or avatar is realistically life-size. These configurations are also part of the invention. The communication unit used exhibits a wide range of configuration options.The communication unit can be, for example, a standard laptop or a customized consultation island or kiosk solution. The physical design of the communication unit is not limited to the variants presented, but is highly variable.
[0021] The system provides a communication workstation for the user's contact person. This workstation also consists of a computer, as well as a camera, a microphone, a monitor, and a speaker, all connected to the computer. The camera and microphone are input devices for the computer, while the monitor and speaker are output devices. Such equipment is available, for example, on a standard laptop or is part of video call center systems. Therefore, it is also intended that the communication workstation be part of such a system, or that the proposed system is integrated into such a system. Such a system has at least one central server to which numerous communication workstations are connected. This system also provides the software necessary for distributing calls and inquiries.In such applications, the speaker at the communication station is often integrated into a headset or headphones. A central component of the system is the digital communicator. The digital communicator (a virtual video agent) is an AI-supported system component that acts as a visually present avatar and includes at least the following modules:
[0022] a) Language analysis unit
[0023] b) Communication module
[0024] c) Image and sound generation unit
[0025] This is a digital assistant designed to engage the user in the most entertaining way possible, both to keep the user in the queue and, potentially, to obtain information that is important or interesting for the business model of the contact person's company. The communicator is a system component, at least partially driven by artificial intelligence, that is capable of responding to the user interactively and appropriately to the situation.
[0026] Demonstrates at least one speech analysis unit that employs AI. This unit is capable of understanding the user's spoken language and independently responding to and answering the user's questions or providing information. It therefore encompasses the functions of language models (such as NLP - Natural Language Processing, STT - Speech-to-Text Transcription, and / or LLM - Large Language Model, capable of understanding complex texts and contexts) and a generative AI (generative AI generates personalized responses based on data, thus creating a more human-like interaction). The speech analysis unit analyzes user input in real time. It possesses language understanding and semantic interpretation capabilities.
[0027] Where appropriate, the speech analysis unit uses company-specific content (from the contact person's company) to respond to the user's questions or input / information. As a result, the speech model generates a response addressed to the user in response to their question, input, or information. Where this application refers to image information, this also includes a sequence of images, such as in a video, and thus also image or video information.
[0028] The customer conversation process is stored in the communication module. A conversation is the chronological sequence of information exchanged between at least two participants. A customer conversation process includes a structured conversation flow designed to gather specific information from the user (customer). This is referred to as pre-qualification. The goal of pre-qualification is to ascertain the user's specific wishes or (potential) needs and to make this information available for further consultation, whether by the communicator or a human agent. The communicator, or the communication module, prepares the conversation according to the criteria of the pre-qualification process (as defined by the system operator's business model) and stores it as a user conversation record.
[0029] The customer conversation process is preferably implemented as an algorithm (a pre-qualification process). Alternatively, it can also be designed as AI. Within the digital communicator, there is a high degree of interaction between the speech analysis unit and the communication module. This ensures targeted and competent communication between the communicator (a digital assistant) and a user (a human).
[0030] The communication module manages the conversation content, conversation history, and context. This content is preferably stored in a user conversation record and saved in the corresponding database.
[0031] Furthermore, the communication module is characterized by its access to a knowledge database, which contains both the customer conversation process and the expert knowledge of the system operator, from which the AI-supported communication module of the communicator draws to answer the user's question or communicate with the user.
[0032] The image and sound generation unit generates an at least acoustically perceptible response message from the response information addressed to the user. This message is then transmitted to the loudspeaker of the communication unit and output there. It therefore possesses text-to-speech (TTS) functionality. Text-to-speech converts flowing text into acoustic speech output. Preferably, in addition to the acoustic response messages, it also generates visual response messages, which appear on a screen, for example, as an animated avatar. The image and sound generation unit or the rendering module synchronizes the audio (the acoustic response message) with the avatar animation. Preferably, this unit is also AI-driven.
[0033] The image and sound generation unit is preferably modular in design and consists in particular of an image generation unit and a sound generation unit, i.e., of two parts.
[0034] Avatars are digital entities with an anthropomorphic appearance, controlled by humans or software, and capable of interaction. An avatar has a human-like appearance and can be controlled by humans or software. Based on real-time 3D engines, interactions can be graphically represented and the avatar animated.
[0035] The typical use of the proposed system involves a user wanting to initiate a video call with a company employee (contact person), for example, to conclude a contract or require information. Such requests are placed in a queue. This is where the communicator, acting as a digital assistant and contact person, comes into play. Therefore, at the start of the video call, the communicator is connected to the user's communication device via the control unit. The user's interaction with the system thus initially results in communication between the user and the digital communicator / avatar.The control unit allows the communication unit to be connected to the communication workstation via data transmission, thus enabling a video call between the user at the communication unit and the contact person at the communication workstation. The control unit is designed to be actively controlled by the user, allowing the user to switch the call from the communicator to the contact person at the communication workstation (provided, of course, that a contact person is available). Alternatively, this switching process can also be automated, based on the progress of the call between the user and the communicator, by sending a corresponding control command from the communication module to the control unit.The proposal stipulates that the control unit either switches the communication station to the communication unit, thus terminating the connection to the communicator, or adds the communication unit to the data connection between the communication station and the communicator. The latter can be advantageous, for example, for the communicator's training purposes.
[0036] The communication unit and the communication workstation are primarily computers (with additional components, as described) that are interconnected via a server. This connection is facilitated by the programs installed on the computers and / or the server. The communicator and the control unit are also installed on this server. For practical or capacity reasons, the individual components of the communicator, particularly the speech analysis unit, the communication module, and / or the audio-visual generation unit, may be located on additional servers. Both the communicator and the control unit are implemented as computer programs running on computers. The various system components are interconnected via data technology, particularly network technology, for example, via the internet and, for data protection, via a virtual private network (VPN) or other network protocols for encrypted connections, such as...SSL, SSH or similar.
[0037] This is a first example of the actual implementation of the system according to the invention, without intending to limit or define the invention to this specific design. It is clear to those skilled in the art that the arrangement and data-technical implementation, division and connection of the communication unit, the communication station, the communicator and the control unit can also be carried out in other ways in accordance with the invention; all these solutions are part of the invention.
[0038] The proposed system consists of the synergistic interaction of several geographically dispersed system components. As previously mentioned, the communication unit can be implemented in a highly flexible manner for the user. Using a standard laptop, the user can access the system from home. The described kiosk variant, on the other hand, can be set up, for example, in the premises of a business (such as a bank or a government office) and can be visited and used there by the user.
[0039] Alternatively, it is also possible to use a mobile device, e.g. a mobile phone or a tablet, as a communication unit.
[0040] The system proposed according to the invention convincingly solves the initial problem. The proposed communicator engages the user in at least a pleasant conversation, or acquires further information valuable to the contact person, and, through the entertaining conversation, ensures a significant reduction in the dropout rate of users waiting in line at video call centers.
[0041] Surprisingly, this is not the only advantage of the inventive proposal. The system also helps to reduce the shortage of skilled workers in call centers or consultant pools, since a highly qualified contact person does not need to be assigned to handle routine conversation content, but rather this is done by the communicator.
[0042] Furthermore, the system according to the invention also ensures a better, more comprehensive customer experience. The customer / user no longer spends their time listening to music while waiting in line, but can instead use the offered conversation with the operator to present and pre-qualify their request. Thus, when using the system according to the invention, the user saves time when dealing with their request and their contact person.
[0043] Furthermore, the initial task is also solved by a method for audio-visual communication between a user and a contact person of the user. The method comprises a communication unit, a communication workstation, a control unit, and a digital communicator.
[0044] The communication unit, the communication workstation, and the communicator are all connected to the control unit via data transmission. The communication unit includes a computer, a camera, a microphone, a screen, and a speaker for the user, while the communication workstation includes a computer, a camera, a microphone, a screen, and a speaker for the person communicating. The communicator comprises a speech analysis unit, a communication module, and an image and sound generation unit. Based on the user's interaction with the system, the communicator produces visual information in the form of images of an artificial human and transmits this information to the screen of the communication unit. The image and sound generation unit then generates an audible message from the information addressed to the user, which is then transmitted to the speaker of the communication unit and played back.The method according to the invention is characterized in that the communication unit is first connected to the communicator by the control unit, thus enabling a conversation between the user and the communicator, and then, depending on a set communication parameter, the control unit connects the communication unit to the communication station, thus transferring the conversation between the user and the communicator to the communication station with the contact person.
[0045] The system also allows users to request a connection with a specific contact person by providing appropriate voice input, for example, to schedule an appointment with a particular contact. The user communicates this request to the system. Ideally, the system has access to a contact database. Upon request from the communication unit, the system searches for the contact person's name and returns an availability signal and information about the contact's workstation. Then, as soon as the contact person is free and available, the system establishes a connection between the communication unit and the user, and the desired contact person's workstation.
[0046] This method also provides that the communication is switched by the control unit from the communication unit to the communication station, or that the communication station is added to the communication between the communication unit and the communicator. The method is preferably implemented with the same elements as the system described above. Preferably, the proposed method is carried out with the system also proposed, without, however, limiting the invention to this. This method also solves the initial problem of the invention and, surprisingly, exhibits the same additional advantages as the system according to the invention.
[0047] The system according to the invention is further characterized by the fact that, in a further development, the communicator includes an image or video analysis unit. The proposed image or video analysis unit is an AI-driven element capable of analyzing the image, image sequences, or video of the user, captured by the camera of the communication unit, with regard to the user's gestures.
[0048] Gestures are the entirety of actions that serve as movements in interpersonal communication. In particular, movements of the arms, hands, and head accompany or replace verbal information. Gestures are signs of nonverbal communication.
[0049] The image and video analysis unit includes various functionalities that can be used individually or in combination:
[0050] - Analysis of the user's gestures, especially their head movements, hand movements or body posture.
[0051] - It also includes facial recognition for user identification. - The unit also performs facial expression and emotion analysis. The latter is used particularly in the emotion monitoring unit, which will be presented below.
[0052] - The unit also performs attention monitoring (e.g.
[0053] (Viewing direction) of the user.
[0054] The image and video analysis unit includes a model that uses machine learning to perform object recognition and computer vision in images and videos. The result is the generation of image content information in text form.
[0055] Furthermore, the proposal advantageously includes a rendering module for the communicator. In the context of image or video production, rendering refers to an automated process that calculates a finished image or video from raw data using 3D, animation, or editing software. It transforms abstract digital information into a visually perceptible form, often with a focus on photorealism, through the simulation of light, colors, and textures.
[0056] One aim of the proposal is to make the unavoidable waiting time in a queue feel shorter for the user through the interplay of acoustic and visual information. Therefore, the invention provides that the image of an artificial person, for example an avatar, which is appropriately animated, appears on the screen of the communication unit.
[0057] The rendering module interacts extensively with the image and sound generation unit and the communication module, delivering or producing an image or...
[0058] Video information, which also integrates appropriate gestures to the speech information prepared simultaneously by the communicator, is modulated onto the avatar, which is displayed as a video or video stream on the communication unit's screen. This significantly enhances the customer experience. By using images of an artificial person who remains clearly recognizable as such, the user is informed that they are (not yet) communicating with a real person, but with the digital communicator. The rendering module generates a photorealistic or stylized avatar representation. The rendering module animates lip movements (lip-sync) to the speech, synchronized with the output acoustic response information. It generates context-appropriate facial expressions and gestures, particularly taking into account any detected heightened emotion parameters.For example, the digital communicator (an avatar) calms the agitated user. The rendering module generates the video stream of the virtual agent (the avatar) in real time.
[0059] Advantageously, the rendering module is designed to generate a photorealistic or stylized avatar that can express emotions through facial expressions and gestures, with the emotions being specified by the speech analysis unit (possibly in conjunction with the communication module) based on the conversational context and the analyzed user mood.
[0060] The rendering module utilizes, among other things, a game engine to generate synchronized lip movements with the support of an audio file. In a preferred embodiment of the proposal, the communication unit is equipped with a user presence sensor that sets a presence parameter in the system while the user is near the unit. When the presence sensor detects a user's presence, this is recorded in the system (the system's software) by setting the presence parameter (for example, in a corresponding database). The operator then initiates the customer call.
[0061] Presence sensors can take many forms. For example, they can be motion sensors, heat sensors, or touch sensors. A motion sensor reacts to the movement of a person in front of the communication unit. A heat sensor registers the heat emitted by a person. Touch sensors, such as a switch in the floor panel the user stands on, a touch-sensitive surface the user places their hand on, or a light barrier or light curtain (with visible or invisible light), are further examples of presence sensors. The presence sensor is connected to the computer of the communication unit via data transmission and is monitored by this computer.If the presence sensor detects the presence of a user in front of the communication unit, the communication unit's computer transmits this information to the server, where it is recorded as a presence parameter. It is clear to those skilled in the art that there are numerous implementation options for this feature.
[0062] A further improvement involves continuously monitoring the presence sensor. This allows for the ongoing monitoring of the user's presence in front of the communication unit. If the user moves away from the communication unit, this is registered by the presence sensor and also reported to the server. The presence parameter is then deleted on the server. This makes it possible to terminate communication between the communication unit and the communicator, or between the user and the system, and to perform other activities as needed. Furthermore, the communication unit is envisaged to have an additional touchscreen and / or the communication unit's screen is designed as a touchscreen and serves as an input device.The proposed touchscreen allows for the implementation of a keyboard, mouse-like pointer control, or similar functions. It also enables the user to enter their signature on the touchscreen. The touchscreen thus serves as an input device for the computer of the communication unit and allows information to be transferred to the system, for example, to the server. Several options are proposed for implementing the touchscreen. In the first option, the existing screen has at least a touch-sensitive area or is entirely designed as a touchscreen.Since this screen is often vertically oriented, it may be more ergonomically advantageous to provide a second touch-sensitive screen in or on the communication unit, either horizontally or tilted at an acute angle to the horizontal.
[0063] In an advantageous embodiment, the communication unit includes a signature panel that can be monitored, controlled, and / or read from either the communication workstation and / or the communicator. Documents can be signed electronically using the signature panel. For example, the documents are displayed on the communication unit's screen by the user's contact person, and the user can then sign them using the signature panel. The signature panel also serves as an input device within the communication unit. It is not merely a replacement for a ballpoint pen that digitally captures handwriting, but also incorporates the user's writing pressure as biometric information. Ideally, the signature panel can be monitored, controlled, and read by the contact person from their communication workstation. This is accomplished via the communication unit's computer.It is also possible for the communicator to monitor, control, or read the signature panel. This proposal thus opens up the possibility that the final steps in a consultation or commissioning process no longer necessarily require the contact person to be present, but rather the digital communicator can be used again. This saves resources on the part of the contact person.
[0064] The communication unit is cleverly designed to include a card reader. This card reader also serves as an input device for the communication unit's computer. The user can hold their customer card to the card reader to identify themselves. This information is then transmitted from the communication unit's computer to the server, where it is recorded that an identified user is now present at a specific communication unit. This identification information is beneficial for the further use of the system or process.
[0065] In a further preferred embodiment, the system includes a communication station availability monitoring unit, which sets an availability parameter when a communication station is available and transmits it to the control unit or the communicator. The user's waiting time in the queue, even when using the proposed communicator, should not be longer than absolutely necessary. This requires that the availability of the contacts at their communication stations be monitored. Therefore, the communication stations are equipped to interact with the proposed communication station availability monitoring unit. For example, the availability parameter can be set by the contact at their communication station (e.g.,A switch, keyboard command, or button in the program running on the communication unit's computer (displayed on the communication unit's screen) can be activated, or this can be done automatically via the communication station control (often part of a video call center solution). For example, the availability parameter of a specific communication station is set when the previous call is completed. If a call between the user and the communicator needs to be transferred to a contact person, the system searches for and finds a communication station with a set availability parameter and then establishes a connection from the communication unit with the user to the communication station with the contact person via the control unit.Furthermore, it is advantageously provided that the microphone of the communication unit serves as an input device for speech information from the user and transmits the speech information to the speech analysis unit of the communicator, and the speech analysis unit includes a speech model that analyzes the speech information and, in conjunction with the communication module, outputs a reaction information to the speech information, and the image and sound generation unit forms an acoustically perceptible reaction message from the reaction information and transmits this to the loudspeaker of the communication unit as an output device for the user.
[0066] The speech information is transmitted via the microphone and computer of the communication unit to the server of the system on which the communicator is installed. The communicator's speech analysis unit analyzes the speech information using a speech model. For this purpose, the speech model or the speech analysis unit may utilize another server on which the speech model is installed. Through the interaction of the speech analysis unit and the communication module, a response information is formulated and output. This response is then converted by the image and sound generation unit into an acoustically perceptible response message, which is then played back on the communication unit's loudspeaker (via the communication unit's computer).
[0067] Advantageously, the microphone of the communication unit serves as an input device for voice information from the user and transmits the voice information to the communicator's voice analysis unit, and the camera of the communication unit serves as an input device for image information from the user and transmits the image information to the communicator's image or video analysis unit.The video analysis unit analyzes the image information and sends the resulting image content information to the speech analysis unit. The speech analysis unit includes a speech model that analyzes the image content information together with the simultaneously received speech information and, in conjunction with the communication module, outputs a reaction information to the combined image content-speech information. The image and sound generation unit, in conjunction with the rendering module (55), forms an acoustically and visually perceptible reaction message from the reaction information and transmits the acoustic reaction message to the loudspeaker of the communication unit as an output device and the visual reaction message as an animated avatar on the screen of the communication unit as an output device for the user.
[0068] In the speech analysis unit, a text is first generated from the speech information, an audio signal, using a speech-to-text (STS - speech recognition) module, which is then more easily processed by a language model that is also part of the speech analysis unit.
[0069] The image or video analysis unit (preferably AI-supported) analyzes the image information and creates image content information from it, preferably in text form, which is then forwarded to the speech analysis unit, the central organ of the communicator for conducting communication with the user and answering their questions.
[0070] To ensure that the image information is processed with the appropriate speech information in the speech analysis unit, both the image information and the speech information contain time information provided by a system clock (an element of the system).
[0071] The language model of the speech analysis unit analyzes both the speech information and the image content information and, in conjunction with the communication module (in which the use cases of the system operator and also a knowledge database are stored), formulates a response information that is sent to the image and sound generation unit.
[0072] This unit prepares the response information, often in text form, for the various output media (in this case, acoustic and visual). The image and sound generation unit therefore also comprises several submodules.
[0073] The audio generation unit essentially comprises a text-to-speech (TTS) module. This module converts text into an audio output. The resulting audio response is played back as an audio file through the communication unit's speaker. The image generation unit receives textual instructions specifying which gestures and facial expressions the avatar should adopt at each stage of its response (based on the output audio response) when responding to the user's voice or image information (the user's request) and displaying its audio and visual response. The image and audio generation unit is supported by the rendering module in creating the video stream that displays the avatar. This video stream, as the visual response, is then displayed on the communication unit's screen.
[0074] The synchronicity of the acoustic and optical response message (for example, lip movement) is established by the rendering module and / or the image and sound generation unit.
[0075] In one embodiment, the image or video analysis unit is designed to analyze the user's gestures and recognize at least the following gestures:
[0076] Head shaking (rejection), head nodding (agreement), hand raising (attention), looking away (disinterest), whereby this gesture information is transmitted to the speech analysis unit.
[0077] In a preferred embodiment, the image or video analysis unit includes a gesture and facial expression analysis module designed to evaluate the user's gestures and / or facial expressions. In addition to the aforementioned gestures, this variant also allows for the evaluation of facial movements and expressions.
[0078] Gestures and facial expressions are an important part of nonverbal communication. Communicators are trained to use gestures and facial expressions and to consider and process these elements when preparing a response message.
[0079] Gestures are the entirety of actions that serve as movements in interpersonal communication. In particular, movements of the arms, hands, and head accompany or replace messages in a given spoken language. Gestures are signs of nonverbal communication.
[0080] Facial expressions are defined as the visible movements of the face. Furthermore, the system is designed to include an emotion monitoring unit that monitors the user's emotional response, determines an emotion parameter from this, and transmits the emotion parameter to the communicator.
[0081] Emotion or emotional state refers to a psychophysical state of being triggered by the conscious or unconscious perception of an event or situation. Various signals indicate whether a user is emotionally tense or relaxed. The proposed emotion monitoring unit monitors these signals and transmits this information, either directly or as an emotion parameter (which may be determined separately for different signals), to the communicator, and in particular to the speech analysis unit, and considers it when preparing the communicator's response. The emotion monitoring unit can, for example, also utilize the result or output of the gesture and facial expression analysis module, as described above. This information is communicated, in particular, when the conversation is handed off to the human agent. The emotion monitoring unit is preferably AI-driven.The emotion parameter is derived from various signals, for example, in speech, from word choice and volume. The emotion monitoring unit is also designed to evaluate the user's gestures and facial expressions. Therefore, the emotion monitoring unit accesses both the acoustic information from the microphone and the visual information from the communication unit's camera.
[0082] Furthermore, the proposal advantageously provides that the communication module or speech analysis unit, depending on the course of the conversation, the analyzed speech information, and / or the analyzed image information, sets a communication parameter and transmits this parameter to the control unit. The control unit then connects the camera-microphone unit of the communication unit to the screen-speaker unit of the communication station, and vice versa, based on the set communication parameter. The above feature describes the technical requirements for transferring the conversation between the user and the communicator to the contact person (a human being, for example, a bank advisor).The set communication parameter is the signal for the control unit to connect the communication unit, where the contact person is located, to the communication station and, if necessary (but not necessarily in an alternative version of the invention), to simultaneously disconnect the connection to the communicator. The communication parameter is set by various processes.
[0083] It is initially proposed that the conversation history be used for this purpose. The conversation history, which is monitored by the communication module and the customer conversation process stored therein, can trigger the setting of the communication parameter. For example, conversation scenarios may be stored in the customer conversation process that are no longer handled by the communicator but by the contact person. If such a conversation scenario is detected, the communication parameter is set. The same applies if the conversation becomes too complex and the communication module no longer has any communication content available for the current conversation. In this case, too, the conversation should switch to the contact person, and the communication parameter should be set.For example, if the user answers "No" to the communicator's question of whether the user wishes to continue speaking with the communicator, the communication parameter is set, with the result that the conversation is transferred from the communicator to the human contact person as quickly as possible.
[0084] If the system detects that the user has unprompted expressed a desire not to speak to the communicator, the communication parameter is set by the communicator (or the communication module) based on the analyzed speech information.
[0085] If the user shakes their head in response to the request for further communication with the communicator, the communication parameter is set based on the analyzed image information. Additionally, it is possible for the user to set the communication parameter using a button (for example, on the touchscreen) or a separate switch, thus enabling them to establish direct communication with the contact person as quickly as possible.For example, it is provided that the touch-sensitive screen of the communication unit serves as an input device and that a button can be displayed on the touch-sensitive screen, the touching of which sets a communication parameter and transmits this to the control unit, and then the control unit connects the camera-microphone unit of the communication unit with the screen-speaker unit of the communication station and the camera-microphone unit of the communication station with the screen-speaker unit of the communication unit.
[0086] The camera-microphone unit describes the unit of input devices at the communication unit or communication station. The screen-speaker unit describes the unit of output devices at the communication unit or communication station.
[0087] For documentation purposes or for training the system, the system has a storage module designed to record and store the conversation between the user and the communicator, as well as the conversation between the user and the contact person.
[0088] In a preferred implementation of the proposal, the system is designed to delete the recorded conversation if the presence parameter is cleared before the conversation ends. This deletion serves as a data protection measure. The presence parameter continuously monitors the user's presence in front of the communication unit. This parameter is set when the user is in front of the communication unit. It is cleared when the presence parameter no longer detects the user, for example, because the user has left the room where the communication unit is located. The end of the conversation can be determined either by the communicator or the recipient. Specifically, the conversation ends, for example, when the user explicitly confirms this or implicitly. If no such indication of conversation termination is given, the recorded conversation is worthless and can be discarded.Furthermore, the control unit is designed so that it only connects the camera-microphone unit of the communication unit with the screen-speaker unit of the communication station, and vice versa, once the availability parameter has been set. This ensures that a contact person from the pool of contacts is actually available for the user.
[0089] Furthermore, the proposal advantageously provides that the computer of the communication unit and / or the communication station performs computationally intensive operations such as the rendering module, the speech analysis unit, the image-video analysis unit, the communication module, the emotion monitoring unit and / or the image-sound generation unit.
[0090] The aforementioned modules are often AI-driven and therefore require considerable computing power. Therefore, a preliminary system architecture proposal envisions that the server connected to the computers of the communication unit and the communication workstation be appropriately powerful, and that the computationally intensive AI-driven program modules be installed and run on this server. The hardware used in the computers of the communication unit and the communication workstation can then be less powerful and thus more cost-effective. However, this will increase the operating costs for the server and the associated data center.
[0091] In another proposed system architecture, at least the computer of the communication unit is designed to be correspondingly powerful, and one or more of the aforementioned AI-driven modules are installed and operated on it.
[0092] This can reduce the ongoing operating costs of the data center / server.
[0093] In a third proposed system architecture, the two aforementioned variants are combined, and depending on the available computing power of the computer at the communication unit, the (at least one, several or all) AI-driven modules are operated on the computer or the data center / server.
[0094] For example, if a user employs a smartphone or tablet with limited processing power as a communication unit, then the AI-driven applications will run on the data center / server. If the communication unit is located in a kiosk and can be used by anyone, it makes sense to equip this hardware with sufficient power so that some or ideally all of the aforementioned AI applications can be installed and run on it.
[0095] In a preferred embodiment of the proposal, a local processing unit is installed on or in the communication unit, which performs the following processing steps locally:
[0096] • Extraction of facial features and gesture descriptors from the user's camera video, whereby only these abstracted features, but not the raw video, are transmitted to remote servers,
[0097] • Extraction of speech features (e.g., MFCC, Mel spectrograms) from the microphone signal, whereby only these features, but not the raw audio signal, are transmitted to remote servers,
[0098] • Local identity verification using facial recognition or data card reading, • Rendering of the avatar video stream locally on the computer of the communication unit.
[0099] Preferably, at least the modules for identity verification, feature extraction from video / audio, and avatar rendering should be executed locally. This approach also improves data privacy.
[0100] Furthermore, the communication unit is to be equipped with a screen the size of a person or a video wall, primarily for displaying a life-size avatar. Preferably, the screen is vertically oriented. This design allows the avatar, or even the person being addressed, to be shown in life-size. The use of a video wall opens up the possibility, for example, of the life-size avatar moving from an entrance area to a conversation table, thus creating an even more realistic scenario. This measure also contributes to improving the atmosphere of the conversation.
[0101] In a further preferred embodiment, the microphone of the communication unit serves as an input device for voice information from the user, and the voice information is transmitted to the speech analysis unit of the communicator; the camera of the communication unit serves as an input device for image or video information from the user, and the image / video analysis unit of the communicator receives and analyzes the image or video information and generates image content information from it; the speech analysis unit converts the voice information and evaluates the image content information and, in conjunction with the communication module, generates a response information on the combined image content-speech information (from the speech-image information); and the communication module,Optionally, in conjunction with the knowledge database, it retrieves pre-stored knowledge and takes it into account to generate the reaction information; a sound generation unit generates an acoustic reaction message from the reaction information and transmits it to the loudspeaker of the communication unit as an output device; and an image generation unit, in conjunction with a rendering module and in conjunction with the sound generation unit and taking the reaction information into account, generates a visually perceptible reaction message as an animated avatar and displays this on the screen of the communication unit as an output device for the user, wherein the acoustic and the visual reaction messages together form an acoustically and visually perceptible reaction message.
[0102] Both speech and image / video information are processed as text in the speech analysis unit. The speech analysis unit has a speech-to-text module that converts the speech information into text. On the output side, the image / video analysis unit also provides text information based on the image information. Both sets of text information are processed simultaneously, preferably in real time, to ensure synchronization in this information processing. Real-time processing enables realistic communication. Potential asynchronicities are avoided, in particular, by using time information from a system clock, which is appropriately integrated into the speech and image information.
[0103] In a preferred variant, the audio-visual generation unit consists of two modules or parts, namely the audio generation unit and the image generation unit, wherein the audio generation unit generates an audio signal and the image generation unit generates a video.
[0104] The proposed procedure is characterized in particular by the fact that
[0105] furthermore, this is achieved by establishing the user's identity as identity information during the course of the conversation, preferably at the beginning of the conversation, between the user and the communicator. This is done in particular by either analyzing the user's relevant speech information or by using an image or video analysis unit to analyze image information of the user's face to create a facial image, and then using a search module to compare this facial image with images in an image database to determine the identity information.
[0106] A data card reading unit of the communication unit reads the user data card to determine the identity information, then, with the help of the identity information, feeds the user-specific user conversation data record belonging to the identity into the communication module of the communicator from a user database.
[0107] This pre-conditions the conversation between the communicator and the user, raising it to a higher quality level from the outset, as it allows the communicator to connect to and build upon previous conversations. This benefits the contact person, as the conversation is conducted at a higher level, for example, towards a conclusion. For the user, the perceived time spent on the platform is reduced due to the higher quality of the conversation.
[0108] The proposal also stipulates that identity verification is performed redundantly, meaning that several different identification methods, as described above, are used. This increases the likelihood that the correct person is actually in front of the communication unit. Furthermore, the proposal advantageously includes a rendering module in the communicator, which synchronizes the image output on the communication unit's screen with the audio output on the communication unit's speaker in real time.
[0109] This design enables a very natural expression from the avatar and enhances the communication experience for the user. Not only is the avatar's head animated, but, if the avatar is depicted at life size, its entire movement, gestures, and facial expressions are animated in real time. The rendering also includes emotional adjustments for the avatar.
[0110] In an advantageous embodiment, the system includes a communication station availability monitoring unit, which monitors the availability of the communication station and is connected to a contact person database in such a way that when an availability parameter is set, the name of the contact person is also transmitted to the communicator, and before the call is transferred to the communication station with the contact person, the communication module issues a farewell and transfer information message, and the image and sound generation unit creates an acoustically perceptible farewell and transfer message from the farewell and transfer information and transmits this to the screen speaker unit of the communication unit as the output device for the user.This ensures a professional handover of the conversation from the digital communicator to the personal contact person.
[0111] Cleverly, the system allows the communicator to record and save the conversation content with the user in a user-specific conversation record before transferring the call to the communication station with the contact person. The communication station then opens this conversation record, or parts thereof (such as a summary), upon taking over the call and displays or plays it back on the screen and / or through the speaker. This feature significantly enhances the professionalism of the communication between the user and the system. The communicator is no longer limited to engaging in trivial small talk but, guided by the customer conversation model, can gather information relevant to the user's request and then pass it on to the personal contact person.
[0112] Advantageously, this approach provides for the communicator, or the speech analysis unit / language model used, to create a summary of the conversation and store this summary in the user conversation record. The communicator prepares the conversation according to the criteria of the pre-qualification process (based on the system operator's business model – specific pre-qualification algorithms) and stores this summary as a user conversation record. This user conversation record is stored in a user database, preferably located on the system operator's server. The proposed method is therefore characterized in particular by the fact that the previous conversation with the communicator is summarized, and recognized user intentions and emotional states are also stored in the user conversation record.
[0113] Prior user identification is not mandatory for implementing this feature. New customers who are not yet registered in the system and therefore cannot yet be identified by the system can be treated in the same way. If a contract is concluded, the necessary identity information is added to the provisionally created user-specific call record. The content of the call record is intended to be displayed either on the screen of the communication workstation, which the contact person then reads, or it can be read aloud by the communication workstation (automatically via text-to-speech) so that the contact person hears the information.
[0114] The proposal stipulates that either the communication station automatically opens the user call record upon transfer of the call, or the contact person at the communication station actively opens the record. In both cases, the user call record is opened via or by the computer at the communication station. As a result, this feature also ensures higher productivity for the contact person on the one hand, and on the other hand increases the user's trust and sense of well-being regarding competent, targeted advice and information.
[0115] In another preferred embodiment, the control unit connects the communication unit to the communication station only when an availability parameter is set. If no contact person is yet available, the availability parameter is not yet set. If the communication parameter is already set, meaning the user wants the call transferred from the communicator to the contact person, the communicator informs the user that a contact person is not yet available. A corresponding message is played through the loudspeaker of the communication unit. This also enhances the professionalism of the conversation.
[0116] Furthermore, it is advantageously provided that a return signal can be sent from the communication station to the control unit, and, after a return signal has been sent from the communication station, the control unit reconnects the communication unit to the communicator, thus transferring the conversation between the user and the contact person back to the communicator. Such a design further reduces the workload for the contact person.
[0117] Often, the consultation between the user and the contact person concludes with the signing of a contract and the associated formalities. These processes can also be mapped through the customer consultation process and therefore automated using the communicator. It is therefore also intended that the communicator, in this case, will also monitor, control, and / or read the signature panel.
[0118] The set return signal simultaneously clears the communication parameter. If the user then requests a personal consultation with the contact person again, the communication parameter is reset with the (optional) additional condition that the contact person from the first part of the conversation should continue the consultation. The communication location or the name of this contact person was already recorded in the previous part of the conversation. In a preferred implementation of the proposal, it is provided that after the communication unit is reconnected to the communicator, the communication parameter is reset, and if the communication parameter is set again during the conversation between the user and the communicator, the control unit reconnects the communication unit to the communication location.
[0119] This design enables flexible, dynamic communication between the virtual agent (the communicator) and the human agent (the contact person). A flexible ping-pong workflow is possible, in which the conversation can be handed off from the virtual video agent to the human contact person (forward handoff) when a communication parameter is set, for example, when a defined complexity threshold is exceeded or the user requests it. This concept also allows the conversation to be handed back from the human contact person to the virtual video agent (backward handoff) when routine tasks (e.g.,
[0120] Data query, form completion, and signature process are to be carried out. If, during this communication, the need for consultation with the contact person arises again, the system switches back to the contact person. This process can be repeated as often as necessary. With each transfer, a structured contextual data set is transmitted, for example, as part of the user conversation record, which includes at least: previous conversation history, recognized user intention, user's emotional state, open questions, tasks already completed by the communicator and / or contact person, and information already gathered. The procedure is preferably designed so that when a contact person is contacted repeatedly, the system always contacts the same contact person, provided they are still available and their availability parameter is still set.
[0121] In a preferred embodiment of the proposal, the communication module is capable of generating an image message as a response from analyzed speech information, which is then displayed as a visually perceptible response message on the communication unit. It is also possible for the communication module to generate a response message from analyzed image information, which is then displayed as an audibly perceptible response message on the loudspeaker of the communication unit.
[0122] In a further preferred embodiment of the proposed system, the system is provided for to have more than one communication unit and more than one communication station. In particular, the system in this embodiment includes a more complex algorithm in the control unit, which is capable of managing and organizing this multitude of communication units and communication stations, as well as the necessary digital communicators. Specifically, the control unit interacts with the control system of the video call center equipment.
[0123] The proposed system can be used in the following areas:
[0124] 1. Banks and financial service providers
[0125] o Consultation and contract conclusion at kiosk terminals or via home banking
[0126] o Pre-qualification of loan applications by communicator
[0127] Routine tasks (transfers, standing orders) are handled by a communications specialist, while complex advice is provided by staff / contact persons.
[0128] 2. Authorities and public institutions
[0129] o Data protection-compliant processing through local architecture
[0130] 3. Telecommunications and utilities
[0131] Customer support with a tiered model: Simple questions — communicator, escalation — contact person
[0132] 4. Healthcare
[0133] o Initial consultation / administrative tasks by the communicator o Handover to doctor / contact person for medical questions 5. E-commerce and retail
[0134] o Virtual sales
[0135] o rater in online shops by commuikateur
[0136] o Kiosk systems in stores for product advice by communicators
[0137] Furthermore, the invention comprises a computer program or a plurality of computer programs comprising program code to cause a computer to execute the steps or part of the steps of the method when the computer program is executed on the computer.
[0138] In this context, it is particularly emphasized that all features and properties described in relation to the device, i.e., the system, as well as all procedures, are analogously transferable to the formulation of the method according to the invention and can be used within the scope of the invention, and are considered to be jointly disclosed. The same applies in reverse, meaning that only structural, i.e., device-related, features mentioned in relation to the method can also be considered and claimed for the claimed system and are likewise part of the disclosure.
[0139] Further features, details and advantages of the invention will become apparent from the wording of the claims and from the following description of exemplary embodiments with reference to the drawings. The drawings show:
[0140] Fig. 1 shows the structure of the system according to the invention in a block diagram.
[0141] Fig. 2 shows in a block diagram the structure of the digital communicator according to a first embodiment of the invention.
[0142] Fig. 3 shows in a further block diagram the structure of the digital communicator according to a second embodiment of the invention.
[0143] Fig. 4 shows in a further block diagram the structure of the digital communicator according to a third embodiment of the invention.
[0144] In the figures, identical or corresponding elements are designated with the same reference numerals and are therefore not described again unless expedient. The disclosures contained in the entire description are transferable analogously to identical parts with the same reference numerals or the same component designations. Furthermore, individual features or combinations of features from the different embodiments shown and described can also represent independent, inventive, or inventive solutions. Figure 1 shows an embodiment of the system 1 according to the invention. The system 1 consists of a communication unit 2 for the user. The user, for example, a customer, can interact with the system 1 using the communication unit 2. The communication unit 2 consists of a computer or central processing unit, a screen, a loudspeaker, a camera, and a microphone.In the embodiment shown here, the communication unit 2 (actually the computer) is additionally connected to a data card reader 7 and a signature panel 8. The data connection to these elements is indicated by the double arrow.
[0145] The communication unit 2, more precisely the computer of the communication unit 2, is connected to the control unit 4 via data transmission.
[0146] Furthermore, the control unit 4 is connected via data link to the digital communicator 5 and the communication station 3 for the user's contact person. The communication station 3 is similarly structured to the communication unit 2. It preferably consists of a computer or central processing unit, a monitor, a speaker, a camera, and a microphone.
[0147] In the embodiment shown here, a storage module 6 is connected to the control unit 4 via data transmission. This is advantageous because all information to and from the user flows through the control unit 4. Alternatively, the storage module 6 can also be connected to the communicator 5 and / or the communication station 3.
[0148] A central component of System 1 is the digital communicator 5. A first embodiment of the digital communicator 5, or digital assistant, is shown in detail in Figure 2. The digital communicator 5 is designed to relieve the contact person at communication station 3 of the burden of communicating with the user, and to conduct brief, engaging communication with the user until a contact person at communication station 3 becomes available and the conversation can then be transferred to them. The capabilities of the communicator 5 are not limited to verbal communication; it is also able to display a person, an avatar, on the screen that communicates with the user using gestures.
[0149] Figure 2 schematically illustrates the structure of the digital communicator 5 based on the processing sequence of incoming speech and / or image information 50. In this case, the control unit 4 is configured in the first section of the inventive method such that the communication unit 2 is connected to the user via the communicator 5. The speech and / or image information 50 is recorded by the microphone or camera of the communication unit 2 and sent to the communicator 5 via the computer of the communication unit 2 and the control unit 4. This digital or analog audio or video information from the user is first processed in the communicator 5 so that it can be reacted to and responded to appropriately.
[0150] Speech information is transmitted via arrow 50a to the speech analysis unit 51.
[0151] The speech analysis unit 51 is capable of understanding the user's spoken language and independently responding to and answering the user's questions or providing information. It includes at least one AI-driven speech-to-text (STT) module that converts the spoken word into searchable text. The transcribed speech message is then fed to an AI-driven large language model (LLM) module, which generates information addressed to the user in response to their request or information.
[0152] Communication module 53 contains a customer conversation process. This process is designed to structure the conversation with the user and to gather relevant information from the user that is important for further consultation with the contact person. Therefore, communication module 53 uses the results of the speech analysis unit 51 (using the provided computer) to analyze the content of the spoken words and the customer conversation process to formulate a response message. This response message is then converted into an acoustically perceptible response message by the audio-visual generation unit 54. This acoustic response message is then transmitted via the computer of communication unit 2 to the loudspeaker of communication unit 2 and output.
[0153] Image information is transmitted via arrow 50b to the image or video analysis unit 52.
[0154] The image / video analysis unit 52 is capable of analyzing user-captured images, image sequences, or videos with regard to the user's gestures and, particularly in interaction with the communication module 53 and the customer conversation process stored therein, formulating (or designing) an image message as a response. This message is then transformed (translated) into a visually perceptible reaction message using the rendering module 55 and the audio-visual generation unit 54. This involves close interaction between the rendering module 55 and the communication module 53 (especially the customer conversation process) and, on the other hand, with the audio-visual generation unit 54. These interactions are indicated by the double arrows.
[0155] This visual response message is then transmitted via the computer of communication unit 2 to the screen of communication unit 2 and displayed or played back. It is clear that this visual response message is either just an image or a video sequence lasting several seconds, preferably displayed as an avatar on the screen of communication unit 2.
[0156] Figure 3 schematically shows an alternative configuration of the digital communicator 5. The functionality of this variant is also illustrated by the process of handling incoming speech and / or image information 50.
[0157] The acoustic or linguistic content of the speech-image information 50 is fed to the speech analysis unit 51. This is indicated by arrow 50a. The speech analysis unit 51 is the central module of the digital communicator 5.
[0158] In Speech Analysis Unit 51, a text is first generated from the speech information, an audio signal, using a Speech-to-Text (STS) module. This text is then more easily processed by a language model, which is also part of Speech Analysis Unit 51. The language model, for example, the Large Language Model (LLM), is capable of understanding and responding to complex texts and relationships.
[0159] The image information contained in the speech / image information 50 is forwarded to the image / video analysis unit 52. This is indicated by arrow 50b.
[0160] The image / video analysis unit 52 (preferably AI-supported) analyzes the image information and generates image content information, preferably in text form, which is then forwarded to the speech analysis unit 51, the communicator's central hub for communication with the user. This is illustrated by arrow 521. This ensures that the speech analysis unit 51 receives all available user-related information. A system clock (not shown here) synchronizes the information with the respective data to prevent asynchronicity.
[0161] The speech analysis unit 51 is closely linked to the communication module 53. This is represented by the double arrow 513. The communication module 53, in turn, accesses a knowledge database 57. This is represented by the double arrow 537. Both the communication module 53 and the knowledge database 57 contain specific information regarding the system operator's use case, which is used by the speech analysis unit 51 for communication and to answer user requests.
[0162] The language model of the language analysis unit 51 analyzes both the language information and the image content information and, in conjunction with the communication module 53 and the knowledge database 57, formulates a response information which is sent to the image and sound generation unit 54, see arrow 514.
[0163] The image and sound generation unit 54 prepares the response information, which is in text form, for the various output media. The image and sound generation unit 54 therefore has several sub-modules (sub-unit sound generation unit and sub-unit image generation unit). The sub-unit sound generation unit essentially comprises a TTS (text-to-speech) module. This generates an acoustic response message 56a, which is played as an audio file on the speaker of the communication unit 2 (not shown), see arrow 546.
[0164] The subunit image generation unit 54b receives instructions in text form regarding which gestures and facial expressions the avatar should adopt at which point in its response (depending on the output acoustic response message).
[0165] The image and sound generation unit 54 is supported by the rendering module 55 for the creation of the video stream that displays the avatar. This is represented by the arrow 545.
[0166] The video stream, as optical response message 56b, is displayed by the rendering module 55 (see arrow 556) on the screen (not shown) of the communication unit 2.
[0167] The synchronicity of the acoustic (56a) and optical (56b) response message (for example, the lip movement) is established by the rendering module 55 (represented by the arrow 554) and / or the image and sound generation unit 54 (as the controlling element).
[0168] Figure 4 schematically shows another configuration of the digital communicator 5. The functionality of this variant is also illustrated by the process of processing incoming speech and / or image information 50.
[0169] The acoustic or linguistic content of the speech-image information 50 is fed to the speech analysis unit 51. This is indicated by arrow 50a. The speech analysis unit 51 is the central module of the digital communicator 5.
[0170] In speech analysis unit 51, a text is generated from the speech information, an audio signal, using a speech-to-text module. This text is then processed by a language model, which is also part of speech analysis unit 51. The language model, for example, LLM (Large Language Model), is capable of understanding and responding to complex texts and relationships. The image or video information contained in speech / image information 50 is fed to image / video analysis unit 52. This is indicated by arrow 50b.
[0171] The image and video analysis unit 52 comprises a model that uses machine learning to perform object recognition and computer vision in images and videos. The result is the generation of image content information in text form. This also includes emotion recognition, person recognition, and other similar functions.
[0172] The image or video analysis unit 52 (preferably AI-supported) analyzes the image information and creates image content information from it, which is then also forwarded to the speech analysis unit 51, the central organ of the communicator 5 for conducting communication with the user.
[0173] This is illustrated by arrow 521. This ensures that the speech analysis unit 51 receives all available user-side information. A system clock (not shown here) adds a time stamp to the respective information to prevent asynchronicity.
[0174] Alternatively, short latency periods or latency blocks are provided within which the speech analysis unit 51 waits until it receives image content information along with speech information, or vice versa. The first incoming signal triggers the latency period, thus ensuring reliable communication. This approach is not limited to the embodiment shown here, but applies to all embodiments presented here.
[0175] The speech analysis unit 51 is closely linked to the communication module 53. This is represented by the double arrow 513. The communication module 53, in turn, accesses a knowledge database 57. This is represented by the double arrow 537. Both the communication module 53 and the knowledge database 57 contain specific information regarding the system operator's use case, which the speech analysis unit 51 uses when communicating with and responding to user requests. The language model of the speech analysis unit 51 accesses the pre-stored knowledge in the knowledge database 57. The language model of the speech analysis unit 51 analyzes both the speech information and the image content information and, in conjunction with the communication module 53 and the knowledge database 57, formulates a response.
[0176] The image and sound generation unit 54 is modularly constructed and consists of an image generation unit 54b and a sound generation unit 54a, i.e., two parts.
[0177] The response information includes an acoustically output component and a visually output component. The response information is therefore fed to the tone generation unit 54a, as indicated by arrow 401.
[0178] The sound generation unit 54a converts the acoustically output portion of the response information into an audio file; a text-to-speech model is provided for this purpose in the sound generation unit 54a. The generated acoustic response message 56a is output as an audio file to the image generation unit 54b, see arrow 403.
[0179] The reaction information is simultaneously sent to the image generation unit 54b, as indicated by arrow 402. The image generation unit 54b interacts extensively with the rendering module 55, as indicated by double arrow 405. The image generation unit 54b animates the avatar.
[0180] The rendering module 55 / image generation unit 54b features a game engine that uses the audio file from the sound generation unit 54a to animate the avatar's lip movements. The context of the animation (greeting, speech gesture, etc.) is determined from the reaction information to ensure the avatar is animated correctly.
[0181] The acoustic / optical response message 56 consists of the generated audio file, which is output along with the visually generated avatar video.
[0182] The customer conversation process includes a protocol that sets the communication parameter when, for example, the user answers relevant key questions or provides specific, unambiguous input. This set communication parameter signals the control unit 4 to transfer the conversation between the user and communicator 5 to the contact person at communication station 3.
[0183] Ideally, the call is only transferred to communication station 3 once the availability parameter is set, meaning a contact person is actually available at communication station 3. If the communication parameter is set, the communicator 5 initiates the transfer of the call to communication station 3 with the contact person. This transfer includes, for example, a farewell from the digital communicator 5 and an announcement of the contact person's name at communication station 3. It also includes the transfer of information gathered (and possibly summarized) during the conversation between the user and the communicator 5, particularly information relevant to the conversation. The control unit 4 then activates the selected communication station 3 with the contact person and, if necessary, deactivates the digital communicator 5.Now the contact person can competently build upon the outcome of the first part of the conversation and work considerably more productively. At the same time, the user experiences significantly more competent communication and advice. If the conversation between the user at communication unit 2 and the contact person at communication station 3 ends with a contract or other formality necessary in the customer process, the call is transferred back from control unit 4 to digital communicator 5 and disconnected from communication station 3. The remaining formal tasks are then completed by digital communicator 5.
[0184] The invention is not limited to one of the embodiments described above, but can be modified in many different ways.
[0185] All features and advantages arising from the claims, the description, and the drawing, including design details, spatial arrangements, and process steps, can be essential to the invention, both individually and in various combinations. Reference numerals:
[0186] 1 system
[0187] 2 Communication unit
[0188] 3 Communication space
[0189] 4 Control unit
[0190] 5 Communicator
[0191] 6 memory modules
[0192] 7 Data card reader unit
[0193] 8 Signature panel
[0194] 50 Language / Image Information
[0195] 50a, 50b Arrow
[0196] 51 Language Analysis Unit
[0197] 52 Image or video analysis units
[0198] 53 Communication module
[0199] 54 Image and sound generation unit
[0200] 54a Tone generation unit
[0201] 54b Image generation unit
[0202] 55 Rendering module
[0203] 56 Acoustic / visual response message 56a Acoustic response message
[0204] 56b Optical response message
[0205] 57 Knowledge Database
[0206] 401, 402 Arrow
[0207] 403 Arrow
[0208] 404 Arrow
[0209] 405 Double Arrow
[0210] 513 Double Arrow
[0211] 514 Arrow
[0212] 521 Arrow
[0213] 537 Double Arrow
[0214] 545 Arrow
[0215] 546 Arrow
[0216] 554 Arrow
[0217] 556 Arrow
Claims
Patent claims 1. System for audio-visual communication between a user and a contact person of the user, wherein the system (1) comprises a communication unit (2), a communication station (3), a control unit (4) and a digital communicator (5), wherein - the communication unit (2), the communication station (3) and the communicator (5) are connected to the control unit (4) via data technology and - the communication unit (2) includes a computer, a camera, a microphone, a screen and a speaker for the user, - the communication station (3) includes a computer, a camera, a microphone, a screen and a speaker for the contact person and - the communicator (5) has a speech analysis unit (51), a communication module (53) and an image and sound generation unit (54). and the communicator (5) produces image information in the form of images of an artificial human as a result of the user's interaction with the system and transmits this information to the screen of the communication unit, and the control unit (4) establishes a connection either between the communication unit (2) and the communicator (5) or between the communication unit (2) and the communication location (3) as a result of the user's interaction with the system (1).
2. System according to the preceding claim, characterized in that the communicator (5) comprises a rendering module (55).
3. System according to one of the preceding claims, characterized in that the microphone of the communication unit (2) serves as an input device for voice information from the user and transmits the voice information to the voice analysis unit (51) of the communicator (5), and the camera of the communication unit (2) serves as an input device for image information from the user and transmits the image information to the image or video analysis unit (52) of the communicator (5), and the image or video analysis unit (52) analyzes the image information and sends the resulting image content information to the voice analysis unit (51). and the language analysis unit (51) comprises a language model that analyzes the image content information, together with the simultaneously received language information, and, in conjunction with the communication module (53), outputs a reaction information on the combined image content-language information and the image and sound generation unit (54), in conjunction with the rendering module (55), forms an acoustically and visually perceptible response message from the response information and transmits the acoustic response message to the loudspeaker of the communication unit (2) as an output device and the optical response message as an animated avatar on the screen of the communication unit (2) as an output device for the user.
4. System according to one of the preceding claims, characterized in that the system has an emotion monitoring unit which monitors the emotional reaction of the user and determines an emotion parameter from this and transmits the emotion parameter to the communicator (5).
5. System according to one of the preceding claims, characterized in that the communication module (53) or the speech analysis unit (51) sets a communication parameter depending on the course of the conversation, the analyzed speech information and / or the analyzed image information and transmits this to the control unit (4) and then, based on the set communication parameter, the control unit (4) connects the camera-microphone unit of the communication unit (2) with the screen-speaker unit of the communication station (3) and the camera-microphone unit of the communication station (3) with the screen-speaker unit of the communication unit (2). 6.System according to one of the preceding claims, characterized in that the control unit (4) connects the camera-microphone unit of the communication unit (2) with the screen-speaker unit of the communication station (3) and the camera-microphone unit of the communication station (3) with the screen-speaker unit of the communication unit (2) only when an availability parameter is set.
7. System according to one of the preceding claims, characterized in that the computer of the communication unit (2) and / or the communication station (3) performs computationally intensive operations such as the rendering module (55), the speech analysis unit (51), the image-video analysis unit (52), the communication module (53), the emotion monitoring unit and / or the image-sound generation unit (54).
8. System according to one of the preceding claims, characterized in that a screen in the dimensions of a person or a screen wall is provided on the communication unit (2), which serves in particular to display a life-size avatar.
9. System according to one of the preceding claims, characterized in that the microphone of the communication unit (2) serves as an input device for speech information from the user and the speech information is transmitted to the speech analysis unit (51) of the communicator (5) and that the camera of the communication unit (2) serves as an input device for image or video information from the user and the image-video analysis unit (52) of the communicator (5) receives and analyzes the image or video information and generates image content information from it, and that the speech analysis unit (51) evaluates the speech information and the image content information and, in conjunction with the communication module (53), generates a response information on the combined image content-speech information, and that the communication module (53), optionally in conjunction with the knowledge database (57), retrieves pre-stored knowledge and takes it into account to generate the response information, and that a sound generation unit (54a) generates an acoustic response message from the response information and transmits it to the loudspeaker of the communication unit (2) as an output device, and that an image generation unit (54b) in conjunction with a rendering module (55) and in conjunction with the sound generation unit (54a) and taking into account the reaction information generates an optically perceptible reaction message (56b) as an animated avatar and outputs this on the screen of the communication unit (2) as an output device for the user, wherein the acoustic and the optical reaction message together form an acoustically and optically perceptible reaction message (56).
10. Method for image and sound communication between a user and a contact person of the user, wherein the method comprises a communication unit (2), a communication station (3), a control unit (4) and a digital communicator (5), wherein - the communication unit (2), the communication station (3) and the communicator (5) are connected to the control unit (4) via data technology and - the communication unit (2) comprises a computer, a camera, a microphone, a screen and a speaker for the user, - the communication station (3) comprises a computer, a camera, a microphone, a screen and a speaker for the contact person, - the communicator (5) comprises a speech analysis unit, a communication module (53) and an image and sound generation unit (54). and the communication unit (2) is first connected to the communicator (5) by the control unit (4), thus enabling a conversation between the user and the communicator (5), and the communicator (5), as a result of the user's interaction with the system, produces image information in the form of images of an artificial human and transmits this to the screen of the communication unit, and the image and sound generation unit (54) generates an acoustically perceptible message from the information addressed to the user, which is then transmitted to the loudspeaker of the communication unit and output there and Then, depending on a set communication parameter, the control unit (4) connects the communication unit (2) with the communication station (3) and thus the conversation between the user and the communicator (5) is transferred to the communication station (3) with the contact person.
11. Method according to claim 10, characterized in that during the course of the conversation, preferably at the beginning of the conversation, the identity of the user is determined as identity information between the user and the communicator (5), for which purpose in particular either - the user's language information is analyzed or - an image or video analysis unit (52) analyzes image information of the user's face to create a facial image and a search module compares the facial image with the images in an image database to determine the identity information or - a data card reading unit (7) of the communication unit (2) reads the user data card to determine the identity information, then, using the identity information, feeds the user-specific user conversation record belonging to the identity from a user database into the communication module (53) of the communicator (5).
12. A method according to any one of claims 10 to 11, characterized in that the communicator (5) comprises a rendering module (55) and the rendering module (55) synchronizes the image output on the screen of the communication unit with the sound output to the loudspeaker of the communication unit in real time.
13. A method according to any one of claims 10 to 12, characterized in that the communicator (5) stores and saves the conversation content with the user in a user-specific user conversation data record with the contact person before the conversation is transferred to the communication station (3), and the communication station (3) opens this user conversation data record, or parts thereof, for example, a summary, upon taking over the conversation and displays or outputs it at least on the screen and / or the loudspeaker.
14. Method according to one of claims 10 to 13, characterized in that a return signal can be sent from the communication station (3) to the control unit (4), and, after a return signal has been sent from the communication station (3), the control unit (4) reconnects the communication unit (2) to the communicator (5) and thus transfers the conversation between the user and the contact person to the communicator (5).
15. Method according to one of claims 10 to 14, characterized in that, after the communication unit (2) is reconnected to the communicator (5), the communication parameter is reset and, if the communication parameter is set again during the conversation between the user and the communicator, the control unit (4) reconnects the communication unit (2) to the communication location (3).