Training a conversation system using video / audio analysis
Patent Information
- Application Number
- EP2024702243
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-07
- Filing Date
- 2024-01-12
- Publication Date
- 2025-11-05
AI Technical Summary
Current chatbot training methods require laborious manual data entry and lack the ability to capture emotional nuances, limiting the efficiency and quality of conversation systems.
A computer-implemented method that analyzes video/audio recordings to automatically identify and label questions and answers, including emotional cues from body language and intonation, to provide additional training data for conversation systems.
This approach eliminates the need for manual data entry and enhances chatbot training by incorporating emotional context, leading to improved and accelerated qualitative and quantitative training of conversation systems.
Smart Images

Figure EP2024050730_15082024_PF_FP
Abstract
Description
[0001] Description
[0002] Training a conversation system through video / audio analysis
[0003] FIELD OF THE INVENTION
[0004] The invention relates to a computer-implemented method and an associated arrangement for training an automated conversation system.
[0005] Regardless of the grammatical gender of a particular term, persons with male, female or other gender identity are included.
[0006] BACKGROUND OF THE INVENTION
[0007] A chatbot is a computer-based system capable of conducting human-like conversations with users. Chatbots are commonly used in social media, websites, or instant messaging applications to provide users with information, answer questions, or offer services.
[0008] Chatbots can be programmed in a variety of ways to conduct human-like conversations. Some chatbots use rule-based systems, which use predefined rules and patterns to respond to user input. Other chatbots use machine learning to conduct human-like conversations and improve based on user input.
[0009] Chatbots can also be used in a variety of industries, such as customer service, e-commerce, and the entertainment industry. They can help improve the efficiency of business processes and provide users with a faster and easier way to respond to their inquiries and needs. There are different ways a chatbot can be trained, depending on how it was programmed.
[0010] A rule-based chatbot is typically programmed by a developer using predefined rules and patterns to respond to user input. These rules and patterns are built into the chatbot from the start and do not require further training.
[0011] A chatbot that uses machine learning to conduct human-like conversations is typically trained by providing it with large amounts of data, known as "training." This data includes typical user inputs and the chatbot's corresponding responses. The machine learning model is then used to identify patterns in this data and learn how to respond to user input.
[0012] There are different types of machine learning models that can be used in chatbots, such as neural networks and decision trees. Each model has its own strengths and weaknesses and may be better suited to certain use cases than others.
[0013] It is important to note that training a chatbot is a continuous process and that the chatbot may need further training to improve its performance and adapt to new user inputs and needs.
[0014] Currently, chatbot administrators offer two main functionalities. One is the input of at least 20 versions of the same question to trigger what is known as NLU (Natural Language Understanding) training from intents, e.g., "Can I have some ice cream?", "Do you have any ice cream?", "I really want ice cream," etc., to train the intent "The user wants ice cream." The other is hard-coded answers that the chatbot can provide when such an intent is recognized, e.g., "Here you can find the nearest ice cream parlor."
[0015] Typically, this content is handed over to the data scientist as unstructured prose. It's also common for the content to be collected from multiple sources in various formats before this one-stop handover.
[0016] A chatbot administrator is someone responsible for managing and maintaining a chatbot. This may include programming, training, and updating the chatbot, depending on how the chatbot is configured and what type of functions it performs.
[0017] A chatbot administrator might also be responsible for monitoring the chatbot's performance and analyzing user inputs and responses to ensure the chatbot provides accurate and helpful answers. They might also be responsible for integrating the chatbot with other systems or applications, ensuring it functions properly and meets the company's needs.
[0018] In some cases, a chatbot administrator might also be responsible for communicating with users and customers when the chatbot is unable to answer their queries or assist them. In this case, the chatbot administrator might interact directly with users to answer their queries or help them resolve their issues.
[0019] Overall, the chatbot administrator is responsible for managing and maintaining the chatbot, ensuring it functions properly and provides helpful information and services to users.
[0020] A data scientist is someone who analyzes and understands large amounts of data. Data scientists use mathematics, statistics, machine learning, and other tools and techniques to identify and understand patterns and relationships in data. They often work with large amounts of structured and unstructured data, using specialized tools and techniques to process and analyze the data.
[0021] The work of a data scientist often involves analyzing data to support business decisions or improve processes. They can also make predictions or forecasts by identifying patterns in the data and presenting the results in a visual format. Data scientists often work closely with other professionals, such as software developers and business analysts, using their knowledge of data analysis to solve complex problems and make decisions.
[0022] Data collection for chatbot systems is laborious and typically requires manual input by a programmer. Questions and corresponding answers must be entered into a chatbot CMS (e.g., Botpress), and at least 10 alternatives must be found for NLU training.
[0023] A user must be identified to collect the required topic content and derive questions and answers themselves. These are then entered manually. The variants required to start an NLU training are based solely on the creativity of this user and often lack a qualitative data basis.
[0024] SUMMARY OF THE INVENTION
[0025] It is an object of the invention to provide a solution with the aid of which the training of communication systems, also referred to as "chatbots", can be improved qualitatively and quantitatively and accelerated. The invention results from the features of the independent claims. Advantageous developments and refinements are the subject of the dependent claims. Further features, possible applications and advantages of the invention result from the following description.
[0026] One aspect of the invention concerns the analysis of a video and / or audio source by a computer-based analysis system, separating questions and answers and placing them in context. It then automatically expands the dialogue capabilities of a conversation system with the questions and answers thus obtained.
[0027] If, in a further aspect, several video and / or audio interviews with the same interview structure are analyzed, existing questions in the conversation system can be expanded to include new answer options. The conversation system thus automatically receives additional training data.
[0028] A key difference compared to the state of the art is that manual data entry is no longer required.
[0029] However, analyzing recorded structured video interviews offers several additional advantages: by analyzing facial and verbal expressions and gestures, the analysis system can recognize which emotion is associated with a question or answer. It can learn to ask questions with a specific emotion (e.g., sarcasm) or formulate answers that emphasize individual words.
[0030] The invention claims a computer-implemented method for training an automated conversation system, comprising the steps of:
[0031] - Video / audio recording of a question-and-answer session between a person asking a question and a person answering a question, - Analysis of the video / audio recording with regard to questions and corresponding answers with regard to their usability for the conversation system and determination of the usable questions and corresponding usable answers,
[0032] - Transferring the questions and answers thus determined to the conversation system, and
[0033] - Training the conversation system with the identified questions and answers.
[0034] In a further development of the procedure, the procedure can be carried out with additional questioning persons and additional answering persons.
[0035] In further training, the procedure includes the following further steps:
[0036] - Analysis of the video / audio recording in relation to the body language of the person asking and / or answering,
[0037] - Determination of the corresponding first emotions,
[0038] - Marking the questions and / or answers with the identified first emotions, and
[0039] - when training the conversation system, the first emotions identified are taken into account.
[0040] Body language includes gestures, facial expressions, posture and habitus.
[0041] In further training, the procedure includes the following further steps:
[0042] - Analysis of the video / audio recording in relation to the intonation of the person asking and / or answering,
[0043] - Determination of the corresponding second emotions,
[0044] - Marking the questions and / or answers with the identified second emotions, and
[0045] - when training the conversation system, the identified second emotions are taken into account.
[0046] Intonation includes, for example, accent, pitch progression, and pause structure. As an alternative to the previous paragraphs, the procedure may include the following additional steps.
[0047] - Analysis of the video / audio recording in relation to the body language and intonation of the person asking and / or answering,
[0048] - Determination of the corresponding third emotions,
[0049] - Marking the questions and / or answers with the identified third emotions, and
[0050] - when training the conversation system, the identified third emotions are taken into account.
[0051] The invention also claims an arrangement for training an automated conversation system, comprising:
[0052] - a recording unit comprising a microphone unit and an image recording unit, wherein the recording unit is designed and configured to record a question-and-answer session between a questioning person and a responding person,
[0053] - an analysis unit that is trained and configured to analyze the video / audio recording with regard to questions and associated answers with regard to their usability for the conversation system and to determine the usable questions and associated usable answers, and to transmit the questions and answers thus determined to the conversation system, and
[0054] - the conversation system, which is trained and equipped to train itself with the identified questions and answers.
[0055] In a further embodiment of the arrangement, the analysis unit can be designed and configured
[0056] - to analyze the video / audio recording in relation to the body language of the person asking and / or answering,
[0057] - to identify the associated first emotions,
[0058] - to mark the questions and / or answers with the identified first emotions, and that the conversation system is trained and set up,
[0059] - to take into account the first emotions identified when training the conversation system.
[0060] In a further training, the analysis unit can be trained and set up,
[0061] - to analyze the video / audio recording in relation to the intonation of the person asking and / or answering,
[0062] - to identify the corresponding second emotions, and
[0063] - to mark the questions and / or answers with the identified second emotions, and the conversation system must be trained and set up,
[0064] - to take the identified second emotions into account when training the conversation system.
[0065] In a further training, the analysis unit can be trained and set up,
[0066] - to analyze the video / audio recording in relation to the body language and intonation of the person asking and / or answering,
[0067] - to identify the corresponding third emotions, and
[0068] - to mark the questions and / or answers with the identified third emotions, and the conversation system must be trained and set up,
[0069] - to take the identified third emotions into account when training the conversation system.
[0070] In further training, the analysis unit can be implemented in the conversation system.
[0071] This means that both components form a unit and can, for example, share computing resources and data storage.
[0072] Further features and advantages of the invention will become apparent from the following explanations of an exemplary embodiment based on schematic drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] It shows :
[0074] FIG. 1 is a flow diagram of the method for training a conversation system by video / audio analysis, and
[0075] FIG. 2 is a block diagram of the arrangement for training a conversation system by video / audio analysis.
[0076] DETAILED DESCRIPTION OF THE INVENTION
[0077] An exemplary embodiment of the invention is shown in FIGS. 1 and 2 and is described in detail below.
[0078] FIG. 1 shows the flowchart of the computer-implemented method for training a conversation system through video / audio analysis. In the first step 101, a video / audio recording of a question-and-answer session between a questioning person 3 and an answering person 4 is made. In the following step 102, the video / audio recording is analyzed with regard to questions and corresponding answers with regard to their usability for the conversation system, and the usable questions and corresponding usable answers are determined.
[0079] Additionally, in step 102, the video / audio recording can optionally be analyzed with respect to the body language of the person asking and / or answering, with the associated first emotions then being determined in step 104. Subsequently, in step 105, the questions and / or answers are identified, for example, labeled, with the determined first emotions. Optionally, second or third emotions can be determined with respect to intonation or with respect to body language and intonation.
[0080] In step 106, the questions and answers thus determined, as well as optionally also the labels for the first, second, or third emotions, are transmitted to the conversation system 1. In the final step 107, the conversation system is trained with the determined questions and answers as well as the optional labels according to known methods from the prior art.
[0081] In the final step 107, the conversation system 1 is trained with the questions and answers transmitted in step 105 as well as the first and / or second or third emotions characterizing the questions and / or answers.
[0082] The described procedure can be repeated as often as desired with additional questioning persons 3 and / or answering persons 4.
[0083] FIG. 2 shows a block diagram of an exemplary arrangement of the invention for training a conversation system 1. In a question and answer session, the questioning person 3 asks the answering person 4 questions. The questions and answers are recorded via the microphone unit 2. 2. The question and answer session is also filmed with the image recording unit 2. 1.
[0084] All recorded data are analyzed by the analysis unit 5 and processed according to the method shown in FIG. 1. For this purpose, the analysis unit 5 has a computing unit or uses resources from a cloud.
[0085] Using the determined questions and answers, and optionally using the emotion labels, the conversation system 1 is trained according to known methods. Although the invention has been illustrated and described in more detail by the exemplary embodiments, the invention is not limited by the disclosed examples, and other variations can be derived therefrom by those skilled in the art without departing from the scope of the invention.
[0086] Reference symbol list
[0087] 1 conversation system
[0088] 2 Recording unit
[0089] 2 . 1 image acquisition unit
[0090] 2 . 2 Microphone unit
[0091] 3 Questioning person
[0092] 4 Respondent
[0093] 5 Analysis unit
[0094] 101 Video / Audio on Drawing
[0095] 102 Analysis of the video / audio recording
[0096] 103 Determination of questions and answers
[0097] 104 Determination of emotions
[0098] 105 Marking of questions and / or answers
[0099] 106 Transferring the questions and answers
[0100] 107 Training the conversation system
Claims
Patent claims 1. A computer-implemented method for training an automated conversation system (1) , characterized by: - video / audio recording (101) of a question-and-answer session between a questioning person (3) and a responding person (4), - analysis (102) of the video / audio recording with regard to questions and corresponding answers with regard to their usability for the conversation system (1) and determination (103) of the usable questions and corresponding usable answers, - transmitting (106) the questions and answers thus determined to the conversation system, and - Training (107) the conversation system with the determined questions and answers.
2. The computer-implemented method according to claim 1, characterized in that the method is carried out with additional questioning persons and additional answering persons.
3. The computer-implemented method according to claim 1 or 2, characterized by - Analysis (102) of the video / audio recording in relation to the body language of the person asking and / or answering 3, 4. - Determination (104) of the corresponding first emotions, - marking (105) the questions and / or answers with the identified first emotions, and - when training (107) the conversation system (1) taking into account the first emotions identified.
4. The computer-implemented method according to one of the preceding claims, characterized by - Analysis (102) of the video / audio recording in relation to the intonation of the person asking and / or answering 3, 4, - Determination (104) of the corresponding second emotions, - marking (105) the questions and / or answers with the identified second emotions, and - when training (107) the conversation system (1) taking into account the determined second emotions.
5. The computer-implemented method according to claim 1 or 2, characterized by - Analysis (102) of the video / audio recording in relation to the body language and intonation of the person asking and / or answering 3, 4, - Determination (104) of the corresponding third emotions, - marking (105) the questions and / or answers with the identified third emotions, and - when training (107) the conversation system (1) taking into account the identified third emotions.
6. An arrangement for training an automated conversation system (1), characterized by: - a recording unit (2) comprising a microphone unit (2.2) and an image recording unit (2.1), wherein the recording unit (2) is designed and configured to record a question-and-answer session between a questioning person (3) and a responding person (4), - an analysis unit (5) which is designed and configured to analyze the video / audio recording with regard to questions and associated answers with regard to their usability for the conversation system and to determine the usable questions and associated usable answers, and to transmit the questions and answers thus determined to the conversation system, and - the conversation system (1) which is trained and equipped to train itself with the identified questions and answers. 7 . The arrangement according to claim 6 , characterized in that the analysis unit ( 5 ) is designed and arranged , - to analyze the video / audio recording in relation to the body language of the person asking and / or answering, - to identify the associated first emotions, - to mark the questions and / or answers with the identified first emotions, and that the conversation system is trained and set up, - to take into account the first emotions identified when training the conversation system. 8 . The arrangement according to claim 5 or 6 , characterized in that the analysis unit ( 5 ) is designed and arranged , - to analyze the video / audio recording in relation to the intonation of the person asking and / or answering, - to identify the corresponding second emotions, and - to mark the questions and / or answers with the identified second emotions, and that the conversation system is trained and set up, - to take the identified second emotions into account when training the conversation system.
9. The arrangement according to claim 5 or 6, characterized in that the analysis unit (5) is designed and arranged, - to analyze the video / audio recording in relation to the body language and intonation of the person asking and / or answering, - to identify the corresponding third emotions, and - to mark the questions and / or answers with the identified third emotions, and that the conversation system is trained and set up, - to take the identified third emotions into account when training the conversation system.
10. The arrangement according to one of claims 6 to 8, characterized in that the analysis unit (5) is formed in the conversation system (1).