Computer-implemented dialogue method, communication system, computer product and computer readable storage medium
By neutralizing user input and adding personalized elements to the output, the problem of chatbots' lack of personalized experience is solved, emotional and empathetic conversation effects are achieved, and the user experience is improved.
Patent Information
- Application Number
- CN202480014418.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-02-23
- Filing Date
- 2024-01-12
- Publication Date
- 2025-10-03
AI Technical Summary
Existing chatbots lack a personalized and individualized user experience, are unable to convey style and express empathy, and as a result, the more emotional the user input, the worse the output.
An interactive communication system implemented by a computer uses artificial intelligence to neutralize and anonymize user input, remove style elements, and add personalization and style elements as needed during output, using training data such as audio, text interviews, dramas, and movies to generate emotional and empathetic responses.
This enables the chatbot's output to be personalized and stylized according to user needs, improving the user experience and making conversations more open, in-depth, and comfortable.
Smart Images

Figure CN120752640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a computer-implemented dialog method with improved user experience, a communication system for this purpose, a computer product for implementing an automated dialog system and a computer-readable storage medium. Background Art
[0002] A conversational system (also referred to as a "chatbot") is a computer-assisted system capable of engaging in human-like conversations with users. Chatbots are often used on social media, websites, or instant messaging applications to provide users with information, answer questions, or offer services.
[0003] Chatbots can be programmed in various ways to engage in human-like conversations. Some chatbots use rule-based systems, where predefined rules and patterns are used to respond to user input. Other chatbots use machine learning to engage in human-like conversations and improve and / or adapt themselves based on user input.
[0004] Chatbots can also be applied in a wide range of fields, such as customer service, e-commerce, and the entertainment industry. They help improve the efficiency of business processes and provide users with faster and simpler solutions to respond to their inquiries and needs.
[0005] Chatbots can be trained in a variety of ways, depending on how they are programmed.
[0006] Chatbots that operate based on a ruleset are typically programmed by developers in such a way that they respond to user input using predefined rules and patterns. These rules and patterns are installed in the chatbot from the outset and require no further training.
[0007] Chatbots that use machine learning to conduct human-like conversations are typically trained by providing them with large amounts of data, known as "training data." This data typically includes user input and the chatbot's corresponding responses. The machine learning model is then used to identify patterns in this data and learn how to respond to user input.
[0008] There are various types of machine learning models that can be used in chatbots, such as neural networks and decision trees. Each model has its own strengths and weaknesses and may be better suited to specific use cases than others.
[0009] It is important to note that training a chatbot is an ongoing process, and further training of the chatbot may be required to improve its performance and adapt to new user input and needs.
[0010] The GPT3 chatbot has been around for some time and has been piloted by approximately 100 million users. The system, trained on massive amounts of text and with the help of thousands of human "trainers," is capable of delivering seemingly well-thought-out responses. However, numerous flaws remain, ranging from mispronunciations to misleading numbers and even explicit insults delivered by the chatbot. The more empathetic the user's input, the worse the chatbot's output. For example, a user's input of an emotional message to end a relationship will not produce the desired result, as the chatbot was not trained to express empathy.
[0011] Artificial intelligence (KI), on the other hand, is perfectly capable of separating the objective content of input from personal style; for example, some processors have been configured to compose classical music, even directly imitating Mozart. If properly trained, AI can also paint in the style of Edward Hopper and write in the style of Shakespeare. Chatbots, on the other hand, can be trained to respond in the typical voice of Donald Trump or Jedi Master Yoda.
[0012] Currently, chatbot administrators offer two main features. One is to input at least 20 versions of the same question to trigger so-called NLU (Natural Language Understanding) intent training, such as "Can I have ice cream?", "Do you have ice cream?", "I really want ice cream," etc., to train the "user wants ice cream" intent. The other is to provide a hard-coded response to the chatbot when such an intent is recognized, such as "The nearest ice cream shop is here."
[0013] Often, this content is handed off to data scientists in the form of unstructured prose. It’s also common to gather said content from different sources in a variety of formats before handing it off in one place.
[0014] A chatbot administrator is someone who manages and maintains a chatbot. This may include programming, training, and updating the chatbot, depending on how the chatbot is configured and the types of functions it performs.
[0015] Chatbot administrators are also responsible for monitoring the performance of the chatbot and analyzing user input and responses to ensure that the chatbot provides accurate and helpful responses. They are also responsible for integrating the chatbot into other systems or applications and ensuring that it operates as expected and meets the requirements of the business.
[0016] In some cases, chatbot administrators may also be responsible for communicating with users and customers when the chatbot is unable to answer their queries or help them. In such cases, the chatbot administrator may interact directly with users to answer their queries or help them resolve their issues.
[0017] Overall, a chatbot administrator is responsible for managing and maintaining the chatbot and is responsible for ensuring it functions as expected and provides helpful information and services to users.
[0018] A data scientist is someone responsible for analyzing and understanding vast amounts of data. Data scientists apply mathematics, statistics, machine learning, and other tools and techniques to identify and understand patterns and connections within data. They typically work with large amounts of structured and unstructured data, using specialized tools and techniques to process and analyze it.
[0019] Data scientists' work typically involves analyzing data to support business decisions or improve processes. They may also create predictions or forecasts by identifying patterns in data and presenting the results in an intuitive form. Data scientists often collaborate with other specialists, such as software developers and business analysts, and apply their knowledge of data analysis to solve complex problems and make decisions.
[0020] Until now, chatbots have operated in a neutral manner, regardless of whether text and / or voice input is present; that is, their responses could not be associated with a specific type or style of human or emotion.
[0021] However, this is disadvantageous because it means a lack of personalization and individualization for the user, the so-called "user experience." Summary of the Invention
[0022] The purpose of the present invention is to provide a dialogue system that, in addition to the functions of conveying user intentions and correct answers, is also equipped with the possibility of conveying style, expressing empathy and imparting a personality format that is independent of both the chatbot administrator and the data scientist.
[0023] This object is achieved by the subject matter of the invention as defined in the description, the drawings and the claims.
[0024] The subject matter of the present invention is therefore a computer-implemented dialog method comprising the following method steps:
[0025] -1) providing a computer-implemented interactive communication system having: an input device, a processor for text and / or speech recognition, a neural network including artificial intelligence, said artificial intelligence being used for text and / or speech evaluation, and generating first results from said text and / or speech recognition and storing them,
[0026] -2) user input into a computer-implemented interactive communication system,
[0027] -3) processing the user's input by artificial intelligence for text and / or speech evaluation and generating a first result from the text and / or speech recognition,
[0028] -4) performing a first processing on the first result by means of an artificial intelligence provided and configured for this purpose, in the form of neutralizing and / or removing stylistic elements in the voice and / or text input obtained from the user, in order to formalize the input and / or anonymize the user who made the input,
[0029] -5) generating a second result by means of said artificial intelligence, which is based on the content of the user input in an anonymized and neutralized form,
[0030] -6) processing the second results by the artificial intelligence to generate appropriate outputs, for example in the form of answers, and outputting them via a suitable output device.
[0031] The subject of the invention is also an exemplary communication system comprising the following components:
[0032] - Input devices such as keyboard, microphone, camera,
[0033] a first processor comprising a memory unit, the first processor being configured and arranged to perform text and / or speech recognition, the first processor having a neural network and artificial intelligence and being able to run them and thereby generate a first result that can be stored in the connected first memory unit,
[0034] - a second processor, the second processor being configured to, first, retrieve the first result from the first storage unit, process the first result, and provide the first result to a configured and trained artificial intelligence so as to neutralize the result, remove rhetoric, and / or remove personalized style elements, and / or insert personalized style elements, wherein the result input to the artificial intelligence is either the first neutralized input from the user or the third neutral answer from the chatbot,
[0035] - wherein the artificial intelligence is set and configured to assign predeterminable or artificial intelligence-selectable stylistic and / or rhetorical and / or personality attributes to the obtained neutral text and / or speech content, and / or convert it accordingly and forward it to the third processor,
[0036] - wherein the third processor is configured and arranged to perform speech and / or text generation by artificial intelligence and provide it in acoustic, optical or both form as output in response to input obtained from the user via an output device of the communication system.
[0037] According to an advantageous embodiment of the invention, between the penultimate step of searching for a suitable output and the final step of outputting a suitable answer, reply and / or reaction via the communication system, a process step 7) also occurs, by which the output generated by the artificial intelligence is supplemented with personalization and / or other stylistic elements.
[0038] For example, process step 7 can be performed by processing the output by the artificial intelligence so that the output performed by the artificial intelligence is rhetorically stylized according to known rules and / or is typicalized by wording, sentence length, sentence structure and / or is performed in a typical personality profile stored in the artificial intelligence, for example in the style of Donald Trump.
[0039] The present invention is based on the general understanding that artificial intelligence can separate universal and individual stylistic elements from the core of user input—that is, it can remove these stylistic elements from the input and "unmask" the content in a neutralized form. The present invention exploits this fact and implements the method during user input by having the artificial intelligence neutralize and anonymize the input and / or add any stylistic elements, typical features, and / or personalization and / or individualization elements to the output, such as a subsequent response. This makes the user feel comfortable and allows the conversation between the machine and the user to proceed in a significantly more open, in-depth, and / or thorough manner. Thus, the chatbot is provided with not only the intonation of, for example, Jedi Master Yoda (because it is very unique), but also his thoughts, phrasing, typical sentence structure, style, and personal approach to specific topics, questions, content, etc.
[0040] The training possibilities for stylizing the output of the processing performed by the artificial intelligence in process step 7) are practically unlimited. Any interview with a user can serve as a basis for generating training data. However, interviews, plays, and / or films stored in audio and / or text form in archives can also be evaluated as a data pool.
[0041] Data used to train AI for style and / or personalization can be extracted from any context involving empathetic conversations, writing, and / or responses. This data enables AI trained accordingly and connected to a chatbot to impart emotional, empathetic, personalized, and / or any other personalized effects to the communication system's output, thereby resulting in an improved user experience for the corresponding computer-implemented communication system.
[0042] For example, some text-to-speech machines use Speech Synthesis Markup Language (SSMML). This technology makes it possible to add emotional effects to the audio output of chatbots, but it does not mean that the resulting machine responses can display any empathy or emotional style.
[0043] The present invention eliminates the previously customary process step of having data scientists individually use algorithms to apply personalization to the chatbot's responses in voice and / or text form, since the artificial intelligence is trained in such a way that it automatically selects a personality and / or style for the user and applies this individualization to its results, which are output, for example, in the form of the chatbot's voice responses.
[0044] Thus, chatbots in computer-implemented communication systems can also be trained to develop empathy, as the emotional style of their responses conveys and / or suggests greater empathy and understanding to the user. For example, in another process step optionally implemented in the method, the user can even enter a desired style profile in their input regarding the desired output style. This desired profile can be generated consciously or unconsciously, as the user's preferred style and personality type can be communicated to the AI, for example, through the wording of the input or simply through questions directed to this aspect.
[0045] However, empathy in the response can be generated by expressing the response in the style of Shakespeare, who himself has empathetic language due to the abundance of flowery descriptions.
[0046] For example, linguistic and / or rhetorical stylistic devices that have been scientifically studied enough can also be used to enhance the expressiveness of written or spoken texts used to train artificial intelligence, and their impact on readers / listeners can be precisely correlated. Examples of these are: metaphor, alliteration, hyperbole, oxymoron, irony, archaisms, word-shifting, euphemisms, tautology, progression, paraphrase, verbosity, etc.
[0047] There are various rhetorical stylistic devices that can be used to train AI and associate them with different personality profiles, thus influencing the chatbot's phrasing, which can result in responses in a specific style.
[0048] In addition to wording, the AI can also be trained to use a specific voice, intonation, and / or speech rhythm, thereby also giving the chatbot's output a specific style.
[0049] Thus, virtually any output content and any type of output, whether acoustic, optical and / or visual, can be automatically given a literary style, a rhetorical style and / or a typical personality style by a suitably trained artificial intelligence.
[0050] According to an advantageous embodiment of the invention, the user's reactions are recorded and input into the system, so that the artificial intelligence can learn from the reactions which type of chatbot is suitable for which type of user.
[0051] According to an advantageous embodiment of the invention, the speech recognition in process step A) is performed by artificial intelligence with NLP capabilities.
[0052] Natural language processing (NLP) is a subfield of artificial intelligence that studies how computers interact with humans through natural language. It involves developing algorithms and models that can analyze, understand, and generate human language, including text and speech.
[0053] NLP capabilities refer to the ability of a computer system or software to perform tasks in the field of natural language processing, including, for example, language translation, text classification, sentiment analysis, and speech generation.
[0054] In general, the goal of NLP is to enable computers to process, analyze, and understand human language in a manner and method similar to that of humans, so that they can interact with humans more effectively and perform tasks that are difficult or impossible for humans to complete. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The invention is explained in more detail below with reference to a flow chart which reproduces the flow of an exemplary embodiment of the disclosed method. DETAILED DESCRIPTION
[0056] - At the top, process step 1) should be seen: providing a computer-implemented interactive communication system having: input devices, such as a microphone, a keyboard, a camera; a processor for text and / or speech recognition; a neural network containing artificial intelligence for text and / or speech evaluation and for generating first results from the text and / or speech recognition and storing them,
[0057] - and the following process step 5): user inputs to the computer-implemented interactive communication system via an input device;
[0058] - Process step 3): processing the user input by a processor suitable for this purpose and configured accordingly, wherein the processor comprises a neural network, an artificial intelligence running on the neural network for text and / or speech evaluation, and generating a first result from the text and / or speech evaluation and / or text and / or speech recognition;
[0059] - then proceeding to process step 4): performing a first processing of the first result by means of an artificial intelligence provided and configured for this purpose in the form of neutralizing and / or removing stylistic elements from the voice and / or text input obtained from the user, in order to formalize the input and / or anonymize the user who made the input; and the following process steps
[0060] -5) generating a second result by said artificial intelligence, which is based on the content of the user input in an anonymized and neutralized form; and the final process step
[0061] -6) processing the second result by the artificial intelligence to generate appropriate outputs, for example in the form of answers, and outputting them via a suitable output device, monitor, speaker, etc.
[0062] Advantageously, an intermediate step 7 is also carried out between process step 5 and process step 6, in which the communication system is executed by means of the corresponding artificial intelligence trained as described above: the output is processed by the artificial intelligence so that the output is stylized by the artificial intelligence and / or is in the form of a personality profile stored in the artificial intelligence.
[0063] The present invention makes it possible for the first time to equip a chatbot with a user experience in which the chatbot's output can be given a personalized and / or rhetorical and / or individualized stylistic approach as required.
Claims
1. A computer-implemented conversation method, comprising the following steps: -1) providing a computer-implemented interactive communication system having: an input device, a processor for text and / or speech recognition, a neural network including artificial intelligence, the artificial intelligence being used for text and / or speech evaluation, and generating first results from the text and / or speech recognition and storing them, -2) user input to the computer-implemented interactive communication system, -3) processing the user's input by artificial intelligence for text and / or speech evaluation and generating a first result from the text and / or speech recognition, -4) performing a first processing on the first result by means of an artificial intelligence provided and configured for this purpose, in the form of removing, formatting and / or neutralizing personalized style elements and / or other style elements in the voice and / or text input obtained from the user, so as to formalize the input and / or anonymize the user who made the input, -5) generating a second result by means of said artificial intelligence, which is based on the objective content of the user input in an anonymized and neutralized form, -6) processing the second results by the artificial intelligence to generate appropriate outputs, for example in the form of answers, and outputting them via a suitable output device.
2. The method according to claim 1, wherein Prior to process step 6), i.e., before the neutral second result is processed by the artificial intelligence and a corresponding output is generated by the artificial intelligence, there is an intermediate step 7), by which the output generated by the artificial intelligence is supplemented with personalized style elements and / or other style elements.
3. The method according to claim 1 or 2, wherein: Assign rhetorical stylistic means to the output produced by the correspondingly trained artificial intelligence.
4. A method according to claim 3, wherein metaphors are incorporated into the text as stylistic devices.
5. A method according to claim 3 or 4, wherein word repetition is incorporated into the text as a stylistic device.
6. Method according to any one of claims 3 to 5, wherein alliteration, hyperbole and / or oxymoron are incorporated into the text as stylistic devices.
7. Method according to any one of claims 3 to 6, wherein irony, part-of-speech translation and / or tautology are incorporated as stylistic means.
8. A method according to any one of claims 3 to 7, wherein archaisms, euphemisms and / or progressions are incorporated into the text as stylistic devices.
9. A method according to any one of claims 3 to 8, wherein paraphrases and / or verbiage are incorporated into the text as stylistic means.
10. A communication system, comprising the following components: - Input devices such as keyboard, microphone, camera, a first processor comprising a memory unit, the first processor being configured and arranged to perform text and / or speech recognition, the first processor having a neural network and artificial intelligence and being able to execute them and thereby generate a first result that can be stored in the connected first memory unit, - a second processor, which is configured to, first, retrieve the first result from the first storage unit, process it and provide it to a configured and trained artificial intelligence so as to neutralize the result, remove rhetoric and / or remove personalized style elements and / or insert personalized style elements, wherein, The result input to the artificial intelligence is either a first neutralized input from the user or a third neutral answer from the chatbot, wherein the artificial intelligence is configured and arranged to impart predetermined or artificial intelligence-selectable stylistic and / or rhetorical and / or personality attributes to the obtained neutral text and / or speech content, and / or to convert the obtained neutral text and / or speech content accordingly and forward it to a third processor, - wherein the third processor is configured and arranged to perform speech and / or text generation by the artificial intelligence and provide it in the form of acoustic, optical or both as an output in response to input obtained from the user via an output device of the communication system.
11. A computer product adapted to implement the method according to any one of claims 1 to 9.
12. A computer-readable storage medium comprising the computer program product according to claim 11.