Dialogue response creator for generating dialogue response based on dialogue characteristics and relationships using language model

By training a language model to reflect the dialogue characteristics between users and interlocutors, natural and context-sensitive dialogue responses are generated, solving the problem of insufficient handling of relationships and characteristics in existing dialogue systems and improving the intelligence and naturalness of the dialogue system.

CN121958459APending Publication Date: 2026-05-01林耀焕
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
林耀焕
Filing Date
2024-11-14
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing dialogue systems are unable to effectively handle the relationship and dialogue characteristics between users and interlocutors, resulting in unnatural responses and a failure to maintain context in dialogue.

Method used

By extracting dialogue feature data from past conversations to train a language model, response data that reflects the relationship, dialogue content, emotions, and habits between users and interlocutors is generated, and dialogue responses are generated using dialogue features and relationships.

Benefits of technology

It enables the rapid and timely generation of natural responses that reflect the relationship, dialogue content, emotions, and dialogue habits between users and interlocutors, thereby enhancing the intelligence and context processing capabilities of the dialogue system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958459A_ABST
    Figure CN121958459A_ABST
Patent Text Reader

Abstract

The present invention provides a dialogue response creator for generating a dialogue response based on dialogue characteristics and relationships using a language model, the dialogue response creator storing one language model, response data generated by the language model reflecting dialogue characteristics of a speaker who has a dialogue with a user and a past dialogue made with the user, the data represents a response that the user can communicate to the speaker; and a processor that inputs, as first input data, real-time dialogue data including an ongoing real-time dialogue statement between the dialogue person and the user into the language model, and outputs, as first output data, the response data from the language model.
Need to check novelty before this filing date? Find Prior Art

Description

A dialogue response creator that uses language models to generate dialogue responses based on dialogue features and relationships. Technical Field

[0001] This invention relates to a dialogue response generation apparatus that uses a language model to generate dialogue responses based on dialogue characteristics and relationships. More specifically, it trains a language model by using dialogue characteristic data extracted from past dialogues between the interlocutor and the user as learning data to generate a language model that reflects the responses that the interlocutor and the user can convey to the interlocutor in real-time dialogue, thereby generating a response that reflects the relationship between the interlocutor and the user, the dialogue content, emotions, and habits. Background Technology

[0002] Large-scale language models (LLMs) and other super-large AI are artificial intelligence models that learn using large amounts of text data to perform natural language processing tasks and can be used for various language modeling tasks.

[0003] Most LLMs can be learned as datasets consisting of hundreds of billions of sentences, including various web documents and text data such as the internet, books, newspaper articles, and blogs, which can be used for a variety of applications such as natural language understanding, sentence generation, machine translation, chatbots, and automatic summarization.

[0004] Furthermore, LLM-based deep learning models are an artificial intelligence technique that can learn human-used languages ​​and perform various tasks using those languages. A prime example is understanding human-written text through natural language and then using that understanding to engage in dialogue.

[0005] Traditional dialogue systems store expected questions and answers in a large database, search the database for a question that matches the user's input, and then provide the answer. However, the traditional approach doesn't perform language processing, so for example, "Give me a TV." Wow, because "Give me a TV" is recognized as another string, so it's "Give me a TV." Even if such questions and answers exist in the database, there are still questions that cannot be answered if the user says "Give me a TV."

[0006] Furthermore, traditional dialogue systems only handle question-and-answer dialogues that cannot maintain context, thus failing to handle goal-oriented dialogues. Dialogues between people flow according to the topic, with variations in the words, predicates, and sentence endings depending on the participants. However, because existing dialogue systems cannot adaptively generate responses based on the dialogue participants, they cannot handle word endings or sentence variations, resulting in unnatural sentences being generated as responses.

[0007] Patent No. 10-1359718 is currently in the process of developing a dialogue management system and method. Patent No. 10-1359718 relates to a model for constructing a chat system that allows for free-flowing, non-purpose-specific dialogue with an AI agent, in order to effectively manage dialogue. Patent No. 10-1359718 provides dialogue services based on pre-defined firing pairs within the system, without considering the relationship between users and interlocutors or firing styles, or other dialogue characteristics.

[0008] Patent application No. 10-2015-0086534 describes a device and method for managing the dialogue order based on the dialogue context and topic. Patent application No. 10-2015-0086534 relates to how to extract keywords from a user's fire information, determine the dialogue topic, and, based on explicit signals including the user's gaze, gestures, and touch, consider the state of floor movement, and adjust the dialogue order according to a dialogue pattern stored in the dialogue context. Patent application No. 10-2015-0086534 does not consider dialogue characteristics such as the relationship between the user and the interlocutors or the dialogue style.

[0009] Patent application No. 10-2013-0124534 is initiating an interactive service apparatus and method based on a user's speaking style. Patent application No. 10-2013-0124534 relates to an interactive service apparatus and method based on a user's speaking style, specifically how the interactive service apparatus analyzes the meaning of user statements, understands the user's speaking intent, and generates response statements based on the user's speaking style. Patent application No. 10-2013-0124534 does not consider dialogue characteristics such as the relationship between the user and the interlocutor or dialogue style.

[0010] Patent No. 10-1497411 is currently in the process of developing a text style conversion device, text style conversion method, storage medium, and automatic dialogue service system and method. Patent No. 10-1497411 relates to a text style conversion device, text style conversion method, storage medium, and automatic dialogue service system and method, aiming to provide users with articles in multiple text styles to improve user familiarity. Patent No. 10-1497411 does not consider the relationship between the user and the interlocutor, or dialogue style and other dialogue characteristics.

[0011] Existing technical documents

[0012] Patent documents

[0013] (Patent Document 1) Korean Patent No. 10-1359718 Summary of the Invention

[0014] The problem that the invention aims to solve

[0015] The problem this invention aims to solve is to use dialogue feature data extracted from past conversations between speakers and users as learning data to train a language model and create a language model to represent the responses that users can convey to speakers in real-time conversations between speakers and users. This provides a dialogue response generation device that can quickly and timely generate responses that reflect the relationship, dialogue content, emotions, and dialogue habits between speakers and users.

[0016] The objectives of this invention are not limited to those described above. Other objectives and advantages of this invention not mentioned can be understood from the following description and can also be more clearly understood from the embodiments of this invention. Furthermore, it is readily apparent that the objectives and advantages of this invention can be achieved by the means and combinations thereof as described in the claims of the patent application.

[0017] means for solving problems

[0018] To address the aforementioned problem, a dialogue response generation device using the language model specified in this invention generates dialogue responses based on dialogue characteristics and relationships. This device stores a language model, whose generated response data reflects the dialogue characteristics of the speaker engaging in dialogue with the user and the past dialogues between the speaker and the user, and represents responses that the user can send to the speaker. The device also includes a processor that inputs real-time dialogue data containing ongoing real-time dialogue statements between the speaker and the user as first input data into the language model and outputs the response data as first output data from the language model.

[0019] It is worth mentioning that the language model is a dialogue feature extraction model that generates dialogue feature data by extracting dialogue features from past dialogue data containing past dialogue statements.

[0020] It is worth mentioning that the dialogue feature extraction model is the dialogue feature data, which can generate one or more contextual data representing the relationship between the interlocutor and the user, interlocutor feature data representing the dialogue sentence end type characteristics of the interlocutor's dialogue sentence, user feature data representing the dialogue sentence end type characteristics of the user's dialogue sentence, emotional feature data representing the user's emotions in the past dialogue, and contextual features representing the context of the past dialogue.

[0021] It is worth mentioning that the processor can input the historical dialogue data as the second input data into the language model, and output the dialogue feature data as the second output data from the language model.

[0022] Ideally, the language model is a response generation model that generates the response data from the real-time dialogue data.

[0023] It is worth mentioning that the processor can use the dialogue feature data as learning data to train the language model, so that the response generation model reflects the features of the past dialogue and generates the response data.

[0024] Invention Effects

[0025] According to the present invention, dialogue features extracted from past dialogues between speakers and users—dialogue feature data—are used as learning data to train a language model and generate a language model that reflects the responses that users can convey to speakers in real-time dialogues between speakers and users. This allows for the rapid and timely generation of responses that reflect the relationship, dialogue content, emotions, and dialogue habits between speakers and users.

[0026] The effects of the present invention are not limited to those mentioned above, and other effects not mentioned can be clearly understood by those skilled in the art from the following description. Attached Figure Description

[0027] Figure 1 shows the connection configuration between the dialogue response generator and the dialogue device, which uses a language model from one of the examples in this guide to generate dialogue responses based on dialogue features and relationships.

[0028] Figure 2 is a block diagram of the dialogue response generation device, which uses a language model in one example of this launch to generate dialogue responses based on dialogue features and relationships.

[0029] Figure 3 is a diagram illustrating the process by which a dialogue response generation device, using a language model to generate dialogue responses based on dialogue features and relationships in an example given at the beginning, generates dialogue feature data using a language model.

[0030] Figure 4 is a diagram showing a screen example of a dialogue response generation device that uses an example language model from this launch to generate dialogue responses based on dialogue characteristics and relationships.

[0031] Figure 5 is a diagram intended to illustrate the process by which the dialogue response generation device uses a language model in an example of this startup to generate dialogue responses based on dialogue features and relationships, and uses the dialogue feature data as learning data to train the language model.

[0032] Figure 6 is a diagram illustrating the process by which a dialogue response generation device uses a language model to generate response data. This device uses a language model from an example in this guide to generate dialogue responses based on dialogue features and relationships.

[0033] Figure 7 shows another example of the screen of the dialogue response generation device, which uses the language model of an example of this launch to generate dialogue responses based on dialogue characteristics and relationships.

[0034] Explanation of reference numerals in the attached figures

[0035] 100: A dialogue response generation device that uses a language model to generate dialogue responses based on dialogue features and relationships.

[0036] 200: Dialogue device;

[0037] 110: Address Book;

[0038] 120: Display;

[0039] 130: Input section;

[0040] 140: Memory;

[0041] 150: Processor. Detailed Implementation

[0042] The advantages and features of this invention, as well as the methods for implementing them, will become clear from the accompanying drawings and detailed implementation examples. However, this invention is not limited to the embodiments described below and may be embodied in different forms. These embodiments are provided only to complete the introduction of the invention and to fully inform those skilled in the art of its scope. The invention is defined solely by the scope of the claims.

[0043] The terminology used in this listing is for illustrative purposes and not for limiting the invention. In this listing, singular forms also include plural forms unless specifically mentioned herein. The use of "comprising" and / or "comprising" in the listing does not exclude the presence or addition of one or more other components. Throughout the description, the same graphic symbol refers to the same component, and "and / or" includes each and one or more combinations of the mentioned components. While "first," "second," etc., are used to describe various components, these components are not limited by these terms. These terms are used merely to distinguish one component from another. Therefore, a first component mentioned below may be a second component in the technical concept of the invention.

[0044] Unless otherwise defined, all terms used in this listing (including technical and scientific terms) are to be commonly understood by one of ordinary skill in the art to which this invention pertains. Furthermore, terms defined in commonly used dictionaries are not to be ideally or over-interpreted unless explicitly defined.

[0045] The terms "parent" or "module" used in this listing refer to hardware components, such as software, FPGAs, or ASICs, where the "parent" or "module" plays a certain role. However, the meaning of "parent" or "module" is not limited to software or hardware. A "parent" or "module" can be configured in addressable storage media or configured to play one or more processors. Thus, for example, a "parent" or "module" includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, properties, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functionality provided in components and "parents" or "modules" can be combined into fewer components and "parents" or "modules," or further separated from other components and "parents" or "modules."

[0046] Spatially relative terms such as "below," "beneath," "lower," "above," and "upper" can be used to describe the relationship between a component and other components, as shown in the figure. Spatially relative terms should be understood as terms that, in addition to the directions shown in the figure, also include the different orientations of the component when used or operated. For example, if the components shown in the figure are flipped, the component described as "below" or "beneath" to other components can be placed "above" to other components. Therefore, the example term "below" can encompass both the below and above directions simultaneously. Components can be assigned to other directions, so spatially relative terms can be interpreted according to their assignment.

[0047] In this description, "computer" refers to all types of hardware devices that contain at least one processor, and, depending on the implementation example, can be understood to include the software configuration running on that hardware device. For example, a computer can be understood to include servers, smartphones, tablets, desktops, laptops, and user clients and applications running on each device, but is not limited thereto.

[0048] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0049] Figure 1 is a connection configuration diagram between the dialogue response generation device and the dialog device. The device uses the language model of an instance of this startup to generate dialogue responses based on dialogue characteristics and relationships. Figure 2 is a block diagram of the dialogue response generation device that uses the language model of an instance of this startup to generate dialogue responses based on dialogue characteristics and relationships.

[0050] As shown in Figures 1 and 2, the dialogue response generation device 100, which uses a language model from one example in this guide to generate dialogue responses based on dialogue features and relationships, can be a user-controlled electronic device.

[0051] Thus, the dialogue response generation device 100, which uses a language model from one example in this guide to generate dialogue responses based on dialogue characteristics and relationships, can be an electronic device that provides a chat function for users to converse with interlocutors via text through wireless communication.

[0052] To this end, using a language model in one of the examples in this guide, a dialogue response generation device 100 that generates dialogue responses based on dialogue characteristics and relationships can display user dialogue statements input by the user on the screen, and if a transmission request is input by the user, the input user dialogue statements can be sent to an electronic dialogue device 200 controlled by the interlocutor.

[0053] Subsequently, using the language model in one example of this startup, a dialogue response generation device 100 can receive and display dialogue statements input by the dialogue device 200, based on dialogue characteristics and relationships, and generate dialogue responses based on dialogue characteristics and relationships.

[0054] To this end, using a language model in one example of this guide, a dialogue response generation device 100 that generates dialogue responses based on dialogue features and relationships may include an address book 110, a display 120, an input component 130, a memory 140, and a processor 150.

[0055] The communication unit 110 can send user conversations to the conversational device 200 and receive conversations from the conversational device 200, as described above.

[0056] In addition, the address book 110 can send and receive conversation devices 200 and various information and data.

[0057] In addition, the address book 110 can also send and receive various information and data with other electronic devices besides the conversational device 200.

[0058] Therefore, the communication unit 110 can include various communication chips, such as Wi-Fi chips, Bluetooth chips, wireless communication chips, NFC chips, and Bluetooth Low Energy (BLE) chips. In this case, the Wi-Fi chip, Bluetooth chip, and NFC chip communicate via LAN, Wi-Fi, Bluetooth, and NFC respectively. When using a Wi-Fi or Bluetooth chip, various connection information, such as SSID and session keys, can be sent and received first. This information is then used to establish a communication connection before sending and receiving various messages. Wireless communication chips refer to chips that perform communication according to various communication standards such as IEEE, Zigby, 3G (3rd Generation), 3GPP (3rd Generation Partnership Project), LTE (Long Term Evolution), and 5G (5th Generation).

[0059] As described above, the display 120 can display user dialogue text and conversation text.

[0060] In addition, the display 120 can display various driving applications and show various information and data on the screen.

[0061] For this purpose, the display 120 may be equipped with a display panel and a display panel control circuit.

[0062] Input section 130 can receive various inputs from the user.

[0063] Specifically, the input section 130 can receive various requests, selections, and user dialogues from the user.

[0064] Thus, the input section 130 can be configured in various ways as long as the user inputs, but it is best to be a touch panel combined with the display panel.

[0065] Additionally, using a language model in one example of this launch, a dialogue response generation device 100 that generates dialogue responses based on dialogue characteristics and relationships can extract dialogue characteristic data, which are the dialogue characteristics of past dialogues between the interlocutor and the user.

[0066] Furthermore, the dialogue response generation device 100, which uses a language model in one example of this launch to generate dialogue responses based on dialogue features and relationships, can train the language model as learning data using dialogue feature data, and use the language model to generate response data that reflects the responses that the user can convey to the interlocutor in a real-time dialogue between the interlocutor and the user.

[0067] In the following sections, please describe the dialogue feature data, the training of the language model, and the generation of the response data.

[0068] Figure 3 is a diagram illustrating the process of a dialogue response generation device that uses a language model to generate dialogue response based on dialogue features and relationships in an example of this launch to generate dialogue feature data. Figure 4 is a diagram showing an example of a screen of a dialogue response generation device that uses a language model to generate dialogue response based on dialogue features and relationships in an example of this launch.

[0069] Referring further to Figures 3 and 4, memory 140 can store a language model (LLM) to generate dialogue feature data representing dialogue features of past conversations between the user and the speaker.

[0070] To this end, a language model (LLM) can include a dialogue feature extraction model (EM), which extracts dialogue features from past dialogue data containing past dialogue statements to generate dialogue feature data.

[0071] At this time, the processor 150 can store past dialogue data in memory 140, including user dialogue sent and received from the start of the dialogue between the user and the interlocutor to the end of the dialogue between the user and the interlocutor, as well as dialogue text (first dialogue text) composed of interlocutor dialogues.

[0072] In other words, past conversations may be conversations that ended in the past.

[0073] Therefore, if, after the start time, the processor 150 inputs a dialogue end time from the user, 1) the dialogue end time from the user input, 2) the time of inputting the user dialogue from the input time or the time of receiving the dialogue does not exceed a preset reference time, or 3) the time of inputting the dialogue from the input time or the time of receiving the dialogue does not exceed a preset reference time, then it can be determined that the dialogue has ended and the dialogue is classified as a past dialogue.

[0074] In addition, the processor 150 can input past dialogue data as second input data into the language model (LLM) and output dialogue feature data from the language model (LLM) as second output data.

[0075] Specifically, if the processor 150 inputs historical data as second input data into the language model (LLM), the historical dialogue data input into the language model (LLM) will be input into the dialogue feature extraction model (EM). The dialogue feature extraction model (EM) will use the input historical dialogue data to generate dialogue feature data, and the generated dialogue feature data can be output from the language model (LLM) to the second output data.

[0076] At this point, the dialogue feature extraction model (EM) can classify user dialogue statements and interlocutor dialogue statements from past dialogues after inputting past dialogue data.

[0077] Subsequently, the Dialogue Feature Extraction Model (EM) can generate relational feature data representing the relationship between the interlocutor and the user as the dialogue feature data.

[0078] Furthermore, the Dialogue Feature Extraction Model (EM) can generate relational feature data representing the relationship between the speaker and the user as the dialogue feature data. For example, the relationship between the speaker and the user can be one of the following: teacher and parent, romantic relationship, workload and boss relationship, friend and parent relationship, but it can also be other relationships.

[0079] In addition, the Dialogue Feature Extraction Model (EM) can generate dialogue feature data, which represents the characteristics of the end type of interlocutor dialogue statements, as the dialogue feature data.

[0080] The dialogue sentence termination type can include affirmative type (indicating the dialogue sentence is affirmative), negative type (indicating the dialogue sentence is negative), interrogative type (indicating the dialogue sentence is interrogative), imperative type (indicating the dialogue sentence is imperative), exclamatory type (indicating the dialogue sentence is exclamatory), and request type (indicating the dialogue sentence is request).

[0081] Among them, the characteristics of the dialogue ending type of the interlocutor's dialogue can be the features of the sentences used by the interlocutor in the dialogue ending type.

[0082] In addition, the Dialogue Feature Extraction Model (EM) can generate user feature data for the dialogue feature data, which represents the features of the end type of the user's dialogue statement.

[0083] Among them, the characteristics of the user's dialogue sentence ending type can refer to the features of the sentences used by the user, or the characteristics of the dialogue sentence ending type.

[0084] In addition, the dialogue feature extraction model (EM) can generate the dialogue feature data by using emotional feature data that reflects the user's emotions in past dialogues.

[0085] Furthermore, the Dialogue Attribute Extraction Model (EM) can generate contextual attribute data representing past dialogue context as the dialogue attribute data.

[0086] On the other hand, before generating dialogue feature data through the language model (LLM), the user device 100 can obtain the dialogue information I1 of the interlocutor from the user through the input section 130.

[0087] The speaker information I1 may include one or more values ​​representing the speaker's gender, age, and the level of intimacy between the speaker and the user.

[0088] Subsequently, the processor 150 serves as the second input data, which can not only input past dialogue data, but also input dialogue information I1 into the language model (LLM), and output dialogue feature data from the language model (LLM) as the second output data.

[0089] This method can improve the accuracy of dialogue features.

[0090] Figure 5 is a diagram illustrating how to generate dialogue responses based on dialogue features and relationships using a language model from one example in this guide. Figure 6 is a diagram illustrating the process of a dialogue response generator generating response data using an example from this guide, which generates dialogue responses based on dialogue features and relationships using a language model from one example in this guide. Figure 7 is a diagram illustrating how to generate dialogue responses based on another example from one example in this guide.

[0091] Referring again to Figure 4 or Figure 7, the processor 150 can, without ending the ongoing real-time dialogue between the interlocutor and the user, if it receives a dialogue from the interlocutor, input the real-time dialogue data containing the real-time dialogue as the first input data into the language model (LLM), and output the response data as the first output into the language model (LLM).

[0092] At this point, the Language Model (LLM) can generate response data that reflects the conversational characteristics of past dialogues between the speaker and the user, thus representing the responses that the user can convey to the speaker.

[0093] Thus, a Language Model (LLM) can include a Response Generation Model (GM) that generates response data from real-time dialogue data, which contains dialogue statements from the ongoing real-time conversation between the speaker and the user.

[0094] To this end, processor 150 can train a language model (LLM). More specifically, processor 150 can train the language model (LLM) using learning data so that the response generation model (GM) can generate response data based on the characteristics of past dialogues.

[0095] More specifically, the processor 150 can train a response generation model (GM) by learning data to reflect the dialogue characteristics of past dialogues between the user and the speaker, thereby generating response data that represents the responses that the user can convey to the speaker in real-time dialogues between the user and the speaker.

[0096] Subsequently, each time the processor 150 receives the dialogue text I2 from the interlocutor, it can input the real-time dialogue data containing the ongoing real-time dialogue between the interlocutor and the user as the first input data into the language model (LLM), and output the response data as the first output data from the language model (LLM).

[0097] Specifically, if the processor 150 inputs real-time dialogue data as the first input data into the language model (LLM), the real-time dialogue data input into the language model (LLM) will be input into the response generation model (GM). The response generation model (GM) will use the input real-time dialogue data to generate response data, and the generated response data can be output from the response generation model (GM) to the first output data.

[0098] At this point, multiple response data I3 can be generated according to the response end type. Specifically, response data representing affirmative sentences, negative sentences, interrogative sentences, command sentences, exclamatory sentences, and request sentences can be output as the first output data.

[0099] Subsequently, the processor 150 can control the display 120 to display response data I3 on ​​the screen as a candidate response to the last received speaker dialogue that the user can send to the speaker, as the first output data output.

[0100] Next, the processor 150 can control the address book 110 to send the selected response data to the dialogue device 200 via the user dialogue statement if any one of the multiple response data I3 is selected from the user through the input terminal 130.

[0101] User device 100 and dialogue device 200 can directly send and receive dialogue messages, but they can also send and receive dialogue messages between user device 100 and dialogue device 200 through a chat server.

[0102] Additionally, if, after a real-time conversation begins, the processor 150 1) inputs a conversation end time from the user, 2) inputs a conversation time from the user's input time or the time the conversation is received that does not exceed a preset baseline time, and 3) inputs a conversation time from the input time or the time the conversation is received that does not exceed the preset baseline time, then it can determine that the real-time conversation has ended and classify the real-time conversation as a past conversation.

[0103] Subsequently, after the real-time dialogue is classified as a past dialogue, the processor 150 can again input the past dialogue data as the second input data into the language model (LLM) and output the dialogue feature data from the language model (LLM) as the second output data.

[0104] In this way, users and interlocutors can generate dialogue feature data that reflects the latest dialogue characteristics.

[0105] On the other hand, according to other embodiments, the processor 150 can calculate a response selection ratio, which represents the ratio of the number of times the user selects response data to the number of times it is displayed on the screen.

[0106] Subsequently, if the selected response ratio is lower than the standard ratio, the processor 150 may not display the response data on the screen while the dialog box is engaged in real-time dialogue with the user.

[0107] In other words, if the selected percentage of response is lower than the standard percentage, the processor 150 can control the display 120 so that the response data is not displayed on the screen when having a real-time conversation with that interlocutor.

[0108] In another embodiment, the processor 150 can calculate the number of response end types selected from the user, and calculate the selection type ratio, which is the ratio of the number of times response data is selected from the user to the number of each response end type.

[0109] The processor 150 can control the display 120 to determine the response termination type with the highest selection type ratio, and among the multiple response data generated for a conversational dialogue, the response data with the highest selection type ratio will be highlighted compared to other response data.

[0110] For example, processor 150 can control display 120 to make the response type with the highest selection type ratio appear larger than other response data, to make the response type with the highest selection type ratio appear higher on the screen than other response data, or to make the response type with the highest selection type ratio appear in a different color than other response data.

[0111] This method allows you to highlight response data for response end types that are selected more frequently.

[0112] Memory 140 can utilize a language model to store various programs and data required for the actions of the dialogue response generation device 100, which generates dialogue responses based on dialogue characteristics and relationships. Memory 140 can be implemented as non-volatile memory, volatile memory, flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.

[0113] The processor 150 can utilize various programs stored in the memory 140 to control the overall operation of the dialogue response generation device 100 using a language model. This device generates dialogue responses based on dialogue characteristics and relationships. The processor 150 may consist of RAM, ROM, a graphics processing unit, a main CPU, first to n interfaces, and a bus. In this case, the RAM, ROM, graphics processing unit, main CPU, first to n interfaces, etc., can be interconnected via the bus.

[0114] RAM stores the operating system and applications. Specifically, if a dialogue response generation device 100 that uses a language model to generate dialogue responses based on dialogue features and relationships is launched, the O / S will be stored in RAM, and various application data selected by the user can also be stored in RAM.

[0115] The ROM stores the command set used to boot the system. If a boot command is entered and power is applied, the main CPU will copy the operating system (O / S) stored in memory 140 to RAM according to the commands stored in the ROM, and then run the O / S to boot the system. After booting, the main CPU copies various application programs stored in memory 140 to RAM and runs the copied applications to perform various operations.

[0116] The main CPU accesses memory 140 and uses the operating system stored in memory 140 to perform booting. Furthermore, the main CPU utilizes various programs, content, and data stored in memory 140 to perform various operations.

[0117] The first to n interfaces are connected to the various components. One of the first or n interfaces may be a network interface that connects to external devices via a network.

[0118] Additionally, processor 150 can control artificial intelligence models (the language model (LLM), dialogue feature extraction model (EM), and response generation model (GM)). In this case, processor 150 may of course include a graphics dedicated processor (e.g., a GPU) for controlling the artificial intelligence models (the language model (LLM), dialogue feature extraction model (EM), and response generation model (GM)).

[0119] Furthermore, the artificial intelligence model according to the present invention (the language model (LLM), dialogue feature extraction model (EM), and response generation model (GM)) can be a model based on supervised learning or unsupervised learning. In addition, the artificial intelligence model according to the present invention (the language model (LLM), dialogue feature extraction model (EM), and response generation model (GM)) can include SVM (support vector machine), decision tree, neural network, and the methodologies used therein.

[0120] As an example, the artificial intelligence model according to the present invention (the language model (LLM), dialogue feature extraction model (EM), and response generation model (GM)) can be an artificial intelligence model based on a synthetic multiplicative neural network (CNN) learned from input learning data. However, it is not limited to this; various artificial intelligence models can also be applied to the present invention. For example, models such as DNN (Deep Neural Network), RNN (Recurrent Neural Network), and BRDNN (Bidirectional Recurrent Deep Neural Network) can be used as artificial intelligence models, but are not limited to this.

[0121] At this point, Convolutional Deep Neural Networks (CNNs) are a type of multiplayer perceptron designed to use minimal preprocessing. A CNN consists of one or more convolutional layers and ordinary artificial neural network layers placed on top of them, further utilizing weights and pooling layers. Thanks to this structure, CNNs can fully utilize two-dimensional input data. Furthermore, CNNs can be trained using standard inverse propagation. CNNs are easier to train than other feedforward artificial neural network techniques and have the advantage of using fewer parameters.

[0122] Furthermore, a deep neural network (DNN) is an artificial neural network (ANN) composed of multiple hidden layers between the input layer and the output layer.

[0123] At this point, the structure of a deep neural network can be composed of perceptrons. A perceptron consists of multiple input values, a processor, and an output value. The processor multiplies each input value by a weight, then sums all the input values ​​multiplied by the weights. The processor then substitutes the synthesized value into an activation function and outputs a single value. If a specific value is desired from the output of the activation function, the weights multiplied by each input value can be modified, and the output value can be recalculated using the modified weights. Each perceptron can use a different activation function. Furthermore, each perceptron accepts the output from the previous layer as input and then uses the activation function to find the output. The found output is then passed as the input to the next layer. Through this process, several output values ​​can ultimately be obtained.

[0124] A Recurrent Neural Network (RNN) is a neural network in which the connections between the units that make up an artificial neural network form a directed cycle. Unlike feedforward neural networks, RNNs can utilize the internal memory of the neural network to process arbitrary inputs.

[0125] Deep Belief Networks (DBNs) are generative graphical models used in machine learning. In deep learning, they refer to deep neural networks composed of multiple layers with latent variables. A key characteristic is that while there are connections between layers, there are no connections between units within a single layer.

[0126] Deep trust neural networks, due to the characteristics of their generative model, can be used for pre-learning. After learning the initial weights through pre-learning, the weights can be fine-tuned through backpropagation or other discriminative algorithms. These characteristics are very useful when there is limited training data, because the less training data available, the greater the impact of the initial weight values ​​on the final model. Compared to arbitrarily set initial weight values, pre-learned initial weight values ​​are closer to the optimal weights, which can improve the performance and speed of the unadjusted phase.

[0127] The description of artificial intelligence and its learning methods is for illustrative purposes only, and the artificial intelligence and its learning methods used in the implementation examples are not limited. For example, systems that can be used to implement all types of artificial intelligence technologies and their learning methods applicable to those typically skilled in the art to solve the same problem have been initiated.

[0128] On the other hand, the processor 150 may include one or more cores (not shown) and graphics processing units (not shown) and / or connection channels (such as buses) for sending and receiving signals with other components.

[0129] According to one example, processor 150 executes the illustrative method related to the present invention by running one or more instructions stored in memory 140.

[0130] For example, processor 150 can acquire new learning data by running one or more instruments stored in memory 140, test the acquired new learning data using the learned model, extract the first learning data whose test results and label information accuracy exceed a specified first standard value, delete the extracted first learning data from the new learning data, delete the new learning data that has been used for learning, and thus reuse the new learning model for learning.

[0131] Additionally, processor 150 may also include RAM (Random Access Memory) and ROM (Read-Only Memory) for temporary and / or permanent storage of signals (or data) processed internally by processor 150. Furthermore, processor 150 may be implemented as a system on chip (SoC), which includes at least one graphics processing unit, RAM, and ROM.

[0132] Memory 140 can store programs (one or more instructions) that process and control processor 150. The programs stored in memory 140 can be divided into multiple modules according to their functions.

[0133] Furthermore, different embodiments of the present invention can be complementary or combined.

[0134] The components of this invention can be implemented by a program (or application program) and stored in a medium for operation in conjunction with a hardware computer. The components of this invention can be executed using software programming or software elements; similarly, instantiation can be implemented using programming or scripting languages ​​such as C, C++, Java, assembly, and Python, including various algorithms composed of data structures, processes, routines, or other programming components. Functionality can be implemented through algorithms that run on one or more processors.

[0135] While embodiments of the present invention have been described above with reference to the accompanying drawings, those skilled in the art will understand that the present invention can be implemented in other specific forms without altering its technical concept or essential features. Therefore, the embodiments described above are exemplary in all respects and should be understood as non-limiting.

Claims

1. A dialogue response creator that uses a language model to generate dialogue responses based on dialogue features and relationships, characterized in that, The dialogue response creator that uses a language model to generate dialogue responses based on dialogue characteristics and relationships includes: storing a language model to reflect dialogue characteristics of past dialogues between a user and a speaker, generating response data reflecting responses that the user can convey to the speaker; and a processor that inputs real-time dialogue data containing ongoing real-time dialogue statements between the speaker and the user as first input data into the language model and outputs the response data as first output data from the language model.

2. The dialogue response creator according to claim 1, which uses a language model to generate dialogue responses based on dialogue features and relationships, is characterized in that, The language model includes a dialogue feature extraction model that extracts dialogue features from past dialogue data containing past dialogue statements and generates dialogue feature data.

3. The dialogue response creator that uses a language model to generate dialogue responses based on dialogue features and relationships according to claim 2, characterized in that, The dialogue feature extraction model generates at least one of the following data: relationship attribute data representing the relationship between the interlocutor and the user, dialogue attribute data representing the dialogue statement end type characteristics of the interlocutor's dialogue statement, user characteristic data representing the dialogue statement end type characteristics of the user's dialogue statement, emotional characteristic data representing the user's emotions in the past dialogue, and context characteristic data representing the context of the past dialogue.

4. The dialogue response creator according to claim 2, which uses a language model to generate dialogue responses based on dialogue features and relationships, is characterized in that... The processor inputs the historical dialogue data as second input data into the language model and outputs the dialogue feature data as second output data from the language model.

5. The dialogue response creator according to claim 2, which uses a language model to generate dialogue responses based on dialogue features and relationships, is characterized in that... The language model includes a response generation model that generates the response data from the real-time dialogue data.

6. The dialogue response creator according to claim 5, which uses a language model to generate dialogue responses based on dialogue features and relationships, is characterized in that... In order to enable the response generation model to reflect the characteristics of the past dialogue, the processor generates the response data and uses the dialogue characteristic data as learning data to train the language model.

Citation Information

Patent Citations

  • Conversation Managemnt System and Method Thereof

    KR101359718B1