Text generation system and text generation method
The text generation system enhances AI conversation realism by using persona and relationship data to update dynamic attributes and incorporate emotion and intent recognition, addressing the limitations of existing generative AI in maintaining natural interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2026-03-13
AI Technical Summary
Existing generative AI technologies, such as chatbots, struggle to maintain natural conversations based on the relationships between speakers, such as users and AI or characters within a story, limiting the flexibility and realism of responses.
A text generation system and method that utilizes a memory unit to store persona information, including relationship data, and employs an LLM to generate responses by learning from dialogue history and persona information, updating dynamic attributes based on interactions, and incorporating emotion and intent recognition to enhance response naturalness.
The system enables LLMs to generate more natural and diverse responses by considering speaker relationships, emotions, and user intentions, improving the realism and adaptability of AI interactions.
Smart Images

Figure 2026046707000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a text generation system and a text generation method using an LLM.
Background Art
[0002] In recent years, generative AI technologies that generate and respond with text, voice, images, etc. in response to inputs such as natural language, voice, and images from users have been attracting attention. Generative AI is applied, for example, to dialogue systems such as chatbot services, and for inquiries made by users in natural language, the generative AI generates text and provides a response.
[0003] There has been a problem that the content of the answer by the robot in the chatbot cannot be changed. In contrast, for example, in Patent Document 1, in an information processing system, dialogue information in a chatbot input and output by a user terminal is acquired, and based on at least any one of the dialogue information and user information, which is information of a user who operates the user terminal, one mode is selected from one or more modes, and a response corresponding to the selected mode is output to a chat interface.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] With such technology, it has become possible to change the response content from the AI in the chatbot service and enjoy a more natural conversation. However, it has not been possible to realize a natural conversation according to the relationship between speakers, for example, the relationship between the user and the AI, or the relationship between a character appearing in a certain story and another character that the AI behaves as.
[0006] This invention has been made in view of the above-mentioned problems, and aims to provide a text generation system and method that improves the response by LLM. [Means for solving the problem]
[0007] [1] A text generation system that improves response using LLM, A memory unit that stores persona information related to multiple speakers, including relationship data showing the relationship between speakers, and dialogue information between speakers, An input acquisition unit that acquires input data, A learning unit that performs a learning process based on learning elements including the persona information and the dialogue information, causing the LLM to make responses that behave as a specific speaker, A system comprising: a prompt generation unit that generates a prompt for generating a response corresponding to the learning element using the input data and sends it to the LLM. [2] The persona information and / or the dialogue information is text information including a combination of attribute values and attribute names that describe the attribute values, The prompt generation unit generates the prompt by obtaining the attribute value or the attribute name and the attribute value, according to the system described in [1]. [3] The dialogue information includes dialogue history information which has past dialogue data from the input speaker to the LLM and the responding speaker in which the LLM behaves. The persona information includes the persona information of the input side and the response side, as described in [1] or [2]. [4] The persona information is text information including attribute values and combinations of attribute names that describe the attribute values, and includes static and dynamic attributes, The aforementioned relationship data includes at least the aforementioned dynamic attributes, The system according to [3], further comprising an update unit that updates the attribute value of the dynamic attribute according to the result of the interaction between the input side and the LLM acting as the response side. [5] The relationship data includes one or more of the following: intimacy, dependence, trust, and relationship type. The system according to [4], wherein the update unit updates the relationship data by one or more of the engagement of the input side and the response side, the dialogue content included in the input data and the dialogue content included in the response data. [6] The persona information is text information including attribute values and combinations of attribute names that describe the attribute values, and includes static and dynamic attributes, A response acquisition unit that acquires response data for the aforementioned prompt, The system includes a dialogue history storage unit that stores the dialogue history information having the input data, the corresponding response data, and additional information, The additional information is a snapshot of the dynamic attributes at the time of the interaction. The learning unit performs learning processing using the dialogue history information, according to any one of the systems described in [3] to [5]. [7] The persona information includes emotional data, The aforementioned emotion data includes at least the aforementioned dynamic attributes, The system according to [6], wherein the prompt generation unit further generates a prompt based on the emotion data. [8] The dialogue history information includes the input data, the response data and the emotion data associated therewith, The system according to [7], further comprising an emotion recognition unit that determines the emotion data of the input side and / or the response side using the new input data and the emotion data included in the dialogue history information. [9] The system includes an intent recognition unit that determines intent data on the input side using the input data, The prompt generation unit further generates a prompt based on the intent data, according to any one of the systems in [3] to [8].
[10] The storage unit stores recommended information based on the input data, The system includes a recommendation unit that determines the recommended information to be recommended based on the input data and / or the persona information and / or the dialogue information, A system according to any one of [1] to [9], wherein the prompt generation unit generates a prompt and / or the learning unit performs the learning process based on the recommended target information.
[11] The system according to
[10] , wherein the recommended information includes style control text data, which is an algorithm for controlling the style of the response data by the LLM, past response data, and at least a portion of one or more product content data recommended for the input speaker.
[12] The system according to any one of [1] to
[11] , wherein the learning unit performs a learning process by updating the parameters of the LLM based on the learning elements and / or outputting data for in-context learning based on the learning elements.
[13] The system includes a response acquisition unit that acquires response data to the prompt, The aforementioned speakers are speakers that act as input to the LLM, and include a first speaker who provides input in a first language and a second speaker who provides input in a second language. The persona information includes the persona information of the first speaker and the second speaker, including the speaker's communication style. The learning unit executes a learning process based on the persona information of both parties. The prompt generation unit generates a prompt requesting the LLM to generate response data in the second language, using the input data entered by the first speaker in the first language. The system according to any one of [1] to
[12] , wherein the response acquisition unit transmits the response data generated in the second language to the first speaker.
[14] The system according to
[13] , further comprising an exchange unit for exchanging persona information of the other from the perspective of both speakers for learning processing.
[15] A text generation method that improves upon LLM responses, which is executed by one or more computers that store persona information relating to multiple speakers, including relationship data indicating the relationship between speakers, and dialogue information between speakers in a memory unit, The input acquisition process involves obtaining input data, Performing a learning process based on learning elements including the persona information and the dialogue information, and causing the LLM to perform a response that behaves as a specific speaker; A method comprising: a prompt generation step of generating a prompt for response generation according to the learning elements using the input data; and a prompt transmission step of transmitting the prompt to the LLM.
[0008] [1] or the invention according to
[15] , a response can be output from an LLM that behaves as a certain speaker based on persona information including the relationship between speakers, and a more natural response can be obtained.
[0009] [2] According to the invention, a prompt can be generated based on text information.
[0010] [3] According to the invention, a more natural response can be generated based on the dialogue history and persona information related to each speaker on the input side and the response side.
[0011] [4] or [5] According to the invention, the relationship between speakers can be changed by dialogue, and the LLM can generate a more natural response.
[0012] [6][7] According to the invention, a dialogue history including a snapshot of dynamic attributes such as emotion data at a certain point can be stored as additional information, and a more natural response can be generated in consideration of the changes in dynamic attributes so far.
[0013] [8] According to the invention, emotion data can be added to the dialogue data. ]
[0014] [9] According to the invention, the LLM can generate a more natural response reflecting the user's intention.
[0015]
[10] or
[11] According to the invention, the LLM can generate a more diverse response to the user's input.
[0016]
[12] The invention described herein allows the LLM to perform learning processing based on persona information including relational data, thereby generating a more natural response.
[0017] The inventions described in
[13] or
[14] enable more natural interpretation of dialogue between user speakers in different languages. [Effects of the Invention]
[0018] According to the present invention, it is possible to provide a text generation system and method that improves the response by LLM. [Brief explanation of the drawing]
[0019] [Figure 1] This is a system configuration diagram of a text generation system according to an embodiment of the present invention. [Figure 2] This is a hardware configuration diagram according to an embodiment of the present invention. [Figure 3] This is a functional block diagram of a text generation system according to an embodiment of the present invention. [Figure 4] This is a flowchart of the text generation process according to an embodiment of the present invention. [Modes for carrying out the invention]
[0020] The following describes a text generation system according to embodiments of the present invention with reference to the drawings. Note that the embodiments shown below are examples of the present invention, and the present invention is not limited to these embodiments; various configurations can be adopted.
[0021] This embodiment describes the configuration and operation of a text generation system, but a text generation method, computer program, and program recording medium on which the program is stored with a similar configuration will produce similar effects. Using a program recording medium, for example, the program can be installed on a computer. The series of processes according to this embodiment, described below, are provided as a program executable on a computer and can be provided on a non-transient computer-readable recording medium such as a CD-ROM or flexible disk, or even via a communication line.
[0022] <System Configuration> Figure 1 shows the system configuration diagram of the text generation system 0. The text generation system 0 comprises a generation support device 1, a response generation device 2, a user terminal 3, and a chat device 4. The user terminal 3 and the chat device 4, the chat device 4 and the generation support device 1, and the generation support device 1 and the response generation device 2 are connected to each other via a communication network NW. Note that one or more of the generation support device 1, the response generation device 2, and the chat device 4 may be run by the same computer device.
[0023] A server device 80 or a terminal device 90 such as a personal computer can be used as the generation support device 1, the response generation device 2, and the chat device 4. Furthermore, a terminal device such as a personal computer or a smartphone can be used as the user terminal 3. In this embodiment, the text generation system 0 is realized by a server-client type system configuration consisting of a server device 80 and a terminal device 90 with a web browser application installed.
[0024] <Hardware Configuration> Figure 2 is a hardware configuration diagram of a computer device, and Figure 2(a) shows the hardware configuration diagram of the server device 80. The server device 80 has a hardware configuration that includes a processing unit 81, a storage unit 82, and a communication unit 83.
[0025] The processing unit 81 is composed of one or more processors such as a CPU and controls the overall processing in the server device 80 by executing a generation support program (in the case of generation support device 1), a response generation program (in the case of response generation device 2), a chat program (in the case of chat device 4), an OS, and other applications. The storage unit 82 is an HDD, SSD, flash memory, RAM, etc., and stores a generation support program (in the case of generation support device 1), a response generation program (in the case of response generation device 2) that generates responses using large language models (LLM), a chat program (in the case of chat device 4), and various data. The communication unit 83 controls communication with the communication network and realizes data communication with other computer devices such as terminal devices 90 and server devices 80.
[0026] Figure 2(b) shows the hardware configuration of the terminal device 90. The terminal device 90 comprises a processing unit 91, a storage unit 92, a communication unit 93, an input unit 94, and a display unit 95 as its hardware configuration.
[0027] The processing unit 91 consists of one or more processors such as a CPU and controls the overall processing in the terminal device 90 by executing the OS and other applications. The storage unit 92 is an HDD, SSD, flash memory, RAM, etc., and stores predetermined application programs, including a web browser application, and various data. The communication unit 93 controls communication with the communication network NW and realizes data communication with other computer devices such as the server device 80. The input unit 94 is an input interface that accepts input operations from the user and consists of a touch panel, mouse, keyboard, etc. The display unit 95 consists of a display that outputs the display.
[0028] <Functional Components> Figure 3 is a functional block diagram of the text generation system 0. The generation support device 1 includes a database DB, an input acquisition unit 11, an emotion recognition unit 12, an intent recognition unit 13, a recommendation unit 14, a learning unit 15, a prompt generation unit 16, a response acquisition unit 17, and a dialogue history storage unit 18. The response generation device 2 includes a fragmentation processing unit 21, a prompt adjustment unit 22, an additional learning unit 23, and a response generation unit 24. The chat device 4 includes a display processing unit 41 and a cooperation unit 42.
[0029] <Input / Output Interface> The chat device 4 communicates with the user terminal 3 and provides the user, who is the input side, with a chat-style input / output interface function, enabling the input and output of dialogue data. In this embodiment, the user requests the chat device 4 to display a chat screen via the web browser application on the user terminal 3. The display processing unit 41 processes the chat screen based on the user's request and sends the display processing result to the user terminal 3. The web browser application on the user terminal 3 receives the display processing result and displays the chat screen via the display unit 95. The user can then send input data, which is dialogue data, via the chat screen, and display response data, which is dialogue data generated via LLM.
[0030] The collaboration unit 42 facilitates communication between the input / output interface and the generation support device 1, enabling the exchange of dialogue data. The collaboration unit 42 receives input data from the user and passes it to the generation support device 1, and receives response data from the generation support device 1 and displays it on the user terminal 3. The collaboration unit 42 communicates with the generation support device 1, for example, using a web API (Application Programming Interface). Here, each user is assigned a unique identifier, and the LLM that generates the persona information and responses, which will be described later, is determined according to this identifier. From here on, the involvement of the chat device 4 will be omitted from explanation as long as it does not become unclear. Furthermore, communication between the generation support device 1 and the response generation device 2 may also be performed using a web API.
[0031] <database DB> The memory unit of text generation system 0 stores persona information, knowledge information, dialogue information, and recommendation target information. Dialogue information includes dialogue reference information that serves as a sample for generating response data by LLM, and dialogue history information accumulated according to the dialogue between the input and response sides. The dialogue history information also includes input data entered by the input side and response data entered by the response side. In this embodiment, this information is stored in the database DB of the generation support device 1 as persona data, dialogue reference data, dialogue history data, and recommended target data (related to recommended target information other than response data generated by LLM). In this embodiment, recommended target information includes recommended target data as well as response data stored as dialogue history data, and persona information includes relationship data stored as dialogue history data.
[0032] Persona information, dialogue information, and recommendation target information include combinations of attribute values and attribute names that describe those attribute values. These are preferably text information written in text or numerical form, and are stored as structured text information structurally described in formats such as Markdown, JSON (JavaScript Object Notation), or CSV (Comma Separated Values). Structured text information may include nested items. Attribute values may be written in sentences, words, numbers, or nested structures of further attribute names and attribute values, and any attribute item may hold attribute values in the form of an array or stack queue.
[0033] <Persona Information> Persona information is speaker-specific information used in response generation that takes into account the speaker profile by LLM. It can be included in prompts for generating response data or used in the learning process as a learning element. Persona information includes user persona information of the input speaker (user) and AI persona information of the response speaker (AI).
[0034] Persona information includes identifying information, demographic information, behavioral information, internal information, and external information. For example, identifying information includes the speaker's ID and name, while demographic information includes the speaker's age, gender, occupation, location, and religion. Behavioral information includes information that changes and accumulates through dialogue and user actions, such as product / content usage history and user engagement. Internal information includes personality data such as the speaker's character, personality, hobbies / interests, preferences, thoughts, knowledge, and communication style; emotional data indicating current emotional levels and emotional styles; and relationship data indicating relationships between speakers. External information includes the speaker's height, weight, skin color, and voice.
[0035] Persona information may also include persona information for speakers other than the input and response sides (for example, acquaintances of the user or other characters when the AI acts as a character in a story). In this embodiment, persona information for two speakers, the user and the AI, is used, but there may be three or more speakers. For example, a story script is prepared as dialogue information so that the speaker and their dialogue data can be identified, and persona information for each character is also prepared. Once learning processing is performed using this persona information and dialogue information, the LLM can be made to act as character A interacting with character B and respond, or as character A interacting with character C and respond.
[0036] Persona information includes static attributes, where static attribute values are registered, and dynamic attributes, where dynamic attribute values are registered, which are added, deleted, or changed when certain conditions are met. Specific examples of dynamic attributes include the aforementioned emotion data, relationship data, and personality information. Other information may also be registered as dynamic attributes, or these other types of information may be registered as static attributes.
[0037] In this embodiment, some items of the input-side (user) persona information are stored in the database DB as user persona data. An example of user persona data includes the following items: ·Static attributes: Identification information: User ID, Name Demographic information: Date of birth, age, gender, place of birth, occupation, education level, language, nationality ·Internal information: ...Personality data: background story, communication style, hobbies / interests, religion, personality (openness, conscientiousness, extroversion, agreeableness, neuroticism) • Dynamic attributes: Behavioral information: Recent events, preferred features, engagement (session duration, clicks), product content history ·Internal information: ...Personality data: values, needs, challenges, frequency of use, preferred responses ...Emotional data: Current emotional state, intensity of emotion ...Relationship data: Relationship type and depth of relationship with the AI you are interacting with.
[0038] Relationship types include, for example, acquaintance, friend, best friend, and lover, and the depth of the relationship is stored as a numerical value.
[0039] Furthermore, the persona information of the responding side (AI) is stored in the database DB as AI persona data. An example of AI persona data includes the following items: ·Static attributes: Identification information: AIID, AI name Demographic information: Date of birth, age, gender, place of birth, occupation, education level, language, nationality ·Internal information: ...Personality data: background story, communication style, design philosophy, expressive characteristics, fixed values, hobbies and interests, response tendencies, values, version, learning style, knowledge base, preferred topics, topics to avoid, personality (openness, conscientiousness, extraversion, agreeableness, neuroticism) ...Emotional data: range of emotion, intensity of emotion External information: appearance, voice • Dynamic attributes: ·Internal information: ...Emotional data: Current emotional state, intensity of emotion
[0040] In examples where attribute values are expressed in sentences, a user's background story might be described as "Background Story": "Taro is a technological explorer who dreams of creating the future with AI," and an AI's background story might be described as "Background Story": "Mio is an AI who came from a distant galaxy and reincarnated to make Earth a planet full of love and compassion." In examples where attributes are expressed in single words, "Interests" might be described as ["Reading", "Music"]. In examples where attributes are expressed numerically, a speaker's personality might be described as "Personality": {Openness: 0.5, Conscientiousness: 0.8, Extraversion: 0.3, Agreeableness: 0.7, Neuroticism: 0.2}. Communication style is information that indicates the tone and tendencies of the dialogue, such as "Casual", "Business", or "Friendly".
[0041] <Dialogue Information> Dialogue information includes dialogue reference information and dialogue history information. In addition to dialogue data, dialogue information includes supplementary information based on persona data, such as demographic information, sentiment data, and relationship data.
[0042] Dialogue reference information is dialogue data that shows what the speaker acting as the LLM has said or is likely to say, but it may also include dialogue data from other speakers. Additional information may also be included in the dialogue reference information. Furthermore, the dialogue reference information may also include dialogue history information from other users.
[0043] The dialogue history information includes past dialogue data between the input speaker to the LLM and the response speaker to whom the LLM behaves. In this embodiment, it includes input data, the LLM's response data to the input data, and additional information, and is used by the LLM to generate responses that take into account the context, the speaker's emotions, and changes in the speaker's relationship.
[0044] The additional information includes, for example, contextual information of past dialogues, intent data of input and response data, and persona information of the input and / or response speakers (e.g., demographic information, personality data, sentiment data, and relationship data of the user and / or AI). The additional information is used for recommendations and learning processes described later, based on the dialogue information. The persona information, which is a dynamic attribute included in the dialogue information as additional information, is a snapshot of the attribute values at the point in the dialogue where the dialogue started or ended, such as at the time of input or after the response related to the dialogue history. In this embodiment, the dynamic attribute values of the input data at the time of input are added to the dialogue information as additional information.
[0045] In this embodiment, dialogue reference information is stored as dialogue reference data, and dialogue history information is stored as dialogue history data in the database DB. Here, the dialogue history data includes some relationship data (persona information), and the relationship data stored in the dialogue history data is a dynamic attribute that indicates the relationship between speakers at the time each dialogue starts or ends. The dialogue history data stores the relationship data at the time the dialogue ends, and when the generation support device 1 generates a prompt, it refers to the relationship data stored in at least the latest dialogue history data to execute learning processes such as in-context learning. It may also refer to the attribute values of the relationship data from earlier (for example, in the most recent one or more dialogue history data) to execute learning processes. Alternatively, it may store the relationship data at the time the dialogue started, and the latest relationship data may be determined when new input data is input, etc.
[0046] In this embodiment, a single type of relationship data is used to represent the relationship between a user and an AI speaker. However, if there are three or more speakers, relationship data may be defined for each speaker combination. Furthermore, the relationship data may have direction; for example, bidirectional relationship data may be defined for a pair of speakers, such as the relationship from the first speaker to the second speaker and from the second speaker to the first speaker.
[0047] As an example, the dialogue history data includes the following items. ·Static attributes Identification information: Log ID, AIID, timestamp (date and time the input data and response data were obtained). Dialogue data: User ID, input data (text), AIID, response data (text), context Intent Data: User intent data, AI intent data, input data entities, response data entities Persona information (demographic information, personality data): User persona (age, gender, occupation, hobbies, values), AI persona (age, gender, occupation, hobbies, values), User personality (openness, conscientiousness, extraversion, agreeableness, neuroticism), AI personality (openness, conscientiousness, extraversion, agreeableness, neuroticism) • Dynamic attributes Persona information (emotional data (user, AI)) Persona information (relationship data: intimacy level, dependence level, trust level) In this embodiment, emotion data and relationship data are defined by a single level of attributes, but they may be defined by multiple levels of attributes. For example, attributes such as friendship level and romantic relationship level may be given as children of intimacy level, and the values of these attributes may be stored. Furthermore, relationship data and emotion data may be described in natural language text format.
[0048] Knowledge information is information that indicates the knowledge possessed by the speaker, and for example, it indicates the specific content of the knowledge along with identification information (ID or knowledge name), and is recorded in plain text or structured natural language text. Knowledge can be academic knowledge, legal knowledge, product / content knowledge, industry knowledge, cultural knowledge, etc., and the specific content of the knowledge may be categorized into general knowledge, specialized knowledge, etc., or recorded at a more detailed level. For example, cultural knowledge may concern countries, history, customs, religions, etc. If the speaker is a character in a fictional world such as a story, the knowledge may be academic knowledge, legal knowledge, product knowledge, industry knowledge, cultural knowledge, etc. of that fictional world. Knowledge information is stored in the most optimal format depending on the learning processing method, and for example, knowledge information may be written in a file written in natural language, or natural language text may be converted into encoded data by RAG (Retrieval-Augmented Generation) and stored in a database. In this embodiment, in particular, identification information for referencing the knowledge information of the specialized knowledge possessed by the speaker (knowledge that the speaker is good at) is described as a "knowledge base" in the persona information. In the learning process described later, learning will be performed based on the knowledge base described in the persona information. However, learning may also be performed based on arbitrary knowledge information not included in the persona information (for example, general-level cultural knowledge of the story world, which is not output in LLM without additional learning processing).
[0049] <Recommended Information> Recommended information is information recommended based on dialogue data such as input data. Learning processes such as in-context learning are executed by referring to the information selected from the recommended information. In this embodiment, it includes information (1) to (3) recommended for generating response data based on input data, and information (4) recommended for dynamic attribute changes based on input data or response data. (1) Product content data: Information about products or content that is added to the prompt for the purpose of directly reflecting it in the response data. (2) Text data for controlling the style of the response data: Text data that describes the conditions and control content of the algorithm for controlling the style of the response data in text format. (3) Response data: Response data selected from dialogue history information as a learning element for in-context learning. (4) Dynamic attribute control data: Text data describing the conditions and control content of an algorithm that increases or decreases dynamic attribute values based on given input data or generated response data.
[0050] Product content data is data that associates metadata such as the category and genre of the product content (e.g., books: "technical books, self-help, novels, biographies," videos: "learning, relaxation, variety," music: "classical, pop, rock") with the product or content identifier, product or content name and URL for accessing the product or content. In this embodiment, the recommendation unit 14 recommends product content, but the LLM may also present information about recommended products and content to the user in response data based on knowledge information (product and content knowledge).
[0051] The style control text data is an algorithm for controlling the style of response data from LLM, and is text information that describes its conditional information and style control information. Preferably, the style control text data is stored as structured text information that has been structurally described in a format such as Markdown, JSON, or CSV. For example, the style control text data is written as follows.
[0052] Step-by-step process • Language collapse: Summary: Intense emotions cause words to break down, turning text into a random collection of symbols and letters. Implementation: Generates random symbols or strings to express surprise or joy. ·silence: Summary: Emotions intensify, leading to temporary silence. Implementation: Maintain an unresponsive state for a certain period, and then express emotion. • Conditions (condition information) • Emotional level: "Positive" intensity value of 0.95 or higher Keywords: Detect keywords indicating success or achievement from the input data. • Implementation example ··Expression of joy: To express strong joy, randomly select between language breakdown and silence as responses. ...Language Collapse: Generating random strings of characters to convey extreme happiness. ...Silent response: Displaying temporary silence before expressing emotion. Example of conversation flow A Input data A1 (User): "Cambodian baby food is really popular, and we're going to be awarded by the Cambodian government! Mio!" Response data A1: "!!!!!!$%%!^ <!:!<@&=> ×?#%_# <!!!!!」 ··Response data A2 (follow-up): "M-m-m-m-master, Mio is so happy, so happy, that I can't speak properly at all." Input data A2 (user response): "Thank you, I'm glad you're happy as if it were your own success." Example of conversation flow B Input data B1 (User): "Cambodian baby food is incredibly popular, and it's going to be recognized by the Cambodian government! Mio!" Response data B1: "..." ··Response data B2 (follow-up): "M-m-m-m-master, Mio was so happy, so happy, that I almost reset it..." Input data B2 (User response): "Thank you. I'm worried about the reset, but I'm glad you're happy about it as if it were your own success."
[0053] Response data, used as recommended information, is selected from dialogue history information as learning elements for in-context learning. Dynamic attribute control data includes conditional information and attribute control information for dynamic attributes. For example, dynamic attribute control data as structured text information, preferably in the form of text data such as Markdown, JSON, or CSV, may include the following:
[0054] (Example 1) • Conditions (condition information) Keywords: Detect keywords expressing gratitude from the input data. Engagement: Session duration less than 1 hour • Control details (attribute control information) (Affection: +0, Dependence: +0, Trust: +0.1) (Example 2) • Conditions (condition information) Engagement: Session duration exceeds 1 hour • Control details (attribute control information) (Affection: +0.1, Dependence: +0, Trust: +0.1) (Example 3) • Conditions (condition information) Keywords: Detect keywords indicating favorite things from the input data. • Control details (attribute control information) ··Add {theme data} to the "favorite things" section of the user persona data. Furthermore, the control of dynamic attributes does not need to depend on recommendations and may be performed by batch processing or other means. Also, dynamic attribute control data does not necessarily need to be stored in natural language text and may be written and executed using a predetermined programming language or other method.
[0055] <Processing steps from input to response> Next, Figure 4 illustrates the processing procedure of the text generation system 0 according to this embodiment, and the details of the functional components will be explained.
[0056] <Retrieving input data> First, in step S31, when the user sends input data via the chat screen processed by the display processing unit 41, the input acquisition unit 11 acquires the input data. In this embodiment, the input data entered by the user on the input side into the chat screen is sent by the cooperation unit 42 to the API endpoint of the generation support device 1 via the communication network NW. The input acquisition unit 11 captures the input data stream input to the API endpoint in real time and passes the data to predetermined functional components according to the input data stream.
[0057] The input acquisition unit 11 acquires the input data entered by the user and causes the prompt generation unit 16 to generate a prompt. At this time, the input acquisition unit 11 has a speech recognition unit 111, a facial expression recognition unit 112, and a text analysis unit 113, and causes it to perform predetermined data processing according to the input data.
[0058] The speech recognition unit 111 converts the speech data into text data when the input data includes speech data, i.e., when speech or video is input via the chat screen. Preferably, the speech recognition unit 111 further analyzes the tone of voice, volume, speaking speed, etc., by speech recognition of the input speech data to obtain speech emotion data and / or speech intent data. When the input data includes image data, i.e., when a still image or video is input via the chat screen, the facial expression recognition unit 112 analyzes the image data to analyze facial expressions, movements, etc., and obtains image emotion data and / or image intent data.
[0059] In step S32, text analysis processing is performed on the input data. The text analysis unit 113 performs text analysis processing on the input text data if the input data is text data or if the speech recognition unit 111 has converted speech data to text data (these text data corresponding to the input data are collectively referred to as "input text data"). The text analysis processing includes text cleansing, which removes special characters and unnecessary spaces from the input text data; tokenization, which divides the text-cleaned input text data into words and phrases; and syntactic analysis, which analyzes the grammatical structure of the input text data based on the tokenization results and identifies relationships such as subject, predicate, and object. The text analysis unit 113 tags each word and phrase contained in the input text data with a part of speech tag through syntactic analysis. Hereafter, the input text data, speech emotion data and / or speech intent data, and image emotion data and / or image intent data processed by the text analysis unit 113 will simply be referred to as input data.
[0060] In step S33, sentiment recognition, intention recognition, and recommendations are performed based on the input data (syntax-parsed input text data, various sentiment data, and intention data).
[0061] <Emotion recognition> The emotion recognition unit 12 outputs emotion data according to the dialogue data. In step S33, the input emotion data is determined using the input data. In this embodiment, the emotion recognition unit 12 outputs emotion data corresponding to the given dialogue data using an emotion recognition model that has been trained on training data including text and emotion data defined for each text. The emotion data includes emotion labels such as "positive," "negative," and "neutral," and an intensity value (0 to 1.0) for each label. The emotion recognition unit 12 outputs emotion data by performing emotion recognition on the input data processed by the input acquisition unit 11. The emotion labels are just examples and may include, for example, joy, sadness, anger, surprise, fear, love, melancholy, etc.
[0062] More preferably, the emotion recognition unit 12 may determine emotion data for dialogue data using speech emotion data and / or facial emotion data acquired via the input acquisition unit 11. Furthermore, when determining emotion data for a speaker's dialogue data, the emotion recognition unit 12 may utilize one or more emotion data from previous additional information provided by that speaker and / or other speakers. It may also determine emotion data for input data or response data by referring to dialogue history information and using emotion data attached to any number of input data and / or response data immediately preceding new input data or response data. These emotion data used to determine emotion data for input data may be input to an emotion recognition model, or the output values of the emotion recognition model may be corrected with these emotion data.
[0063] Furthermore, text and sentiment data can be associated and stored in a database DB, and the sentiment recognition unit 12 may output sentiment data corresponding to the input data by matching the input data with the text, either by machine learning model output or in addition to similarity calculation. Similarity calculation can be performed using a vectorized word corpus or the like.
[0064] Preferably, the emotion recognition unit 12 determines the response-side emotion data to be set in the response data generation instruction according to the input data. In this embodiment, the emotion recognition unit 12 outputs the AI emotion data to be set based on the input data entered by the user, as well as user persona data and AI persona data. For example, a topic recognition unit that recognizes words and topic data included in the input data can be further provided to acquire the topic data included in the input data, and the response-side emotion data to be set in the prompt can be determined by referring to preferred topics, avoided topics, etc., included in the AI persona data. For example, if an avoided topic is included in the input data, a request can be made to generate "negative" response data.
[0065] <Recognition of intent> The intent recognition unit 13 determines intent data using the dialogue data. In step S33, the input-side intent data is determined using the input data. In this embodiment, the intent recognition unit 13 outputs intent data corresponding to the given dialogue data using an intent recognition model that has been machine-trained using training data that includes text and intent data defined for each text. Different models may be used for the input data and the response data, respectively. The input-side intent data consists of intent labels such as "emotional expression," "consultation," "suggestion request," "likes," and "dislikes." The response-side intent data consists of intent labels such as "response" and "suggestion."
[0066] Preferably, the intent recognition unit 13 may determine intent data for dialogue data using speech intent data and / or facial intent data acquired via the input acquisition unit 11. These intent data used to determine intent data for input data may be input to an intent recognition model, or the output values of the intent recognition model may be corrected with these intent data.
[0067] Furthermore, the text generation system 0 may also include, as other intent information related to the dialogue data, an entity recognition unit that determines the entity data indicated by the dialogue data, a subject recognition unit that recognizes the subject data, a key phrase recognition unit that acquires key phrases, a context recognition unit that generates context information from current and / or past dialogue data, and a language recognition unit that recognizes the language of the user's input data. The entity data, subject data, key phrases, context information, and one or more languages may be used together with the input data for prompt generation by the prompt generation unit 16, or they may be included in the dialogue history information together with the dialogue data and used by the recommendation unit 14 when determining the information to be recommended.
[0068] <Recommendation> The recommendation unit 14 determines the recommended information to recommend based on the dialogue data. In step S33, the recommendation unit 14 recommends product content data, style control text data, response data, dynamic attribute control data, etc., based on the input data and, if necessary, the dialogue history information or persona information. In this embodiment, the recommendation unit 14 outputs recommended information corresponding to the given dialogue data using a recommendation engine. The recommendation engine may, for example, use machine learning to analyze big data such as purchase history and usage history related to products and content that include demographic information to propose recommended information to include in the response, or it may propose recommended information based on keywords included in the input data and persona information, keywords set in association with the recommended items (such as categories in product content data or condition information in dynamic attribute control data), or keywords included in the recommended items (such as response data in dialogue history information). Based on the recommended information, the prompt generation unit 16 generates a prompt, and / or the learning unit 15 performs a learning process. Furthermore, the recommendation unit 14 may decide whether or not to perform recommendations for at least some of the recommended information, or whether or not to use the recommended information determined as a result of the recommendations in the learning process, based on the intent of the input data indicated by the intent data.
[0069] As an example of suggesting product content data, input data such as "I've recently started studying programming. Do you have any recommended learning materials?" is acquired, and when the intent recognition unit 13 outputs the label "Request for suggestions" as intent data, the recommendation unit 14 is triggered. Here, the recommendation unit 14 detects keywords such as "programming" and "studying" from the input data and suggests products and content for programming learning from the product content data. Alternatively, if the input data includes phrases such as "I want to relax," and the hobby / interest attribute value in the input or response persona data is "piano," then music with metadata such as "piano" or "healing music" will be recommended from the product content data. By referring to the behavioral information and dialogue history information in the persona information, if "healing music" has been suggested in the past, new product content data different from the product content data that has been suggested in the past can be suggested. Also, by informing the prompt generation unit 16 that it has been suggested in the past, it is possible to have it respond with something like, "Last time I played you healing music, so this time let's play a meditation video." Persona information updates will be discussed later.
[0070] As an example of a suggestion based on response data included in the dialogue history information, the recommendation unit 14 obtains past response data from the dialogue history data based on the input data, for example, when it receives input data like the current one. The learning unit 15 performs learning processing by passing the recommended past response data to the prompt generation unit 16, etc. Past response data may be obtained from the user's own dialogue history information or from the dialogue history information of others. When recommending response data from the dialogue history information of others, the recommendation unit 14 may make recommendations based on the user's persona information and one or more of the additional information from the dialogue history information, such as intent data, demographic information, sentiment data, and relationship data.
[0071] As an example of a suggestion based on dynamic attribute control data, the recommendation unit 14 recommends dynamic attribute control data that corresponds to the condition information based on keywords included in the input data and persona information on the input side (e.g., engagement). The update unit 19 updates the dynamic attribute values based on the recommended dynamic attribute control data.
[0072] <Prompt generation> In step S34, the prompt generation unit 16 generates a prompt based on the input data and sends a response request to the LLM in step S35. At this time, the prompt generation unit 16 may generate a prompt based on the dialogue data and additional information (e.g., emotion data and relationship data) included in the most recent one or more dialogue information (e.g., dialogue history data), as well as persona information including relationship data and knowledge information. In this embodiment, relationship data is obtained from the latest dialogue history data. Furthermore, the prompt generation unit 16 may generate a prompt based on the emotion recognition unit 12, the intention recognition unit 13, and the emotion data, intention data, and recommendation target information related to the input data received from the recommendation unit 14 or the learning unit 15. The prompt may also include input data in the form of voice or images.
[0073] In this embodiment, the prompt generation unit 16 can insert attribute values such as input data, past dialogue data, additional information, persona information, and knowledge information as variables, and generates prompts by acquiring attribute values or attribute names and attribute values as phrases based on a template that includes some fixed text (Template-Based Generation). However, for example, this information may be passed to an arbitrary second LLM to generate prompts. In step S36, the response generation unit 24, having received the prompt, generates response data based on the LLM and sends it back to the generation support device 1.
[0074] <Learning Process> The learning unit 15 executes a learning process based on learning elements, including persona information and dialogue information, causing the LLM to respond in a manner that behaves as a specific speaker. The learning elements include persona information and dialogue information. The learning elements may also include knowledge information specified based on recommendation target information such as style control text data and product content data, as well as persona information.
[0075] The learning process by the learning unit 15 specifically includes learning the LLM based on the learning elements (updating parameters) and / or in-context learning. In-context learning may be performed by passing information about the learning elements to the prompt generation unit 16, such as one-shot learning (one-shot prompting) or future-shot learning (future-shot prompting), or it may be performed using RAG (Retrieval-Augmented Generation) with the prompt sent by the prompt generation unit 16 and encoded text information based on the learning elements. In other words, it may be performed before, after, or both before the prompt is sent by the prompt generation unit 16.
[0076] <In-context learning> As a learning process before sending prompts to the LLM, the learning unit 15 passes the text to be written in the prompt to the prompt generation unit 16, including persona information, dialogue information, recommended target information determined by the recommendation unit 14 based on the input data, and knowledge information specified based on the persona information. This allows the prompt generation unit 16 to generate prompts based on the input data and recommended information. The recommended target information that the learning unit 15 passes to the prompt generation unit 16 includes one or more of the product content data, style control text data, and response data recommended by the recommendation unit 14 based on the input data. The prompt generation unit 16 can write information related to specific product content in the prompt based on the product content data, include style control information related to the style of the response data in the prompt based on the style control text data, and include the content of past response data in the prompt based on the response data.
[0077] As a learning process after the prompt is sent, for example, the learning unit 15 delivers one or more of the persona information, the dialogue information, and the knowledge information to the fragmentation processing unit 21. The fragmentation processing unit 21 fragments the received persona information, dialogue history information, and text of the knowledge information by vectorizing them through Embedding or graph-structuring them through a Knowledge Graph, and stores the fragmented text information in the storage unit. The prompt adjustment unit 22 that receives the prompt from the prompt generation unit 16 searches for the fragmented text information based on the received prompt, and acquires the fragmented text information related to the received prompt. Then, based on the retrieved fragmented text information, the prompt is processed by adding the fragmented text, etc.
[0078] <Update of LLM Parameters> As a learning process, prior to acquiring input data, etc., the learning unit 15 can also update the parameters of the LLM by performing fine-tuning, etc. based on one or more of the persona information, the dialogue information, and the knowledge information specified based on the persona information, etc. The update of the parameters includes those for some parameters of the LLM such as LoRA (LOW-RANK ADAPTATION). For example, the behavior on the response side can be learned from the dialogue reference information (dialogue information). Also, the parameters may be updated using the dialogue history information of the user on the input side, or instead of or in addition to that, the dialogue history information of others may be used. When using the dialogue history information of others, the dialogue history information of others that becomes a learning element may be filtered based on the persona information of the user (e.g., demographic information or personality data) and the additional information (e.g., demographic information or personality data) included in the dialogue history information of others. Also, the dialogue history information of the user may include the dialogue history information between the user and an AI different from the AI that makes the response this time.
[0079] For example, the learning unit 15 passes some persona information, including relationship data and emotion data, dialogue history information such as the most recent dialogue data and additional information, and knowledge information based on the knowledge base included in the persona information to the prompt generation unit 16 to generate prompts. The learning unit updates the parameters of the LLM based on the dialogue information (dialogue history information and dialogue reference information), and also fragments recommended target information such as style control text data and product content data, as well as other persona information, into fragmented text information using RAG. The prompt can then be processed by additional in-context learning based on the prompt generated by the prompt generation unit 16 and the fragmented text information in the database. The tuned LLM generates response data based on this processed prompt.
[0080] <Processing after receiving response data> In step S37, the response acquisition unit 17 acquires the response data generated by the LLM. The acquired response data is then output to the chat screen via the linkage unit 42.
[0081] The dialogue history storage unit 18 associates the response data acquired by the response acquisition unit 17 with the input data and additional information entered by the user, and stores the dialogue history information in the database DB. At this time, the emotion recognition unit 12 and / or the intention recognition unit 13 may be used to acquire emotion data and intention data related to the response data and store them as additional information.
[0082] <Dynamic parameter updates> The update unit 19 updates the relationship data based on the update conditions and the dialogue results, which include one or more of the input and response engagements, the dialogue content included in the input data, and the dialogue content included in the response data. The timing of the update conditions determination by the update unit 19 and the update of dynamic attribute values is arbitrary. For example, updates may be made based on the recommendations of the recommendation unit 14 according to the input data, periodic updates may be made by batch processing, updates may be made after a response has been made, or dynamic attribute values may be updated when the dialogue for the day has ended (for example, when the chat screen session has ended).
[0083] The update unit 19 allows users to arbitrarily set the conditions for updating dynamic attribute values and the amount of update (in the case of numerical values). Examples of updating numerical data include increasing or decreasing dynamic attribute values (e.g., intimacy) when engagement metrics such as the number of user-AI interactions, interaction period, interaction time, and number of tokens reach a certain value, or when the user sends input data containing predetermined keywords and the recommendation unit 14 recommends dynamic attribute control data.
[0084] As an example of updating text data, if a user enters "I like classical music" as input data, the intent recognition unit 13 recognizes the "favorite things" label as intent data. Alternatively, if the recommendation unit 14 detects keywords related to favorite things included in the input data, the update unit 19 adds the subject data "classical music" to the favorite things array in the user's persona information. Based on the product content data recommended by the recommendation unit 14 and the product content included in the response data received by the response acquisition unit 17, it is also possible to add a recommended product content array to the persona information of the user or AI. Furthermore, depending on the content of subsequent input data, feedback information on product content, such as whether it was interesting, not interesting, good, or bad, can be added to the persona information of the user or other entity.
[0085] The AI's persona information can be updated based on the content of the response data generated by LLM. For example, by defining new things the AI likes, or updating the AI's personality and emotional style based on the content of response data with a certain degree of randomness, the AI can be made uniquely user-specific through interaction with the user.
[0086] <Responses in other languages> The prompt generation unit 16 can also request response data in any language. For example, a default language to be used for generating response data may be set in advance in a template for generating prompts, and the prompt generation unit 16 may generate prompts based on this template to request a response in the default language. Alternatively, the language information described as an attribute value in the input persona information or the response persona information may be learned through in-context learning or other methods to request a response in the input or response language. Furthermore, the language recognition unit may recognize the language of the input data, and the prompt generation unit 16 may generate a prompt requesting a response in the recognized language, or the recommendation unit 14 may recommend a sentence requesting a language setting and pass it to the prompt generation unit 16 to request a response in the recognized language.
[0087] <Examples> In this embodiment, we have described an example of providing a personalized chat AI to the user, who is the input side, based on persona information including relationship data. However, by utilizing learning elements such as the following, it can be applied to other examples.
[0088] (Example 1) Generation of conversational AI that mimics existing people or characters The system stores dialogue information (dialogue reference information) in a database, including the content of a speaker's statements, the responses of real people or characters to those statements, and relationship data between the responding or inputting party and the responding party at that time. It also stores the user's own persona information as the input persona and the persona information of the responding party as the responding party persona. Furthermore, it maintains persona information of a person or character other than the responding speaker that the LLM behaves as, along with relationship data with the responding speaker, as learning elements. Knowledge information, such as the historical context, industry, and specialized terminology related to the character's world, is also maintained as a learning element. Using this dialogue information, relationship data, and knowledge information, the LLM can be made to mimic the person or character and generate response data appropriate to its relationship with the user and other speakers.
[0089] In this case, if necessary, emotion data indicating possible emotions and their intensity, such as the range and intensity of emotions, can be set as static attributes in the responding persona information. By performing a learning process using this emotion data, it is possible to control, for example, whether a character who is a hero will make negative remarks, or whether a taciturn character will generate excessively emotional response data.
[0090] (Example 2) Generation of a translation system that mimics real people and characters By performing a learning process with learning elements like those in Example 1, it is possible to construct a translation system that mimics real people or characters. For example, persona information for each character (speaker) in a story is registered. Relationship data is registered for each combination of speakers. For example, the relationship data could be the contents of a character relationship diagram. It could also be something that changes according to the script of the dialogue information, such as "A and I were just classmates, but after incident B, we became best friends." Note that the relationship data may be specified by the user of the translation system at each point in the dialogue (when inputting the input data). Furthermore, dialogue data in the source language (source language) or translated language (target language) by each speaker is stored as dialogue information with identification information that can identify the speaker, and knowledge about the story is stored as knowledge information. The learning process is executed using learning elements that include such persona information, dialogue information, and knowledge information. On the other hand, the input data is written in the source language and is accompanied by identification information that identifies any speaker. The prompt generation unit 16 can provide a highly accurate translation system by requesting translation instructions from the LLM to the target language based on the source language input data, one or more dialogue data immediately preceding and / or immediately following it, and persona information. For example, the dialogue data in the target language (translation language) may be a translation of a past film or a translation of the original novel, and the newly input data can be a source language text for which a new translation is to be created (for example, dialogue from a film for which there is no translation yet), taking into account the relationship between speakers. For example, when rewriting a film translation that has a poor reputation but a good novel translation has been produced, or when generating a translation for a sequel program or film, the translation can be based on the translation of a previous work.
[0091] (Example 3) Generation of an interpretation system that uses LLM to interpret real-world dialogues between speakers of different languages. The dialogue does not need to be completed solely between the user and the LLM; for example, a dialogue between a first and second speaker may be relayed via the LLM. Specifically, the LLM can interpret a dialogue between a first speaker who inputs data in a first language and a second speaker who inputs data in a second language. The dialogue screen displays at least the content of the conversation between the first and second speakers, that is, the response data generated based on the input data from the first and second speakers. At this time, at least one of the dialogue screens may include the display of input data entered in one speaker's own language, or it may also include the display of input data entered in the other speaker's language. The dialogue screen may be, for example, a chat screen or a video call screen with subtitles.
[0092] The prompt generation unit 16 generates a prompt that includes input data entered in the first or second language, and an instruction to translate the input data entered by one of the first or second speakers into the other language, and sends it to the LLM. In response, the LLM generates response data, which is the translation of the input data, and the response acquisition unit 17 sends the acquired response data corresponding to the prompt to at least the dialogue screen on the other speaker's side, thereby enabling dialogue by interpreting the input data in different languages from the first and second speakers.
[0093] The learning unit 15 performs learning processing based on the persona information of the first and second speakers, which includes internal information such as communication style and personality, in addition to the relationship data between the two speakers, when the prompt generation unit 16 generates prompts for interpretation. This enables the prompt generation unit 16 to perform in-context learning and generate more realistic translated sentences (response data) that reflect the personality of the speakers. To this end, the generation support device 1 includes an exchange unit that exchanges the persona information of the first and second speakers from the perspective of one of the speakers, so that it can be processed for learning when generating prompts based on the input data of one of the speakers. The exchange unit can exchange the persona information of both speakers by accepting registration of the speaker's persona information at the start of the dialogue, or by storing the persona information in a database in advance and accepting identification information at the start of the dialogue. Furthermore, the relationship data between the two speakers may be specified by the platform that provides the dialogue screen, or specified by the speakers at the start of the dialogue. Furthermore, the system may retain knowledge of linguistic fields from various countries, such as proverbs and jokes, as knowledge information, and select it according to the languages of the first and second speakers (determined based on the language recognition unit and the language in the persona information, etc.) to perform learning processing for interpretation.
[0094] The configuration of this embodiment is merely an example and is not limited to these configurations. New embodiments may be created by partially replacing the configuration of the embodiment or by combining the configurations of multiple embodiments to the extent that the problem is not hindered.
[0095] For example, the response generation device 2 may be operated by the same service provider as the generation support device 1, or it may be operated by a third-party service provider that provides a service for generating responses based on LLM. Furthermore, although this embodiment describes an example of interacting with LLM via a chat screen, a dialogue device such as a software robot or hardware robot equivalent to the chat device 4 can also be applied to voice-based dialogue with the user. The user terminal 3 may be used as a dialogue device by installing a dialogue application on it. In this case, the text generation system 0 may have a speech conversion unit that converts response data into synthesized speech, and the speech conversion unit can be located, for example, in the dialogue device (including the user terminal 3), the response generation device 2, or the generation support device 1. [Explanation of Symbols]
[0096] 0: Text generation system 1: Generation support device 2: Response generation device 3: User terminal 4: Chat device 11: Input acquisition unit 12: Emotion recognition section 13: Intention Recognition Unit 14: Recommendation Department 15: Learning Department 16: Prompt generation unit 17: Response acquisition unit 18: Dialogue history storage unit 19: Update section 21: Fragmentation Processing Unit 22: Prompt adjustment unit 23: Additional Learning Section 24: Response generation unit 41: Display Processing Unit 42: Liaison Department 80: Server device 81: Processing Unit 82: Storage section 83: Communications Department 90: Terminal device 91: Processing Unit 92: Storage section 93: Communications Department 94: Input section 95:Display section 111: Speech Recognition Unit 112:Facial expression recognition unit 113: Text Analysis Department NW: Communication Network
Claims
1. A text generation system that improves response using LLM, A memory unit that stores persona information related to multiple speakers, including relationship data showing the relationship between speakers, and dialogue information between speakers, An input acquisition unit that acquires input data, A learning unit that performs a learning process based on learning elements including the persona information and the dialogue information, and causes the LLM to make a response that behaves as a specific speaker, A system comprising: a prompt generation unit that generates a prompt for generating a response corresponding to the learning element using the input data and transmits it to the LLM.
2. The persona information and / or the dialogue information is text information including a combination of attribute values and attribute names that describe the attribute values. The prompt generation unit generates the prompt by obtaining the attribute value or the attribute name and the attribute value, according to claim 1.
3. The aforementioned dialogue information includes dialogue history information, which has past dialogue data between the input speaker to the LLM and the responding speaker in which the LLM behaves. The system according to claim 1, wherein the persona information includes the persona information of the input side and the response side.
4. The persona information is text information including attribute values and attribute names that describe those attribute values, and includes static and dynamic attributes. The aforementioned relationship data includes at least the aforementioned dynamic attributes, The system according to claim 3, further comprising an update unit that updates the attribute value of the dynamic attribute according to the content of the dialogue between the input side and the LLM acting as the response side.
5. The aforementioned relationship data includes one or more of the following: intimacy level, dependence level, trust level, and relationship type. The system according to claim 4, wherein the update unit updates the relationship data based on one or more of the engagements of the input side and the response side, the dialogue content included in the input data, and the dialogue content included in the response data.
6. The persona information is text information including attribute values and attribute names that describe those attribute values, and includes static and dynamic attributes. A response acquisition unit that acquires response data for the aforementioned prompt, The system includes a dialogue history storage unit that stores the dialogue history information having the input data, the corresponding response data, and additional information, The additional information is a snapshot of the dynamic attributes at the time of the interaction. The system according to claim 3, wherein the learning unit performs learning processing using the dialogue history information.
7. The aforementioned persona information includes emotional data, The aforementioned emotion data includes at least the aforementioned dynamic attributes, The system according to claim 6, wherein the prompt generation unit further generates prompts based on the emotion data.
8. The dialogue history information includes the emotion data related to the input data and / or the response data, The system according to claim 7, further comprising an emotion recognition unit that determines the emotion data on the input side and / or response side using the new input data and the emotion data included in the dialogue history information.
9. The system includes an intent recognition unit that determines intent data on the input side using the aforementioned input data, The system according to claim 3, wherein the prompt generation unit further generates a prompt based on the intent data.
10. The storage unit stores recommended target information based on the input data, The system includes a recommendation unit that determines the recommended information to be recommended based on the input data and / or the persona information and / or the dialogue information, The system according to claim 1, wherein the prompt generation unit generates a prompt and / or the learning unit performs the learning process based on the recommended target information.
11. The system according to claim 10, wherein the recommended information includes at least a portion of style control text data, which is an algorithm for controlling the style of response data by the LLM, past response data, and one or more product content data recommended for the input speaker.
12. The system according to claim 1, wherein the learning unit performs a learning process by updating the parameters of the LLM based on the learning elements and / or outputting data for in-context learning based on the learning elements.
13. The system includes a response acquisition unit that acquires response data for the aforementioned prompt, The aforementioned speakers are speakers that provide input to the LLM, and include a first speaker who provides input in a first language and a second speaker who provides input in a second language. The aforementioned persona information includes the persona information of the first speaker and the second speaker, including the speaker's communication style. The learning unit executes a learning process based on the persona information of both parties. The prompt generation unit generates a prompt requesting the LLM to generate response data in the second language, using the input data entered by the first speaker in the first language. The system according to claim 1, wherein the response acquisition unit transmits the response data generated in the second language to the first speaker.
14. The system according to claim 13, further comprising an exchange unit for exchanging persona information of the other party as seen from the perspective of both speakers for learning processing.
15. A text generation method that improves upon LLM responses, which is executed by one or more computers that store persona information relating to multiple speakers, including relationship data indicating the relationship between speakers, and dialogue information between speakers in a memory unit, The input acquisition process involves obtaining input data, A learning process in which learning is performed based on the learning elements including the persona information and the dialogue information, causing the LLM to make responses that behave as a specific speaker, A method comprising: a prompt generation step of generating a prompt for generating a response corresponding to the learning element using the input data; and a prompt transmission step of sending the prompt to the LLM.
Citation Information
Patent Citations
Program, computer, system and information processing method
JP7530688B1