Method and system for providing persona chatbot

By dividing text content and generating prompt templates for persona chatbots, the method addresses the high cost and resource intensity of existing LLMs, enabling immersive and realistic character interactions with reduced computational load.

KR102997471B1Active Publication Date: 2026-07-29AINE BLUME CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
AINE BLUME CO LTD
Filing Date
2023-12-18
Publication Date
2026-07-29

AI Technical Summary

Technical Problem

The high cost and resource-intensive process of creating character-specific Large Language Models (LLMs) for persona chatbots, which require significant human effort to reflect character personalities and relationships, limits the scalability and affordability of entertainment-oriented chatbot services.

Method used

A method and system for generating persona chatbots that involve dividing text content into subtext contents, creating prompt templates based on these subtexts, and using a language model to generate responses considering character relationships, including gender, age, and personality, with summary information stored in a vector database for contextually accurate responses.

Benefits of technology

This approach enables a more immersive and realistic persona chatbot service by accurately reflecting character relationships and reducing the computational burden, allowing for extensive text content processing and diverse conversational scenarios, including group chats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 112023142044936-PAT00001_ABST
    Figure 112023142044936-PAT00001_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method for providing a persona chatbot. The method for providing a persona chatbot includes the steps of receiving text content, generating a prompt template for a first persona associated with the text content based on the text content, receiving user input, and inputting the user input and the prompt template into a language model to generate a response of the first persona to the user input.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present disclosure relates to a method and system for providing a persona chatbot, and specifically to a method and system for providing a persona chatbot that generates a response of a persona associated with text content in response to user input. Background Technology

[0002] The chatbot service market is growing rapidly with the recent emergence of Large Language Models (LLMs). As a result, there is increasing demand for entertainment-oriented persona chatbot services that allow users to converse with celebrities or cartoon characters.

[0003] For persona chatbots, it is important to effectively reflect the characteristics of characters familiar to the user. Since this process requires analyzing the character's personality, speech patterns, background, and relationships with other characters using vast amounts of basic data, significant human resources and time are required in the initial stages of implementing a persona chatbot. Consequently, the cost of building a character-specific LLM model directly can be very high. The problem to be solved

[0004] To solve the above problems, various embodiments of the present disclosure provide a method for providing a persona chatbot, a computer program stored on a recording medium, and a device (system). means of solving the problem

[0005] The present disclosure may be implemented in various ways, including a method, an apparatus (system), or a computer program stored on a readable storage medium.

[0006] According to one embodiment of the present disclosure, a persona chatbot providing method is provided, which is performed by at least one processor. The persona chatbot providing method includes the steps of receiving text content, generating a prompt template for a first persona associated with text content based on the text content, receiving user input, and inputting the user input and the prompt template into a language model to generate a response of the first persona to the user input.

[0007] According to one embodiment of the present disclosure, user input is an utterance of a second persona associated with text content, the first persona and the second persona are different from each other, and the response of the first persona is a response generated by considering the relationship between the first persona and the second persona.

[0008] According to one embodiment of the present disclosure, the step of generating a prompt template for a first persona includes the step of dividing text content to generate a plurality of subtext contents and the step of generating a prompt template based on the plurality of subtext contents.

[0009] According to one embodiment of the present disclosure, text content is divided into a plurality of subtext contents based on token limitations of a language model.

[0010] According to one embodiment of the present disclosure, text content is divided based on episode information.

[0011] According to one embodiment of the present disclosure, text content is divided into scenes.

[0012] According to one embodiment of the present disclosure, a plurality of subtext contents include a first subtext content and a second subtext content, and the step of generating a prompt template based on the plurality of subtext contents includes the step of generating a first prompt template for a first persona based on the first subtext content and the step of updating the first prompt template based on the second subtext content to generate a second prompt template.

[0013] According to one embodiment of the present disclosure, the step of generating a prompt template based on a plurality of subtext contents further includes the step of generating first summary information associated with a first subtext content, the step of generating second summary information associated with a second subtext content, and the step of storing the first and second summary information in a vector database.

[0014] According to one embodiment of the present disclosure, the step of generating a response of a first persona includes the step of retrieving summary information associated with user input within a vector database and the step of inputting the user input, the summary information associated with user input, and a prompt template into a language model to generate a response of the first persona to the user input.

[0015] According to one embodiment of the present disclosure, a prompt template for a first persona includes information associated with at least one of the first persona's name, gender, age, personality, or relationship information with a character.

[0016] According to one embodiment of the present disclosure, user input is a utterance of a second persona associated with text content, the first persona and the second persona are different from each other, and the response of the first persona to the user input includes the utterance of the first persona and a fingerprint associated with the first persona or the second persona.

[0017] According to one embodiment of the present disclosure, a computer program stored on a computer-readable recording medium is provided for executing a persona chatbot providing method on a computer.

[0018] According to one embodiment of the present disclosure, the system comprises a communication module, a memory, and at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, and the at least one program comprises instructions for receiving text content, generating a prompt template for a first persona associated with the text content based on the text content, receiving user input, and inputting the user input and the prompt template into a language model to generate a response of the first persona to the user input. Effects of the invention

[0019] According to various embodiments of the present disclosure, a persona chatbot can provide a response that is faithful to the relationships between characters within text content by generating a response that considers not only the gender, age, and personality of the persona, but also the relationship with the characters. In addition, by generating and providing fingerprint information to the user in addition to utterances, the persona chatbot can provide a more immersive chatbot service.

[0020] According to various embodiments of the present disclosure, prompt templates and content summary information for specific characters within text content can be generated even when text content is extensive, despite the token limitations of the language model. Additionally, since the inference accuracy of the language model decreases when input data becomes long, the accuracy of the generated results can be improved by dividing the text content into multiple sub-text contents to generate prompt templates and content summary information.

[0021] According to various embodiments of the present disclosure, since a prompt template for a specific persona includes relationship information with a character, it is possible to generate different attitudes / responses depending on which character the specific persona converses with. Accordingly, a more realistic persona chatbot service can be provided. In addition, by generating and storing summary information for each session and using it when providing the persona chatbot service, contextually appropriate text and responses to user utterances can be provided.

[0022] According to various embodiments of the present disclosure, a persona prompt including at least one of a name, age, gender, personality, and relationship with a character is generated in advance for each character appearing in text content and utilized to generate a persona response, thereby providing a persona chatbot service to a user that faithfully reflects the tendencies of the character within the text content.

[0023] According to various embodiments of the present disclosure, since the user can select a character appearing in the text content as their own character as well as a character to converse with, that is, a response can be provided that takes into account the relationship between the first persona and the second persona. Accordingly, since it is possible to go beyond simple chatting and have an experience where the user feels as if they themselves are a character in the text content, a persona chatbot service can be provided that allows for more active immersion in the text content.

[0024] According to various embodiments of the present disclosure, summary information summarizing the plot of text content (or summary information for each sub-text content) is generated in advance and stored in a vector database, thereby allowing the retrieval of summary information associated with user dialogue. Additionally, by generating a persona response using the retrieved summary information, a persona response that also takes into account related events within the text content can be provided.

[0025] According to various embodiments of the present disclosure, not only the utterance of the first persona associated with user input but also the response of the first persona, including contextually relevant text, background photos, sounds, flashes, etc., can be provided. Additionally, the user can select multiple personas as conversation partner characters. In this case, not only one-on-one chatting but also group chatting becomes possible, thereby enabling the realization of a scene where multiple characters appearing in text content communicate. Accordingly, a persona chatbot service with greater realism, depth, and immersion can be provided.

[0026] The effects of the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by a person skilled in the art to which the present disclosure pertains (referred to as "person skilled in the art") from the description in the claims. Brief explanation of the drawing

[0027] FIG. 1 is a drawing showing an example in which a persona chatbot is provided according to one embodiment of the present disclosure. FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to communicate with a plurality of user terminals in order to provide a persona chatbot service according to one embodiment of the present disclosure. FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure. FIG. 4 is an example of a method for generating a prompt template and content summary information for a specific persona from text content according to one embodiment of the present disclosure. FIG. 5 is an example in which, according to one embodiment of the present disclosure, a prompt template for a specific persona is generated / updated from a plurality of subtext contents and summary information is generated. FIG. 6 is an example of a method in which a processor generates a persona's response to a user's input according to one embodiment of the present disclosure. FIG. 7 is an example in which, according to one embodiment of the present disclosure, a processor generates a response of the first persona by considering the relationship between the first persona and the second persona. FIG. 8 is an example of a persona chatbot service according to one embodiment of the present disclosure. FIG. 9 is a flowchart illustrating an example of a method for providing a persona chatbot service to a user according to one embodiment of the present disclosure. Specific details for implementing the invention

[0028] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions regarding well-known functions or configurations will be omitted if there is a risk that the gist of the present disclosure may be unnecessarily obscured.

[0029] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Additionally, in the description of the following embodiments, the description of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.

[0030] The advantages and features of the disclosed embodiments and the methods for achieving them will become clear by referring to the embodiments described below in conjunction with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various different forms, and the embodiments provided are merely to make the present disclosure complete and to fully inform those skilled in the art of the scope of the invention.

[0031] The terms used in this specification will be briefly explained, and the disclosed embodiments will be described in detail. The terms used in this specification have been selected to be as generally used as possible, taking into account their functions in this disclosure; however, these terms may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should be defined not merely by their names, but based on their meanings and the content throughout this disclosure.

[0032] In this specification, singular expressions include plural expressions unless the context clearly specifies them as singular. Additionally, plural expressions include singular expressions unless the context clearly specifies them as plural. Throughout the specification, when a part is described as including a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0033] Additionally, the terms 'module' or 'part' as used in the specification refer to software or hardware components, and the 'module' or 'part' performs certain roles. However, the meaning of 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside in an addressable storage medium or configured to run on one or more processors. Thus, as an example, the 'module' or 'part' may include components such as software components, object-oriented software components, class components, and task components, and at least one of processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, or variables. The components and the functions provided within the 'module' or 'part' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0034] According to one embodiment of the present disclosure, a ‘module’ or ‘part’ may be implemented as a processor and memory. The term ‘processor’ should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, the term ‘processor’ may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. The term ‘processor’ may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a DSP core, or any other combination of such configurations. Additionally, the term ‘memory’ should be broadly interpreted to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as Random Access Memory (RAM), Read-Only Memory (ROM), Non-Volatile Random Access Memory (NVRAM), Programmable Read-Only Memory (PROM), Erasable-Programmable Read-Only Memory (EPROM), Electrically Erasable PROM (EEPROM), Flash Memory, Magnetic or Optical Data Storage Devices, Registers, etc. If a processor can read information from memory and / or write information to memory, the memory is said to be in an electronic communication state with the processor. Memory integrated into a processor is in an electronic communication state with the processor.

[0035] In the present disclosure, "text content" may refer to stories expressed in text, such as light novels, web sophies, and online novels. The genres of text content may include, but are not limited to, romance, thriller, action, detective, and adult content. Furthermore, the language in which the text content is written is not limited to Korean but may include English, Japanese, Chinese, German, French, etc. Additionally, the size or length of the text content is not limited to a specific range. The text content may include not only character speech but also fingerprint information.

[0036] In the present disclosure, 'persona' may refer to a character that reflects the name, gender, age, personality, region, linguistic characteristics, and relationship with other characters of a person or character in text content, and possesses characteristics that distinguish it from other characters or people by manifesting unique linguistic features through the use of specific vocabulary, interjections, tone of voice, etc. In the present disclosure, a persona may be referred to as a person or character.

[0037] In the present disclosure, "chatbot" may refer to software that provides information associated with a specific service or provides a response (e.g., including utterances and / or fingerprints) to a user's utterance, which is conversed in natural language using a computer program or AI (Artificial Intelligence) technology.

[0038] In the present disclosure, "utterance" may refer to a linguistic act of speaking aloud or a written description of said linguistic act (e.g., text). In addition, it may also refer to a written description of the thoughts of a character or person appearing in text content.

[0039] In the present disclosure, a "super-large language model" may be a language model having more than 10 times as many parameters as a conventional general language model (e.g., more than 100 billion parameters). Examples may include GPT-4 developed by OpenAI and BERT developed by Google. In the present disclosure, a super-large language model may be referred to as a language model.

[0040] In the present disclosure, 'scene' may refer to a dialogue scene in which at least one character appearing in the text content participates and one or more utterances are composed.

[0041] FIG. 1 is a diagram illustrating an example in which a persona chatbot is provided according to one embodiment of the present disclosure. As illustrated, a user can converse with a character (e.g., a persona chatbot of a character) of specific text content (e.g., a novel titled "Break Down More, To Me") through a conversation screen (100). Episode information (e.g., title) (110) of the text content may be displayed at the top of the conversation screen (100).

[0042] In one embodiment, the user can select a person or character within the text content to converse with. For example, the user can select "Yoon Si-hoo" from "Break Down More, To Me" as a conversation partner. Additionally, the user can choose to take on the role of a person or character within the text content. For example, the user can choose to take on the role of "Cha Ju-ha" from "Break Down More, To Me". In this case, the "Yoon Si-hoo" persona chatbot can recognize the utterance entered by the user as the utterance of "Cha Ju-ha" and generate the utterance and / or text of "Yoon Si-hoo".

[0043] In one embodiment, the user may input a user utterance "(Yoon Si-hoo..?)" into an input field on the conversation screen (100). In this case, the character name "Cha Ju-ha" selected by the user and the "(Yoon Si-hoo..?)" entered by the user may be displayed as a message on the conversation screen (100) as user input (120). Afterward, the persona chatbot may generate a response (130, 140) of "Yoon Si-hoo" to the user input (120) based on the user input (120) and a pre-generated prompt template for "Yoon Si-hoo". For example, the user input (120) and a pre-generated prompt template for "Yoon Si-hoo" may be input into a language model to generate a response (130, 140) of "Yoon Si-hoo" to the user input (120). The response (130, 140) of "Yoon Si-hoo" generated by the persona chatbot can be displayed as a message on the conversation screen (100).

[0044] In one embodiment, the response generated by the persona chatbot may include fingerprints and / or persona utterances. For example, the response (130, 140) of "Yoon Si-hoo" may include a fingerprint (130) stating "Yoon Si-hoo answered Ju-ha with a frown" and a persona utterance (140) of "Yoon Si-hoo" stating "Uh." The response generated by the persona chatbot may be a response generated by considering "Yoon Si-hoo's" gender, age, personality, and relationship with characters, etc., included in a pre-generated prompt template. By this configuration / method, the persona chatbot can provide a response that is faithful to the character relationships within the text content by generating a response that considers not only the persona's gender, age, and personality but also the relationship with characters. Additionally, by generating and providing fingerprint information to the user in addition to utterances, the persona chatbot can provide a more immersive chatbot service.

[0045] FIG. 2 is a schematic diagram showing a configuration in which an information processing system (230) is connected to communicate with a plurality of user terminals (210_1, 210_2, 210_3) to provide a persona chatbot service according to one embodiment of the present disclosure. The information processing system (230) may include a system capable of providing a persona chatbot service. In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer-executable programs (e.g., downloadable applications) and data related to the persona chatbot service, or one or more distributed computing devices and / or distributed databases based on cloud computing services. For example, the information processing system (230) may include separate systems (e.g., servers) for the persona chatbot service. The persona chatbot service, etc. provided by the information processing system (230) can be provided to the user through an instant messaging application, artificial intelligence-based communication software, a web browser, etc. installed on each of the plurality of user terminals (210_1, 210_2, 210_3). In one embodiment, the information processing system (230) can provide the persona chatbot service to the user terminal using a super-large language model (240).

[0046] Multiple user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) through a network (220). The network (220) can be configured to enable communication between the multiple user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) may be configured as a wired network such as Ethernet, Power Line Communication, telephone line communication device and RS-serial communication, a mobile communication network, a Wireless LAN (WLAN), Wi-Fi, Bluetooth and ZigBee, or a combination thereof. The communication method is not limited and may include not only communication methods utilizing communication networks that the network (220) may include (e.g., mobile communication network, wired internet, wireless internet, broadcasting network, satellite network, etc.) but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).

[0047] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are illustrated as examples of user terminals, but are not limited thereto, and the user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication. For example, user terminals may include smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (Virtual Reality) devices, AR (Augmented Reality) devices, etc. Additionally, FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with an information processing system (230) through a network (220), but is not limited thereto, and may be configured so that a different number of user terminals communicate with an information processing system (230) through a network (220).

[0048] In one embodiment, the information processing system (230) may receive user utterances from user terminals (210_1, 210_2, 210_3). Although the super-large language model (240) is depicted as existing outside the information processing system (230) in FIG. 2, it is not limited thereto, and the super-large language model (240) may be stored and used inside the information processing system (230). Additionally, although FIG. 2 depicts the information processing system (230) generating a character utterance and providing it to the user terminal after receiving user utterances from the user terminal, it is not limited thereto, and hardware / software for providing a persona chatbot service may be provided in the user terminal.

[0049] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure. The user terminal (210) may refer to any computing device capable of running an instant messaging application, artificial intelligence-based communication software, a web browser, etc., and capable of wired / wireless communication, and may include, for example, the mobile phone terminal (210_1), tablet terminal (210_2), PC terminal (210_3) of FIG. 2. As illustrated, the user terminal (210) may include a memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include a memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data through the network (220) using their respective communication modules (316, 336). Additionally, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) through the input / output interface (318).

[0050] The memory (312, 332) may include any non-transient computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as random access memory (RAM), read-only memory (ROM), a disk drive, a solid state drive (SSD), or flash memory. As another example, a permanent mass storage device such as ROM, an SSD, flash memory, or a disk drive may be included in the user terminal (210) or information processing system (230) as a separate permanent storage device distinct from the memory. Additionally, an operating system and at least one program code may be stored in the memory (312, 332).

[0051] These software components may be loaded from a computer-readable recording medium separate from memory (312, 332). This separate computer-readable recording medium may include a recording medium that can be directly connected to the user terminal (210) and the information processing system (230), for example, a computer-readable recording medium such as a floppy drive, disk, tape, DVD / CD-ROM drive, or memory card. As another example, the software components may be loaded into memory (312, 332) via a communication module (316, 336) rather than a computer-readable recording medium. For example, at least one program may be loaded into memory (312, 332) based on a computer program installed by files provided through a network (220) by developers or a file distribution system that distributes installation files for the application.

[0052] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a recording device such as memory (312, 332).

[0053] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system). For example, a request or data generated by the processor (314) of the user terminal (210) according to program code stored in a recording device such as memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, a control signal or command provided under the control of the processor (334) of the information processing system (230) may be received by the user terminal (210) via the communication module (316) of the user terminal (210) through the communication module (336) and the network (220).

[0054] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera including an audio sensor and / or an image sensor, a keyboard, a microphone, or a mouse, and the output device may include a device such as a display, a speaker, or a haptic feedback device. As another example, the input / output interface (318) may be a means for interfacing with a device in which the configuration or function for performing input and output is integrated into one, such as a touchscreen. Although the input / output device (320) is depicted in FIG. 3 as not being included in the user terminal (210), it is not limited thereto and may be configured as a single device with the user terminal (210). Additionally, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (230) or that the information processing system (230) may include. In FIG. 3, the input / output interface (318, 338) is shown as an element configured separately from the processor (314, 334), but is not limited thereto, and the input / output interface (318, 338) may be configured to be included in the processor (314, 334).

[0055] The user terminal (210) and the information processing system (230) may include more components than those of FIG. 3. However, it is not necessary to clearly illustrate most of the prior art components. In one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. Additionally, the user terminal (210) may further include other components such as a transceiver, a GPS (Global Positioning System) module, a camera, various sensors, a database, etc. For example, if the user terminal (210) is a smartphone, it may include components that are generally included in a smartphone, and may be implemented to include various components such as an accelerometer, a gyroscope, a microphone module, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration.

[0056] FIG. 4 is an example of a method (400) for generating a prompt template and content summary information for a specific persona from text content according to one embodiment of the present disclosure. The method (400) may be performed by a processor (e.g., at least one processor of a user terminal or an information processing system). As illustrated, the method (400) may be disclosed by receiving specific text content (S410).

[0057] According to one embodiment, the processor can generate a plurality of subtext contents by dividing the received text content according to predetermined criteria (S420). For example, the predetermined criteria for dividing the text content may be the token size of the language model (240), the number of times the text content is written, the scene, etc. Additionally, multiple criteria may be used in combination to divide the text content. Also, the size / length of the subtext contents may all be within the allowed token range of the language model (e.g., 100,000 tokens or less).

[0058] According to one embodiment, the processor can generate a prompt template for a specific persona (or character) based on a divided subtext content (S430). Additionally, the processor can generate content summary information based on the subtext content and store it in a vector database (S440). In one example, the content summary information may be encoded by an encoder and stored in the vector database in a vector form. The processor can generate the prompt template and / or content summary information using a language model. Then, the processor can determine whether the subtext content corresponds to the last subtext content (S450).

[0059] If it is determined that the corresponding subtext content does not correspond to the last subtext content (NO in S450), the processor (234) can generate / update a prompt template for a specific persona based on the next subtext content (S430), generate content summary information and save it to a vector database (S440), and repeat the process of determining whether it is the last subtext content (S450).

[0060] On the other hand, if the corresponding subtext content is determined to be the last subtext content (YES in S450), the method (400) may be terminated. That is, if the step of creating / updating a prompt template of a specific persona (or character) for all of the divided multiple subtext contents (S430) and the step of creating content summary information and saving to a vector database (S440) are completed, the method (400) may be terminated.

[0061] In FIG. 4, text content is shown as being divided into multiple sub-text contents, but is not limited thereto. For example, without dividing the received text content, a prompt template for a specific character or content summary information can be generated.

[0062] With this configuration, prompt templates and content summary information for specific characters within the text content can be generated, even when the text content is extensive and despite the token limitations of the language model. Furthermore, since the inference accuracy of the language model decreases when the input data becomes long, the accuracy of the generated output can be improved by dividing the text content into multiple sub-text contents to generate prompt templates and content summary information.

[0063] FIG. 5 is an example in which, according to one embodiment of the present disclosure, a prompt template (522, 542) for a specific persona is generated / updated from a plurality of subtext contents and summary information (524, 544) is generated. FIG. 5 is illustrated as dividing text content into a plurality of subtext contents, then generating / updating a prompt template and generating summary information, but is not limited thereto.

[0064] According to one embodiment, a plurality of subtext contents may include a first subtext content (514), a second subtext content (534), etc. Here, the first subtext content (514) and the second subtext content (534) may correspond to different subtext contents among the plurality of subtext contents. For example, if the text content is divided by episode, the first subtext content (514) may correspond to the first episode subtext content and the second subtext content (534) may correspond to the second episode subtext content, but is not limited thereto.

[0065] According to one embodiment, the processor may receive a prompt template (512) and a first-round subtext content (or first-round content, 514) for a specific persona "Kwak Ki-hoon" as a first-round input (510). As illustrated, the prompt template (512) may include an instruction such as "Update the information about the character 'Kwak Ki-hoon' according to the form below based on the given round content." Additionally, since the prompt template (512) within the first-round input (510) is the initial prompt template, the remaining information (e.g., gender, age, personality, relationship with characters, round summary, etc.) excluding the target persona "Name: Kwak Ki-hoon" may be left blank.

[0066] According to one embodiment, the processor may encode a first input (510) and input the encoded input into a language model. The language model may generate a result (520) based on the encoded input. As illustrated, the result (520) generated based on the first subtext content (514) may include a prompt template (522) and summary information (524) for a specific persona. For example, the prompt template (522) for a specific persona may include the gender (e.g., male), age (e.g., information not provided), personality (e.g., a character who appears outwardly smiling and gentle but is actually cold and good at acting. Cold attitude toward Juha~~~), relationship with a character (e.g., cousin of Cha Juha, who dislikes going to school~~~), etc. Summary information (524) may include information summarizing the content of an episode, such as "Juha, who returned home late after playing with friends after school...". Afterward, the processor may store the summary information (524) in a vector database (590). Here, the vector database (590) may be included in an internal and / or external storage device of the information processing system.

[0067] After the result (520) for the first input (510) is generated, the processor may receive the prompt template (522) and second subtext content (or second content, 534) for the previously generated specific persona "Kwak Ki-hoon" as the second input (530). As illustrated, the prompt template (512) may include instructions such as "Update the information about the character 'Kwak Ki-hoon' according to the format below based on the given episode content." Additionally, since the prompt template (532) within the second input (530) is the previously generated prompt template, in addition to the target persona "Name: Kwak Ki-hoon," the gender, age, personality, relationship with the character, etc., may be listed. However, the episode summary section may be left blank to generate summary information for each subtext content.

[0068] According to one embodiment, the processor may encode a second input (530) and input the encoded input into a language model. The language model may generate / update a result (540) based on the encoded input. As illustrated, the result (540) generated based on the second subtext content (534) may include an updated prompt template (542) and summary information (544) for a specific persona. For example, an updated prompt template (542) for a specific persona may include the gender (e.g., male), age (e.g., information not provided), personality (e.g., outwardly smiling and displaying a gentle attitude to deceive people; in reality, having a cool and cold personality, ~~~), and relationship with characters (e.g., known as a famous actor at school and very popular; appears to know Im Seoyoon, meeting and talking with her at school; in the group with Juha, people are unaware of Kwak Kihun's two faces, ~~~). Summary information (544) may include information summarizing the content of the episode, such as "Kwak Kihun is famous and popular as an actor, and went to the rooftop to see Juha~~~". Afterward, the processor may store the summary information (544) in a vector database (590).

[0069] Subsequently, the processor can update the prompt template for a specific persona for the third input, fourth input, etc., in the same or similar manner as described above, and generate summary information for each input. With this configuration, since the prompt template for a specific persona includes relationship information with the characters, it is possible to generate different attitudes / responses depending on which character the specific persona is conversing with. Accordingly, a more realistic persona chatbot service can be provided. Additionally, by generating and storing summary information for each input and using it when providing the persona chatbot service, text and responses appropriate to the context of the user's utterance can be provided.

[0070] FIG. 6 is an example of a method (600) in which a processor generates a response of a persona to a user's input according to one embodiment of the present disclosure. The method (600) may be performed by a processor (e.g., at least one processor of a user terminal or an information processing system). The method (600) may be initiated by the user selecting a conversational partner character (or a first persona) among the characters appearing in text content (S610). In this case, the processor may receive information about the conversational partner character (or first persona) selected by the user. In this case, the processor may retrieve a prompt template associated with a previously generated first persona.

[0071] Next, the user can input user dialogue (S620). Additionally, the user can select a user character (or a second persona) from among the characters appearing in the text content. In this case, the user dialogue may be the dialogue of the second persona. Here, the first persona and the second persona may be different from each other. The processor can input information about the second persona selected by the user and user dialogue.

[0072] According to one embodiment, the processor can retrieve associated summary information within a vector database based on at least one of information about a first persona selected by the user, information about a second persona, or user dialogue entered by the user (S630). In one example, at least one of information about a first persona selected by the user, information about a second persona, or user dialogue entered by the user is encoded by an encoder and converted into a vector, and associated summary information can be retrieved through vector search processing for the vector database.

[0073] Next, the processor inputs the input user dialogue (including information about the second persona), summary information retrieved from a vector database, and the prompt template of the first persona into a language model to generate a response of the first persona to the input of the second persona (user) entered by the user (S640). Here, the response of the first persona may include a passage and / or utterance in which the first persona converses with the second persona.

[0074] FIG. 7 is an example in which a processor generates a response of a first persona by considering the relationship between a first persona and a second persona according to an embodiment of the present disclosure. FIG. 7 specifically specifies a prompt template of the first persona ("Please perform the role of the character below. For the answer, please generate dialogue along with the text. {Generated Persona Prompt}") (710) and a directive ("Please refer to the following content for answer generation.") among the results of a vector database query (730), but is not limited thereto. Additionally, although the prompt template of the first persona (710) is shown as including ({Generated Persona Prompt}), this is for convenience of explanation, and in reality, the full text of the prompt template generated according to the method described in FIG. 4 and FIG. 5 may be included.

[0075] In one example, the user may select "Kwak Ki-hoon" as the first persona, which is the character to converse with, from among the characters appearing in specific text content, and select "Cha Ju-ha" as the second persona, which is the user's own character. In this case, the prompt template (710) may include a pre-generated prompt template for "Kwak Ki-hoon." Additionally, the user may input a user input (720) which is the utterance of the second persona. For example, the user input (720) may include information about the second persona (Cha Ju-ha) and utterance information ("Why did you say that to me on the rooftop?").

[0076] According to one embodiment, the processor may retrieve and retrieve summary information (or vector database query result (730)) associated therewith within a vector database based on at least some of the first persona (Kwak Ki-hoon) selected by the user and user input (Cha Ju-ha: "Why did you say that to me on the rooftop?"). For example, summary information associated with the first persona (Kwak Ki-hoon) selected by the user and / or user input (Cha Ju-ha: "Why did you say that to me on the rooftop?") may be retrieved as "Kwak Ki-hoon is famous and popular as an actor and appears on the rooftop to see Ju-ha. He reveals his true nature to Ju-ha, throws threatening remarks, and displays dangerous-looking behavior by suggesting to Ju-ha that they sleep together all night."

[0077] According to one embodiment, the processor inputs a pre-generated prompt template (710) of the first persona, user input (720), and summary information associated with the user input (vector database query result (730)) into a language model to generate a response of the first persona as an inferred result (740) considering the relationship between the first persona and the second persona. Here, the response of the first persona may include not only the utterance of the first persona but also fingerprints associated with the first persona and / or the second persona. For example, the processor may generate a response (or answer) of the first persona that includes the utterance (or dialogue) of the first persona (Kwak Ki-hoon: "I want to play with you all night long haha. Why? Are you dissatisfied?") along with fingerprints (Kwak Ki-hoon smiled artificially, looked at Cha Ju-ha, and spoke.).

[0078] Through this configuration, a persona prompt (710) containing at least one of name, age, gender, personality, and relationship with characters is generated in advance for each character appearing in the text content and utilized to generate a persona response, thereby providing a persona chatbot service to the user that faithfully reflects the character's tendencies within the text content.

[0079] In addition, since users can select a character within the text content as their own character as well as a character to converse with, they can provide a response that takes into account the relationship between the first persona and the second persona. Accordingly, since it is possible to go beyond simple chatting and have an experience where the user feels as if they themselves have become a character in the text content, it is possible to provide a persona chatbot service that allows for more active immersion in the text content.

[0080] Additionally, by generating summary information that summarizes the plot of the text content (or summary information for each sub-text content) in advance and storing it in a vector database, the summary information associated with the user input (720) can be retrieved. Furthermore, by generating a persona response using the retrieved summary information, a persona response that also takes into account related events within the text content can be provided.

[0081] FIG. 8 is an example of a persona chatbot service according to one embodiment of the present disclosure. According to one embodiment, a user may select a first persona and / or a second persona. For example, the first persona may be "Jeonghyeon" as a character with which the user converses. Additionally, the second persona may be "Seonah" as the user's own character.

[0082] The first operation (810) may represent an example in which a fingerprint (812) and a background photo (814) are output on a display. For example, the fingerprint (812) and the background photo (814) may be generated based on user input and / or a second persona response. In another example, the fingerprint (812) and the background photo (814) may be for any event (or event selected by the user) within the text content and may be output on the display. Here, the fingerprint (812) may include information such as the location, time, surrounding buildings, and surrounding circumstances of the event / scene associated with the text content. For example, the fingerprint (812) is "[Near a certain new town] [PM 11:55]" A hideous abandoned building standing alone in a deserted place. 'Woosung Obstetrics and Gynecology' It may include ". In this case, the user can initiate a conversation by referring to the fingerprint (812) and background photo (814). In one embodiment, in addition to the fingerprint and background photo, background music, sound effects, etc. may be provided.

[0083] In one embodiment, at least one of a prompt template, user input (816), and summary information for a first persona ("Jung-hyun") selected by a user is input into a language model to generate a first persona response including a first persona utterance (818). For example, if the user inputs "Haha, we must be really crazy to come here" as a line for Seona, the utterance "Seeing it at night" of "Jung-hyun" may be output.

[0084] The second operation (820) represents an example in which a fingerprint (822) and a background photo (824) are generated and displayed on a display based on user input (816) and a first persona utterance (818). For example, the fingerprint may output "[1st floor lobby] Friends, turn on the mobile phone flash. Jeonghyeon, wandering around and taking pictures, is very excited." Additionally, a background photo (824) associated with each persona's conversation, such as "a photo of a camera flash turned on in a dark space," may be displayed on the display. In one embodiment, the flash of the user terminal may be controlled together.

[0085] The third action (830) illustrates an example where the user converses with two or more personas. For example, the user may select “Hyejeong, Jeonghyeon, Taeseong” as the multiple first personas who are the conversation partners, and “Seona” as the second persona who is their own character. In this case, the prompt templates for the three pre-generated multiple first personas “Hyejeong, Jeonghyeon, Taeseong” may be used.

[0086] As described above, when a user inputs a utterance (832) of a second persona, a response from each of the first personas to the utterance (832) of the second persona can be generated / provided using a prompt template for each of the first personas. For example, if the user inputs "Well, nothing special" as the utterance of the second persona "Seon-a," "That's true" can be provided as the response (834) of "Hye-jeong."

[0087] In one embodiment, not only a conversation between a second persona, which is the user's own character, and one of a plurality of first personas, which is the conversation partner character, but also a conversation between the plurality of first personas may be generated / provided. For example, based on the utterance (832) "Well, nothing special" by the second persona "Sun-ah" and the response (834) "Yeah" by the first persona "Hye-jeong," a response (836) by another first persona "Jeong-hyeon" "For some reason, now that I've come in here, I feel like my head hurts, don't you, Tae-seong?" may be provided. Additionally, a response (838) "Huh? I don't know" by another first persona "Tae-seong" may be provided.

[0088] Through this configuration, it is possible to provide not only the utterance of the first persona associated with user input but also the first persona's response, which includes contextually relevant text, background images, sounds, flashes, etc. Additionally, the user can select multiple personas as conversation partner characters. In this case, group chatting becomes possible in addition to simple one-on-one chatting, enabling the realization of scenes where multiple characters appearing in text content communicate. Accordingly, a persona chatbot service with greater realism, depth, and immersion can be provided.

[0089] FIG. 9 is a flowchart illustrating an example of a method (900) for providing a persona chatbot service to a user according to one embodiment of the present disclosure. In one embodiment, the persona chatbot providing method (900) may be performed by a processor (e.g., at least one processor of a user terminal or an information processing system). In another embodiment, the information processing system and the user terminal may divide and perform the steps of the persona chatbot providing method (900).

[0090] In one embodiment, the persona chatbot provision method (900) may be initiated by receiving text content (S910). Then, the processor may generate a prompt template for a first persona associated with the text content based on the text content (S920). In one embodiment, the prompt template for the first persona may include information associated with at least one of the first persona's name, gender, age, personality, or relationship information with a character.

[0091] After generating a prompt template for the first persona, the processor can receive user input (S930). Then, the processor can input the user input and the prompt template into a language model to generate a response of the first persona to the user input (S940). Here, the user input may be an utterance of the second persona associated with text content. The first persona and the second persona may be different from each other.

[0092] In one embodiment, the response of the first persona may be a response generated by considering the relationship between the first persona and the second persona. In another embodiment, the response of the first persona to user input may include the utterance of the first persona and fingerprints associated with the first persona or the second persona.

[0093] According to one embodiment, a prompt template for a first persona can be generated by dividing text content. For example, text content can be divided into multiple sub-text contents based on token limitations of a language model. In another example, text content can be divided based on episode information. In yet another example, text content can be divided into scenes.

[0094] Specifically, the processor can divide text content to generate multiple subtext contents. Then, the processor can generate a prompt template based on the multiple subtext contents.

[0095] According to one embodiment, the prompt template may be updated after the initial template is generated. For example, a plurality of subtext contents may include a first subtext content and a second subtext content. In this case, the processor may generate a first prompt template for a first persona based on the first subtext content. Then, the processor may generate a second prompt template by updating the first prompt template based on the second subtext content.

[0096] Additionally, the response of the first persona can be generated based on summary information of the subtext content. Specifically, the processor can generate summary information of the subtext content and store it in a vector database, and generate the response of the first persona using user input and summary information of the subtext content.

[0097] For example, the processor can generate first summary information associated with first subtext content. Additionally, the processor can generate second summary information associated with second subtext content. The processor can store the first and second summary information in a vector database. Subsequently, the processor can retrieve the summary information associated with user input within the vector database. The processor can input the user input, the summary information associated with the user input, and the prompt template into a language model to generate a response of the first persona to the user input.

[0098] The method described above may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may continuously store a program executable by a computer, or temporarily store it for execution or download. Additionally, the medium may be various recording or storage means in the form of a single or multiple hardware components, and may not be limited to a medium directly connected to a computer system but may exist distributed over a network. Examples of media may include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and media configured to store program instructions, including ROM, RAM, and flash memory. Furthermore, other examples of media may include recording or storage media managed by app stores that distribute applications or sites and servers that supply or distribute various other software.

[0099] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will understand that the various exemplary logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein may be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate such interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been generally described above in terms of their functional aspects. Whether such functions are implemented in hardware or in software depends on the design requirements imposed on the specific application and the overall system. Those skilled in the art may implement the functions described in various ways for each specific application, but such implementations should not be construed as departing from the scope of the present disclosure.

[0100] In a hardware implementation, the processing units used to perform the techniques may be implemented in one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described in this disclosure, computers, or a combination thereof.

[0101] Accordingly, the various exemplary logic blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of general-purpose processors, DSPs, ASICs, FPGAs or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, for example, a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors coupled with a DSP core, or any other combination of configurations.

[0102] In firmware and / or software implementations, techniques may be implemented as instructions stored on a computer-readable medium such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, compact disc (CD), magnetic or optical data storage devices, etc. The instructions may be executable by one or more processors, and may cause the processor(s) to perform specific aspects of the functions described in this disclosure.

[0103] Where implemented in software, the techniques may be stored on a computer-readable medium as one or more instructions or code, or transmitted through a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transmission of a computer program from one place to another. Storage media may be any available medium accessible by a computer. As a non-limiting example, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium accessible by a computer that can be used to transfer or store desired program code in the form of instructions or data structures. Additionally, any connection is appropriately referred to as a computer-readable medium.

[0104] For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, coaxial cable, fiber optic cable, twisted pair cable, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of a medium. As used herein, disk and disc include CD, laser disc, optical disc, DVD (digital versatile disc), floppy disk, and Blu-ray disc, wherein disks usually play data magnetically, whereas discs play data optically using a laser. The above combinations should also be included within the scope of computer-readable media.

[0105] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other known form of storage medium. An exemplary storage medium may be connected to a processor so that the processor can read information from the storage medium or write information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and the storage medium may exist within an ASIC. The ASIC may exist within a user terminal. Alternatively, the processor and the storage medium may exist as separate components within the user terminal.

[0106] Although the embodiments described above have been described as utilizing aspects of the subject matter disclosed herein in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or a distributed computing environment. Furthermore, aspects of the subject matter in the present disclosure may be implemented in a plurality of processing chips or devices, and storage may be similarly affected across a plurality of devices. Such devices may include PCs, network servers, and portable devices.

[0107] Although the present disclosure has been described in relation to some embodiments, various modifications and changes may be made without departing from the scope of the present disclosure as understood by a person skilled in the art to which the invention of the present disclosure pertains. Furthermore, such modifications and changes should be considered to fall within the scope of the claims appended to this specification. Explanation of the symbols

[0108] 100: Dialogue screen 110: Episode Information 120: User utterance 130: Fingerprint 140: Persona Chatbot Response

Claims

Claim 1 A method for providing a persona chatbot performed by at least one processor, comprising: receiving text content; generating a prompt template for a first persona associated with the text content based on the text content; receiving user input; and inputting the user input and the prompt template into a language model to generate a response of the first persona to the user input, wherein the step of generating a prompt template for the first persona comprises: dividing the text content to generate a plurality of subtext content; and generating the prompt template based on the plurality of subtext content, wherein the text content is divided into a plurality of subtext content based on the token limit of the language model. Claim 2 A method for providing a persona chatbot according to claim 1, wherein the user input is an utterance of a second persona associated with the text content, the first persona and the second persona are different from each other, and the response of the first persona is a response generated by considering the relationship between the first persona and the second persona. Claim 3 delete Claim 4 delete Claim 5 In claim 1, the method of providing a persona chatbot, wherein the text content is divided based on episode information. Claim 6 A method for providing a persona chatbot according to claim 1, wherein the text content is divided into scene units. Claim 7 A method for providing a persona chatbot according to claim 1, wherein the plurality of subtext contents includes a first subtext content and a second subtext content, and the step of generating the prompt template based on the plurality of subtext contents includes: a step of generating a first prompt template for the first persona based on the first subtext content; and a step of updating the first prompt template based on the second subtext content to generate a second prompt template. Claim 8 A method for providing a persona chatbot according to claim 7, wherein the step of generating the prompt template based on the plurality of subtext contents further comprises: generating first summary information associated with the first subtext content; generating second summary information associated with the second subtext content; and storing the first and second summary information in a vector database. Claim 9 A method for providing a persona chatbot according to claim 8, wherein the step of generating a response of the first persona comprises: a step of searching for summary information associated with the user input within the vector database; and a step of inputting the user input, the summary information associated with the user input, and the prompt template into a language model to generate a response of the first persona to the user input. Claim 10 A method for providing a persona chatbot according to claim 1, wherein the prompt template for the first persona includes information associated with at least one of the name, gender, age, personality, or relationship information with a character of the first persona. Claim 11 A method for providing a persona chatbot according to claim 1, wherein the user input is a utterance of a second persona associated with the text content, the first persona and the second persona are different from each other, and the response of the first persona to the user input includes a utterance of the first persona and a fingerprint associated with the first persona or the second persona. Claim 12 A computer program stored on a computer-readable recording medium for executing a method according to any one of paragraphs 1, 2, 5 through 11 on a computer. Claim 13 A system comprising: a communication module; a memory; and at least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory, wherein the at least one program comprises instructions for receiving text content, generating a prompt template for a first persona associated with the text content based on the text content, receiving user input, inputting the user input and the prompt template into a language model to generate a response of the first persona to the user input, wherein generating a prompt template for the first persona comprises dividing the text content to generate a plurality of subtext content, and generating the prompt template based on the plurality of subtext content, wherein the text content is divided into a plurality of subtext content based on the token limit of the language model.