system
By acquiring and processing dialogue data to generate character-like responses, the system addresses the lack of immersion in conventional chatbots, improving user engagement and interaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-16
- Publication Date
- 2026-04-28
AI Technical Summary
Conventional chatbot systems struggle to provide immersive experiences by generating individualized responses in the tone and language pattern of a specific character, leading to reduced user engagement and inquiry frequency.
The system acquires dialogue data from the internet or databases, preprocesses it to learn a character's tone and language patterns, and uses a generative model to generate character-like responses, followed by post-processing to emphasize these elements, thereby providing an immersive experience.
This approach enhances user engagement and inquiry frequency by allowing users to interact with a specific character, increasing interest and interaction with the company.
Smart Images

Figure 2026070891000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Conventional chatbot systems have a problem that it is difficult to provide an immersive experience for users because they can only generate standard answers. Therefore, in scenes where individualized responses using the tone and language pattern of a specific character are required, they cannot attract users' interest, and the improvement of inquiry frequency and the enhancement of customer engagement have not been fully realized.
Means for Solving the Problems
[0005] This invention provides a means for acquiring dialogue data based on a specified character from the internet or a database, and for learning the character's tone of voice and language patterns by preprocessing the acquired data and using it as training data for a generative model. Furthermore, it proposes a system that generates character-like responses using the generative model in response to user inquiries, and then performs post-processing on these responses to emphasize character-specific elements before responding to the user. Through this means, users can experience an immersive feeling as if they are interacting with a specific character, and companies can increase the number of inquiries and strengthen customer engagement.
[0006] "Users" refers to individuals or organizations that make inquiries through this system.
[0007] A "character" is a fictional person or entity specified by the user, and the focus is on their distinctive speech patterns and language styles.
[0008] "Dialogue data" refers to text data collected from works in which characters appear, and is used as reference material to understand the characters' speech patterns and expressions.
[0009] A "generative model" refers to a program built on machine learning algorithms that learns a character's tone of voice and language patterns, and generates responses in response to inquiries.
[0010] "Preprocessing" refers to a series of processing operations to convert dialogue data into a format suitable for the generation model.
[0011] "Post-processing" refers to additional processing operations performed on the generated response to emphasize character-specific elements.
[0012] A "terminal" refers to an information processing device used by a user to access and receive the generated response. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] The present invention provides a platform for users to obtain responses that mimic the speech patterns of a specific character. The natural language processing functions of this system are described below.
[0035] First, the system is triggered when the user specifies their favorite character. The server receives this information and collects dialogue data from works and databases related to the specified character. The collected data is used as material to understand the character's speech patterns and characteristics. This data is preprocessed and converted into a format optimal for training a generative model.
[0036] The server then uses this preprocessed data to train a generative model, learning the language patterns and emotional expressions of the specified character. Once trained, the model retains the generated patterns and makes them available in response to actual queries.
[0037] Next, when a user makes a request, the device sends the content to the server. The server uses a generative model to generate a response to the user's request that has a character-like tone. At this stage, the response undergoes post-processing to reflect the character's characteristics and enable realistic interaction.
[0038] For example, if a user asks, "What products do you recommend?" and they are a fan of the dragon character, the system will generate a response like, "We have items that are as powerful as a dragon and will support you every day!" This allows the user to have an experience as if they were interacting with that character.
[0039] This system allows companies to personalize inquiry responses and increase user engagement. Because the system generates responses that include character-specific elements, users become more interested and more likely to actively engage in communication with the company.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] Users specify their favorite character via LINE and begin a conversation with a chatbot. The character information entered by the user acts as a trigger to retrieve related data about that character.
[0043] Step 2:
[0044] The server receives character information from users and collects dialogue data for those characters via the internet or an internal database. This collected data serves as material for learning the characters' language styles and characteristics.
[0045] Step 3:
[0046] The server preprocesses the collected dialogue data. This includes removing noise, normalizing the text, and tokenizing it. This preprocessing creates a dataset that allows the generative model to train more effectively.
[0047] Step 4:
[0048] The server uses pre-processed data to train a generative model. During this process, machine learning algorithms are used to teach the model the character's speech patterns and mannerisms. Once the model is trained, it is ready to generate responses from that character.
[0049] Step 5:
[0050] When a user enters their inquiry, the terminal sends that information to the server. This user input forms the basis for generating the response.
[0051] Step 6:
[0052] The server uses a generative model to generate character-like responses to user questions. These responses are designed to reflect the unique language patterns and emotional expressions of the characters.
[0053] Step 7:
[0054] The server performs post-processing on the generated response. This post-processing involves adding specific phrases or sentence endings to emphasize character-specific elements.
[0055] Step 8:
[0056] The server sends the final response to the device, allowing the user to view it on LINE. Seeing the displayed response, the user can experience what it's like to be having a conversation with their favorite character.
[0057] (Example 1)
[0058] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0059] There is a need to reproduce the unique speech patterns and language styles of characters, providing users with a more relatable and engaging communication experience. However, existing technologies have difficulty generating expressions that adequately mimic the characteristics of characters. This hinders increased user satisfaction and the promotion of deeper engagement.
[0060] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0061] In this invention, when a user specifies a preferred character, the server includes means for an information processing device to collect speech data of that character from a communication network or information aggregation device; means for initial processing the collected speech data and optimizing it as training data for a learning algorithm; and means for training the learning algorithm with the training data to enable it to understand the character's expression style and language structure. This makes it possible to provide realistic and personalized communication using the character's unique language expressions.
[0062] A "user" is a person or organization that uses this system with the aim of obtaining a response from a specific character.
[0063] An "information processing device" is a device that processes digital data and has the function of acquiring, processing, and transmitting specific information based on inquiries from users.
[0064] A "communication network" is the infrastructure used for sending and receiving data, and includes the internet and other network technologies.
[0065] An "information aggregation device" is any system or platform for collecting, storing, and utilizing large amounts of data as needed.
[0066] "Speech data" refers to data about a character's lines and linguistic expressions, and is text information that reflects the character's characteristics.
[0067] "Initial processing" refers to processes performed after data collection, such as data formatting, cleaning, and format conversion, which prepare the data for input into the AI model that generates it.
[0068] A "learning algorithm" is a technique that aims to learn patterns and characteristics by processing large amounts of data based on mathematical models, and includes machine learning techniques.
[0069] "Training data" refers to the dataset provided to a learning algorithm, which serves as the material for the algorithm to learn specific patterns and features.
[0070] "Style of expression" refers to the unique language style, tone of voice, and emotional expressions used by a character, and is an element that forms the character's individuality.
[0071] "Language structure" refers to the grammatical and structural characteristics of the language used by a character, encompassing elements related to the flow and structure of words.
[0072] The system of this invention starts operating when the user specifies a particular character. At this time, the user uses a character selection interface using a smartphone or computer application. Once the user selects a character, the specified information is transmitted to the server via the terminal.
[0073] Based on this information, the server collects speech data from databases associated with the specified character via the communication network. This process involves web scraping and API calls to obtain appropriate text data from video content and books in which the character appears.
[0074] The acquired speech data is initially processed by the server. This initial processing uses natural language processing libraries such as NLTK and spaCy to remove unnecessary symbols and normalize the text, followed by tokenization and part-of-speech tagging. This process ensures the data is in a format optimized for the learning algorithm.
[0075] Next, the server trains a generative AI model using the processed data. At this stage, a machine learning framework such as Tensorflow® or PyTorch is used to build a model that learns the character's expressive style and language structure. The model reflects the character's characteristics and enables the generation of responses to user inquiries.
[0076] When a user enters a specific question or request using their device, that information is sent to the server. The server utilizes a trained generative AI model to generate character-like responses to the user's inquiries. For example, if a user enters "What products do you recommend?", the server generates a response that reflects the character's personality, such as "We have items that are as powerful as a dragon and will support your everyday life!" The generated response is then given a finishing touch to emphasize the character's characteristics before being sent back to the device.
[0077] An example of a prompt message is given to the generating AI model as, "Please respond to the question 'What products do you recommend?' in a dragon's voice." This system allows users to have a personalized conversational experience with a selected character, and is expected to improve customer satisfaction and engagement for companies and services.
[0078] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0079] Step 1:
[0080] The user uses their device to specify a particular character. As input, the user selects or enters the character name within the application's interface. This information, reflecting the user's preference, is sent to the server.
[0081] Step 2:
[0082] The server uses the received character information to collect relevant dialogue data from information aggregation devices via the communication network. As input, queries based on character names are executed, and the output is the character's dialogue data. Database access and API calls are performed to aggregate the dialogue data.
[0083] Step 3:
[0084] The server performs initial processing on the collected speech data. The input data is text information in a non-standardized format, and natural language processing libraries (such as NLTK and spaCy) are used to cleanse and normalize the text. The output is tokenized text data.
[0085] Step 4:
[0086] The server trains a generative AI model using an initialized dataset. The input is formatted character data, and tools such as TensorFlow and PyTorch are used. The model learns the character's representational style and language structure, and the output is a trained AI model.
[0087] Step 5:
[0088] When a user uses a terminal to ask and answer questions, the user inputs questions and requests from the terminal. Specifically, the user enters text through the interface, which is then sent to the server. As output, a character-like response based on this input is ready to be generated.
[0089] Step 6:
[0090] The server uses a character-like response model based on the user's question as input. The output is a response text that highlights the character's characteristics. The server then processes this response to emphasize the character's unique elements.
[0091] Step 7:
[0092] The terminal receives the completed response sent back from the server and displays it to the user. The input includes the response data from the server. This text is visualized on the terminal's screen, and a character-like dialogue is established as the final result provided to the user.
[0093] (Application Example 1)
[0094] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0095] Modern commercial platforms require providing personalized experiences in interactions with users to enhance customer engagement. However, typical systems struggle to provide personalized responses based on user interests and preferences, making it challenging to deliver individual user experiences, such as generating responses in the tone of a specific character.
[0096] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0097] In this invention, the server includes means for acquiring speech data of a character specified by the user from an information network or information storage device, means for preprocessing the acquired speech data and structuring it as training data for a generative model, and means for introducing products and information using the generated response. This makes it possible to introduce products and information in the tone of voice of the character set by the user.
[0098] "User" refers to a person who uses a system or application.
[0099] A "character" refers to a fictional person or creature that appears in a specific work or story, possessing its own unique language patterns and emotions.
[0100] An "information network" refers to the entire system for acquiring or transmitting information through networks such as the internet.
[0101] "Information storage device" refers to physical or virtual media used to store information using databases or other storage devices.
[0102] "Word data" refers to information expressed in text format, including data such as lines and dialogues spoken by specific characters.
[0103] "Acquisition" refers to the act of collecting data from external sources.
[0104] "Preprocessing" refers to the process of organizing acquired data into an appropriate format and processing it to a state suitable for analysis and training machine learning models.
[0105] A "generative model" refers to a machine learning algorithm that automatically generates new information or responses based on input data.
[0106] "Training data" refers to the dataset used to create a machine learning model; it is the fundamental data that the model learns from.
[0107] "Introducing products or information" refers to the act of explaining or suggesting specific content or products to users in detail.
[0108] A "generative AI model" refers to a model that uses artificial intelligence technology to generate new text or responses from data.
[0109] The program to implement this system primarily consists of data exchange between the server and the user's terminal, and natural language processing. The server retrieves the corresponding character's speech data from an information network or data storage device based on the character specified by the user. The retrieved data is preprocessed and structured as training data for a generative AI model. This process utilizes programming languages such as Python for cleaning and formatting text data, and natural language processing APIs (e.g., OpenAI's GPT-3).
[0110] The server trains a generative model based on pre-processed data, learning the character's tone of voice and language patterns. Based on inquiries from the user's terminal, the server utilizes the generative AI model to generate character-like responses. These responses are further post-processed to emphasize character-specific elements. The generated responses are sent to the user's terminal, providing a unique purchasing experience by introducing products and information.
[0111] For example, if a user inquires via smartphone, "I want a new laptop," the system will generate a response in the tone of their favorite character saying, "A laptop that opens the door to the future—unleash your creativity with this!" An example of a prompt to the generating AI model would be, "Please give me a recommendation about [topic] in the tone of the character's name."
[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0113] Step 1:
[0114] The user selects a specific character on the terminal and enters their inquiry. The input data includes the character name set by the user and the specific inquiry (e.g., "I want a new laptop"). This triggers the system to begin generating a character-like response to the inquiry.
[0115] Step 2:
[0116] The terminal sends the user's selected character information and inquiry details to the server. At this time, the terminal sends the character identifier and inquiry details to the server as JSON data. The input data is received by the server and forms the basis for proceeding to the next step.
[0117] Step 3:
[0118] The server collects the corresponding character's speech data from the information storage device based on the character's identifier. The collected data is text data containing the character's lines and characteristic expressions. This data is then pre-processed.
[0119] Step 4:
[0120] The server preprocesses the collected word data and structures it as training data for a generative AI model. This preprocessing involves text cleaning (removing unnecessary information and converting it to a unified format). The preprocessed data is then in a format suitable for training the model.
[0121] Step 5:
[0122] The server trains a generative AI model using preprocessed data to learn character speech patterns and language styles. The input is structured word data, and the output is the learned model parameters, which are used in the next step.
[0123] Step 6:
[0124] The server inputs the user's inquiry into a generating AI model and generates a character-like response. This prompt is in the form of "Please give me a recommendation about XX in the tone of the character's name," and the input is transformed into a specific response, enabling dynamic dialogue.
[0125] Step 7:
[0126] The generated responses undergo post-processing on the server to emphasize character-specific elements. This post-processing adds specific tones and emotions to the lines, resulting in more realistic character responses.
[0127] Step 8:
[0128] Finally, the server sends a post-processed response back to the terminal. This response is displayed on the user's terminal, and the user receives personalized product and information recommendations from their favorite character. The output data allows the user to have a unique purchasing experience.
[0129] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0130] This invention is an interactive system that allows users to receive responses based on the tone of a specific character, while simultaneously recognizing the user's emotions and customizing the response content to suit those emotions. The operation of this system is described below in natural language.
[0131] First, the system activates when a user specifies their favorite character via LINE and initiates an inquiry. Based on the entered character information, the server collects dialogue data for that character from the internet or an internal database.
[0132] The collected dialogue data is preprocessed to understand the character's tone of voice and linguistic characteristics, and then structured as training data for a generative model. The server uses this data to train the generative model and teach it character-specific language patterns. At this stage, the generative model has the foundation to reproduce the character's tone of voice.
[0133] Next, when the user enters their inquiry, the terminal sends that information to the server. At this point, the system's built-in emotion engine is utilized. The server uses this emotion engine to recognize the emotions from the user's input text.
[0134] Based on the recognized emotion, the server uses a generative model to generate a character-like response corresponding to that emotion. This response reflects pre-trained character speech patterns and emotion-based expressions. The generated response is then post-processed to take emotion information into account, further emphasizing the character's unique emotional expressions.
[0135] For example, if a user types "I'm feeling a bit down today," the system will work to have the designated character respond with positive words to cheer them up. Let's say the generated response is "A fantastic adventure awaits you that will cheer you up!" In this way, users can enjoy more personalized interactions tailored to the character and their emotions.
[0136] This system can strengthen the relationship between companies and users and improve the usefulness of inquiries by simultaneously achieving realistic character reproduction and response generation adapted to the user's emotions.
[0137] The following describes the processing flow.
[0138] Step 1:
[0139] Users specify their favorite character on LINE and initiate an inquiry. This necessitates a conversational experience tailored to the user's preferences.
[0140] Step 2:
[0141] The server receives information about the specified character and collects dialogue data from works and databases in which that character appears. This includes the character's characteristic speech patterns and phrases.
[0142] Step 3:
[0143] The server preprocesses the collected dialogue data, performing tasks such as text normalization and tokenization, and structures it into a data format suitable for training generative models. This processing improves the accuracy of the learning process.
[0144] Step 4:
[0145] The server uses the pre-processed data to train the generative model. The model learns the characters' speech patterns and language styles, and is then ready to go.
[0146] Step 5:
[0147] When a user enters and submits an inquiry, the device transmits that data to the server. At that time, sentiment analysis of the user's input is performed.
[0148] Step 6:
[0149] The emotion engine installed on the server recognizes emotions from the user's input text. For example, basic emotions such as joy, sadness, and anger can be identified.
[0150] Step 7:
[0151] The server considers the recognized emotions and uses a generative model to generate character-like responses that correspond to those emotions. The responses not only reflect the character's tone of voice but also fit the user's emotions.
[0152] Step 8:
[0153] The server then performs post-processing on the generated response to enhance its emotional expression. This involves adding additional character-specific phrases and nuances depending on the specific emotion.
[0154] Step 9:
[0155] The server sends the completed character-like response to the user's device, allowing the user to view it on LINE. This allows users to enjoy a richer experience through dialogue that resonates with their own emotions.
[0156] (Example 2)
[0157] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0158] Conventional dialogue systems have problems accurately reproducing the tone of voice of a character specified by the user, and are unable to generate flexible responses that respond to the user's emotions. This limits the quality of the dialogue provided by the system and hinders the improvement of the user experience. In addition, the lack of processing to effectively emphasize character-specific expressions makes it difficult to provide responses that fully utilize the characteristics of the character.
[0159] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0160] In this invention, the server includes means for acquiring language data of a symbol specified by the user from a network or information recording device, means for preprocessing the acquired language data and structuring it as training information for a generation algorithm, and means for training the generation algorithm with the training information to learn the language features and expression patterns of the symbol. This makes it possible to faithfully reproduce the tone of voice of a character specified by the user and generate appropriate responses that correspond to the user's emotions.
[0161] A "symbol" is an element designated by the user to represent a specific character or its characteristics.
[0162] A "network" is a communication infrastructure used to transmit digital information and connect multiple devices and systems.
[0163] An "information recording device" is a technology or platform for storing data and accessing it as needed.
[0164] "Linguistic data" refers to information about texts and speeches, including the tone of voice and linguistic characteristics of a specific character.
[0165] "Preprocessing" is the process of normalizing acquired data, removing noise, and preparing it for training generative algorithms.
[0166] A "generative algorithm" is a computational method for generating new data or responses based on presented information.
[0167] "Training information" refers to a structured dataset used by a generative algorithm for learning.
[0168] "Linguistic features" refer to the unique expressive styles and grammatical characteristics of the language spoken by a particular character or individual.
[0169] A "pattern of expression" is a characteristic that indicates a specific form or flow in language or communication.
[0170] "Emotion analysis" is a technology that recognizes a user's emotional state from text or audio.
[0171] "Control variables" are parameters that are adjusted to optimize the performance of a generation algorithm.
[0172] The specific implementation methods for this system are described below.
[0173] First, the user uses a network-enabled device to specify their preferred character and access the system via a messaging service such as LINE. Based on the specified character, the server collects language data related to that character. This data is stored on the network or in a database. The collected language data is then structured for processing by the server.
[0174] The server runs software developed using programming languages such as Python and Java (registered trademark) to preprocess the collected language data. This process includes data normalization and noise reduction, preparing the data for use as training data for a generative AI model. Next, the server uses natural language processing libraries to train the generative AI model with this training data. In this way, the model learns the language features and expression patterns of a specified character.
[0175] Subsequently, when a user submits a query to the system, the terminal sends the query to the server. Upon receiving the query, the server uses natural language processing technology to analyze the text sent by the user and recognize its sentiment. Sentiment analysis is performed using sentiment analysis libraries and APIs.
[0176] When the generative AI model generates responses based on the trained data, the server generates character-like responses that take the user's emotions into account, based on the analyzed emotions. Finally, the generated responses undergo further post-processing to enhance character-specific expressions. This process generates emotionally natural responses that are sent to the user's device.
[0177] For example, if a user sends a message such as "I'm a little tired today," the system will generate a response that resonates with the user, while maintaining the character's tone of voice. "For instance, it might respond with something like, 'Please rest well today!' In this way, the character provides the user with an emotionally resonant dialogue."
[0178] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0179] Step 1:
[0180] A user initiates a query by specifying a particular character through the LINE application. The input is the character's name, and this information is sent to the server as output. Based on the received character name, the server collects relevant language data using the network or database. Specifically, this involves issuing API queries to the character or performing database searches.
[0181] Step 2:
[0182] The server preprocesses the collected language data. At this stage, raw language data is provided as input, and preprocessed data is generated as output. Data processing is performed, such as normalization, removal of unnecessary information, and identification of frequently occurring words and phrases. Specifically, text tokenization and morphological analysis are performed to build a training dataset.
[0183] Step 3:
[0184] The server uses preprocessed data to train a generative AI model. It receives preprocessed data as input and generates a trained model as output. Here, the model uses a machine learning algorithm to learn the language patterns of characters. Specifically, the model's parameters are updated using a deep learning framework.
[0185] Step 4:
[0186] The user enters their inquiry into the system and sends it to the server via their terminal. The input includes the inquiry text, and this information is passed to the server as output. Specifically, the text data entered on the user's terminal is encrypted and securely transmitted to the server.
[0187] Step 5:
[0188] The server analyzes the received inquiry and uses an emotion analysis engine to evaluate the user's emotions. The user's inquiry is the input for analysis, and the recognized emotion information is the output. The server extracts emotions such as positive, negative, and neutral from the text. Specifically, an emotion score is calculated based on an analysis of keywords and phrases.
[0189] Step 6:
[0190] The server uses a generative AI model to generate character-like responses corresponding to the recognized emotions. Emotional information and query text are used as input, and response text is generated as output. The generated response reflects the character's tone and emotions. For example, when emphasizing positive emotions, phrases expressing encouragement and joy are included.
[0191] Step 7:
[0192] The server performs post-processing on the generated response to emphasize character-specific expressions. The input is a generated response, and the output is an even more character-driven response. Specifically, the post-processing involves adjustments to reflect the character's unique tone of voice and expression patterns.
[0193] Step 8:
[0194] The server sends the final response to the user's terminal. The server has the post-processed response text as input, and it is displayed on the user's terminal as output. Specifically, the response text is delivered to the user in real time over the network.
[0195] (Application Example 2)
[0196] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0197] Modern virtual interactions require effectively providing users with personalized experiences. However, current systems struggle to reproduce a character's tone of voice while simultaneously generating responses that adapt to the user's emotions. To solve this problem, a system is needed that faithfully reproduces a character's tone of voice while generating responses that align with the user's feelings.
[0198] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0199] In this invention, the server includes means for analyzing the user's emotional state, means for generating a character-like response based on the user's emotional state using a generative model, and means for post-processing the generated response to emphasize character-specific elements. This makes it possible to generate a personalized response that is appropriate to the user's emotions while faithfully reproducing the tone of voice of the character specified by the user.
[0200] A "user" is the entity that operates the system and enjoys interacting with a specific character.
[0201] A "character" is an entity specified by the user that generates responses based on its tone of voice and language patterns.
[0202] A "network" is a communication infrastructure used to acquire and transmit information.
[0203] A "recording medium" is hardware used to store and retain data and information.
[0204] "Dialogue information" refers to data related to a character's language expression and tone of voice.
[0205] "Preprocessing" refers to the process of formally preparing acquired data for use in a generative model.
[0206] A "generative model" is an algorithm that generates responses that reproduce a character's tone of voice and writing style through learning.
[0207] "Training information" refers to the dataset used to train a generative model.
[0208] "Emotional state" refers to information that indicates the user's psychological state or mood.
[0209] "Analysis" is the process of interpreting data and information and extracting useful insights.
[0210] "Post-processing" refers to the process of further adjusting the generated response and emphasizing the tone and characteristics of the specified character.
[0211] A "device" refers to hardware used to perform interactions with users.
[0212] This system is an interactive system that reproduces the speech patterns of specific characters and generates responses that match the user's emotional state. The specific implementation method is described below.
[0213] First, the user selects a specific character through their device and begins interacting with that character. Suitable devices include smartphones and smart glasses. The selected character's dialogue information is retrieved by a server using a network or recording medium. The server preprocesses this dialogue information and structures it as training data for a generative AI model. During the preprocessing stage, the dialogue information is formatted to allow the generative model to easily learn from it.
[0214] Next, the server uses a generative model to learn the character's tone of voice and language patterns from the training data. For example, an open-source generative AI model could be used.
[0215] When a user enters an inquiry from their device, the server uses an emotion analysis engine to analyze the user's emotions based on that inquiry. This analysis can utilize a cloud-based emotion recognition service, and based on the results, a generative model generates a character-like response that is adapted to the user's emotional state.
[0216] The generated responses are then post-processed to further emphasize character-specific elements. This post-processing ensures that the responses more clearly reflect the character's unique phrasing and emotional expressions.
[0217] For example, if the user types "I want to do something fun today," the character will generate a response such as "A wonderful day awaits you, what would you like to do?" and an encouraging dialogue will unfold. An example of a corresponding prompt would be: "You are an assistant that analyzes the user's emotions and faithfully reproduces the tone of the character the user likes to continue the conversation. When the user says, 'I want a new jacket, but I'm not sure,' what would your response be?"
[0218] This configuration enables personalized intef interactions for users.
[0219] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0220] Step 1:
[0221] The user selects a specific character via the terminal. The input is the user's character selection data, which the terminal sends to the server. The terminal constructs the appropriate request to transmit the user's selection to the server over the network and sends the input information through the interface.
[0222] Step 2:
[0223] The server retrieves dialogue information related to the selected character via a network or storage medium. The input is the character's identification information, and the output is the dialogue information as raw data. In this process, the server collects necessary information from external resources or an internal database via an API and stores it in a cache.
[0224] Step 3:
[0225] The server preprocesses the acquired dialogue information and structures it as training data for a generative AI model. The input is raw dialogue data, and the output is in a format suitable for training. Specifically, the server prepares the data by normalizing character codes and removing meaningless strings. This process also includes data cleaning and tokenization techniques.
[0226] Step 4:
[0227] The server trains a generative model, learning the character's speech patterns and language. The input is the training data prepared in step 3, and the output is the trained generative model. The server uses computing nodes to input data into the model and learns while adjusting parameters. A deep learning framework, for example, is suitable for use as the generative AI model.
[0228] Step 5:
[0229] The user enters their inquiry into a terminal and sends it to the server. The input is the user's inquiry text, and the output is raw data for analysis. The terminal uses a secure communication protocol to send the user's input to the server.
[0230] Step 6:
[0231] The server analyzes the user's inquiry using an emotion analysis engine and generates emotion data. The input is the inquiry text received in step 5, and the output is the emotion analysis result. The server utilizes natural language processing and emotion analysis algorithms to score the emotions within the text.
[0232] Step 7:
[0233] The server generates character-like responses using a generative model based on sentiment analysis results. The input is the analyzed sentiment data and the user's original query, and the output is the response text. The server inputs prompts to the generative AI model to create the optimal response, which reflects learned character expression patterns and the user's emotional needs.
[0234] Step 8:
[0235] The server performs post-processing on the generated response to emphasize character-specific elements. The input is the generated response text, and the output is the final response text. The server adds specific words or phrases or adjusts the style to emphasize the expression.
[0236] Step 9:
[0237] The server sends the final, post-processed response back to the user's terminal. The input is the post-processed response text, and the output is the data displayed on the user's terminal. The server sends the response data from the server to the terminal and presents the response in a formatted form for visualization by the user.
[0238] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0239] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0240] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0241] [Second Embodiment]
[0242] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0243] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0244] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0245] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0246] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0247] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0248] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0249] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0250] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0251] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0252] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0253] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0254] The present invention provides a platform for users to obtain responses that mimic the speech patterns of a specific character. The natural language processing functions of this system are described below.
[0255] First, the system is triggered when the user specifies their favorite character. The server receives this information and collects dialogue data from works and databases related to the specified character. The collected data is used as material to understand the character's speech patterns and characteristics. This data is preprocessed and converted into a format optimal for training a generative model.
[0256] The server then uses this preprocessed data to train a generative model, learning the language patterns and emotional expressions of the specified character. Once trained, the model retains the generated patterns and makes them available in response to actual queries.
[0257] Next, when a user makes a request, the device sends the content to the server. The server uses a generative model to generate a response to the user's request that has a character-like tone. At this stage, the response undergoes post-processing to reflect the character's characteristics and enable realistic interaction.
[0258] For example, if a user asks, "What products do you recommend?" and they are a fan of the dragon character, the system will generate a response like, "We have items that are as powerful as a dragon and will support you every day!" This allows the user to have an experience as if they were interacting with that character.
[0259] This system allows companies to personalize inquiry responses and increase user engagement. Because the system generates responses that include character-specific elements, users become more interested and more likely to actively engage in communication with the company.
[0260] The following describes the processing flow.
[0261] Step 1:
[0262] Users specify their favorite character via LINE and begin a conversation with a chatbot. The character information entered by the user acts as a trigger to retrieve related data about that character.
[0263] Step 2:
[0264] The server receives character information from users and collects dialogue data for those characters via the internet or an internal database. This collected data serves as material for learning the characters' language styles and characteristics.
[0265] Step 3:
[0266] The server preprocesses the collected dialogue data. This includes removing noise, normalizing the text, and tokenizing it. This preprocessing creates a dataset that allows the generative model to train more effectively.
[0267] Step 4:
[0268] The server uses pre-processed data to train a generative model. During this process, machine learning algorithms are used to teach the model the character's speech patterns and mannerisms. Once the model is trained, it is ready to generate responses from that character.
[0269] Step 5:
[0270] When a user enters their inquiry, the terminal sends that information to the server. This user input forms the basis for generating the response.
[0271] Step 6:
[0272] The server uses a generative model to generate character-like responses to user questions. These responses are designed to reflect the unique language patterns and emotional expressions of the characters.
[0273] Step 7:
[0274] The server performs post-processing on the generated response. This post-processing involves adding specific phrases or sentence endings to emphasize character-specific elements.
[0275] Step 8:
[0276] The server sends the final response to the device, allowing the user to view it on LINE. Seeing the displayed response, the user can experience what it's like to be having a conversation with their favorite character.
[0277] (Example 1)
[0278] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0279] There is a need to reproduce the unique speech patterns and language styles of characters, providing users with a more relatable and engaging communication experience. However, existing technologies have difficulty generating expressions that adequately mimic the characteristics of characters. This hinders increased user satisfaction and the promotion of deeper engagement.
[0280] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0281] In this invention, when a user specifies a preferred character, the server includes means for an information processing device to collect speech data of that character from a communication network or information aggregation device; means for initial processing the collected speech data and optimizing it as training data for a learning algorithm; and means for training the learning algorithm with the training data to enable it to understand the character's expression style and language structure. This makes it possible to provide realistic and personalized communication using the character's unique language expressions.
[0282] A "user" is a person or organization that uses this system with the aim of obtaining a response from a specific character.
[0283] An "information processing device" is a device that processes digital data and has the function of acquiring, processing, and transmitting specific information based on inquiries from users.
[0284] A "communication network" is an infrastructure for transmitting and receiving data, including the Internet and other network technologies.
[0285] An "information integration device" is any system or platform for collecting, storing, and using large amounts of data as needed.
[0286] "Utterance data" is data related to a character's lines and language expressions, and is text information reflecting the characteristics of the character.
[0287] "Initial processing" is processing such as data shaping, cleaning, and format conversion performed after data collection, and is a process of preparing data in a form suitable for input to a generation AI model.
[0288] A "learning algorithm" is a technology aimed at processing large amounts of data based on a mathematical model and learning patterns and characteristics, and includes machine learning technology.
[0289] "Learning data" is a data set provided to a learning algorithm and is a material for the algorithm to learn specific patterns and features.
[0290] "Expression style" refers to the unique language style, tone, emotional expression, etc. used by a character, and is an element that forms the personality of the character.
[0291] "Language structure" refers to the grammatical and structural features of the language used by a character, and is an element related to the flow and composition method of words.
[0292] The system of this invention starts operating when the user specifies a particular character. At this time, the user uses a character selection interface using a smartphone or computer application. Once the user selects a character, the specified information is transmitted to the server via the terminal.
[0293] Based on this information, the server collects speech data from databases associated with the specified character via the communication network. This process involves web scraping and API calls to obtain appropriate text data from video content and books in which the character appears.
[0294] The acquired speech data is initially processed by the server. This initial processing uses natural language processing libraries such as NLTK and spaCy to remove unnecessary symbols and normalize the text, followed by tokenization and part-of-speech tagging. This process ensures the data is in a format optimized for the learning algorithm.
[0295] Next, the server trains a generative AI model using the processed data. At this stage, a machine learning framework such as TensorFlow or PyTorch is used to build a model that learns the character's expressive style and language structure. The model reflects the character's characteristics and enables the generation of responses to user inquiries.
[0296] When a user enters a specific question or request using their device, that information is sent to the server. The server utilizes a trained generative AI model to generate character-like responses to the user's inquiries. For example, if a user enters "What products do you recommend?", the server generates a response that reflects the character's personality, such as "We have items that are as powerful as a dragon and will support your everyday life!" The generated response is then given a finishing touch to emphasize the character's characteristics before being sent back to the device.
[0297] An example of a prompt message is given to the generating AI model as, "Please respond to the question 'What products do you recommend?' in a dragon's voice." This system allows users to have a personalized conversational experience with a selected character, and is expected to improve customer satisfaction and engagement for companies and services.
[0298] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0299] Step 1:
[0300] The user uses their device to specify a particular character. As input, the user selects or enters the character name within the application's interface. This information, reflecting the user's preference, is sent to the server.
[0301] Step 2:
[0302] The server uses the received character information to collect relevant dialogue data from information aggregation devices via the communication network. As input, queries based on character names are executed, and the output is the character's dialogue data. Database access and API calls are performed to aggregate the dialogue data.
[0303] Step 3:
[0304] The server performs initial processing on the collected speech data. The input data is text information in a non-standardized format, and natural language processing libraries (such as NLTK and spaCy) are used to cleanse and normalize the text. The output is tokenized text data.
[0305] Step 4:
[0306] The server trains a generative AI model using the preprocessed dataset. The input is the formatted character data, and tools such as TensorFlow and PyTorch are used. The model learns the character representation style and language structure, and the output is the trained AI model.
[0307] Step 5:
[0308] When the user uses the terminal to conduct Q&A, questions and requests are input from the terminal as the user's input. As a specific operation, the user inputs text from the interface, which is then sent to the server. As the output, preparations are made to generate a character-style response based on this.
[0309] Step 6:
[0310] The server uses the generative AI model with the user's question as the input to generate a character-style response. The output obtained is the response text that makes use of the character's features. The server performs finishing processing on this response to emphasize the character-specific elements.
[0311] Step 7:
[0312] The terminal receives the finished response sent back from the server and displays it to the user. The input includes the response data from the server. This text is visualized on the terminal screen, and a character-style conversation is established as the final result provided to the user.
[0313] (Application Example 1)
[0314] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0315] Modern commercial platforms require providing personalized experiences in interactions with users to enhance customer engagement. However, typical systems struggle to provide personalized responses based on user interests and preferences, making it challenging to deliver individual user experiences, such as generating responses in the tone of a specific character.
[0316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0317] In this invention, the server includes means for acquiring speech data of a character specified by the user from an information network or information storage device, means for preprocessing the acquired speech data and structuring it as training data for a generative model, and means for introducing products and information using the generated response. This makes it possible to introduce products and information in the tone of voice of the character set by the user.
[0318] "User" refers to a person who uses a system or application.
[0319] A "character" refers to a fictional person or creature that appears in a specific work or story, possessing its own unique language patterns and emotions.
[0320] An "information network" refers to the entire system for acquiring or transmitting information through networks such as the internet.
[0321] "Information storage device" refers to physical or virtual media used to store information using databases or other storage devices.
[0322] "Word data" refers to information expressed in text format, including data such as lines and dialogues spoken by specific characters.
[0323] "Acquisition" refers to the act of collecting data from external sources.
[0324] "Preprocessing" refers to the process of organizing acquired data into an appropriate format and processing it to a state suitable for analysis and training machine learning models.
[0325] A "generative model" refers to a machine learning algorithm that automatically generates new information or responses based on input data.
[0326] "Training data" refers to the dataset used to create a machine learning model; it is the fundamental data that the model learns from.
[0327] "Introducing products or information" refers to the act of explaining or suggesting specific content or products to users in detail.
[0328] A "generative AI model" refers to a model that uses artificial intelligence technology to generate new text or responses from data.
[0329] The program to implement this system primarily consists of data exchange between the server and the user's terminal, and natural language processing. The server retrieves the corresponding character's speech data from an information network or data storage device based on the character specified by the user. The retrieved data is preprocessed and structured as training data for a generative AI model. This process utilizes programming languages such as Python for cleaning and formatting text data, and natural language processing APIs (e.g., OpenAI's GPT-3).
[0330] The server trains a generative model based on pre-processed data, learning the character's tone of voice and language patterns. Based on inquiries from the user's terminal, the server utilizes the generative AI model to generate character-like responses. These responses are further post-processed to emphasize character-specific elements. The generated responses are sent to the user's terminal, providing a unique purchasing experience by introducing products and information.
[0331] For example, if a user inquires via smartphone, "I want a new laptop," the system will generate a response in the tone of their favorite character saying, "A laptop that opens the door to the future—unleash your creativity with this!" An example of a prompt to the generating AI model would be, "Please give me a recommendation about [topic] in the tone of the character's name."
[0332] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0333] Step 1:
[0334] The user selects a specific character on the terminal and enters their inquiry. The input data includes the character name set by the user and the specific inquiry (e.g., "I want a new laptop"). This triggers the system to begin generating a character-like response to the inquiry.
[0335] Step 2:
[0336] The terminal sends the user's selected character information and inquiry details to the server. At this time, the terminal sends the character identifier and inquiry details to the server as JSON data. The input data is received by the server and forms the basis for proceeding to the next step.
[0337] Step 3:
[0338] The server collects the corresponding character's speech data from the information storage device based on the character's identifier. The collected data is text data containing the character's lines and characteristic expressions. This data is then pre-processed.
[0339] Step 4:
[0340] The server preprocesses the collected word data and structures it as training data for a generative AI model. This preprocessing involves text cleaning (removing unnecessary information and converting it to a unified format). The preprocessed data is then in a format suitable for training the model.
[0341] Step 5:
[0342] The server trains a generative AI model using preprocessed data to learn character speech patterns and language styles. The input is structured word data, and the output is the learned model parameters, which are used in the next step.
[0343] Step 6:
[0344] The server inputs the user's inquiry into a generating AI model and generates a character-like response. This prompt is in the form of "Please give me a recommendation about XX in the tone of the character's name," and the input is transformed into a specific response, enabling dynamic dialogue.
[0345] Step 7:
[0346] The generated responses undergo post-processing on the server to emphasize character-specific elements. This post-processing adds specific tones and emotions to the lines, resulting in more realistic character responses.
[0347] Step 8:
[0348] Finally, the server sends a post-processed response back to the terminal. This response is displayed on the user's terminal, and the user receives personalized product and information recommendations from their favorite character. The output data allows the user to have a unique purchasing experience.
[0349] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0350] This invention is an interactive system that allows users to receive responses based on the tone of a specific character, while simultaneously recognizing the user's emotions and customizing the response content to suit those emotions. The operation of this system is described below in natural language.
[0351] First, the system activates when a user specifies their favorite character via LINE and initiates an inquiry. Based on the entered character information, the server collects dialogue data for that character from the internet or an internal database.
[0352] The collected dialogue data is preprocessed to understand the character's tone of voice and linguistic characteristics, and then structured as training data for a generative model. The server uses this data to train the generative model and teach it character-specific language patterns. At this stage, the generative model has the foundation to reproduce the character's tone of voice.
[0353] Next, when the user enters their inquiry, the terminal sends that information to the server. At this point, the system's built-in emotion engine is utilized. The server uses this emotion engine to recognize the emotions from the user's input text.
[0354] Based on the recognized emotion, the server uses a generative model to generate a character-like response corresponding to that emotion. This response reflects pre-trained character speech patterns and emotion-based expressions. The generated response is then post-processed to take emotion information into account, further emphasizing the character's unique emotional expressions.
[0355] For example, if a user types "I'm feeling a bit down today," the system will work to have the designated character respond with positive words to cheer them up. Let's say the generated response is "A fantastic adventure awaits you that will cheer you up!" In this way, users can enjoy more personalized interactions tailored to the character and their emotions.
[0356] This system can strengthen the relationship between companies and users and improve the usefulness of inquiries by simultaneously achieving realistic character reproduction and response generation adapted to the user's emotions.
[0357] The following describes the processing flow.
[0358] Step 1:
[0359] Users specify their favorite character on LINE and initiate an inquiry. This necessitates a conversational experience tailored to the user's preferences.
[0360] Step 2:
[0361] The server receives information about the specified character and collects dialogue data from works and databases in which that character appears. This includes the character's characteristic speech patterns and phrases.
[0362] Step 3:
[0363] The server preprocesses the collected dialogue data, performing tasks such as text normalization and tokenization, and structures it into a data format suitable for training generative models. This processing improves the accuracy of the learning process.
[0364] Step 4:
[0365] The server uses the pre-processed data to train the generative model. The model learns the characters' speech patterns and language styles, and is then ready to go.
[0366] Step 5:
[0367] When a user enters and submits an inquiry, the device transmits that data to the server. At that time, sentiment analysis of the user's input is performed.
[0368] Step 6:
[0369] The emotion engine installed on the server recognizes emotions from the user's input text. For example, basic emotions such as joy, sadness, and anger can be identified.
[0370] Step 7:
[0371] The server considers the recognized emotions and uses a generative model to generate character-like responses that correspond to those emotions. The responses not only reflect the character's tone of voice but also fit the user's emotions.
[0372] Step 8:
[0373] The server then performs post-processing on the generated response to enhance its emotional expression. This involves adding additional character-specific phrases and nuances depending on the specific emotion.
[0374] Step 9:
[0375] The server sends the completed character-like response to the user's device, allowing the user to view it on LINE. This allows users to enjoy a richer experience through dialogue that resonates with their own emotions.
[0376] (Example 2)
[0377] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0378] Conventional dialogue systems have problems accurately reproducing the tone of voice of a character specified by the user, and are unable to generate flexible responses that respond to the user's emotions. This limits the quality of the dialogue provided by the system and hinders the improvement of the user experience. In addition, the lack of processing to effectively emphasize character-specific expressions makes it difficult to provide responses that fully utilize the characteristics of the character.
[0379] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0380] In this invention, the server includes means for acquiring language data of a symbol specified by the user from a network or information recording device, means for preprocessing the acquired language data and structuring it as training information for a generation algorithm, and means for training the generation algorithm with the training information to learn the language features and expression patterns of the symbol. This makes it possible to faithfully reproduce the tone of voice of a character specified by the user and generate appropriate responses that correspond to the user's emotions.
[0381] A "symbol" is an element designated by the user to represent a specific character or its characteristics.
[0382] A "network" is a communication infrastructure used to transmit digital information and connect multiple devices and systems.
[0383] An "information recording device" is a technology or platform for storing data and accessing it as needed.
[0384] "Linguistic data" refers to information about texts and speeches, including the tone of voice and linguistic characteristics of a specific character.
[0385] "Preprocessing" is the process of normalizing acquired data, removing noise, and preparing it for training generative algorithms.
[0386] A "generative algorithm" is a computational method for generating new data or responses based on presented information.
[0387] "Training information" refers to a structured dataset used by a generative algorithm for learning.
[0388] "Linguistic features" refer to the unique expressive styles and grammatical characteristics of the language spoken by a particular character or individual.
[0389] A "pattern of expression" is a characteristic that indicates a specific form or flow in language or communication.
[0390] "Emotion analysis" is a technology that recognizes a user's emotional state from text or audio.
[0391] "Control variables" are parameters that are adjusted to optimize the performance of a generation algorithm.
[0392] The specific implementation methods for this system are described below.
[0393] First, the user uses a network-enabled device to specify their preferred character and access the system via a messaging service such as LINE. Based on the specified character, the server collects language data related to that character. This data is stored on the network or in a database. The collected language data is then structured for processing by the server.
[0394] The server runs software developed using programming languages such as Python and Java to preprocess the collected language data. This process includes data normalization and noise reduction, preparing the data for use as training data for a generative AI model. Next, the server uses natural language processing libraries to train the generative AI model with this training data. In this way, the model learns the language features and expression patterns of a specified character.
[0395] Subsequently, when a user submits a query to the system, the terminal sends the query to the server. Upon receiving the query, the server uses natural language processing technology to analyze the text sent by the user and recognize its sentiment. Sentiment analysis is performed using sentiment analysis libraries and APIs.
[0396] When the generative AI model generates responses based on the trained data, the server generates character-like responses that take the user's emotions into account, based on the analyzed emotions. Finally, the generated responses undergo further post-processing to enhance character-specific expressions. This process generates emotionally natural responses that are sent to the user's device.
[0397] For example, if a user sends a message such as "I'm a little tired today," the system will generate a response that resonates with the user, while maintaining the character's tone of voice. "For instance, it might respond with something like, 'Please rest well today!' In this way, the character provides the user with an emotionally resonant dialogue."
[0398] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0399] Step 1:
[0400] A user initiates a query by specifying a particular character through the LINE application. The input is the character's name, and this information is sent to the server as output. Based on the received character name, the server collects relevant language data using the network or database. Specifically, this involves issuing API queries to the character or performing database searches.
[0401] Step 2:
[0402] The server preprocesses the collected language data. At this stage, raw language data is provided as input, and preprocessed data is generated as output. Data processing is performed, such as normalization, removal of unnecessary information, and identification of frequently occurring words and phrases. Specifically, text tokenization and morphological analysis are performed to build a training dataset.
[0403] Step 3:
[0404] The server uses preprocessed data to train a generative AI model. It receives preprocessed data as input and generates a trained model as output. Here, the model uses a machine learning algorithm to learn the language patterns of characters. Specifically, the model's parameters are updated using a deep learning framework.
[0405] Step 4:
[0406] The user enters their inquiry into the system and sends it to the server via their terminal. The input includes the inquiry text, and this information is passed to the server as output. Specifically, the text data entered on the user's terminal is encrypted and securely transmitted to the server.
[0407] Step 5:
[0408] The server analyzes the received inquiry and uses an emotion analysis engine to evaluate the user's emotions. The user's inquiry is the input for analysis, and the recognized emotion information is the output. The server extracts emotions such as positive, negative, and neutral from the text. Specifically, an emotion score is calculated based on an analysis of keywords and phrases.
[0409] Step 6:
[0410] The server uses a generative AI model to generate character-like responses corresponding to the recognized emotions. Emotional information and query text are used as input, and response text is generated as output. The generated response reflects the character's tone and emotions. For example, when emphasizing positive emotions, phrases expressing encouragement and joy are included.
[0411] Step 7:
[0412] The server performs post-processing on the generated response to emphasize character-specific expressions. The input is a generated response, and the output is an even more character-driven response. Specifically, the post-processing involves adjustments to reflect the character's unique tone of voice and expression patterns.
[0413] Step 8:
[0414] The server sends the final response to the user's terminal. The server has the post-processed response text as input, and it is displayed on the user's terminal as output. Specifically, the response text is delivered to the user in real time over the network.
[0415] (Application Example 2)
[0416] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0417] Modern virtual interactions require effectively providing users with personalized experiences. However, current systems struggle to reproduce a character's tone of voice while simultaneously generating responses that adapt to the user's emotions. To solve this problem, a system is needed that faithfully reproduces a character's tone of voice while generating responses that align with the user's feelings.
[0418] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0419] In this invention, the server includes means for analyzing the user's emotional state, means for generating a character-like response based on the user's emotional state using a generative model, and means for post-processing the generated response to emphasize character-specific elements. This makes it possible to generate a personalized response that is appropriate to the user's emotions while faithfully reproducing the tone of voice of the character specified by the user.
[0420] A "user" is the entity that operates the system and enjoys interacting with a specific character.
[0421] A "character" is an entity specified by the user that generates responses based on its tone of voice and language patterns.
[0422] A "network" is a communication infrastructure used to acquire and transmit information.
[0423] A "recording medium" is hardware used to store and retain data and information.
[0424] "Dialogue information" refers to data related to a character's language expression and tone of voice.
[0425] "Preprocessing" refers to the process of formally preparing acquired data for use in a generative model.
[0426] A "generative model" is an algorithm that generates responses that reproduce a character's tone of voice and writing style through learning.
[0427] "Training information" refers to the dataset used to train a generative model.
[0428] "Emotional state" refers to information that indicates the user's psychological state or mood.
[0429] "Analysis" is the process of interpreting data and information and extracting useful insights.
[0430] "Post-processing" refers to the process of further adjusting the generated response and emphasizing the tone and characteristics of the specified character.
[0431] A "device" refers to hardware used to perform interactions with users.
[0432] This system is an interactive system that reproduces the speech patterns of specific characters and generates responses that match the user's emotional state. The specific implementation method is described below.
[0433] First, the user selects a specific character through their device and begins interacting with that character. Suitable devices include smartphones and smart glasses. The selected character's dialogue information is retrieved by a server using a network or recording medium. The server preprocesses this dialogue information and structures it as training data for a generative AI model. During the preprocessing stage, the dialogue information is formatted to allow the generative model to easily learn from it.
[0434] Next, the server uses a generative model to learn the character's tone of voice and language patterns from the training data. For example, an open-source generative AI model could be used.
[0435] When a user enters an inquiry from their device, the server uses an emotion analysis engine to analyze the user's emotions based on that inquiry. This analysis can utilize a cloud-based emotion recognition service, and based on the results, a generative model generates a character-like response that is adapted to the user's emotional state.
[0436] The generated responses are then post-processed to further emphasize character-specific elements. This post-processing ensures that the responses more clearly reflect the character's unique phrasing and emotional expressions.
[0437] For example, if the user types "I want to do something fun today," the character will generate a response such as "A wonderful day awaits you, what would you like to do?" and an encouraging dialogue will unfold. An example of a corresponding prompt would be: "You are an assistant that analyzes the user's emotions and faithfully reproduces the tone of the character the user likes to continue the conversation. When the user says, 'I want a new jacket, but I'm not sure,' what would your response be?"
[0438] This configuration enables personalized intef interactions for users.
[0439] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0440] Step 1:
[0441] The user selects a specific character via the terminal. The input is the user's character selection data, which the terminal sends to the server. The terminal constructs the appropriate request to transmit the user's selection to the server over the network and sends the input information through the interface.
[0442] Step 2:
[0443] The server retrieves dialogue information related to the selected character via a network or storage medium. The input is the character's identification information, and the output is the dialogue information as raw data. In this process, the server collects necessary information from external resources or an internal database via an API and stores it in a cache.
[0444] Step 3:
[0445] The server preprocesses the acquired dialogue information and structures it as training data for a generative AI model. The input is raw dialogue data, and the output is in a format suitable for training. Specifically, the server prepares the data by normalizing character codes and removing meaningless strings. This process also includes data cleaning and tokenization techniques.
[0446] Step 4:
[0447] The server trains a generative model, learning the character's speech patterns and language. The input is the training data prepared in step 3, and the output is the trained generative model. The server uses computing nodes to input data into the model and learns while adjusting parameters. A deep learning framework, for example, is suitable for use as the generative AI model.
[0448] Step 5:
[0449] The user enters their inquiry into a terminal and sends it to the server. The input is the user's inquiry text, and the output is raw data for analysis. The terminal uses a secure communication protocol to send the user's input to the server.
[0450] Step 6:
[0451] The server analyzes the user's inquiry using an emotion analysis engine and generates emotion data. The input is the inquiry text received in step 5, and the output is the emotion analysis result. The server utilizes natural language processing and emotion analysis algorithms to score the emotions within the text.
[0452] Step 7:
[0453] The server generates character-like responses using a generative model based on sentiment analysis results. The input is the analyzed sentiment data and the user's original query, and the output is the response text. The server inputs prompts to the generative AI model to create the optimal response, which reflects learned character expression patterns and the user's emotional needs.
[0454] Step 8:
[0455] The server performs post-processing on the generated response to emphasize character-specific elements. The input is the generated response text, and the output is the final response text. The server adds specific words or phrases or adjusts the style to emphasize the expression.
[0456] Step 9:
[0457] The server sends the final, post-processed response back to the user's terminal. The input is the post-processed response text, and the output is the data displayed on the user's terminal. The server sends the response data from the server to the terminal and presents the response in a formatted form for visualization by the user.
[0458] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0459] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0460] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0461] [Third Embodiment]
[0462] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0463] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0464] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0465] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0466] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0467] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0468] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0469] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0470] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0471] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0472] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0473] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0474] The present invention provides a platform for users to obtain responses that mimic the speech patterns of a specific character. The natural language processing functions of this system are described below.
[0475] First, the system is triggered when the user specifies their favorite character. The server receives this information and collects dialogue data from works and databases related to the specified character. The collected data is used as material to understand the character's speech patterns and characteristics. This data is preprocessed and converted into a format optimal for training a generative model.
[0476] The server then uses this preprocessed data to train a generative model, learning the language patterns and emotional expressions of the specified character. Once trained, the model retains the generated patterns and makes them available in response to actual queries.
[0477] Next, when a user makes a request, the device sends the content to the server. The server uses a generative model to generate a response to the user's request that has a character-like tone. At this stage, the response undergoes post-processing to reflect the character's characteristics and enable realistic interaction.
[0478] For example, if a user asks, "What products do you recommend?" and they are a fan of the dragon character, the system will generate a response like, "We have items that are as powerful as a dragon and will support you every day!" This allows the user to have an experience as if they were interacting with that character.
[0479] This system allows companies to personalize inquiry responses and increase user engagement. Because the system generates responses that include character-specific elements, users become more interested and more likely to actively engage in communication with the company.
[0480] The following describes the processing flow.
[0481] Step 1:
[0482] Users specify their favorite character via LINE and begin a conversation with a chatbot. The character information entered by the user acts as a trigger to retrieve related data about that character.
[0483] Step 2:
[0484] The server receives character information from users and collects dialogue data for those characters via the internet or an internal database. This collected data serves as material for learning the characters' language styles and characteristics.
[0485] Step 3:
[0486] The server preprocesses the collected dialogue data. This includes removing noise, normalizing the text, and tokenizing it. This preprocessing creates a dataset that allows the generative model to train more effectively.
[0487] Step 4:
[0488] The server uses pre-processed data to train a generative model. During this process, machine learning algorithms are used to teach the model the character's speech patterns and mannerisms. Once the model is trained, it is ready to generate responses from that character.
[0489] Step 5:
[0490] When a user enters their inquiry, the terminal sends that information to the server. This user input forms the basis for generating the response.
[0491] Step 6:
[0492] The server uses a generative model to generate character-like responses to user questions. These responses are designed to reflect the unique language patterns and emotional expressions of the characters.
[0493] Step 7:
[0494] The server performs post-processing on the generated response. This post-processing involves adding specific phrases or sentence endings to emphasize character-specific elements.
[0495] Step 8:
[0496] The server sends the final response to the device, allowing the user to view it on LINE. Seeing the displayed response, the user can experience what it's like to be having a conversation with their favorite character.
[0497] (Example 1)
[0498] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0499] There is a need to reproduce the unique speech patterns and language styles of characters, providing users with a more relatable and engaging communication experience. However, existing technologies have difficulty generating expressions that adequately mimic the characteristics of characters. This hinders increased user satisfaction and the promotion of deeper engagement.
[0500] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0501] In this invention, when a user specifies a preferred character, the server includes means for an information processing device to collect speech data of that character from a communication network or information aggregation device; means for initial processing the collected speech data and optimizing it as training data for a learning algorithm; and means for training the learning algorithm with the training data to enable it to understand the character's expression style and language structure. This makes it possible to provide realistic and personalized communication using the character's unique language expressions.
[0502] A "user" is a person or organization that uses this system with the aim of obtaining a response from a specific character.
[0503] An "information processing device" is a device that processes digital data and has the function of acquiring, processing, and transmitting specific information based on inquiries from users.
[0504] A "communication network" is the infrastructure used for sending and receiving data, and includes the internet and other network technologies.
[0505] An "information aggregation device" is any system or platform for collecting, storing, and utilizing large amounts of data as needed.
[0506] "Speech data" refers to data about a character's lines and linguistic expressions, and is text information that reflects the character's characteristics.
[0507] "Initial processing" refers to processes performed after data collection, such as data formatting, cleaning, and format conversion, which prepare the data for input into the AI model that generates it.
[0508] A "learning algorithm" is a technique that aims to learn patterns and characteristics by processing large amounts of data based on mathematical models, and includes machine learning techniques.
[0509] "Training data" refers to the dataset provided to a learning algorithm, which serves as the material for the algorithm to learn specific patterns and features.
[0510] "Style of expression" refers to the unique language style, tone of voice, and emotional expressions used by a character, and is an element that forms the character's individuality.
[0511] "Language structure" refers to the grammatical and structural characteristics of the language used by a character, encompassing elements related to the flow and structure of words.
[0512] The system of this invention starts operating when the user specifies a particular character. At this time, the user uses a character selection interface using a smartphone or computer application. Once the user selects a character, the specified information is transmitted to the server via the terminal.
[0513] Based on this information, the server collects speech data from databases associated with the specified character via the communication network. This process involves web scraping and API calls to obtain appropriate text data from video content and books in which the character appears.
[0514] The acquired speech data is initially processed by the server. This initial processing uses natural language processing libraries such as NLTK and spaCy to remove unnecessary symbols and normalize the text, followed by tokenization and part-of-speech tagging. This process ensures the data is in a format optimized for the learning algorithm.
[0515] Next, the server trains a generative AI model using the processed data. At this stage, a machine learning framework such as TensorFlow or PyTorch is used to build a model that learns the character's expressive style and language structure. The model reflects the character's characteristics and enables the generation of responses to user inquiries.
[0516] When a user enters a specific question or request using their device, that information is sent to the server. The server utilizes a trained generative AI model to generate character-like responses to the user's inquiries. For example, if a user enters "What products do you recommend?", the server generates a response that reflects the character's personality, such as "We have items that are as powerful as a dragon and will support your everyday life!" The generated response is then given a finishing touch to emphasize the character's characteristics before being sent back to the device.
[0517] An example of a prompt message is given to the generating AI model as, "Please respond to the question 'What products do you recommend?' in a dragon's voice." This system allows users to have a personalized conversational experience with a selected character, and is expected to improve customer satisfaction and engagement for companies and services.
[0518] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0519] Step 1:
[0520] The user uses their device to specify a particular character. As input, the user selects or enters the character name within the application's interface. This information, reflecting the user's preference, is sent to the server.
[0521] Step 2:
[0522] The server uses the received character information to collect relevant dialogue data from information aggregation devices via the communication network. As input, queries based on character names are executed, and the output is the character's dialogue data. Database access and API calls are performed to aggregate the dialogue data.
[0523] Step 3:
[0524] The server performs initial processing on the collected speech data. The input data is text information in a non-standardized format, and natural language processing libraries (such as NLTK and spaCy) are used to cleanse and normalize the text. The output is tokenized text data.
[0525] Step 4:
[0526] The server trains a generative AI model using an initialized dataset. The input is formatted character data, and tools such as TensorFlow and PyTorch are used. The model learns the character's representational style and language structure, and the output is a trained AI model.
[0527] Step 5:
[0528] When a user uses a terminal to ask and answer questions, the user inputs questions and requests from the terminal. Specifically, the user enters text through the interface, which is then sent to the server. As output, a character-like response based on this input is ready to be generated.
[0529] Step 6:
[0530] The server uses a character-like response model based on the user's question as input. The output is a response text that highlights the character's characteristics. The server then processes this response to emphasize the character's unique elements.
[0531] Step 7:
[0532] The terminal receives the completed response sent back from the server and displays it to the user. The input includes the response data from the server. This text is visualized on the terminal's screen, and a character-like dialogue is established as the final result provided to the user.
[0533] (Application Example 1)
[0534] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0535] Modern commercial platforms require providing personalized experiences in interactions with users to enhance customer engagement. However, typical systems struggle to provide personalized responses based on user interests and preferences, making it challenging to deliver individual user experiences, such as generating responses in the tone of a specific character.
[0536] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0537] In this invention, the server includes means for acquiring speech data of a character specified by the user from an information network or information storage device, means for preprocessing the acquired speech data and structuring it as training data for a generative model, and means for introducing products and information using the generated response. This makes it possible to introduce products and information in the tone of voice of the character set by the user.
[0538] "User" refers to a person who uses a system or application.
[0539] A "character" refers to a fictional person or creature that appears in a specific work or story, possessing its own unique language patterns and emotions.
[0540] An "information network" refers to the entire system for acquiring or transmitting information through networks such as the internet.
[0541] "Information storage device" refers to physical or virtual media used to store information using databases or other storage devices.
[0542] "Word data" refers to information expressed in text format, including data such as lines and dialogues spoken by specific characters.
[0543] "Acquisition" refers to the act of collecting data from external sources.
[0544] "Preprocessing" refers to the process of organizing acquired data into an appropriate format and processing it to a state suitable for analysis and training machine learning models.
[0545] A "generative model" refers to a machine learning algorithm that automatically generates new information or responses based on input data.
[0546] "Training data" refers to the dataset used to create a machine learning model; it is the fundamental data that the model learns from.
[0547] "Introducing products or information" refers to the act of explaining or suggesting specific content or products to users in detail.
[0548] A "generative AI model" refers to a model that uses artificial intelligence technology to generate new text or responses from data.
[0549] The program to implement this system primarily consists of data exchange between the server and the user's terminal, and natural language processing. The server retrieves the corresponding character's speech data from an information network or data storage device based on the character specified by the user. The retrieved data is preprocessed and structured as training data for a generative AI model. This process utilizes programming languages such as Python for cleaning and formatting text data, and natural language processing APIs (e.g., OpenAI's GPT-3).
[0550] The server trains a generative model based on pre-processed data, learning the character's tone of voice and language patterns. Based on inquiries from the user's terminal, the server utilizes the generative AI model to generate character-like responses. These responses are further post-processed to emphasize character-specific elements. The generated responses are sent to the user's terminal, providing a unique purchasing experience by introducing products and information.
[0551] For example, if a user inquires via smartphone, "I want a new laptop," the system will generate a response in the tone of their favorite character saying, "A laptop that opens the door to the future—unleash your creativity with this!" An example of a prompt to the generating AI model would be, "Please give me a recommendation about [topic] in the tone of the character's name."
[0552] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0553] Step 1:
[0554] The user selects a specific character on the terminal and enters their inquiry. The input data includes the character name set by the user and the specific inquiry (e.g., "I want a new laptop"). This triggers the system to begin generating a character-like response to the inquiry.
[0555] Step 2:
[0556] The terminal sends the user's selected character information and inquiry details to the server. At this time, the terminal sends the character identifier and inquiry details to the server as JSON data. The input data is received by the server and forms the basis for proceeding to the next step.
[0557] Step 3:
[0558] The server collects the corresponding character's speech data from the information storage device based on the character's identifier. The collected data is text data containing the character's lines and characteristic expressions. This data is then pre-processed.
[0559] Step 4:
[0560] The server preprocesses the collected word data and structures it as training data for a generative AI model. This preprocessing involves text cleaning (removing unnecessary information and converting it to a unified format). The preprocessed data is then in a format suitable for training the model.
[0561] Step 5:
[0562] The server trains a generative AI model using preprocessed data to learn character speech patterns and language styles. The input is structured word data, and the output is the learned model parameters, which are used in the next step.
[0563] Step 6:
[0564] The server inputs the user's inquiry into a generating AI model and generates a character-like response. This prompt is in the form of "Please give me a recommendation about XX in the tone of the character's name," and the input is transformed into a specific response, enabling dynamic dialogue.
[0565] Step 7:
[0566] The generated responses undergo post-processing on the server to emphasize character-specific elements. This post-processing adds specific tones and emotions to the lines, resulting in more realistic character responses.
[0567] Step 8:
[0568] Finally, the server sends a post-processed response back to the terminal. This response is displayed on the user's terminal, and the user receives personalized product and information recommendations from their favorite character. The output data allows the user to have a unique purchasing experience.
[0569] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0570] This invention is an interactive system that allows users to receive responses based on the tone of a specific character, while simultaneously recognizing the user's emotions and customizing the response content to suit those emotions. The operation of this system is described below in natural language.
[0571] First, the system activates when a user specifies their favorite character via LINE and initiates an inquiry. Based on the entered character information, the server collects dialogue data for that character from the internet or an internal database.
[0572] The collected dialogue data is preprocessed to understand the character's tone of voice and linguistic characteristics, and then structured as training data for a generative model. The server uses this data to train the generative model and teach it character-specific language patterns. At this stage, the generative model has the foundation to reproduce the character's tone of voice.
[0573] Next, when the user enters their inquiry, the terminal sends that information to the server. At this point, the system's built-in emotion engine is utilized. The server uses this emotion engine to recognize the emotions from the user's input text.
[0574] Based on the recognized emotion, the server uses a generative model to generate a character-like response corresponding to that emotion. This response reflects pre-trained character speech patterns and emotion-based expressions. The generated response is then post-processed to take emotion information into account, further emphasizing the character's unique emotional expressions.
[0575] For example, if a user types "I'm feeling a bit down today," the system will work to have the designated character respond with positive words to cheer them up. Let's say the generated response is "A fantastic adventure awaits you that will cheer you up!" In this way, users can enjoy more personalized interactions tailored to the character and their emotions.
[0576] This system can strengthen the relationship between companies and users and improve the usefulness of inquiries by simultaneously achieving realistic character reproduction and response generation adapted to the user's emotions.
[0577] The following describes the processing flow.
[0578] Step 1:
[0579] Users specify their favorite character on LINE and initiate an inquiry. This necessitates a conversational experience tailored to the user's preferences.
[0580] Step 2:
[0581] The server receives information about the specified character and collects dialogue data from works and databases in which that character appears. This includes the character's characteristic speech patterns and phrases.
[0582] Step 3:
[0583] The server preprocesses the collected dialogue data, performing tasks such as text normalization and tokenization, and structures it into a data format suitable for training generative models. This processing improves the accuracy of the learning process.
[0584] Step 4:
[0585] The server uses the pre-processed data to train the generative model. The model learns the characters' speech patterns and language styles, and is then ready to go.
[0586] Step 5:
[0587] When a user enters and submits an inquiry, the device transmits that data to the server. At that time, sentiment analysis of the user's input is performed.
[0588] Step 6:
[0589] The emotion engine installed on the server recognizes emotions from the user's input text. For example, basic emotions such as joy, sadness, and anger can be identified.
[0590] Step 7:
[0591] The server considers the recognized emotions and uses a generative model to generate character-like responses that correspond to those emotions. The responses not only reflect the character's tone of voice but also fit the user's emotions.
[0592] Step 8:
[0593] The server then performs post-processing on the generated response to enhance its emotional expression. This involves adding additional character-specific phrases and nuances depending on the specific emotion.
[0594] Step 9:
[0595] The server sends the completed character-like response to the user's device, allowing the user to view it on LINE. This allows users to enjoy a richer experience through dialogue that resonates with their own emotions.
[0596] (Example 2)
[0597] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0598] Conventional dialogue systems have problems accurately reproducing the tone of voice of a character specified by the user, and are unable to generate flexible responses that respond to the user's emotions. This limits the quality of the dialogue provided by the system and hinders the improvement of the user experience. In addition, the lack of processing to effectively emphasize character-specific expressions makes it difficult to provide responses that fully utilize the characteristics of the character.
[0599] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0600] In this invention, the server includes means for acquiring language data of a symbol specified by the user from a network or information recording device, means for preprocessing the acquired language data and structuring it as training information for a generation algorithm, and means for training the generation algorithm with the training information to learn the language features and expression patterns of the symbol. This makes it possible to faithfully reproduce the tone of voice of a character specified by the user and generate appropriate responses that correspond to the user's emotions.
[0601] A "symbol" is an element designated by the user to represent a specific character or its characteristics.
[0602] A "network" is a communication infrastructure used to transmit digital information and connect multiple devices and systems.
[0603] An "information recording device" is a technology or platform for storing data and accessing it as needed.
[0604] "Linguistic data" refers to information about texts and speeches, including the tone of voice and linguistic characteristics of a specific character.
[0605] "Preprocessing" is the process of normalizing acquired data, removing noise, and preparing it for training generative algorithms.
[0606] A "generative algorithm" is a computational method for generating new data or responses based on presented information.
[0607] "Training information" refers to a structured dataset used by a generative algorithm for learning.
[0608] "Linguistic features" refer to the unique expressive styles and grammatical characteristics of the language spoken by a particular character or individual.
[0609] A "pattern of expression" is a characteristic that indicates a specific form or flow in language or communication.
[0610] "Emotion analysis" is a technology that recognizes a user's emotional state from text or audio.
[0611] "Control variables" are parameters that are adjusted to optimize the performance of a generation algorithm.
[0612] The specific implementation methods for this system are described below.
[0613] First, the user uses a network-enabled device to specify their preferred character and access the system via a messaging service such as LINE. Based on the specified character, the server collects language data related to that character. This data is stored on the network or in a database. The collected language data is then structured for processing by the server.
[0614] The server runs software developed using programming languages such as Python and Java to preprocess the collected language data. This process includes data normalization and noise reduction, preparing the data for use as training data for a generative AI model. Next, the server uses natural language processing libraries to train the generative AI model with this training data. In this way, the model learns the language features and expression patterns of a specified character.
[0615] Subsequently, when a user submits a query to the system, the terminal sends the query to the server. Upon receiving the query, the server uses natural language processing technology to analyze the text sent by the user and recognize its sentiment. Sentiment analysis is performed using sentiment analysis libraries and APIs.
[0616] When the generative AI model generates responses based on the trained data, the server generates character-like responses that take the user's emotions into account, based on the analyzed emotions. Finally, the generated responses undergo further post-processing to enhance character-specific expressions. This process generates emotionally natural responses that are sent to the user's device.
[0617] For example, if a user sends a message such as "I'm a little tired today," the system will generate a response that resonates with the user, while maintaining the character's tone of voice. "For instance, it might respond with something like, 'Please rest well today!' In this way, the character provides the user with an emotionally resonant dialogue."
[0618] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0619] Step 1:
[0620] A user initiates a query by specifying a particular character through the LINE application. The input is the character's name, and this information is sent to the server as output. Based on the received character name, the server collects relevant language data using the network or database. Specifically, this involves issuing API queries to the character or performing database searches.
[0621] Step 2:
[0622] The server preprocesses the collected language data. At this stage, raw language data is provided as input, and preprocessed data is generated as output. Data processing is performed, such as normalization, removal of unnecessary information, and identification of frequently occurring words and phrases. Specifically, text tokenization and morphological analysis are performed to build a training dataset.
[0623] Step 3:
[0624] The server uses preprocessed data to train a generative AI model. It receives preprocessed data as input and generates a trained model as output. Here, the model uses a machine learning algorithm to learn the language patterns of characters. Specifically, the model's parameters are updated using a deep learning framework.
[0625] Step 4:
[0626] The user enters their inquiry into the system and sends it to the server via their terminal. The input includes the inquiry text, and this information is passed to the server as output. Specifically, the text data entered on the user's terminal is encrypted and securely transmitted to the server.
[0627] Step 5:
[0628] The server analyzes the received inquiry and uses an emotion analysis engine to evaluate the user's emotions. The user's inquiry is the input for analysis, and the recognized emotion information is the output. The server extracts emotions such as positive, negative, and neutral from the text. Specifically, an emotion score is calculated based on an analysis of keywords and phrases.
[0629] Step 6:
[0630] The server uses a generative AI model to generate character-like responses corresponding to the recognized emotions. Emotional information and query text are used as input, and response text is generated as output. The generated response reflects the character's tone and emotions. For example, when emphasizing positive emotions, phrases expressing encouragement and joy are included.
[0631] Step 7:
[0632] The server performs post-processing on the generated response to emphasize character-specific expressions. The input is a generated response, and the output is an even more character-driven response. Specifically, the post-processing involves adjustments to reflect the character's unique tone of voice and expression patterns.
[0633] Step 8:
[0634] The server sends the final response to the user's terminal. The server has the post-processed response text as input, and it is displayed on the user's terminal as output. Specifically, the response text is delivered to the user in real time over the network.
[0635] (Application Example 2)
[0636] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0637] Modern virtual interactions require effectively providing users with personalized experiences. However, current systems struggle to reproduce a character's tone of voice while simultaneously generating responses that adapt to the user's emotions. To solve this problem, a system is needed that faithfully reproduces a character's tone of voice while generating responses that align with the user's feelings.
[0638] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0639] In this invention, the server includes means for analyzing the user's emotional state, means for generating a character-like response based on the user's emotional state using a generative model, and means for post-processing the generated response to emphasize character-specific elements. This makes it possible to generate a personalized response that is appropriate to the user's emotions while faithfully reproducing the tone of voice of the character specified by the user.
[0640] A "user" is the entity that operates the system and enjoys interacting with a specific character.
[0641] A "character" is an entity specified by the user that generates responses based on its tone of voice and language patterns.
[0642] A "network" is a communication infrastructure used to acquire and transmit information.
[0643] A "recording medium" is hardware used to store and retain data and information.
[0644] "Dialogue information" refers to data related to a character's language expression and tone of voice.
[0645] "Preprocessing" refers to the process of formally preparing acquired data for use in a generative model.
[0646] A "generative model" is an algorithm that generates responses that reproduce a character's tone of voice and writing style through learning.
[0647] "Training information" refers to the dataset used to train a generative model.
[0648] "Emotional state" refers to information that indicates the user's psychological state or mood.
[0649] "Analysis" is the process of interpreting data and information and extracting useful insights.
[0650] "Post-processing" refers to the process of further adjusting the generated response and emphasizing the tone and characteristics of the specified character.
[0651] A "device" refers to hardware used to perform interactions with users.
[0652] This system is an interactive system that reproduces the speech patterns of specific characters and generates responses that match the user's emotional state. The specific implementation method is described below.
[0653] First, the user selects a specific character through their device and begins interacting with that character. Suitable devices include smartphones and smart glasses. The selected character's dialogue information is retrieved by a server using a network or recording medium. The server preprocesses this dialogue information and structures it as training data for a generative AI model. During the preprocessing stage, the dialogue information is formatted to allow the generative model to easily learn from it.
[0654] Next, the server uses a generative model to learn the character's tone of voice and language patterns from the training data. For example, an open-source generative AI model could be used.
[0655] When a user enters an inquiry from their device, the server uses an emotion analysis engine to analyze the user's emotions based on that inquiry. This analysis can utilize a cloud-based emotion recognition service, and based on the results, a generative model generates a character-like response that is adapted to the user's emotional state.
[0656] The generated responses are then post-processed to further emphasize character-specific elements. This post-processing ensures that the responses more clearly reflect the character's unique phrasing and emotional expressions.
[0657] For example, if the user types "I want to do something fun today," the character will generate a response such as "A wonderful day awaits you, what would you like to do?" and an encouraging dialogue will unfold. An example of a corresponding prompt would be: "You are an assistant that analyzes the user's emotions and faithfully reproduces the tone of the character the user likes to continue the conversation. When the user says, 'I want a new jacket, but I'm not sure,' what would your response be?"
[0658] This configuration enables personalized intef interactions for users.
[0659] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0660] Step 1:
[0661] The user selects a specific character via the terminal. The input is the user's character selection data, which the terminal sends to the server. The terminal constructs the appropriate request to transmit the user's selection to the server over the network and sends the input information through the interface.
[0662] Step 2:
[0663] The server retrieves dialogue information related to the selected character via a network or storage medium. The input is the character's identification information, and the output is the dialogue information as raw data. In this process, the server collects necessary information from external resources or an internal database via an API and stores it in a cache.
[0664] Step 3:
[0665] The server preprocesses the acquired dialogue information and structures it as training data for a generative AI model. The input is raw dialogue data, and the output is in a format suitable for training. Specifically, the server prepares the data by normalizing character codes and removing meaningless strings. This process also includes data cleaning and tokenization techniques.
[0666] Step 4:
[0667] The server trains a generative model, learning the character's speech patterns and language. The input is the training data prepared in step 3, and the output is the trained generative model. The server uses computing nodes to input data into the model and learns while adjusting parameters. A deep learning framework, for example, is suitable for use as the generative AI model.
[0668] Step 5:
[0669] The user enters their inquiry into a terminal and sends it to the server. The input is the user's inquiry text, and the output is raw data for analysis. The terminal uses a secure communication protocol to send the user's input to the server.
[0670] Step 6:
[0671] The server analyzes the user's inquiry using an emotion analysis engine and generates emotion data. The input is the inquiry text received in step 5, and the output is the emotion analysis result. The server utilizes natural language processing and emotion analysis algorithms to score the emotions within the text.
[0672] Step 7:
[0673] The server generates character-like responses using a generative model based on sentiment analysis results. The input is the analyzed sentiment data and the user's original query, and the output is the response text. The server inputs prompts to the generative AI model to create the optimal response, which reflects learned character expression patterns and the user's emotional needs.
[0674] Step 8:
[0675] The server performs post-processing on the generated response to emphasize character-specific elements. The input is the generated response text, and the output is the final response text. The server adds specific words or phrases or adjusts the style to emphasize the expression.
[0676] Step 9:
[0677] The server sends the final, post-processed response back to the user's terminal. The input is the post-processed response text, and the output is the data displayed on the user's terminal. The server sends the response data from the server to the terminal and presents the response in a formatted form for visualization by the user.
[0678] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0679] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0680] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0681] [Fourth Embodiment]
[0682] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0683] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0684] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0685] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0686] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0687] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0688] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0689] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0690] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0691] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0692] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0693] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0694] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0695] The present invention provides a platform for users to obtain responses that mimic the speech patterns of a specific character. The natural language processing functions of this system are described below.
[0696] First, the system is triggered when the user specifies their favorite character. The server receives this information and collects dialogue data from works and databases related to the specified character. The collected data is used as material to understand the character's speech patterns and characteristics. This data is preprocessed and converted into a format optimal for training a generative model.
[0697] The server then uses this preprocessed data to train a generative model, learning the language patterns and emotional expressions of the specified character. Once trained, the model retains the generated patterns and makes them available in response to actual queries.
[0698] Next, when a user makes a request, the device sends the content to the server. The server uses a generative model to generate a response to the user's request that has a character-like tone. At this stage, the response undergoes post-processing to reflect the character's characteristics and enable realistic interaction.
[0699] For example, if a user asks, "What products do you recommend?" and they are a fan of the dragon character, the system will generate a response like, "We have items that are as powerful as a dragon and will support you every day!" This allows the user to have an experience as if they were interacting with that character.
[0700] This system allows companies to personalize inquiry responses and increase user engagement. Because the system generates responses that include character-specific elements, users become more interested and more likely to actively engage in communication with the company.
[0701] The following describes the processing flow.
[0702] Step 1:
[0703] Users specify their favorite character via LINE and begin a conversation with a chatbot. The character information entered by the user acts as a trigger to retrieve related data about that character.
[0704] Step 2:
[0705] The server receives character information from users and collects dialogue data for those characters via the internet or an internal database. This collected data serves as material for learning the characters' language styles and characteristics.
[0706] Step 3:
[0707] The server preprocesses the collected dialogue data. This includes removing noise, normalizing the text, and tokenizing it. This preprocessing creates a dataset that allows the generative model to train more effectively.
[0708] Step 4:
[0709] The server uses pre-processed data to train a generative model. During this process, machine learning algorithms are used to teach the model the character's speech patterns and mannerisms. Once the model is trained, it is ready to generate responses from that character.
[0710] Step 5:
[0711] When a user enters their inquiry, the terminal sends that information to the server. This user input forms the basis for generating the response.
[0712] Step 6:
[0713] The server uses a generative model to generate character-like responses to user questions. These responses are designed to reflect the unique language patterns and emotional expressions of the characters.
[0714] Step 7:
[0715] The server performs post-processing on the generated response. This post-processing involves adding specific phrases or sentence endings to emphasize character-specific elements.
[0716] Step 8:
[0717] The server sends the final response to the device, allowing the user to view it on LINE. Seeing the displayed response, the user can experience what it's like to be having a conversation with their favorite character.
[0718] (Example 1)
[0719] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0720] There is a need to reproduce the unique speech patterns and language styles of characters, providing users with a more relatable and engaging communication experience. However, existing technologies have difficulty generating expressions that adequately mimic the characteristics of characters. This hinders increased user satisfaction and the promotion of deeper engagement.
[0721] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0722] In this invention, when a user specifies a preferred character, the server includes means for an information processing device to collect speech data of that character from a communication network or information aggregation device; means for initial processing the collected speech data and optimizing it as training data for a learning algorithm; and means for training the learning algorithm with the training data to enable it to understand the character's expression style and language structure. This makes it possible to provide realistic and personalized communication using the character's unique language expressions.
[0723] A "user" is a person or organization that uses this system with the aim of obtaining a response from a specific character.
[0724] An "information processing device" is a device that processes digital data and has the function of acquiring, processing, and transmitting specific information based on inquiries from users.
[0725] A "communication network" is the infrastructure used for sending and receiving data, and includes the internet and other network technologies.
[0726] An "information aggregation device" is any system or platform for collecting, storing, and utilizing large amounts of data as needed.
[0727] "Speech data" refers to data about a character's lines and linguistic expressions, and is text information that reflects the character's characteristics.
[0728] "Initial processing" refers to processes performed after data collection, such as data formatting, cleaning, and format conversion, which prepare the data for input into the AI model that generates it.
[0729] A "learning algorithm" is a technique that aims to learn patterns and characteristics by processing large amounts of data based on mathematical models, and includes machine learning techniques.
[0730] "Training data" refers to the dataset provided to a learning algorithm, which serves as the material for the algorithm to learn specific patterns and features.
[0731] "Style of expression" refers to the unique language style, tone of voice, and emotional expressions used by a character, and is an element that forms the character's individuality.
[0732] "Language structure" refers to the grammatical and structural characteristics of the language used by a character, encompassing elements related to the flow and structure of words.
[0733] The system of this invention starts operating when the user specifies a particular character. At this time, the user uses a character selection interface using a smartphone or computer application. Once the user selects a character, the specified information is transmitted to the server via the terminal.
[0734] Based on this information, the server collects speech data from databases associated with the specified character via the communication network. This process involves web scraping and API calls to obtain appropriate text data from video content and books in which the character appears.
[0735] The acquired speech data is initially processed by the server. This initial processing uses natural language processing libraries such as NLTK and spaCy to remove unnecessary symbols and normalize the text, followed by tokenization and part-of-speech tagging. This process ensures the data is in a format optimized for the learning algorithm.
[0736] Next, the server trains a generative AI model using the processed data. At this stage, a machine learning framework such as TensorFlow or PyTorch is used to build a model that learns the character's expressive style and language structure. The model reflects the character's characteristics and enables the generation of responses to user inquiries.
[0737] When a user enters a specific question or request using their device, that information is sent to the server. The server utilizes a trained generative AI model to generate character-like responses to the user's inquiries. For example, if a user enters "What products do you recommend?", the server generates a response that reflects the character's personality, such as "We have items that are as powerful as a dragon and will support your everyday life!" The generated response is then given a finishing touch to emphasize the character's characteristics before being sent back to the device.
[0738] An example of a prompt message is given to the generating AI model as, "Please respond to the question 'What products do you recommend?' in a dragon's voice." This system allows users to have a personalized conversational experience with a selected character, and is expected to improve customer satisfaction and engagement for companies and services.
[0739] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0740] Step 1:
[0741] The user uses their device to specify a particular character. As input, the user selects or enters the character name within the application's interface. This information, reflecting the user's preference, is sent to the server.
[0742] Step 2:
[0743] The server uses the received character information to collect relevant dialogue data from information aggregation devices via the communication network. As input, queries based on character names are executed, and the output is the character's dialogue data. Database access and API calls are performed to aggregate the dialogue data.
[0744] Step 3:
[0745] The server performs initial processing on the collected speech data. The input data is text information in a non-standardized format, and natural language processing libraries (such as NLTK and spaCy) are used to cleanse and normalize the text. The output is tokenized text data.
[0746] Step 4:
[0747] The server trains a generative AI model using an initialized dataset. The input is formatted character data, and tools such as TensorFlow and PyTorch are used. The model learns the character's representational style and language structure, and the output is a trained AI model.
[0748] Step 5:
[0749] When a user uses a terminal to ask and answer questions, the user inputs questions and requests from the terminal. Specifically, the user enters text through the interface, which is then sent to the server. As output, a character-like response based on this input is ready to be generated.
[0750] Step 6:
[0751] The server uses a character-like response model based on the user's question as input. The output is a response text that highlights the character's characteristics. The server then processes this response to emphasize the character's unique elements.
[0752] Step 7:
[0753] The terminal receives the completed response sent back from the server and displays it to the user. The input includes the response data from the server. This text is visualized on the terminal's screen, and a character-like dialogue is established as the final result provided to the user.
[0754] (Application Example 1)
[0755] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0756] Modern commercial platforms require providing personalized experiences in interactions with users to enhance customer engagement. However, typical systems struggle to provide personalized responses based on user interests and preferences, making it challenging to deliver individual user experiences, such as generating responses in the tone of a specific character.
[0757] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0758] In this invention, the server includes means for acquiring speech data of a character specified by the user from an information network or information storage device, means for preprocessing the acquired speech data and structuring it as training data for a generative model, and means for introducing products and information using the generated response. This makes it possible to introduce products and information in the tone of voice of the character set by the user.
[0759] "User" refers to a person who uses a system or application.
[0760] A "character" refers to a fictional person or creature that appears in a specific work or story, possessing its own unique language patterns and emotions.
[0761] An "information network" refers to the entire system for acquiring or transmitting information through networks such as the internet.
[0762] "Information storage device" refers to physical or virtual media used to store information using databases or other storage devices.
[0763] "Word data" refers to information expressed in text format, including data such as lines and dialogues spoken by specific characters.
[0764] "Acquisition" refers to the act of collecting data from external sources.
[0765] "Preprocessing" refers to the process of organizing acquired data into an appropriate format and processing it to a state suitable for analysis and training machine learning models.
[0766] A "generative model" refers to a machine learning algorithm that automatically generates new information or responses based on input data.
[0767] "Training data" refers to the dataset used to create a machine learning model; it is the fundamental data that the model learns from.
[0768] "Introducing products or information" refers to the act of explaining or suggesting specific content or products to users in detail.
[0769] A "generative AI model" refers to a model that uses artificial intelligence technology to generate new text or responses from data.
[0770] The program to implement this system primarily consists of data exchange between the server and the user's terminal, and natural language processing. The server retrieves the corresponding character's speech data from an information network or data storage device based on the character specified by the user. The retrieved data is preprocessed and structured as training data for a generative AI model. This process utilizes programming languages such as Python for cleaning and formatting text data, and natural language processing APIs (e.g., OpenAI's GPT-3).
[0771] The server trains a generative model based on pre-processed data, learning the character's tone of voice and language patterns. Based on inquiries from the user's terminal, the server utilizes the generative AI model to generate character-like responses. These responses are further post-processed to emphasize character-specific elements. The generated responses are sent to the user's terminal, providing a unique purchasing experience by introducing products and information.
[0772] For example, if a user inquires via smartphone, "I want a new laptop," the system will generate a response in the tone of their favorite character saying, "A laptop that opens the door to the future—unleash your creativity with this!" An example of a prompt to the generating AI model would be, "Please give me a recommendation about [topic] in the tone of the character's name."
[0773] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0774] Step 1:
[0775] The user selects a specific character on the terminal and enters their inquiry. The input data includes the character name set by the user and the specific inquiry (e.g., "I want a new laptop"). This triggers the system to begin generating a character-like response to the inquiry.
[0776] Step 2:
[0777] The terminal sends the user's selected character information and inquiry details to the server. At this time, the terminal sends the character identifier and inquiry details to the server as JSON data. The input data is received by the server and forms the basis for proceeding to the next step.
[0778] Step 3:
[0779] The server collects the corresponding character's speech data from the information storage device based on the character's identifier. The collected data is text data containing the character's lines and characteristic expressions. This data is then pre-processed.
[0780] Step 4:
[0781] The server preprocesses the collected word data and structures it as training data for a generative AI model. This preprocessing involves text cleaning (removing unnecessary information and converting it to a unified format). The preprocessed data is then in a format suitable for training the model.
[0782] Step 5:
[0783] The server trains a generative AI model using preprocessed data to learn character speech patterns and language styles. The input is structured word data, and the output is the learned model parameters, which are used in the next step.
[0784] Step 6:
[0785] The server inputs the user's inquiry into a generating AI model and generates a character-like response. This prompt is in the form of "Please give me a recommendation about XX in the tone of the character's name," and the input is transformed into a specific response, enabling dynamic dialogue.
[0786] Step 7:
[0787] The generated responses undergo post-processing on the server to emphasize character-specific elements. This post-processing adds specific tones and emotions to the lines, resulting in more realistic character responses.
[0788] Step 8:
[0789] Finally, the server sends a post-processed response back to the terminal. This response is displayed on the user's terminal, and the user receives personalized product and information recommendations from their favorite character. The output data allows the user to have a unique purchasing experience.
[0790] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0791] This invention is an interactive system that allows users to receive responses based on the tone of a specific character, while simultaneously recognizing the user's emotions and customizing the response content to suit those emotions. The operation of this system is described below in natural language.
[0792] First, the system activates when a user specifies their favorite character via LINE and initiates an inquiry. Based on the entered character information, the server collects dialogue data for that character from the internet or an internal database.
[0793] The collected dialogue data is preprocessed to understand the character's tone of voice and linguistic characteristics, and then structured as training data for a generative model. The server uses this data to train the generative model and teach it character-specific language patterns. At this stage, the generative model has the foundation to reproduce the character's tone of voice.
[0794] Next, when the user enters their inquiry, the terminal sends that information to the server. At this point, the system's built-in emotion engine is utilized. The server uses this emotion engine to recognize the emotions from the user's input text.
[0795] Based on the recognized emotion, the server uses a generative model to generate a character-like response corresponding to that emotion. This response reflects pre-trained character speech patterns and emotion-based expressions. The generated response is then post-processed to take emotion information into account, further emphasizing the character's unique emotional expressions.
[0796] For example, if a user types "I'm feeling a bit down today," the system will work to have the designated character respond with positive words to cheer them up. Let's say the generated response is "A fantastic adventure awaits you that will cheer you up!" In this way, users can enjoy more personalized interactions tailored to the character and their emotions.
[0797] This system can strengthen the relationship between companies and users and improve the usefulness of inquiries by simultaneously achieving realistic character reproduction and response generation adapted to the user's emotions.
[0798] The following describes the processing flow.
[0799] Step 1:
[0800] Users specify their favorite character on LINE and initiate an inquiry. This necessitates a conversational experience tailored to the user's preferences.
[0801] Step 2:
[0802] The server receives information about the specified character and collects dialogue data from works and databases in which that character appears. This includes the character's characteristic speech patterns and phrases.
[0803] Step 3:
[0804] The server preprocesses the collected dialogue data, performing tasks such as text normalization and tokenization, and structures it into a data format suitable for training generative models. This processing improves the accuracy of the learning process.
[0805] Step 4:
[0806] The server uses the pre-processed data to train the generative model. The model learns the characters' speech patterns and language styles, and is then ready to go.
[0807] Step 5:
[0808] When a user enters and submits an inquiry, the device transmits that data to the server. At that time, sentiment analysis of the user's input is performed.
[0809] Step 6:
[0810] The emotion engine installed on the server recognizes emotions from the user's input text. For example, basic emotions such as joy, sadness, and anger can be identified.
[0811] Step 7:
[0812] The server considers the recognized emotions and uses a generative model to generate character-like responses that correspond to those emotions. The responses not only reflect the character's tone of voice but also fit the user's emotions.
[0813] Step 8:
[0814] The server then performs post-processing on the generated response to enhance its emotional expression. This involves adding additional character-specific phrases and nuances depending on the specific emotion.
[0815] Step 9:
[0816] The server sends the completed character-like response to the user's device, allowing the user to view it on LINE. This allows users to enjoy a richer experience through dialogue that resonates with their own emotions.
[0817] (Example 2)
[0818] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0819] Conventional dialogue systems have problems accurately reproducing the tone of voice of a character specified by the user, and are unable to generate flexible responses that respond to the user's emotions. This limits the quality of the dialogue provided by the system and hinders the improvement of the user experience. In addition, the lack of processing to effectively emphasize character-specific expressions makes it difficult to provide responses that fully utilize the characteristics of the character.
[0820] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0821] In this invention, the server includes means for acquiring language data of a symbol specified by the user from a network or information recording device, means for preprocessing the acquired language data and structuring it as training information for a generation algorithm, and means for training the generation algorithm with the training information to learn the language features and expression patterns of the symbol. This makes it possible to faithfully reproduce the tone of voice of a character specified by the user and generate appropriate responses that correspond to the user's emotions.
[0822] A "symbol" is an element designated by the user to represent a specific character or its characteristics.
[0823] A "network" is a communication infrastructure used to transmit digital information and connect multiple devices and systems.
[0824] An "information recording device" is a technology or platform for storing data and accessing it as needed.
[0825] "Linguistic data" refers to information about texts and speeches, including the tone of voice and linguistic characteristics of a specific character.
[0826] "Preprocessing" is the process of normalizing acquired data, removing noise, and preparing it for training generative algorithms.
[0827] A "generative algorithm" is a computational method for generating new data or responses based on presented information.
[0828] "Training information" refers to a structured dataset used by a generative algorithm for learning.
[0829] "Linguistic features" refer to the unique expressive styles and grammatical characteristics of the language spoken by a particular character or individual.
[0830] A "pattern of expression" is a characteristic that indicates a specific form or flow in language or communication.
[0831] "Emotion analysis" is a technology that recognizes a user's emotional state from text or audio.
[0832] "Control variables" are parameters that are adjusted to optimize the performance of a generation algorithm.
[0833] The specific implementation methods for this system are described below.
[0834] First, the user uses a network-enabled device to specify their preferred character and access the system via a messaging service such as LINE. Based on the specified character, the server collects language data related to that character. This data is stored on the network or in a database. The collected language data is then structured for processing by the server.
[0835] The server runs software developed using programming languages such as Python and Java to preprocess the collected language data. This process includes data normalization and noise reduction, preparing the data for use as training data for a generative AI model. Next, the server uses natural language processing libraries to train the generative AI model with this training data. In this way, the model learns the language features and expression patterns of a specified character.
[0836] Subsequently, when a user submits a query to the system, the terminal sends the query to the server. Upon receiving the query, the server uses natural language processing technology to analyze the text sent by the user and recognize its sentiment. Sentiment analysis is performed using sentiment analysis libraries and APIs.
[0837] When the generative AI model generates responses based on the trained data, the server generates character-like responses that take the user's emotions into account, based on the analyzed emotions. Finally, the generated responses undergo further post-processing to enhance character-specific expressions. This process generates emotionally natural responses that are sent to the user's device.
[0838] For example, if a user sends a message such as "I'm a little tired today," the system will generate a response that resonates with the user, while maintaining the character's tone of voice. "For instance, it might respond with something like, 'Please rest well today!' In this way, the character provides the user with an emotionally resonant dialogue."
[0839] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0840] Step 1:
[0841] A user initiates a query by specifying a particular character through the LINE application. The input is the character's name, and this information is sent to the server as output. Based on the received character name, the server collects relevant language data using the network or database. Specifically, this involves issuing API queries to the character or performing database searches.
[0842] Step 2:
[0843] The server preprocesses the collected language data. At this stage, raw language data is provided as input, and preprocessed data is generated as output. Data processing is performed, such as normalization, removal of unnecessary information, and identification of frequently occurring words and phrases. Specifically, text tokenization and morphological analysis are performed to build a training dataset.
[0844] Step 3:
[0845] The server uses preprocessed data to train a generative AI model. It receives preprocessed data as input and generates a trained model as output. Here, the model uses a machine learning algorithm to learn the language patterns of characters. Specifically, the model's parameters are updated using a deep learning framework.
[0846] Step 4:
[0847] The user enters their inquiry into the system and sends it to the server via their terminal. The input includes the inquiry text, and this information is passed to the server as output. Specifically, the text data entered on the user's terminal is encrypted and securely transmitted to the server.
[0848] Step 5:
[0849] The server analyzes the received inquiry and uses an emotion analysis engine to evaluate the user's emotions. The user's inquiry is the input for analysis, and the recognized emotion information is the output. The server extracts emotions such as positive, negative, and neutral from the text. Specifically, an emotion score is calculated based on an analysis of keywords and phrases.
[0850] Step 6:
[0851] The server uses a generative AI model to generate character-like responses corresponding to the recognized emotions. Emotional information and query text are used as input, and response text is generated as output. The generated response reflects the character's tone and emotions. For example, when emphasizing positive emotions, phrases expressing encouragement and joy are included.
[0852] Step 7:
[0853] The server performs post-processing on the generated response to emphasize character-specific expressions. The input is a generated response, and the output is an even more character-driven response. Specifically, the post-processing involves adjustments to reflect the character's unique tone of voice and expression patterns.
[0854] Step 8:
[0855] The server sends the final response to the user's terminal. The server has the post-processed response text as input, and it is displayed on the user's terminal as output. Specifically, the response text is delivered to the user in real time over the network.
[0856] (Application Example 2)
[0857] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0858] Modern virtual interactions require effectively providing users with personalized experiences. However, current systems struggle to reproduce a character's tone of voice while simultaneously generating responses that adapt to the user's emotions. To solve this problem, a system is needed that faithfully reproduces a character's tone of voice while generating responses that align with the user's feelings.
[0859] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0860] In this invention, the server includes means for analyzing the user's emotional state, means for generating a character-like response based on the user's emotional state using a generative model, and means for post-processing the generated response to emphasize character-specific elements. This makes it possible to generate a personalized response that is appropriate to the user's emotions while faithfully reproducing the tone of voice of the character specified by the user.
[0861] A "user" is the entity that operates the system and enjoys interacting with a specific character.
[0862] A "character" is an entity specified by the user that generates responses based on its tone of voice and language patterns.
[0863] A "network" is a communication infrastructure used to acquire and transmit information.
[0864] A "recording medium" is hardware used to store and retain data and information.
[0865] "Dialogue information" refers to data related to a character's language expression and tone of voice.
[0866] "Preprocessing" refers to the process of formally preparing acquired data for use in a generative model.
[0867] A "generative model" is an algorithm that generates responses that reproduce a character's tone of voice and writing style through learning.
[0868] "Training information" refers to the dataset used to train a generative model.
[0869] "Emotional state" refers to information that indicates the user's psychological state or mood.
[0870] "Analysis" is the process of interpreting data and information and extracting useful insights.
[0871] "Post-processing" refers to the process of further adjusting the generated response and emphasizing the tone and characteristics of the specified character.
[0872] A "device" refers to hardware used to perform interactions with users.
[0873] This system is an interactive system that reproduces the speech patterns of specific characters and generates responses that match the user's emotional state. The specific implementation method is described below.
[0874] First, the user selects a specific character through their device and begins interacting with that character. Suitable devices include smartphones and smart glasses. The selected character's dialogue information is retrieved by a server using a network or recording medium. The server preprocesses this dialogue information and structures it as training data for a generative AI model. During the preprocessing stage, the dialogue information is formatted to allow the generative model to easily learn from it.
[0875] Next, the server uses a generative model to learn the character's tone of voice and language patterns from the training data. For example, an open-source generative AI model could be used.
[0876] When a user enters an inquiry from their device, the server uses an emotion analysis engine to analyze the user's emotions based on that inquiry. This analysis can utilize a cloud-based emotion recognition service, and based on the results, a generative model generates a character-like response that is adapted to the user's emotional state.
[0877] The generated responses are then post-processed to further emphasize character-specific elements. This post-processing ensures that the responses more clearly reflect the character's unique phrasing and emotional expressions.
[0878] For example, if the user types "I want to do something fun today," the character will generate a response such as "A wonderful day awaits you, what would you like to do?" and an encouraging dialogue will unfold. An example of a corresponding prompt would be: "You are an assistant that analyzes the user's emotions and faithfully reproduces the tone of the character the user likes to continue the conversation. When the user says, 'I want a new jacket, but I'm not sure,' what would your response be?"
[0879] This configuration enables personalized intef interactions for users.
[0880] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0881] Step 1:
[0882] The user selects a specific character via the terminal. The input is the user's character selection data, which the terminal sends to the server. The terminal constructs the appropriate request to transmit the user's selection to the server over the network and sends the input information through the interface.
[0883] Step 2:
[0884] The server retrieves dialogue information related to the selected character via a network or storage medium. The input is the character's identification information, and the output is the dialogue information as raw data. In this process, the server collects necessary information from external resources or an internal database via an API and stores it in a cache.
[0885] Step 3:
[0886] The server preprocesses the acquired dialogue information and structures it as training data for a generative AI model. The input is raw dialogue data, and the output is in a format suitable for training. Specifically, the server prepares the data by normalizing character codes and removing meaningless strings. This process also includes data cleaning and tokenization techniques.
[0887] Step 4:
[0888] The server trains a generative model, learning the character's speech patterns and language. The input is the training data prepared in step 3, and the output is the trained generative model. The server uses computing nodes to input data into the model and learns while adjusting parameters. A deep learning framework, for example, is suitable for use as the generative AI model.
[0889] Step 5:
[0890] The user enters their inquiry into a terminal and sends it to the server. The input is the user's inquiry text, and the output is raw data for analysis. The terminal uses a secure communication protocol to send the user's input to the server.
[0891] Step 6:
[0892] The server analyzes the user's inquiry using an emotion analysis engine and generates emotion data. The input is the inquiry text received in step 5, and the output is the emotion analysis result. The server utilizes natural language processing and emotion analysis algorithms to score the emotions within the text.
[0893] Step 7:
[0894] The server generates character-like responses using a generative model based on sentiment analysis results. The input is the analyzed sentiment data and the user's original query, and the output is the response text. The server inputs prompts to the generative AI model to create the optimal response, which reflects learned character expression patterns and the user's emotional needs.
[0895] Step 8:
[0896] The server performs post-processing on the generated response to emphasize character-specific elements. The input is the generated response text, and the output is the final response text. The server adds specific words or phrases or adjusts the style to emphasize the expression.
[0897] Step 9:
[0898] The server sends the final, post-processed response back to the user's terminal. The input is the post-processed response text, and the output is the data displayed on the user's terminal. The server sends the response data from the server to the terminal and presents the response in a formatted form for visualization by the user.
[0899] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0900] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0901] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0902] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0903] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0904] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0905] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0906] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0907] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0908] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0909] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0910] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0911] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0912] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0913] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0914] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0915] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0916] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0917] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0918] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0919] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0920] The following is further disclosed regarding the embodiments described above.
[0921] (Claim 1)
[0922] A means of obtaining dialogue data for a character specified by the user from the internet or a database,
[0923] A means of preprocessing the acquired dialogue data and structuring it as training data for a generative model,
[0924] A means for training a generative model with the aforementioned training data to learn the character's tone of voice and language patterns,
[0925] A means for generating character-like responses using the aforementioned generation model in response to inquiries from users,
[0926] A means for post-processing the generated response to emphasize character-specific elements,
[0927] A means for sending the post-processed response back to the user's terminal,
[0928] A system that includes this.
[0929] (Claim 2)
[0930] The system according to claim 1, comprising normalization of pre-processed dialogue data.
[0931] (Claim 3)
[0932] The system according to claim 1, comprising means for evaluating the performance of a generative model and adjusting hyperparameters as necessary.
[0933] "Example 1"
[0934] (Claim 1)
[0935] When a user specifies their preferred character, the information processing device has means to collect the character's speech data from a communication network or information aggregation device.
[0936] A means for initial processing the collected speech data and optimizing it as training data for a learning algorithm,
[0937] A means for training a learning algorithm with the aforementioned training data to enable it to understand the character's expression style and language structure,
[0938] A means for generating character-like responses using the aforementioned learning algorithm in response to information provided by the user,
[0939] A means of performing finishing touches on the generated response to emphasize character-specific elements,
[0940] A means for sending the completed response back to the user's information terminal,
[0941] A system that includes this.
[0942] (Claim 2)
[0943] The system according to claim 1, comprising normalization of initially processed speech data.
[0944] (Claim 3)
[0945] The system according to claim 1, comprising means for measuring the efficiency of a learning algorithm and adjusting control parameters as necessary.
[0946] "Application Example 1"
[0947] (Claim 1)
[0948] A means for obtaining the character's speech data from an information network or information storage device based on the character specified by the user,
[0949] A means of preprocessing acquired word data and structuring it as training data for a generative model,
[0950] A means for training a generative model with the aforementioned training data to learn the character's tone of voice and language patterns,
[0951] A means for generating character-like responses using the aforementioned generation model in response to inquiries from users,
[0952] A means for post-processing the generated response to emphasize character-specific elements,
[0953] Means for sending the post-processed response back to the user's electronic device,
[0954] A means of introducing products and information using the generated response,
[0955] A means of promoting purchases based on the generation of character-like responses using a generative AI model for coded inputs,
[0956] A system that includes this.
[0957] (Claim 2)
[0958] The system according to claim 1, comprising the standardization of preprocessed word data.
[0959] (Claim 3)
[0960] The system according to claim 1, comprising means for evaluating the performance of a generative model and adjusting control parameters as necessary.
[0961] "Example 2 of combining an emotion engine"
[0962] (Claim 1)
[0963] A means for obtaining language data for a symbol specified by a user from a network or information recording device,
[0964] A means for preprocessing acquired language data and structuring it as training information for a generative algorithm,
[0965] A means for training a generation algorithm with the aforementioned training information to learn the linguistic features and representation patterns of symbols,
[0966] A means for generating a symbolic response using the aforementioned generation algorithm in response to an inquiry from a user,
[0967] A means for post-processing the generated response to highlight symbol-specific attributes,
[0968] A means of emotion analysis that recognizes emotions from input text,
[0969] A means of generating a response corresponding to recognized emotions,
[0970] A means for sending the post-processed response back to the user's information terminal,
[0971] A system that includes this.
[0972] (Claim 2)
[0973] The system according to claim 1, comprising normalization of preprocessed language data.
[0974] (Claim 3)
[0975] The system according to claim 1, further comprising means for evaluating the performance of the generation algorithm and adjusting control variables as necessary.
[0976] "Application example 2 when combining with an emotional engine"
[0977] (Claim 1)
[0978] A means for obtaining dialogue information of a character specified by the user from a network or recording medium,
[0979] A means for preprocessing acquired dialogue information and structuring it as training information for a generative model,
[0980] A means for training a generative model with the aforementioned training information to learn the character's tone of voice and language patterns,
[0981] A means of analyzing the emotional state of users,
[0982] A means for generating character-like responses based on the user's emotional state using the aforementioned generation model,
[0983] A means for post-processing the generated response to emphasize character-specific elements,
[0984] Means for sending the post-processed response back to the user's device,
[0985] A system that includes this.
[0986] (Claim 2)
[0987] The system according to claim 1, comprising normalization of pre-processed dialogue information.
[0988] (Claim 3)
[0989] The system according to claim 1, comprising means for evaluating the performance of the generative model and adjusting hyperparameters as necessary. [Explanation of Symbols]
[0990] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining dialogue data for a character specified by the user from the internet or a database, A means of preprocessing the acquired dialogue data and structuring it as training data for a generative model, A means for training a generative model with the aforementioned training data to learn the character's tone of voice and language patterns, A means for generating character-like responses using the aforementioned generation model in response to inquiries from users, A means for post-processing the generated response to emphasize character-specific elements, A means for sending the post-processed response back to the user's terminal, A system that includes this.
2. The system according to claim 1, comprising normalization of pre-processed dialogue data.
3. The system according to claim 1, comprising means for evaluating the performance of a generative model and adjusting hyperparameters as necessary.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A