system
The information processing system addresses the challenge of creating realistic and personalized virtual character interactions by training models with user data and real-time information, ensuring emotionally responsive and timely conversations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-09
- Publication Date
- 2026-06-19
AI Technical Summary
Existing systems struggle to provide realistic and personalized interactions with virtual characters, requiring significant time and resources, and fail to incorporate real-time and time-dependent information in conversations.
An information processing system that collects user preferences, trains a model to mimic a virtual character's personality and speech patterns, updates with real-time information, and generates personalized and time-sensitive responses using generative AI and emotion recognition.
Enables natural, personalized, and emotionally responsive conversations with virtual characters, providing a realistic experience akin to interacting with a real person.
Smart Images

Figure 2026100609000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the real world, it is difficult to realize an individual and realistic conversation with a specific virtual character or idol, and it takes a large amount of time and cost to execute this. Therefore, it is required to provide a means for ordinary users to interact with these characters. Also, it is one of the problems to provide communication that reflects the latest information and time-dependent situations to the user as if it were a human through a virtual character.
Means for Solving the Problems
[0005] The present invention provides an information processing system for realizing natural language communication with a specific virtual character. Specifically, it solves the above problems by providing means for collecting and storing data on the user's preferences and interests, means for training a model using information processing technology to imitate the personality and speaking style of a designated character, means for updating the model based on the latest information obtained from external information sources, means for analyzing input from the user and generating an appropriate response using the trained and updated model, and means for generating an automatic message at an appropriate time using real-world time information. This provides natural and personalized conversations between the character and the user, and enables the user to have an experience as if they were interacting with a real idol or character.
[0006] An "information processing system" is a combination of hardware and software that has the functions of collecting, storing, analyzing, and generating data in order to achieve a specific purpose.
[0007] A "user" is a person who uses this system to communicate with a specific virtual character.
[0008] A "virtual character" is a person or character that does not exist in reality but is represented as if it exists through computer generation and control.
[0009] "Natural language" refers to the language that humans use on a daily basis, including text and audio input by specific users.
[0010] "Data relating to preferences and interests" refers to information that represents the content of a user's interests and concerns regarding a particular subject.
[0011] "Information processing technology" refers to technical methods and algorithms for effectively collecting, analyzing, generating, and managing data.
[0012] "Training a model" is the process of preparing data using information processing technology to mimic the personality and speech patterns of a virtual character, and then building an algorithm that generates the character's responses based on that data.
[0013] "External information sources" refer to means of providing information from third parties, such as news feeds on the internet, APIs, and databases.
[0014] "Generating a response" is the process of creating character dialogue or messages that the system returns as output, based on input from the user.
[0015] An "automated message" is a pre-programmed text or voice message that is generated by a system according to set conditions and sent to the user. [Brief explanation of the drawing]
[0016] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map on which multiple emotions are mapped. [Figure 10] It shows an emotion map on which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0021] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention constructs an information processing system for enabling a user to communicate with a specific virtual character using natural language. This system consists of a server, a terminal, and a user, and each component works together to facilitate interaction with the virtual character.
[0038] First, the user initiates interaction with the system using a terminal. The terminal provides an interface for inputting data about the user's preferences and interests. This information is sent to the server as foundational data for creating the user's profile.
[0039] The server updates its database based on the received user data and generates a user profile. This profile is used to provide personalized character responses. The server also trains a model that mimics the personality and speech patterns of a specified virtual character using AI technology. Publicly available information and historical data about the character are used for training.
[0040] The server also retrieves the latest news and trend information from external sources, keeping the character's knowledge base constantly up-to-date. This latest information is reflected in user conversations in real time.
[0041] When a user sends a message to a virtual character through their device, the device forwards the input to a server. The server analyzes this input and generates an appropriate response based on the user profile, a trained model, and the latest training data. The generated response is then sent to the device and displayed to the user.
[0042] For example, if a user asks a character, "What's the weather like now?", the server retrieves weather information from an external source and generates a response in a style that matches the character's tone, such as, "Today's weather is sunny, and the temperature is 20 degrees Celsius." In this way, users can have an experience that feels like they are having a conversation with a real idol or character.
[0043] Furthermore, the server uses real-world time information to generate automated messages from the character at appropriate times and send them to the user through the device. For example, by sending a message like "Good morning, let's do our best today!" in the morning, it enables intimate communication that is linked to the time of day.
[0044] This type of system allows users to have a personalized and enjoyable experience, making them feel as if their virtual character is actually real.
[0045] The following describes the processing flow.
[0046] Step 1:
[0047] The user uses their device to input information about their preferences and interests, including their favorite virtual characters and topics of interest. The device then sends this information to the server.
[0048] Step 2:
[0049] The server generates a user profile based on the user data it receives. This profile is stored in a database and used in subsequent conversation generation processes.
[0050] Step 3:
[0051] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. The training utilizes publicly available information and historical materials related to the character.
[0052] Step 4:
[0053] The server periodically retrieves data from external sources to obtain the latest news and trend information. Based on this information, it updates the model to ensure that the information users receive is fresh and accurate.
[0054] Step 5:
[0055] The user sends a message to a virtual character via a terminal. The terminal forwards this user input to the server.
[0056] Step 6:
[0057] The server analyzes the user input it receives. It generates an appropriate response by referring to the user profile, the latest training data, and the trained model.
[0058] Step 7:
[0059] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to interact with the virtual character in real time.
[0060] Step 8:
[0061] The server references real-time information to automatically generate messages at the appropriate time. These messages are sent to the user via the terminal, providing personalized communication tailored to the time.
[0062] (Example 1)
[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0064] In recent years, communication technologies with virtual characters have become widely used, but conventional systems have faced challenges in achieving realistic dialogue and personalized responses. Furthermore, real-time dialogue that reflects time and external information is difficult. Current technology requires the ability to provide individualized responses based on user profiles and to achieve natural dialogue that reflects external information.
[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0066] In this invention, the server includes means for collecting data from users via a terminal and transmitting that data to the server; means for updating a database based on the data received by the server and generating individual user profiles; means for training a model using generative AI technology to mimic the personality and conversational style of a virtual character based on the generated user profiles; means for obtaining current news and weather from external sources and reflecting them in the character's knowledge base; means for analyzing messages sent by users and generating responses based on the user profile and the latest information; and means for sending time-sensitive automated messages to users using real-time time information. This enables users to enjoy real-time, personalized, and natural conversations with virtual characters.
[0067] A "terminal" is a device used by users to input information and interact with virtual characters.
[0068] A "server" is a central device that processes data received from terminals and generates user profiles and responses.
[0069] "Generative AI technology" is an artificial intelligence technology that learns the personality and conversational style of virtual characters and imitates them.
[0070] A "database" is an information system used to store and manage collected user data and profile information.
[0071] A "user profile" is a unique set of information that reflects a user's interests, preferences, and past interactions.
[0072] "External information sources" refer to information providers that are obtained from outside the system, such as news and weather information.
[0073] A "knowledge base" is a collection of information that a virtual character holds for use in interacting with the user.
[0074] An "automatic message" is a message that a virtual character sends at a specific time or period based on pre-set conditions.
[0075] The information processing system of this invention enables a user to communicate with a specific virtual character using natural language. This system consists of a server, a terminal, and a user. Each component works together to realize the interaction function with the virtual character.
[0076] Hardware and software usage
[0077] 1. Terminal:
[0078] Users initiate interaction with virtual characters using devices such as smartphones, tablets, or personal computers. These devices collect information about the user's preferences and interests and transmit it to the server. For this purpose, dedicated applications or web interfaces are provided on the devices.
[0079] 2. Server:
[0080] The server is implemented using programming languages such as Python and Java (registered trademark), and receives and analyzes data sent from the terminal. Based on the received data, the server generates a user profile. This profile forms the basis for personalized interactions.
[0081] The server also uses generative AI technology to train models for generating the personalities and conversational styles of virtual characters. This process utilizes machine learning frameworks such as TENSORFLOW® and PyTorch.
[0082] The server retrieves the latest information from external sources such as OpenWeatherMap and news APIs to keep the virtual character's knowledge base up to date.
[0083] Specific example
[0084] When a user asks a virtual character "What's the latest news?" using their device, the server retrieves the latest news information via the NewsAPI, generates a response that matches the virtual character's speaking style, such as "The latest news is that a new museum has opened in the city," and sends it to the device. In this way, the user can have a natural experience as if they were having a conversation with a real person.
[0085] Example of a prompt
[0086] "Please generate a response for when a user asks the character, 'What events do you recommend this weekend?' The user profile should indicate that the character enjoys the outdoors, and the character should use a friendly tone."
[0087] This system provides users with a real-time, personalized, and enjoyable experience, making them feel as if their virtual character is a close friend or acquaintance.
[0088] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0089] Step 1:
[0090] The user initiates interaction with the system using a terminal. The terminal displays an interface for entering data about the user's preferences and interests. This input data includes information about the user's selected areas of interest and hobbies. The entered information is collected by the terminal and sent directly to the server.
[0091] Step 2:
[0092] The server receives user data sent from the terminal and analyzes it. The server accesses the database and generates individual user profiles based on the newly received data. Data processing includes string analysis and data normalization, resulting in updated user profiles as output.
[0093] Step 3:
[0094] The server uses generative AI technology to train the personality and speech patterns of virtual characters. User profiles and official virtual character documentation are used as input data. Based on this data, the AI model is trained to mimic the virtual character's personality. The output is a highly personalized AI model.
[0095] Step 4:
[0096] The server accesses external news and weather APIs to retrieve the latest information. This retrieved information is stored in the virtual character's knowledge base and used in conversations. Specifically, JSON data from external data sources is parsed, the necessary information is extracted, and added to the server's internal database.
[0097] Step 5:
[0098] When a user asks a question to a virtual character via their device, the message is forwarded to the server. The server analyzes this input and generates the optimal response based on the user profile, a trained model, and external information. Text analysis and natural language generation are performed, and the resulting response message to the user is output.
[0099] Step 6:
[0100] The terminal receives a response from the server and displays it to the user. This process visualizes the received text data on the terminal's user interface, allowing the user to continue interacting with the virtual character.
[0101] Step 7:
[0102] The server generates time-sensitive automated messages using real-world time information and sends them to the terminal periodically. Pre-set messages are generated and sent to the user based on the time and date. This allows the user to receive continuous feedback from the virtual character.
[0103] (Application Example 1)
[0104] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0105] In natural language dialogue systems with virtual entities, there is a need to provide personalized responses tailored to each user's preferences while ensuring a unified interactive experience across different devices. Furthermore, it is necessary to update the virtual entity's knowledge in real time using the latest data obtained from external sources, enabling smooth and intimate communication with users.
[0106] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0107] In this invention, the server includes means for collecting and storing data relating to the user's preferences and interests, means for training a computational model using information processing technology to mimic the characteristics and speech patterns of a designated virtual entity, and means for improving the computational model based on the latest data obtained from external sources. This enables personalized content for the user and real-time interaction that reflects the latest information.
[0108] A "virtual entity" is a digital entity that possesses the characteristics and speech patterns of a human or character, created using information processing technology.
[0109] "Natural language" refers to the forms of language that humans use on a daily basis, including texts and conversations processed by information processing devices.
[0110] An "information processing device" is a computing device that manipulates digital data, and through its functions, it collects, analyzes, stores, and displays data.
[0111] "Preferences" refer to data that shows the individual tastes and preferences of users, and are the basis for information processing systems to generate personalized responses.
[0112] A "computational model" is a mathematical model trained using information processing technology, and it is responsible for the process of generating an appropriate response to an input.
[0113] "External information sources" are various data providers that exist outside the system and are used as input to the system, providing information such as news and weather.
[0114] "Personalization" refers to adjusting content to suit the specific needs and preferences of each user, and is the process of optimizing the responses and content provided by information processing devices for each user.
[0115] An "interactive experience" is a user experience that is realized through real-time information exchange and mutual influence between the user and the system.
[0116] This system consists of client terminals owned by the user and servers operating in the cloud. Users can initiate interactions with virtual entities using terminals such as smartphones or smart glasses. The terminals provide an interface for inputting data about the user's preferences and interests, and transmit this data to the server. The server stores and manages the received data in a database and uses it to update the computational model.
[0117] The server utilizes a high-performance cloud platform (e.g., Google Cloud Platform or AWS) to train an AI model that mimics the characteristics and speech patterns of a virtual entity. The trained model generates appropriate responses based on user input. Natural language processing libraries such as SpaCy and BERT are used to train the AI model.
[0118] Furthermore, the server retrieves the latest data (such as news and weather information) from external sources and updates the computational model in real time. This up-to-date information is immediately utilized in conversations with virtual entities, providing realistic and timely answers to user questions.
[0119] For example, if a user asks, "When is the next live event?", the server uses an AI model to generate a response such as, "The next live event is on October 10th! Please book your tickets early if you plan to attend!" In this way, users can enjoy an interactive experience with a virtual entity.
[0120] An example of a prompt for a generative AI model is, "When is the next live event?". In response to this prompt, the model will generate an answer that matches the tone and speech patterns of the specified character.
[0121] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0122] Step 1:
[0123] The user accesses the system using a client terminal and begins interacting with a virtual entity. The terminal interface displays a screen for inputting data about the user's preferences and interests. This input data is processed on the terminal and then sent to the server.
[0124] Step 2:
[0125] The server receives user preference data from the terminal and stores it in a cloud-based database. This data is saved as a user profile, which serves as the basis for generating personalized responses in subsequent interactions.
[0126] Step 3:
[0127] The server uses natural language processing libraries (e.g., SpaCy or BERT) to update and train the generation AI model based on the received user data. The AI model reflects the characteristics and speech patterns of the specified virtual entity, improving the accuracy of response generation according to user preferences. Model parameters are adjusted during this process.
[0128] Step 4:
[0129] The server accesses external databases via APIs to collect the latest news and weather information from external sources. This data is input into the AI model in real time, constantly updating the knowledge base of the virtual entity.
[0130] Step 5:
[0131] The user sends a message to a virtual entity through a terminal. The terminal receives the user's input message and sends it to the server. This message becomes input data for natural language processing.
[0132] Step 6:
[0133] The server uses an AI model to analyze messages received from users. Based on the analysis results, it generates an appropriate response, referencing the user profile and the latest external data. The generated response is expressed in a tone and style appropriate to the virtual character.
[0134] Step 7:
[0135] The server sends the generated response to the terminal. The terminal displays this response to the user, and the interactive dialogue with the virtual entity is completed.
[0136] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0137] This invention provides a more personalized experience in an information processing system where a user communicates with a specific virtual character using natural language, by combining it with an emotion engine that recognizes the user's emotions. This system consists of a server, a terminal, and a user, and each component works in an integrated manner to realize emotion-responsive communication.
[0138] First, the user operates an input interface through a terminal, providing the system with their preferences, interests, and daily emotions. The terminal formats this information and sends it to the server. The server then generates a user profile based on this information and stores it in a database.
[0139] The server uses generative AI technology to train a model that mimics the personality and speech patterns of a specified virtual character, and establishes a mechanism to recognize emotions from user input using an emotion engine. The emotion engine analyzes the content of the input text and voice to identify emotional states such as positive, negative, and neutral.
[0140] Furthermore, the server retrieves the latest news and trend information from external sources to keep the model up-to-date. This updated information, combined with sentiment recognition, provides personalized responses to the user.
[0141] When a user sends a message to a virtual character via their device, the device forwards the input to the server. The server analyzes the received user input and generates an appropriate response based on the user profile, the latest training data, and the trained model. In doing so, it takes into account the emotional state analyzed by the emotion engine and adjusts the tone and content of the response.
[0142] For example, if a user inputs "I'm tired today," the emotion engine recognizes this input as a negative emotion. Based on this emotion information, the server has a virtual character choose encouraging words such as "You had a tough day, well done," to provide a response that is empathetic to the user.
[0143] The server takes real-world time information into account and generates emotionally appropriate automated messages at the right time, sending them to the user through the terminal. For example, by sending a message like "You've had a long day, good night" late at night, it enables intimate communication tailored to the time of day.
[0144] In this way, this system, which combines an emotion engine, allows users to enjoy a more intimate and human-like conversational experience with virtual characters.
[0145] The following describes the processing flow.
[0146] Step 1:
[0147] Users use their device to input their preferences, interests, and daily emotions. This includes their favorite characters, topics of interest, and recent emotional states. The device formats this information and sends it to the server.
[0148] Step 2:
[0149] The server generates a user profile based on the received user data. This profile is stored in a database and, including emotional information, is used in future conversation generation.
[0150] Step 3:
[0151] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. This process utilizes publicly available information and historical data about the character.
[0152] Step 4:
[0153] The server retrieves the latest news and trend information from external sources. This keeps the character's knowledge base up-to-date, allowing it to provide users with realistic information.
[0154] Step 5:
[0155] The user sends a message to a virtual character via a terminal. The message can be entered in text or voice format. The terminal then forwards this user input to the server.
[0156] Step 6:
[0157] The server analyzes the user's input and uses an emotion engine to recognize emotions from the input. The emotion engine identifies states such as positive, negative, and neutral.
[0158] Step 7:
[0159] The server generates an appropriate response based on the user profile, the latest training data, and the trained model, taking into account the recognized emotional state. The tone and content of the response are adjusted according to the emotion.
[0160] Step 8:
[0161] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to experience real-time interaction with the virtual character.
[0162] Step 9:
[0163] The server references real-world time information to generate automated messages that respond to emotions at the appropriate time. These messages are sent to the user via the terminal, providing intimate communication tailored to the time.
[0164] (Example 2)
[0165] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0166] In modern information processing systems, for users to engage in more intimate and personalized interactions with virtual characters, responses that take into account the user's emotions and temporal information are necessary. However, conventional systems have the challenge of not being able to generate responses that fully utilize emotion recognition and temporal information, and thus failing to sufficiently improve user satisfaction.
[0167] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0168] In this invention, the server includes means for collecting and storing data on the user's preferences and interests, means for training a model using generative AI technology, and means for identifying emotional states using an emotion engine. This enables the generation of personalized responses that take into account the user's emotions and temporal information.
[0169] A "user" is an individual who uses an information processing system to communicate with a virtual character.
[0170] "Preference and interest data" refers to information about a user's interests and preferences that they provide to the system.
[0171] "Generative AI technology" is a method that uses artificial intelligence to enable virtual characters to generate conversations, and is a technology used for training models.
[0172] "Methods for training a model" refers to a series of methods for training and updating a model using generative AI technology to mimic the personality and speech patterns of a virtual character.
[0173] An "emotion engine" refers to a technology or program that analyzes user input and identifies their emotional state based on that input.
[0174] "Emotional state" refers to the result of analyzing and classifying the emotional responses exhibited by users, and can be divided into categories such as positive, negative, and neutral.
[0175] "External information sources" are sources of information that are outside the system and provide up-to-date information and trends that are useful for generating responses from virtual characters.
[0176] "Personalized responses" refer to virtual character responses tailored to a specific user, customized based on the user's profile and emotional state.
[0177] "Real-world time information" refers to data that takes into account temporal factors such as actual time and date, and is used to enable responses at the appropriate time.
[0178] An "automated message" is a message that a system automatically generates and sends, taking into account the time and the user's emotional state.
[0179] This invention is an information processing system for users to communicate with a specific virtual character in natural language. The system consists of a server, a terminal, and a user, and each component works in an integrated manner to achieve personalized communication that responds to the user's emotions.
[0180] The user first interacts with the interface through a device, inputting their preferences, interests, and daily emotions. This information is then formatted and sent from the device to the server. The device can be a typical smartphone or computer, and input can utilize touchscreens or voice recognition.
[0181] The server generates user profiles based on the received data and stores them in a database. Here, it utilizes an emotion engine to recognize emotions from user input. The emotion engine uses natural language processing techniques to analyze text and audio content and identify emotional states such as positive, negative, and neutral. Because this entire process requires high-performance computing resources, the server is equipped with the latest processors and large-capacity memory.
[0182] Furthermore, the server retrieves the latest news and trend information from external sources and uses generative AI technology to train models to mimic the personality and speaking style of virtual characters. The updated models are then used to generate responses to the user. The generated responses are personalized, taking into account the user profile and emotional state.
[0183] For example, if a user types "I'm tired today," the emotion engine will identify this as a negative emotion, and the server can use this emotion information to generate an encouraging message from a virtual character such as "You had a tough day, well done." An example of a prompt would be, "How does the server recognize an emotion and generate a response when a user sends an emotional message?"
[0184] In this way, users can enjoy a more intimate and human-like interaction experience with virtual characters through the system.
[0185] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0186] Step 1:
[0187] Users interact with the interface using natural language through their device, inputting their preferences, interests, and daily emotions. The input data is collected on the device as text or audio data. The device then converts this data into a formatted text file and sends it to the server.
[0188] Step 2:
[0189] The server receives user input data sent from the terminal. Based on the received data, the server generates or updates a user profile. The profile stores the user's interests, past conversation history, and emotional tendencies in a database.
[0190] Step 3:
[0191] The server then uses an emotion engine to analyze the user's input data. This analysis process employs natural language processing techniques to identify emotional states such as positive, negative, and neutral from the input. As a result, classification information of the user's emotional state is obtained.
[0192] Step 4:
[0193] The server uses information on emotional states identified by the emotion engine to generate responses from virtual characters using generative AI technology. During this process, the latest news and trend information obtained from external sources is also analyzed, and a model is trained to provide personalized responses to the user. The generated responses are output as customized messages that take into account the user profile and emotional state.
[0194] Step 5:
[0195] The server sends the generated response message to the terminal. The terminal displays or plays the received message audibly to the user. Specifically, if it's late at night, it will display a message appropriate to the time, such as "Thank you for your hard work today, good night," enabling a response at the right time.
[0196] (Application Example 2)
[0197] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0198] The present invention aims to provide users with a more natural and human-like dialogue experience. However, existing dialogue systems with virtual characters do not adequately provide personalized responses that take into account the user's emotions, making it difficult to achieve empathetic communication based on emotions.
[0199] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0200] In this invention, the server includes means for collecting and storing information relating to the user's preferences and interests; means for training a model using data processing techniques to mimic the personality and speech patterns of a designated character; means for updating the model based on the latest information obtained from external sources; and means for identifying the user's emotions using an emotion engine and adjusting the content and expression of responses. This makes it possible to provide the user with empathetic, emotion-based dialogue and offer support and reminders as needed.
[0201] A "specific virtual character" is a digital entity designed to enable natural language interaction with a user, possessing a specific personality and speech pattern.
[0202] "Natural language" refers to the language forms that users use on a daily basis, and the words that are expressed through speech or text.
[0203] An "information processing system" is a technical mechanism that inputs, processes, and stores data, and provides output as needed.
[0204] "Preferences and interests" refers to information that indicates a user's personal preferences and interests, including their hobbies and areas of interest.
[0205] "Data processing technology" refers to a group of technologies for efficiently processing large amounts of data and utilizing it as information.
[0206] "Training a model" is the act of adjusting the parameters of an algorithm using data to improve its ability to perform a particular task.
[0207] "External information sources" refer to databases and online resources that provide information from outside the system.
[0208] An "emotion engine" is a system that identifies emotions from user input and selects or adjusts appropriate responses based on those emotions.
[0209] "Empathetic dialogue" refers to a type of dialogue that is flexible, responsive to the user's emotions and state of mind, and conducted with empathy.
[0210] To implement this invention, the user first accesses the system using a dedicated terminal. The terminal receives natural language input via voice or text and transmits that data to the server. The terminal is equipped with an interface for the user to input information about their preferences and interests.
[0211] Next, the server processes the received data. The server is designed based on a small computer like a Raspberry Pi and uses Python as its main programming language. The server converts the speech data to text using the Google Speech-to-Text API and implements an emotion engine using a custom model based on OpenAI's GPT-3™. This executes an algorithm to identify emotions from user input and generate responses. The server keeps the model up-to-date and provides more appropriate responses by regularly incorporating updates from external sources.
[0212] As a concrete example, suppose a user types "I'm tired today" into their device. The server, upon receiving this sentence, uses an emotion engine to recognize the negative emotion and generates a response such as "You had a tough day, well done." This generated response is sent back to the user via the device in either voice or text, resulting in a natural conversation.
[0213] The generative AI model in this system plays a crucial role in flexibly adjusting responses based on the user's emotions. Examples of prompts include the following:
[0214] example:
[0215] User: I'm really tired today.
[0216] System: That sounds tough. Is there anything I can do to help?
[0217] In this way, the terminal and server can work together to provide personalized conversational services to users.
[0218] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0219] Step 1:
[0220] The user uses a device to input messages in natural language. This input includes voice or text data. The input data is received by the device and sent to the server.
[0221] Step 2:
[0222] The server uses the Google Speech-to-Text API to convert the received audio data into text. This conversion process analyzes the audio waveform data and generates corresponding text data. This text data serves as input for the next step in sentiment analysis.
[0223] Step 3:
[0224] The server processes text data using a sentiment engine based on a custom model from OpenAI's GPT-3. In this step, the user's emotions are classified from the text data as positive, negative, or neutral, and sentiment metadata is generated based on the results.
[0225] Step 4:
[0226] The server uses user profile data and the latest external information to generate an appropriate response for the user. This process utilizes a generative AI model to generate the optimal response text based on the prompt. The tone and context of the response are adjusted based on the user's sentiment metadata.
[0227] Step 5:
[0228] The server sends the generated response data to the terminal. The terminal provides feedback to the user as voice or text. For voice feedback, a speech synthesis tool converts the text data into natural-sounding speech.
[0229] Step 6:
[0230] The user receives feedback through the device, and the interaction is completed. New input or responses from the user restart the process from step 1, enabling continuous interaction.
[0231] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0232] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0233] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0234] [Second Embodiment]
[0235] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0236] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0237] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0238] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0239] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0240] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0241] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0242] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0243] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0244] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0245] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0246] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0247] This invention constructs an information processing system for enabling a user to communicate with a specific virtual character using natural language. This system consists of a server, a terminal, and a user, and each component works together to facilitate interaction with the virtual character.
[0248] First, the user initiates interaction with the system using a terminal. The terminal provides an interface for inputting data about the user's preferences and interests. This information is sent to the server as foundational data for creating the user's profile.
[0249] The server updates its database based on the received user data and generates a user profile. This profile is used to provide personalized character responses. The server also trains a model that mimics the personality and speech patterns of a specified virtual character using AI technology. Publicly available information and historical data about the character are used for training.
[0250] The server also retrieves the latest news and trend information from external sources, keeping the character's knowledge base constantly up-to-date. This latest information is reflected in user conversations in real time.
[0251] When a user sends a message to a virtual character through their device, the device forwards the input to a server. The server analyzes this input and generates an appropriate response based on the user profile, a trained model, and the latest training data. The generated response is then sent to the device and displayed to the user.
[0252] For example, if a user asks a character, "What's the weather like now?", the server retrieves weather information from an external source and generates a response in a style that matches the character's tone, such as, "Today's weather is sunny, and the temperature is 20 degrees Celsius." In this way, users can have an experience that feels like they are having a conversation with a real idol or character.
[0253] Furthermore, the server uses real-world time information to generate automated messages from the character at appropriate times and send them to the user through the device. For example, by sending a message like "Good morning, let's do our best today!" in the morning, it enables intimate communication that is linked to the time of day.
[0254] This type of system allows users to have a personalized and enjoyable experience, making them feel as if their virtual character is actually real.
[0255] The following describes the processing flow.
[0256] Step 1:
[0257] The user uses their device to input information about their preferences and interests, including their favorite virtual characters and topics of interest. The device then sends this information to the server.
[0258] Step 2:
[0259] The server generates a user profile based on the user data it receives. This profile is stored in a database and used in subsequent conversation generation processes.
[0260] Step 3:
[0261] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. The training utilizes publicly available information and historical materials related to the character.
[0262] Step 4:
[0263] The server periodically retrieves data from external sources to obtain the latest news and trend information. Based on this information, it updates the model to ensure that the information users receive is fresh and accurate.
[0264] Step 5:
[0265] The user sends a message to a virtual character via a terminal. The terminal forwards this user input to the server.
[0266] Step 6:
[0267] The server analyzes the user input it receives. It generates an appropriate response by referring to the user profile, the latest training data, and the trained model.
[0268] Step 7:
[0269] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to interact with the virtual character in real time.
[0270] Step 8:
[0271] The server references real-time information to automatically generate messages at the appropriate time. These messages are sent to the user via the terminal, providing personalized communication tailored to the time.
[0272] (Example 1)
[0273] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0274] In recent years, communication technologies with virtual characters have become widely used, but conventional systems have faced challenges in achieving realistic dialogue and personalized responses. Furthermore, real-time dialogue that reflects time and external information is difficult. Current technology requires the ability to provide individualized responses based on user profiles and to achieve natural dialogue that reflects external information.
[0275] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0276] In this invention, the server includes means for collecting data from users via a terminal and transmitting that data to the server; means for updating a database based on the data received by the server and generating individual user profiles; means for training a model using generative AI technology to mimic the personality and conversational style of a virtual character based on the generated user profiles; means for obtaining current news and weather from external sources and reflecting them in the character's knowledge base; means for analyzing messages sent by users and generating responses based on the user profile and the latest information; and means for sending time-sensitive automated messages to users using real-time time information. This enables users to enjoy real-time, personalized, and natural conversations with virtual characters.
[0277] A "terminal" is a device used by users to input information and interact with virtual characters.
[0278] A "server" is a central device that processes data received from terminals and generates user profiles and responses.
[0279] "Generative AI technology" is an artificial intelligence technology that learns the personality and conversational style of virtual characters and imitates them.
[0280] A "database" is an information system used to store and manage collected user data and profile information.
[0281] A "user profile" is a unique set of information that reflects a user's interests, preferences, and past interactions.
[0282] "External information sources" refer to information providers that are obtained from outside the system, such as news and weather information.
[0283] A "knowledge base" is a collection of information that a virtual character holds for use in interacting with the user.
[0284] An "automatic message" is a message sent by a virtual character at a specific time or the like based on pre-set conditions.
[0285] The information processing system of the present invention enables a specific virtual character and a user to communicate in natural language. This system is composed of a server, a terminal, and a user. Each component cooperates to realize the interaction function with the virtual character.
[0286] Use of Hardware and Software
[0287] 1. Terminal:
[0288] The user uses a terminal such as a smartphone, a tablet, or a personal computer to start an interaction with the virtual character. The terminal has the role of collecting information about the user's preferences and interests and sending it to the server. For this purpose, a dedicated application or a web interface is provided on the terminal.
[0289] 2. Server:
[0290] The server is implemented using programming languages such as Python and Java, receives and analyzes the data sent from the terminal. The server generates a user profile based on the received data. This profile serves as the basis for individualized conversations.
[0291] In addition, the server trains a model for generating the personality and conversation style of the virtual character using generative AI technology. Machine learning frameworks such as TensorFlow and PyTorch are used in this process.
[0292] The server obtains the latest information from external information sources such as OpenWeatherMap and news APIs, and keeps the knowledge base of the virtual character up-to-date.
[0293] Specific example
[0294] When a user asks a virtual character "What's the latest news?" using their device, the server retrieves the latest news information via the NewsAPI, generates a response that matches the virtual character's speaking style, such as "The latest news is that a new museum has opened in the city," and sends it to the device. In this way, the user can have a natural experience as if they were having a conversation with a real person.
[0295] Example of a prompt
[0296] "Please generate a response for when a user asks the character, 'What events do you recommend this weekend?' The user profile should indicate that the character enjoys the outdoors, and the character should use a friendly tone."
[0297] This system provides users with a real-time, personalized, and enjoyable experience, making them feel as if their virtual character is a close friend or acquaintance.
[0298] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0299] Step 1:
[0300] The user initiates interaction with the system using a terminal. The terminal displays an interface for entering data about the user's preferences and interests. This input data includes information about the user's selected areas of interest and hobbies. The entered information is collected by the terminal and sent directly to the server.
[0301] Step 2:
[0302] The server receives the user data sent from the terminal and analyzes it. The server accesses the database and generates individual user profiles based on the newly received data. Data processing includes string analysis and data normalization, and an updated user profile is obtained as the output.
[0303] Step 3:
[0304] The server uses generative AI technology to train the personality and speaking style of the virtual character. For this, the user profile and the official materials of the virtual character are used as input data. Based on this data, the AI model is trained to mimic the personality of the virtual character. As an output, a highly personalized AI model is generated.
[0305] Step 4:
[0306] The server accesses external news APIs and weather APIs to obtain the latest information. The obtained information is stored in the knowledge base of the virtual character and utilized in conversations. Specifically, JSON data from external data sources is parsed, the necessary information is extracted, and added to the database inside the server.
[0307] Step 5:
[0308] When the user asks a question to the virtual character via the terminal, the message is transferred to the server. The server analyzes this input and generates an optimal response based on the user profile, the trained model, and external information. Text analysis and natural language generation are performed, and as a result, a response message to the user is output.
[0309] Step 6:
[0310] The terminal receives a response from the server and displays it to the user. This process visualizes the received text data on the terminal's user interface, allowing the user to continue interacting with the virtual character.
[0311] Step 7:
[0312] The server generates time-sensitive automated messages using real-world time information and sends them to the terminal periodically. Pre-set messages are generated and sent to the user based on the time and date. This allows the user to receive continuous feedback from the virtual character.
[0313] (Application Example 1)
[0314] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0315] In natural language dialogue systems with virtual entities, there is a need to provide personalized responses tailored to each user's preferences while ensuring a unified interactive experience across different devices. Furthermore, it is necessary to update the virtual entity's knowledge in real time using the latest data obtained from external sources, enabling smooth and intimate communication with users.
[0316] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0317] In this invention, the server includes means for collecting and storing data relating to the user's preferences and interests, means for training a computational model using information processing technology to mimic the characteristics and speech patterns of a designated virtual entity, and means for improving the computational model based on the latest data obtained from external sources. This enables personalized content for the user and real-time interaction that reflects the latest information.
[0318] A "virtual entity" is a digital entity that possesses the characteristics and speech patterns of a human or character, created using information processing technology.
[0319] "Natural language" refers to the forms of language that humans use on a daily basis, including texts and conversations processed by information processing devices.
[0320] An "information processing device" is a computing device that manipulates digital data, and through its functions, it collects, analyzes, stores, and displays data.
[0321] "Preferences" refer to data that shows the individual tastes and preferences of users, and are the basis for information processing systems to generate personalized responses.
[0322] A "computational model" is a mathematical model trained using information processing technology, and it is responsible for the process of generating an appropriate response to an input.
[0323] "External information sources" are various data providers that exist outside the system and are used as input to the system, providing information such as news and weather.
[0324] "Personalization" refers to adjusting content to suit the specific needs and preferences of each user, and is the process of optimizing the responses and content provided by information processing devices for each user.
[0325] An "interactive experience" is a user experience that is realized through real-time information exchange and mutual influence between the user and the system.
[0326] This system consists of client terminals owned by the user and servers operating in the cloud. Users can initiate interactions with virtual entities using terminals such as smartphones or smart glasses. The terminals provide an interface for inputting data about the user's preferences and interests, and transmit this data to the server. The server stores and manages the received data in a database and uses it to update the computational model.
[0327] The server utilizes a high-performance cloud platform (e.g., Google Cloud Platform or AWS) to train an AI model that mimics the characteristics and speech patterns of a virtual entity. The trained model generates appropriate responses based on user input. Natural language processing libraries such as SpaCy and BERT are used to train the AI model.
[0328] Furthermore, the server retrieves the latest data (such as news and weather information) from external sources and updates the computational model in real time. This up-to-date information is immediately utilized in conversations with virtual entities, providing realistic and timely answers to user questions.
[0329] For example, if a user asks, "When is the next live event?", the server uses an AI model to generate a response such as, "The next live event is on October 10th! Please book your tickets early if you plan to attend!" In this way, users can enjoy an interactive experience with a virtual entity.
[0330] An example of a prompt for a generative AI model is, "When is the next live event?". In response to this prompt, the model will generate an answer that matches the tone and speech patterns of the specified character.
[0331] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0332] Step 1:
[0333] The user accesses the system using a client terminal and begins interacting with a virtual entity. The terminal interface displays a screen for inputting data about the user's preferences and interests. This input data is processed on the terminal and then sent to the server.
[0334] Step 2:
[0335] The server receives user preference data from the terminal and stores it in a cloud-based database. This data is saved as a user profile, which serves as the basis for generating personalized responses in subsequent interactions.
[0336] Step 3:
[0337] The server uses natural language processing libraries (e.g., SpaCy or BERT) to update and train the generation AI model based on the received user data. The AI model reflects the characteristics and speech patterns of the specified virtual entity, improving the accuracy of response generation according to user preferences. Model parameters are adjusted during this process.
[0338] Step 4:
[0339] The server accesses external databases via APIs to collect the latest news and weather information from external sources. This data is input into the AI model in real time, constantly updating the knowledge base of the virtual entity.
[0340] Step 5:
[0341] The user sends a message to a virtual entity through a terminal. The terminal receives the user's input message and sends it to the server. This message becomes input data for natural language processing.
[0342] Step 6:
[0343] The server uses an AI model to analyze messages received from users. Based on the analysis results, it generates an appropriate response, referencing the user profile and the latest external data. The generated response is expressed in a tone and style appropriate to the virtual character.
[0344] Step 7:
[0345] The server sends the generated response to the terminal. The terminal displays this response to the user, and the interactive dialogue with the virtual entity is completed.
[0346] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0347] This invention provides a more personalized experience in an information processing system where a user communicates with a specific virtual character using natural language, by combining it with an emotion engine that recognizes the user's emotions. This system consists of a server, a terminal, and a user, and each component works in an integrated manner to realize emotion-responsive communication.
[0348] First, the user operates an input interface through a terminal, providing the system with their preferences, interests, and daily emotions. The terminal formats this information and sends it to the server. The server then generates a user profile based on this information and stores it in a database.
[0349] The server uses generative AI technology to train a model that mimics the personality and speech patterns of a specified virtual character, and establishes a mechanism to recognize emotions from user input using an emotion engine. The emotion engine analyzes the content of the input text and voice to identify emotional states such as positive, negative, and neutral.
[0350] Furthermore, the server retrieves the latest news and trend information from external sources to keep the model up-to-date. This updated information, combined with sentiment recognition, provides personalized responses to the user.
[0351] When a user sends a message to a virtual character via their device, the device forwards the input to the server. The server analyzes the received user input and generates an appropriate response based on the user profile, the latest training data, and the trained model. In doing so, it takes into account the emotional state analyzed by the emotion engine and adjusts the tone and content of the response.
[0352] For example, if a user inputs "I'm tired today," the emotion engine recognizes this input as a negative emotion. Based on this emotion information, the server has a virtual character choose encouraging words such as "You had a tough day, well done," to provide a response that is empathetic to the user.
[0353] The server takes real-world time information into account and generates emotionally appropriate automated messages at the right time, sending them to the user through the terminal. For example, by sending a message like "You've had a long day, good night" late at night, it enables intimate communication tailored to the time of day.
[0354] In this way, this system, which combines an emotion engine, allows users to enjoy a more intimate and human-like conversational experience with virtual characters.
[0355] The following describes the processing flow.
[0356] Step 1:
[0357] Users use their device to input their preferences, interests, and daily emotions. This includes their favorite characters, topics of interest, and recent emotional states. The device formats this information and sends it to the server.
[0358] Step 2:
[0359] The server generates a user profile based on the received user data. This profile is stored in a database and, including emotional information, is used in future conversation generation.
[0360] Step 3:
[0361] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. This process utilizes publicly available information and historical data about the character.
[0362] Step 4:
[0363] The server retrieves the latest news and trend information from external sources. This keeps the character's knowledge base up-to-date, allowing it to provide users with realistic information.
[0364] Step 5:
[0365] The user sends a message to a virtual character via a terminal. The message can be entered in text or voice format. The terminal then forwards this user input to the server.
[0366] Step 6:
[0367] The server analyzes the user's input and uses an emotion engine to recognize emotions from the input. The emotion engine identifies states such as positive, negative, and neutral.
[0368] Step 7:
[0369] The server generates an appropriate response based on the user profile, the latest training data, and the trained model, taking into account the recognized emotional state. The tone and content of the response are adjusted according to the emotion.
[0370] Step 8:
[0371] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to experience real-time interaction with the virtual character.
[0372] Step 9:
[0373] The server references real-world time information to generate automated messages that respond to emotions at the appropriate time. These messages are sent to the user via the terminal, providing intimate communication tailored to the time.
[0374] (Example 2)
[0375] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0376] In modern information processing systems, for users to engage in more intimate and personalized interactions with virtual characters, responses that take into account the user's emotions and temporal information are necessary. However, conventional systems have the challenge of not being able to generate responses that fully utilize emotion recognition and temporal information, and thus failing to sufficiently improve user satisfaction.
[0377] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0378] In this invention, the server includes means for collecting and storing data on the user's preferences and interests, means for training a model using generative AI technology, and means for identifying emotional states using an emotion engine. This enables the generation of personalized responses that take into account the user's emotions and temporal information.
[0379] A "user" is an individual who uses an information processing system to communicate with a virtual character.
[0380] "Preference and interest data" refers to information about a user's interests and preferences that they provide to the system.
[0381] "Generative AI technology" is a method that uses artificial intelligence to enable virtual characters to generate conversations, and is a technology used for training models.
[0382] "Methods for training a model" refers to a series of methods for training and updating a model using generative AI technology to mimic the personality and speech patterns of a virtual character.
[0383] An "emotion engine" refers to a technology or program that analyzes user input and identifies their emotional state based on that input.
[0384] "Emotional state" refers to the result of analyzing and classifying the emotional responses exhibited by users, and can be divided into categories such as positive, negative, and neutral.
[0385] "External information sources" are sources of information that are outside the system and provide up-to-date information and trends that are useful for generating responses from virtual characters.
[0386] "Personalized responses" refer to virtual character responses tailored to a specific user, customized based on the user's profile and emotional state.
[0387] "Real-world time information" refers to data that takes into account temporal factors such as actual time and date, and is used to enable responses at the appropriate time.
[0388] An "automated message" is a message that a system automatically generates and sends, taking into account the time and the user's emotional state.
[0389] This invention is an information processing system for users to communicate with a specific virtual character in natural language. The system consists of a server, a terminal, and a user, and each component works in an integrated manner to achieve personalized communication that responds to the user's emotions.
[0390] The user first interacts with the interface through a device, inputting their preferences, interests, and daily emotions. This information is then formatted and sent from the device to the server. The device can be a typical smartphone or computer, and input can utilize touchscreens or voice recognition.
[0391] The server generates user profiles based on the received data and stores them in a database. Here, it utilizes an emotion engine to recognize emotions from user input. The emotion engine uses natural language processing techniques to analyze text and audio content and identify emotional states such as positive, negative, and neutral. Because this entire process requires high-performance computing resources, the server is equipped with the latest processors and large-capacity memory.
[0392] Furthermore, the server retrieves the latest news and trend information from external sources and uses generative AI technology to train models to mimic the personality and speaking style of virtual characters. The updated models are then used to generate responses to the user. The generated responses are personalized, taking into account the user profile and emotional state.
[0393] For example, if a user types "I'm tired today," the emotion engine will identify this as a negative emotion, and the server can use this emotion information to generate an encouraging message from a virtual character such as "You had a tough day, well done." An example of a prompt would be, "How does the server recognize an emotion and generate a response when a user sends an emotional message?"
[0394] In this way, users can enjoy a more intimate and human-like interaction experience with virtual characters through the system.
[0395] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0396] Step 1:
[0397] Users interact with the interface using natural language through their device, inputting their preferences, interests, and daily emotions. The input data is collected on the device as text or audio data. The device then converts this data into a formatted text file and sends it to the server.
[0398] Step 2:
[0399] The server receives user input data sent from the terminal. Based on the received data, the server generates or updates a user profile. The profile stores the user's interests, past conversation history, and emotional tendencies in a database.
[0400] Step 3:
[0401] The server then uses an emotion engine to analyze the user's input data. This analysis process employs natural language processing techniques to identify emotional states such as positive, negative, and neutral from the input. As a result, classification information of the user's emotional state is obtained.
[0402] Step 4:
[0403] The server uses information on emotional states identified by the emotion engine to generate responses from virtual characters using generative AI technology. During this process, the latest news and trend information obtained from external sources is also analyzed, and a model is trained to provide personalized responses to the user. The generated responses are output as customized messages that take into account the user profile and emotional state.
[0404] Step 5:
[0405] The server sends the generated response message to the terminal. The terminal displays or plays the received message audibly to the user. Specifically, if it's late at night, it will display a message appropriate to the time, such as "Thank you for your hard work today, good night," enabling a response at the right time.
[0406] (Application Example 2)
[0407] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0408] The present invention aims to provide users with a more natural and human-like dialogue experience. However, existing dialogue systems with virtual characters do not adequately provide personalized responses that take into account the user's emotions, making it difficult to achieve empathetic communication based on emotions.
[0409] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0410] In this invention, the server includes means for collecting and storing information relating to the user's preferences and interests; means for training a model using data processing techniques to mimic the personality and speech patterns of a designated character; means for updating the model based on the latest information obtained from external sources; and means for identifying the user's emotions using an emotion engine and adjusting the content and expression of responses. This makes it possible to provide the user with empathetic, emotion-based dialogue and offer support and reminders as needed.
[0411] A "specific virtual character" is a digital entity designed to enable natural language interaction with a user, possessing a specific personality and speech pattern.
[0412] "Natural language" refers to the language forms that users use on a daily basis, and the words that are expressed through speech or text.
[0413] An "information processing system" is a technical mechanism that inputs, processes, and stores data, and provides output as needed.
[0414] "Preferences and interests" refers to information that indicates a user's personal preferences and interests, including their hobbies and areas of interest.
[0415] "Data processing technology" refers to a group of technologies for efficiently processing large amounts of data and utilizing it as information.
[0416] "Training a model" is the act of adjusting the parameters of an algorithm using data to improve its ability to perform a particular task.
[0417] "External information sources" refer to databases and online resources that provide information from outside the system.
[0418] An "emotion engine" is a system that identifies emotions from user input and selects or adjusts appropriate responses based on those emotions.
[0419] "Empathetic dialogue" refers to a type of dialogue that is flexible, responsive to the user's emotions and state of mind, and conducted with empathy.
[0420] To implement this invention, the user first accesses the system using a dedicated terminal. The terminal receives natural language input via voice or text and transmits that data to the server. The terminal is equipped with an interface for the user to input information about their preferences and interests.
[0421] Next, the server processes the received data. The server is designed based on a small computer like a Raspberry Pi and uses Python as its main programming language. The server converts the speech data to text using the Google Speech-to-Text API and implements an emotion engine using a custom model based on OpenAI's GPT-3. This runs an algorithm to identify emotions from user input and generate responses. The server keeps the model up-to-date and provides more appropriate responses by regularly incorporating updates from external sources.
[0422] As a concrete example, suppose a user types "I'm tired today" into their device. The server, upon receiving this sentence, uses an emotion engine to recognize the negative emotion and generates a response such as "You had a tough day, well done." This generated response is sent back to the user via the device in either voice or text, resulting in a natural conversation.
[0423] The generative AI model in this system plays a crucial role in flexibly adjusting responses based on the user's emotions. Examples of prompts include the following:
[0424] example:
[0425] User: I'm really tired today.
[0426] System: That sounds tough. Is there anything I can do to help?
[0427] In this way, the terminal and server can work together to provide personalized conversational services to users.
[0428] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0429] Step 1:
[0430] The user uses a device to input messages in natural language. This input includes voice or text data. The input data is received by the device and sent to the server.
[0431] Step 2:
[0432] The server uses the Google Speech-to-Text API to convert the received audio data into text. This conversion process analyzes the audio waveform data and generates corresponding text data. This text data serves as input for the next step in sentiment analysis.
[0433] Step 3:
[0434] The server processes text data using a sentiment engine based on a custom model from OpenAI's GPT-3. In this step, the user's emotions are classified from the text data as positive, negative, or neutral, and sentiment metadata is generated based on the results.
[0435] Step 4:
[0436] The server uses user profile data and the latest external information to generate an appropriate response for the user. This process utilizes a generative AI model to generate the optimal response text based on the prompt. The tone and context of the response are adjusted based on the user's sentiment metadata.
[0437] Step 5:
[0438] The server sends the generated response data to the terminal. The terminal provides feedback to the user as voice or text. For voice feedback, a speech synthesis tool converts the text data into natural-sounding speech.
[0439] Step 6:
[0440] The user receives feedback through the device, and the interaction is completed. New input or responses from the user restart the process from step 1, enabling continuous interaction.
[0441] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0442] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0443] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0444] [Third Embodiment]
[0445] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0446] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0447] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0448] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0449] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0450] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0451] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0452] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0453] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0454] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0455] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0456] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0457] This invention constructs an information processing system for enabling a user to communicate with a specific virtual character using natural language. This system consists of a server, a terminal, and a user, and each component works together to facilitate interaction with the virtual character.
[0458] First, the user initiates interaction with the system using a terminal. The terminal provides an interface for inputting data about the user's preferences and interests. This information is sent to the server as foundational data for creating the user's profile.
[0459] The server updates its database based on the received user data and generates a user profile. This profile is used to provide personalized character responses. The server also trains a model that mimics the personality and speech patterns of a specified virtual character using AI technology. Publicly available information and historical data about the character are used for training.
[0460] The server also retrieves the latest news and trend information from external sources, keeping the character's knowledge base constantly up-to-date. This latest information is reflected in user conversations in real time.
[0461] When a user sends a message to a virtual character through their device, the device forwards the input to a server. The server analyzes this input and generates an appropriate response based on the user profile, a trained model, and the latest training data. The generated response is then sent to the device and displayed to the user.
[0462] For example, if a user asks a character, "What's the weather like now?", the server retrieves weather information from an external source and generates a response in a style that matches the character's tone, such as, "Today's weather is sunny, and the temperature is 20 degrees Celsius." In this way, users can have an experience that feels like they are having a conversation with a real idol or character.
[0463] Furthermore, the server uses real-world time information to generate automated messages from the character at appropriate times and send them to the user through the device. For example, by sending a message like "Good morning, let's do our best today!" in the morning, it enables intimate communication that is linked to the time of day.
[0464] This type of system allows users to have a personalized and enjoyable experience, making them feel as if their virtual character is actually real.
[0465] The following describes the processing flow.
[0466] Step 1:
[0467] The user uses their device to input information about their preferences and interests, including their favorite virtual characters and topics of interest. The device then sends this information to the server.
[0468] Step 2:
[0469] The server generates a user profile based on the user data it receives. This profile is stored in a database and used in subsequent conversation generation processes.
[0470] Step 3:
[0471] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. The training utilizes publicly available information and historical materials related to the character.
[0472] Step 4:
[0473] The server periodically retrieves data from external sources to obtain the latest news and trend information. Based on this information, it updates the model to ensure that the information users receive is fresh and accurate.
[0474] Step 5:
[0475] The user sends a message to a virtual character via a terminal. The terminal forwards this user input to the server.
[0476] Step 6:
[0477] The server analyzes the user input it receives. It generates an appropriate response by referring to the user profile, the latest training data, and the trained model.
[0478] Step 7:
[0479] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to interact with the virtual character in real time.
[0480] Step 8:
[0481] The server references real-time information to automatically generate messages at the appropriate time. These messages are sent to the user via the terminal, providing personalized communication tailored to the time.
[0482] (Example 1)
[0483] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0484] In recent years, communication technologies with virtual characters have become widely used, but conventional systems have faced challenges in achieving realistic dialogue and personalized responses. Furthermore, real-time dialogue that reflects time and external information is difficult. Current technology requires the ability to provide individualized responses based on user profiles and to achieve natural dialogue that reflects external information.
[0485] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0486] In this invention, the server includes means for collecting data from users via a terminal and transmitting that data to the server; means for updating a database based on the data received by the server and generating individual user profiles; means for training a model using generative AI technology to mimic the personality and conversational style of a virtual character based on the generated user profiles; means for obtaining current news and weather from external sources and reflecting them in the character's knowledge base; means for analyzing messages sent by users and generating responses based on the user profile and the latest information; and means for sending time-sensitive automated messages to users using real-time time information. This enables users to enjoy real-time, personalized, and natural conversations with virtual characters.
[0487] A "terminal" is a device used by users to input information and interact with virtual characters.
[0488] A "server" is a central device that processes data received from terminals and generates user profiles and responses.
[0489] "Generative AI technology" is an artificial intelligence technology that learns the personality and conversational style of virtual characters and imitates them.
[0490] A "database" is an information system used to store and manage collected user data and profile information.
[0491] A "user profile" is a unique set of information that reflects a user's interests, preferences, and past interactions.
[0492] "External information sources" refer to information providers that are obtained from outside the system, such as news and weather information.
[0493] A "knowledge base" is a collection of information that a virtual character holds for use in interacting with the user.
[0494] An "automatic message" is a message that a virtual character sends at a specific time or period based on pre-set conditions.
[0495] The information processing system of this invention enables a user to communicate with a specific virtual character using natural language. This system consists of a server, a terminal, and a user. Each component works together to realize the interaction function with the virtual character.
[0496] Hardware and software usage
[0497] 1. Terminal:
[0498] Users initiate interaction with virtual characters using devices such as smartphones, tablets, or personal computers. These devices collect information about the user's preferences and interests and transmit it to the server. For this purpose, dedicated applications or web interfaces are provided on the devices.
[0499] 2. Server:
[0500] The server is implemented using programming languages such as Python or Java, and receives and analyzes data sent from the terminal. Based on the received data, the server generates a user profile. This profile forms the basis for personalized interactions.
[0501] The server also uses generative AI technology to train models for generating the personalities and conversational styles of virtual characters. This process utilizes machine learning frameworks such as TensorFlow and PyTorch.
[0502] The server retrieves the latest information from external sources such as OpenWeatherMap and news APIs to keep the virtual character's knowledge base up to date.
[0503] Specific example
[0504] When a user asks a virtual character "What's the latest news?" using their device, the server retrieves the latest news information via the NewsAPI, generates a response that matches the virtual character's speaking style, such as "The latest news is that a new museum has opened in the city," and sends it to the device. In this way, the user can have a natural experience as if they were having a conversation with a real person.
[0505] Example of a prompt
[0506] "Please generate a response for when a user asks the character, 'What events do you recommend this weekend?' The user profile should indicate that the character enjoys the outdoors, and the character should use a friendly tone."
[0507] This system provides users with a real-time, personalized, and enjoyable experience, making them feel as if their virtual character is a close friend or acquaintance.
[0508] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0509] Step 1:
[0510] The user initiates interaction with the system using a terminal. The terminal displays an interface for entering data about the user's preferences and interests. This input data includes information about the user's selected areas of interest and hobbies. The entered information is collected by the terminal and sent directly to the server.
[0511] Step 2:
[0512] The server receives user data sent from the terminal and analyzes it. The server accesses the database and generates individual user profiles based on the newly received data. Data processing includes string analysis and data normalization, resulting in updated user profiles as output.
[0513] Step 3:
[0514] The server uses generative AI technology to train the personality and speech patterns of virtual characters. User profiles and official virtual character documentation are used as input data. Based on this data, the AI model is trained to mimic the virtual character's personality. The output is a highly personalized AI model.
[0515] Step 4:
[0516] The server accesses external news and weather APIs to retrieve the latest information. This retrieved information is stored in the virtual character's knowledge base and used in conversations. Specifically, JSON data from external data sources is parsed, the necessary information is extracted, and added to the server's internal database.
[0517] Step 5:
[0518] When a user asks a question to a virtual character via their device, the message is forwarded to the server. The server analyzes this input and generates the optimal response based on the user profile, a trained model, and external information. Text analysis and natural language generation are performed, and the resulting response message to the user is output.
[0519] Step 6:
[0520] The terminal receives a response from the server and displays it to the user. This process visualizes the received text data on the terminal's user interface, allowing the user to continue interacting with the virtual character.
[0521] Step 7:
[0522] The server generates time-sensitive automated messages using real-world time information and sends them to the terminal periodically. Pre-set messages are generated and sent to the user based on the time and date. This allows the user to receive continuous feedback from the virtual character.
[0523] (Application Example 1)
[0524] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0525] In natural language dialogue systems with virtual entities, there is a need to provide personalized responses tailored to each user's preferences while ensuring a unified interactive experience across different devices. Furthermore, it is necessary to update the virtual entity's knowledge in real time using the latest data obtained from external sources, enabling smooth and intimate communication with users.
[0526] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0527] In this invention, the server includes means for collecting and storing data relating to the user's preferences and interests, means for training a computational model using information processing technology to mimic the characteristics and speech patterns of a designated virtual entity, and means for improving the computational model based on the latest data obtained from external sources. This enables personalized content for the user and real-time interaction that reflects the latest information.
[0528] A "virtual entity" is a digital entity that possesses the characteristics and speech patterns of a human or character, created using information processing technology.
[0529] "Natural language" refers to the forms of language that humans use on a daily basis, including texts and conversations processed by information processing devices.
[0530] An "information processing device" is a computing device that manipulates digital data, and through its functions, it collects, analyzes, stores, and displays data.
[0531] "Preferences" refer to data that shows the individual tastes and preferences of users, and are the basis for information processing systems to generate personalized responses.
[0532] A "computational model" is a mathematical model trained using information processing technology, and it is responsible for the process of generating an appropriate response to an input.
[0533] "External information sources" are various data providers that exist outside the system and are used as input to the system, providing information such as news and weather.
[0534] "Personalization" refers to adjusting content to suit the specific needs and preferences of each user, and is the process of optimizing the responses and content provided by information processing devices for each user.
[0535] An "interactive experience" is a user experience that is realized through real-time information exchange and mutual influence between the user and the system.
[0536] This system consists of client terminals owned by the user and servers operating in the cloud. Users can initiate interactions with virtual entities using terminals such as smartphones or smart glasses. The terminals provide an interface for inputting data about the user's preferences and interests, and transmit this data to the server. The server stores and manages the received data in a database and uses it to update the computational model.
[0537] The server utilizes a high-performance cloud platform (e.g., Google Cloud Platform or AWS) to train an AI model that mimics the characteristics and speech patterns of a virtual entity. The trained model generates appropriate responses based on user input. Natural language processing libraries such as SpaCy and BERT are used to train the AI model.
[0538] Furthermore, the server retrieves the latest data (such as news and weather information) from external sources and updates the computational model in real time. This up-to-date information is immediately utilized in conversations with virtual entities, providing realistic and timely answers to user questions.
[0539] For example, if a user asks, "When is the next live event?", the server uses an AI model to generate a response such as, "The next live event is on October 10th! Please book your tickets early if you plan to attend!" In this way, users can enjoy an interactive experience with a virtual entity.
[0540] An example of a prompt for a generative AI model is, "When is the next live event?". In response to this prompt, the model will generate an answer that matches the tone and speech patterns of the specified character.
[0541] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0542] Step 1:
[0543] The user accesses the system using a client terminal and begins interacting with a virtual entity. The terminal interface displays a screen for inputting data about the user's preferences and interests. This input data is processed on the terminal and then sent to the server.
[0544] Step 2:
[0545] The server receives user preference data from the terminal and stores it in a cloud-based database. This data is saved as a user profile, which serves as the basis for generating personalized responses in subsequent interactions.
[0546] Step 3:
[0547] The server uses natural language processing libraries (e.g., SpaCy or BERT) to update and train the generation AI model based on the received user data. The AI model reflects the characteristics and speech patterns of the specified virtual entity, improving the accuracy of response generation according to user preferences. Model parameters are adjusted during this process.
[0548] Step 4:
[0549] The server accesses external databases via APIs to collect the latest news and weather information from external sources. This data is input into the AI model in real time, constantly updating the knowledge base of the virtual entity.
[0550] Step 5:
[0551] The user sends a message to a virtual entity through a terminal. The terminal receives the user's input message and sends it to the server. This message becomes input data for natural language processing.
[0552] Step 6:
[0553] The server uses an AI model to analyze messages received from users. Based on the analysis results, it generates an appropriate response, referencing the user profile and the latest external data. The generated response is expressed in a tone and style appropriate to the virtual character.
[0554] Step 7:
[0555] The server sends the generated response to the terminal. The terminal displays this response to the user, and the interactive dialogue with the virtual entity is completed.
[0556] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0557] This invention provides a more personalized experience in an information processing system where a user communicates with a specific virtual character using natural language, by combining it with an emotion engine that recognizes the user's emotions. This system consists of a server, a terminal, and a user, and each component works in an integrated manner to realize emotion-responsive communication.
[0558] First, the user operates an input interface through a terminal, providing the system with their preferences, interests, and daily emotions. The terminal formats this information and sends it to the server. The server then generates a user profile based on this information and stores it in a database.
[0559] The server uses generative AI technology to train a model that mimics the personality and speech patterns of a specified virtual character, and establishes a mechanism to recognize emotions from user input using an emotion engine. The emotion engine analyzes the content of the input text and voice to identify emotional states such as positive, negative, and neutral.
[0560] Furthermore, the server retrieves the latest news and trend information from external sources to keep the model up-to-date. This updated information, combined with sentiment recognition, provides personalized responses to the user.
[0561] When a user sends a message to a virtual character via their device, the device forwards the input to the server. The server analyzes the received user input and generates an appropriate response based on the user profile, the latest training data, and the trained model. In doing so, it takes into account the emotional state analyzed by the emotion engine and adjusts the tone and content of the response.
[0562] For example, if a user inputs "I'm tired today," the emotion engine recognizes this input as a negative emotion. Based on this emotion information, the server has a virtual character choose encouraging words such as "You had a tough day, well done," to provide a response that is empathetic to the user.
[0563] The server takes real-world time information into account and generates emotionally appropriate automated messages at the right time, sending them to the user through the terminal. For example, by sending a message like "You've had a long day, good night" late at night, it enables intimate communication tailored to the time of day.
[0564] In this way, this system, which combines an emotion engine, allows users to enjoy a more intimate and human-like conversational experience with virtual characters.
[0565] The following describes the processing flow.
[0566] Step 1:
[0567] Users use their device to input their preferences, interests, and daily emotions. This includes their favorite characters, topics of interest, and recent emotional states. The device formats this information and sends it to the server.
[0568] Step 2:
[0569] The server generates a user profile based on the received user data. This profile is stored in a database and, including emotional information, is used in future conversation generation.
[0570] Step 3:
[0571] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. This process utilizes publicly available information and historical data about the character.
[0572] Step 4:
[0573] The server retrieves the latest news and trend information from external sources. This keeps the character's knowledge base up-to-date, allowing it to provide users with realistic information.
[0574] Step 5:
[0575] The user sends a message to a virtual character via a terminal. The message can be entered in text or voice format. The terminal then forwards this user input to the server.
[0576] Step 6:
[0577] The server analyzes the user's input and uses an emotion engine to recognize emotions from the input. The emotion engine identifies states such as positive, negative, and neutral.
[0578] Step 7:
[0579] The server generates an appropriate response based on the user profile, the latest training data, and the trained model, taking into account the recognized emotional state. The tone and content of the response are adjusted according to the emotion.
[0580] Step 8:
[0581] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to experience real-time interaction with the virtual character.
[0582] Step 9:
[0583] The server references real-world time information to generate automated messages that respond to emotions at the appropriate time. These messages are sent to the user via the terminal, providing intimate communication tailored to the time.
[0584] (Example 2)
[0585] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0586] In modern information processing systems, for users to engage in more intimate and personalized interactions with virtual characters, responses that take into account the user's emotions and temporal information are necessary. However, conventional systems have the challenge of not being able to generate responses that fully utilize emotion recognition and temporal information, and thus failing to sufficiently improve user satisfaction.
[0587] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0588] In this invention, the server includes means for collecting and storing data on the user's preferences and interests, means for training a model using generative AI technology, and means for identifying emotional states using an emotion engine. This enables the generation of personalized responses that take into account the user's emotions and temporal information.
[0589] A "user" is an individual who uses an information processing system to communicate with a virtual character.
[0590] "Preference and interest data" refers to information about a user's interests and preferences that they provide to the system.
[0591] "Generative AI technology" is a method that uses artificial intelligence to enable virtual characters to generate conversations, and is a technology used for training models.
[0592] "Methods for training a model" refers to a series of methods for training and updating a model using generative AI technology to mimic the personality and speech patterns of a virtual character.
[0593] An "emotion engine" refers to a technology or program that analyzes user input and identifies their emotional state based on that input.
[0594] "Emotional state" refers to the result of analyzing and classifying the emotional responses exhibited by users, and can be divided into categories such as positive, negative, and neutral.
[0595] "External information sources" are sources of information that are outside the system and provide up-to-date information and trends that are useful for generating responses from virtual characters.
[0596] "Personalized responses" refer to virtual character responses tailored to a specific user, customized based on the user's profile and emotional state.
[0597] "Real-world time information" refers to data that takes into account temporal factors such as actual time and date, and is used to enable responses at the appropriate time.
[0598] An "automated message" is a message that a system automatically generates and sends, taking into account the time and the user's emotional state.
[0599] This invention is an information processing system for users to communicate with a specific virtual character in natural language. The system consists of a server, a terminal, and a user, and each component works in an integrated manner to achieve personalized communication that responds to the user's emotions.
[0600] The user first interacts with the interface through a device, inputting their preferences, interests, and daily emotions. This information is then formatted and sent from the device to the server. The device can be a typical smartphone or computer, and input can utilize touchscreens or voice recognition.
[0601] The server generates user profiles based on the received data and stores them in a database. Here, it utilizes an emotion engine to recognize emotions from user input. The emotion engine uses natural language processing techniques to analyze text and audio content and identify emotional states such as positive, negative, and neutral. Because this entire process requires high-performance computing resources, the server is equipped with the latest processors and large-capacity memory.
[0602] Furthermore, the server retrieves the latest news and trend information from external sources and uses generative AI technology to train models to mimic the personality and speaking style of virtual characters. The updated models are then used to generate responses to the user. The generated responses are personalized, taking into account the user profile and emotional state.
[0603] For example, if a user types "I'm tired today," the emotion engine will identify this as a negative emotion, and the server can use this emotion information to generate an encouraging message from a virtual character such as "You had a tough day, well done." An example of a prompt would be, "How does the server recognize an emotion and generate a response when a user sends an emotional message?"
[0604] In this way, users can enjoy a more intimate and human-like interaction experience with virtual characters through the system.
[0605] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0606] Step 1:
[0607] Users interact with the interface using natural language through their device, inputting their preferences, interests, and daily emotions. The input data is collected on the device as text or audio data. The device then converts this data into a formatted text file and sends it to the server.
[0608] Step 2:
[0609] The server receives user input data sent from the terminal. Based on the received data, the server generates or updates a user profile. The profile stores the user's interests, past conversation history, and emotional tendencies in a database.
[0610] Step 3:
[0611] The server then uses an emotion engine to analyze the user's input data. This analysis process employs natural language processing techniques to identify emotional states such as positive, negative, and neutral from the input. As a result, classification information of the user's emotional state is obtained.
[0612] Step 4:
[0613] The server uses information on emotional states identified by the emotion engine to generate responses from virtual characters using generative AI technology. During this process, the latest news and trend information obtained from external sources is also analyzed, and a model is trained to provide personalized responses to the user. The generated responses are output as customized messages that take into account the user profile and emotional state.
[0614] Step 5:
[0615] The server sends the generated response message to the terminal. The terminal displays or plays the received message audibly to the user. Specifically, if it's late at night, it will display a message appropriate to the time, such as "Thank you for your hard work today, good night," enabling a response at the right time.
[0616] (Application Example 2)
[0617] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0618] The present invention aims to provide users with a more natural and human-like dialogue experience. However, existing dialogue systems with virtual characters do not adequately provide personalized responses that take into account the user's emotions, making it difficult to achieve empathetic communication based on emotions.
[0619] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0620] In this invention, the server includes means for collecting and storing information relating to the user's preferences and interests; means for training a model using data processing techniques to mimic the personality and speech patterns of a designated character; means for updating the model based on the latest information obtained from external sources; and means for identifying the user's emotions using an emotion engine and adjusting the content and expression of responses. This makes it possible to provide the user with empathetic, emotion-based dialogue and offer support and reminders as needed.
[0621] A "specific virtual character" is a digital entity designed to enable natural language interaction with a user, possessing a specific personality and speech pattern.
[0622] "Natural language" refers to the language forms that users use on a daily basis, and the words that are expressed through speech or text.
[0623] An "information processing system" is a technical mechanism that inputs, processes, and stores data, and provides output as needed.
[0624] "Preferences and interests" refers to information that indicates a user's personal preferences and interests, including their hobbies and areas of interest.
[0625] "Data processing technology" refers to a group of technologies for efficiently processing large amounts of data and utilizing it as information.
[0626] "Training a model" is the act of adjusting the parameters of an algorithm using data to improve its ability to perform a particular task.
[0627] "External information sources" refer to databases and online resources that provide information from outside the system.
[0628] An "emotion engine" is a system that identifies emotions from user input and selects or adjusts appropriate responses based on those emotions.
[0629] "Empathetic dialogue" refers to a type of dialogue that is flexible, responsive to the user's emotions and state of mind, and conducted with empathy.
[0630] To implement this invention, the user first accesses the system using a dedicated terminal. The terminal receives natural language input via voice or text and transmits that data to the server. The terminal is equipped with an interface for the user to input information about their preferences and interests.
[0631] Next, the server processes the received data. The server is designed based on a small computer like a Raspberry Pi and uses Python as its main programming language. The server converts the speech data to text using the Google Speech-to-Text API and implements an emotion engine using a custom model based on OpenAI's GPT-3. This runs an algorithm to identify emotions from user input and generate responses. The server keeps the model up-to-date and provides more appropriate responses by regularly incorporating updates from external sources.
[0632] As a concrete example, suppose a user types "I'm tired today" into their device. The server, upon receiving this sentence, uses an emotion engine to recognize the negative emotion and generates a response such as "You had a tough day, well done." This generated response is sent back to the user via the device in either voice or text, resulting in a natural conversation.
[0633] The generative AI model in this system plays a crucial role in flexibly adjusting responses based on the user's emotions. Examples of prompts include the following:
[0634] example:
[0635] User: I'm really tired today.
[0636] System: That sounds tough. Is there anything I can do to help?
[0637] In this way, the terminal and server can work together to provide personalized conversational services to users.
[0638] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0639] Step 1:
[0640] The user uses a device to input messages in natural language. This input includes voice or text data. The input data is received by the device and sent to the server.
[0641] Step 2:
[0642] The server uses the Google Speech-to-Text API to convert the received audio data into text. This conversion process analyzes the audio waveform data and generates corresponding text data. This text data serves as input for the next step in sentiment analysis.
[0643] Step 3:
[0644] The server processes text data using a sentiment engine based on a custom model from OpenAI's GPT-3. In this step, the user's emotions are classified from the text data as positive, negative, or neutral, and sentiment metadata is generated based on the results.
[0645] Step 4:
[0646] The server uses user profile data and the latest external information to generate an appropriate response for the user. This process utilizes a generative AI model to generate the optimal response text based on the prompt. The tone and context of the response are adjusted based on the user's sentiment metadata.
[0647] Step 5:
[0648] The server sends the generated response data to the terminal. The terminal provides feedback to the user as voice or text. For voice feedback, a speech synthesis tool converts the text data into natural-sounding speech.
[0649] Step 6:
[0650] The user receives feedback through the device, and the interaction is completed. New input or responses from the user restart the process from step 1, enabling continuous interaction.
[0651] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0652] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0653] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0654] [Fourth Embodiment]
[0655] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0656] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0657] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0658] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0659] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0660] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0661] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0662] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0663] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0664] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0665] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0666] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0667] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0668] This invention constructs an information processing system for enabling a user to communicate with a specific virtual character using natural language. This system consists of a server, a terminal, and a user, and each component works together to facilitate interaction with the virtual character.
[0669] First, the user initiates interaction with the system using a terminal. The terminal provides an interface for inputting data about the user's preferences and interests. This information is sent to the server as foundational data for creating the user's profile.
[0670] The server updates its database based on the received user data and generates a user profile. This profile is used to provide personalized character responses. The server also trains a model that mimics the personality and speech patterns of a specified virtual character using AI technology. Publicly available information and historical data about the character are used for training.
[0671] The server also retrieves the latest news and trend information from external sources, keeping the character's knowledge base constantly up-to-date. This latest information is reflected in user conversations in real time.
[0672] When a user sends a message to a virtual character through their device, the device forwards the input to a server. The server analyzes this input and generates an appropriate response based on the user profile, a trained model, and the latest training data. The generated response is then sent to the device and displayed to the user.
[0673] For example, if a user asks a character, "What's the weather like now?", the server retrieves weather information from an external source and generates a response in a style that matches the character's tone, such as, "Today's weather is sunny, and the temperature is 20 degrees Celsius." In this way, users can have an experience that feels like they are having a conversation with a real idol or character.
[0674] Furthermore, the server uses real-world time information to generate automated messages from the character at appropriate times and send them to the user through the device. For example, by sending a message like "Good morning, let's do our best today!" in the morning, it enables intimate communication that is linked to the time of day.
[0675] This type of system allows users to have a personalized and enjoyable experience, making them feel as if their virtual character is actually real.
[0676] The following describes the processing flow.
[0677] Step 1:
[0678] The user uses their device to input information about their preferences and interests, including their favorite virtual characters and topics of interest. The device then sends this information to the server.
[0679] Step 2:
[0680] The server generates a user profile based on the user data it receives. This profile is stored in a database and used in subsequent conversation generation processes.
[0681] Step 3:
[0682] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. The training utilizes publicly available information and historical materials related to the character.
[0683] Step 4:
[0684] The server periodically retrieves data from external sources to obtain the latest news and trend information. Based on this information, it updates the model to ensure that the information users receive is fresh and accurate.
[0685] Step 5:
[0686] The user sends a message to a virtual character via a terminal. The terminal forwards this user input to the server.
[0687] Step 6:
[0688] The server analyzes the user input it receives. It generates an appropriate response by referring to the user profile, the latest training data, and the trained model.
[0689] Step 7:
[0690] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to interact with the virtual character in real time.
[0691] Step 8:
[0692] The server references real-time information to automatically generate messages at the appropriate time. These messages are sent to the user via the terminal, providing personalized communication tailored to the time.
[0693] (Example 1)
[0694] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0695] In recent years, communication technologies with virtual characters have become widely used, but conventional systems have faced challenges in achieving realistic dialogue and personalized responses. Furthermore, real-time dialogue that reflects time and external information is difficult. Current technology requires the ability to provide individualized responses based on user profiles and to achieve natural dialogue that reflects external information.
[0696] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0697] In this invention, the server includes means for collecting data from users via a terminal and transmitting that data to the server; means for updating a database based on the data received by the server and generating individual user profiles; means for training a model using generative AI technology to mimic the personality and conversational style of a virtual character based on the generated user profiles; means for obtaining current news and weather from external sources and reflecting them in the character's knowledge base; means for analyzing messages sent by users and generating responses based on the user profile and the latest information; and means for sending time-sensitive automated messages to users using real-time time information. This enables users to enjoy real-time, personalized, and natural conversations with virtual characters.
[0698] A "terminal" is a device used by users to input information and interact with virtual characters.
[0699] A "server" is a central device that processes data received from terminals and generates user profiles and responses.
[0700] "Generative AI technology" is an artificial intelligence technology that learns the personality and conversational style of virtual characters and imitates them.
[0701] A "database" is an information system used to store and manage collected user data and profile information.
[0702] A "user profile" is a unique set of information that reflects a user's interests, preferences, and past interactions.
[0703] "External information sources" refer to information providers that are obtained from outside the system, such as news and weather information.
[0704] A "knowledge base" is a collection of information that a virtual character holds for use in interacting with the user.
[0705] An "automatic message" is a message that a virtual character sends at a specific time or period based on pre-set conditions.
[0706] The information processing system of this invention enables a user to communicate with a specific virtual character using natural language. This system consists of a server, a terminal, and a user. Each component works together to realize the interaction function with the virtual character.
[0707] Hardware and software usage
[0708] 1. Terminal:
[0709] Users initiate interaction with virtual characters using devices such as smartphones, tablets, or personal computers. These devices collect information about the user's preferences and interests and transmit it to the server. For this purpose, dedicated applications or web interfaces are provided on the devices.
[0710] 2. Server:
[0711] The server is implemented using programming languages such as Python or Java, and receives and analyzes data sent from the terminal. Based on the received data, the server generates a user profile. This profile forms the basis for personalized interactions.
[0712] The server also uses generative AI technology to train models for generating the personalities and conversational styles of virtual characters. This process utilizes machine learning frameworks such as TensorFlow and PyTorch.
[0713] The server retrieves the latest information from external sources such as OpenWeatherMap and news APIs to keep the virtual character's knowledge base up to date.
[0714] Specific example
[0715] When a user asks a virtual character "What's the latest news?" using their device, the server retrieves the latest news information via the NewsAPI, generates a response that matches the virtual character's speaking style, such as "The latest news is that a new museum has opened in the city," and sends it to the device. In this way, the user can have a natural experience as if they were having a conversation with a real person.
[0716] Example of a prompt
[0717] "Please generate a response for when a user asks the character, 'What events do you recommend this weekend?' The user profile should indicate that the character enjoys the outdoors, and the character should use a friendly tone."
[0718] This system provides users with a real-time, personalized, and enjoyable experience, making them feel as if their virtual character is a close friend or acquaintance.
[0719] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0720] Step 1:
[0721] The user initiates interaction with the system using a terminal. The terminal displays an interface for entering data about the user's preferences and interests. This input data includes information about the user's selected areas of interest and hobbies. The entered information is collected by the terminal and sent directly to the server.
[0722] Step 2:
[0723] The server receives user data sent from the terminal and analyzes it. The server accesses the database and generates individual user profiles based on the newly received data. Data processing includes string analysis and data normalization, resulting in updated user profiles as output.
[0724] Step 3:
[0725] The server uses generative AI technology to train the personality and speech patterns of virtual characters. User profiles and official virtual character documentation are used as input data. Based on this data, the AI model is trained to mimic the virtual character's personality. The output is a highly personalized AI model.
[0726] Step 4:
[0727] The server accesses external news and weather APIs to retrieve the latest information. This retrieved information is stored in the virtual character's knowledge base and used in conversations. Specifically, JSON data from external data sources is parsed, the necessary information is extracted, and added to the server's internal database.
[0728] Step 5:
[0729] When a user asks a question to a virtual character via their device, the message is forwarded to the server. The server analyzes this input and generates the optimal response based on the user profile, a trained model, and external information. Text analysis and natural language generation are performed, and the resulting response message to the user is output.
[0730] Step 6:
[0731] The terminal receives a response from the server and displays it to the user. This process visualizes the received text data on the terminal's user interface, allowing the user to continue interacting with the virtual character.
[0732] Step 7:
[0733] The server generates time-sensitive automated messages using real-world time information and sends them to the terminal periodically. Pre-set messages are generated and sent to the user based on the time and date. This allows the user to receive continuous feedback from the virtual character.
[0734] (Application Example 1)
[0735] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0736] In natural language dialogue systems with virtual entities, there is a need to provide personalized responses tailored to each user's preferences while ensuring a unified interactive experience across different devices. Furthermore, it is necessary to update the virtual entity's knowledge in real time using the latest data obtained from external sources, enabling smooth and intimate communication with users.
[0737] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0738] In this invention, the server includes means for collecting and storing data relating to the user's preferences and interests, means for training a computational model using information processing technology to mimic the characteristics and speech patterns of a designated virtual entity, and means for improving the computational model based on the latest data obtained from external sources. This enables personalized content for the user and real-time interaction that reflects the latest information.
[0739] A "virtual entity" is a digital entity that possesses the characteristics and speech patterns of a human or character, created using information processing technology.
[0740] "Natural language" refers to the forms of language that humans use on a daily basis, including texts and conversations processed by information processing devices.
[0741] An "information processing device" is a computing device that manipulates digital data, and through its functions, it collects, analyzes, stores, and displays data.
[0742] "Preferences" refer to data that shows the individual tastes and preferences of users, and are the basis for information processing systems to generate personalized responses.
[0743] A "computational model" is a mathematical model trained using information processing technology, and it is responsible for the process of generating an appropriate response to an input.
[0744] "External information sources" are various data providers that exist outside the system and are used as input to the system, providing information such as news and weather.
[0745] "Personalization" refers to adjusting content to suit the specific needs and preferences of each user, and is the process of optimizing the responses and content provided by information processing devices for each user.
[0746] An "interactive experience" is a user experience that is realized through real-time information exchange and mutual influence between the user and the system.
[0747] This system consists of client terminals owned by the user and servers operating in the cloud. Users can initiate interactions with virtual entities using terminals such as smartphones or smart glasses. The terminals provide an interface for inputting data about the user's preferences and interests, and transmit this data to the server. The server stores and manages the received data in a database and uses it to update the computational model.
[0748] The server utilizes a high-performance cloud platform (e.g., Google Cloud Platform or AWS) to train an AI model that mimics the characteristics and speech patterns of a virtual entity. The trained model generates appropriate responses based on user input. Natural language processing libraries such as SpaCy and BERT are used to train the AI model.
[0749] Furthermore, the server retrieves the latest data (such as news and weather information) from external sources and updates the computational model in real time. This up-to-date information is immediately utilized in conversations with virtual entities, providing realistic and timely answers to user questions.
[0750] For example, if a user asks, "When is the next live event?", the server uses an AI model to generate a response such as, "The next live event is on October 10th! Please book your tickets early if you plan to attend!" In this way, users can enjoy an interactive experience with a virtual entity.
[0751] An example of a prompt for a generative AI model is, "When is the next live event?". In response to this prompt, the model will generate an answer that matches the tone and speech patterns of the specified character.
[0752] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0753] Step 1:
[0754] The user accesses the system using a client terminal and begins interacting with a virtual entity. The terminal interface displays a screen for inputting data about the user's preferences and interests. This input data is processed on the terminal and then sent to the server.
[0755] Step 2:
[0756] The server receives user preference data from the terminal and stores it in a cloud-based database. This data is saved as a user profile, which serves as the basis for generating personalized responses in subsequent interactions.
[0757] Step 3:
[0758] The server uses natural language processing libraries (e.g., SpaCy or BERT) to update and train the generation AI model based on the received user data. The AI model reflects the characteristics and speech patterns of the specified virtual entity, improving the accuracy of response generation according to user preferences. Model parameters are adjusted during this process.
[0759] Step 4:
[0760] The server accesses external databases via APIs to collect the latest news and weather information from external sources. This data is input into the AI model in real time, constantly updating the knowledge base of the virtual entity.
[0761] Step 5:
[0762] The user sends a message to a virtual entity through a terminal. The terminal receives the user's input message and sends it to the server. This message becomes input data for natural language processing.
[0763] Step 6:
[0764] The server uses an AI model to analyze messages received from users. Based on the analysis results, it generates an appropriate response, referencing the user profile and the latest external data. The generated response is expressed in a tone and style appropriate to the virtual character.
[0765] Step 7:
[0766] The server sends the generated response to the terminal. The terminal displays this response to the user, and the interactive dialogue with the virtual entity is completed.
[0767] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0768] This invention provides a more personalized experience in an information processing system where a user communicates with a specific virtual character using natural language, by combining it with an emotion engine that recognizes the user's emotions. This system consists of a server, a terminal, and a user, and each component works in an integrated manner to realize emotion-responsive communication.
[0769] First, the user operates an input interface through a terminal, providing the system with their preferences, interests, and daily emotions. The terminal formats this information and sends it to the server. The server then generates a user profile based on this information and stores it in a database.
[0770] The server uses generative AI technology to train a model that mimics the personality and speech patterns of a specified virtual character, and establishes a mechanism to recognize emotions from user input using an emotion engine. The emotion engine analyzes the content of the input text and voice to identify emotional states such as positive, negative, and neutral.
[0771] Furthermore, the server retrieves the latest news and trend information from external sources to keep the model up-to-date. This updated information, combined with sentiment recognition, provides personalized responses to the user.
[0772] When a user sends a message to a virtual character via their device, the device forwards the input to the server. The server analyzes the received user input and generates an appropriate response based on the user profile, the latest training data, and the trained model. In doing so, it takes into account the emotional state analyzed by the emotion engine and adjusts the tone and content of the response.
[0773] For example, if a user inputs "I'm tired today," the emotion engine recognizes this input as a negative emotion. Based on this emotion information, the server has a virtual character choose encouraging words such as "You had a tough day, well done," to provide a response that is empathetic to the user.
[0774] The server takes real-world time information into account and generates emotionally appropriate automated messages at the right time, sending them to the user through the terminal. For example, by sending a message like "You've had a long day, good night" late at night, it enables intimate communication tailored to the time of day.
[0775] In this way, this system, which combines an emotion engine, allows users to enjoy a more intimate and human-like conversational experience with virtual characters.
[0776] The following describes the processing flow.
[0777] Step 1:
[0778] Users use their device to input their preferences, interests, and daily emotions. This includes their favorite characters, topics of interest, and recent emotional states. The device formats this information and sends it to the server.
[0779] Step 2:
[0780] The server generates a user profile based on the received user data. This profile is stored in a database and, including emotional information, is used in future conversation generation.
[0781] Step 3:
[0782] The server uses generative AI to train a model that mimics the personality and speech patterns of a specified virtual character. This process utilizes publicly available information and historical data about the character.
[0783] Step 4:
[0784] The server retrieves the latest news and trend information from external sources. This keeps the character's knowledge base up-to-date, allowing it to provide users with realistic information.
[0785] Step 5:
[0786] The user sends a message to a virtual character via a terminal. The message can be entered in text or voice format. The terminal then forwards this user input to the server.
[0787] Step 6:
[0788] The server analyzes the user's input and uses an emotion engine to recognize emotions from the input. The emotion engine identifies states such as positive, negative, and neutral.
[0789] Step 7:
[0790] The server generates an appropriate response based on the user profile, the latest training data, and the trained model, taking into account the recognized emotional state. The tone and content of the response are adjusted according to the emotion.
[0791] Step 8:
[0792] The server sends the generated response to the terminal. The terminal displays this response to the user, allowing the user to experience real-time interaction with the virtual character.
[0793] Step 9:
[0794] The server references real-world time information to generate automated messages that respond to emotions at the appropriate time. These messages are sent to the user via the terminal, providing intimate communication tailored to the time.
[0795] (Example 2)
[0796] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0797] In modern information processing systems, for users to engage in more intimate and personalized interactions with virtual characters, responses that take into account the user's emotions and temporal information are necessary. However, conventional systems have the challenge of not being able to generate responses that fully utilize emotion recognition and temporal information, and thus failing to sufficiently improve user satisfaction.
[0798] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0799] In this invention, the server includes means for collecting and storing data on the user's preferences and interests, means for training a model using generative AI technology, and means for identifying emotional states using an emotion engine. This enables the generation of personalized responses that take into account the user's emotions and temporal information.
[0800] A "user" is an individual who uses an information processing system to communicate with a virtual character.
[0801] "Preference and interest data" refers to information about a user's interests and preferences that they provide to the system.
[0802] "Generative AI technology" is a method that uses artificial intelligence to enable virtual characters to generate conversations, and is a technology used for training models.
[0803] "Methods for training a model" refers to a series of methods for training and updating a model using generative AI technology to mimic the personality and speech patterns of a virtual character.
[0804] An "emotion engine" refers to a technology or program that analyzes user input and identifies their emotional state based on that input.
[0805] "Emotional state" refers to the result of analyzing and classifying the emotional responses exhibited by users, and can be divided into categories such as positive, negative, and neutral.
[0806] "External information sources" are sources of information that are outside the system and provide up-to-date information and trends that are useful for generating responses from virtual characters.
[0807] "Personalized responses" refer to virtual character responses tailored to a specific user, customized based on the user's profile and emotional state.
[0808] "Real-world time information" refers to data that takes into account temporal factors such as actual time and date, and is used to enable responses at the appropriate time.
[0809] An "automated message" is a message that a system automatically generates and sends, taking into account the time and the user's emotional state.
[0810] This invention is an information processing system for users to communicate with a specific virtual character in natural language. The system consists of a server, a terminal, and a user, and each component works in an integrated manner to achieve personalized communication that responds to the user's emotions.
[0811] The user first interacts with the interface through a device, inputting their preferences, interests, and daily emotions. This information is then formatted and sent from the device to the server. The device can be a typical smartphone or computer, and input can utilize touchscreens or voice recognition.
[0812] The server generates user profiles based on the received data and stores them in a database. Here, it utilizes an emotion engine to recognize emotions from user input. The emotion engine uses natural language processing techniques to analyze text and audio content and identify emotional states such as positive, negative, and neutral. Because this entire process requires high-performance computing resources, the server is equipped with the latest processors and large-capacity memory.
[0813] Furthermore, the server retrieves the latest news and trend information from external sources and uses generative AI technology to train models to mimic the personality and speaking style of virtual characters. The updated models are then used to generate responses to the user. The generated responses are personalized, taking into account the user profile and emotional state.
[0814] For example, if a user types "I'm tired today," the emotion engine will identify this as a negative emotion, and the server can use this emotion information to generate an encouraging message from a virtual character such as "You had a tough day, well done." An example of a prompt would be, "How does the server recognize an emotion and generate a response when a user sends an emotional message?"
[0815] In this way, users can enjoy a more intimate and human-like interaction experience with virtual characters through the system.
[0816] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0817] Step 1:
[0818] Users interact with the interface using natural language through their device, inputting their preferences, interests, and daily emotions. The input data is collected on the device as text or audio data. The device then converts this data into a formatted text file and sends it to the server.
[0819] Step 2:
[0820] The server receives user input data sent from the terminal. Based on the received data, the server generates or updates a user profile. The profile stores the user's interests, past conversation history, and emotional tendencies in a database.
[0821] Step 3:
[0822] The server then uses an emotion engine to analyze the user's input data. This analysis process employs natural language processing techniques to identify emotional states such as positive, negative, and neutral from the input. As a result, classification information of the user's emotional state is obtained.
[0823] Step 4:
[0824] The server uses information on emotional states identified by the emotion engine to generate responses from virtual characters using generative AI technology. During this process, the latest news and trend information obtained from external sources is also analyzed, and a model is trained to provide personalized responses to the user. The generated responses are output as customized messages that take into account the user profile and emotional state.
[0825] Step 5:
[0826] The server sends the generated response message to the terminal. The terminal displays or plays the received message audibly to the user. Specifically, if it's late at night, it will display a message appropriate to the time, such as "Thank you for your hard work today, good night," enabling a response at the right time.
[0827] (Application Example 2)
[0828] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0829] The present invention aims to provide users with a more natural and human-like dialogue experience. However, existing dialogue systems with virtual characters do not adequately provide personalized responses that take into account the user's emotions, making it difficult to achieve empathetic communication based on emotions.
[0830] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0831] In this invention, the server includes means for collecting and storing information relating to the user's preferences and interests; means for training a model using data processing techniques to mimic the personality and speech patterns of a designated character; means for updating the model based on the latest information obtained from external sources; and means for identifying the user's emotions using an emotion engine and adjusting the content and expression of responses. This makes it possible to provide the user with empathetic, emotion-based dialogue and offer support and reminders as needed.
[0832] A "specific virtual character" is a digital entity designed to enable natural language interaction with a user, possessing a specific personality and speech pattern.
[0833] "Natural language" refers to the language forms that users use on a daily basis, and the words that are expressed through speech or text.
[0834] An "information processing system" is a technical mechanism that inputs, processes, and stores data, and provides output as needed.
[0835] "Preferences and interests" refers to information that indicates a user's personal preferences and interests, including their hobbies and areas of interest.
[0836] "Data processing technology" refers to a group of technologies for efficiently processing large amounts of data and utilizing it as information.
[0837] "Training a model" is the act of adjusting the parameters of an algorithm using data to improve its ability to perform a particular task.
[0838] "External information sources" refer to databases and online resources that provide information from outside the system.
[0839] An "emotion engine" is a system that identifies emotions from user input and selects or adjusts appropriate responses based on those emotions.
[0840] "Empathetic dialogue" refers to a type of dialogue that is flexible, responsive to the user's emotions and state of mind, and conducted with empathy.
[0841] To implement this invention, the user first accesses the system using a dedicated terminal. The terminal receives natural language input via voice or text and transmits that data to the server. The terminal is equipped with an interface for the user to input information about their preferences and interests.
[0842] Next, the server processes the received data. The server is designed based on a small computer like a Raspberry Pi and uses Python as its main programming language. The server converts the speech data to text using the Google Speech-to-Text API and implements an emotion engine using a custom model based on OpenAI's GPT-3. This runs an algorithm to identify emotions from user input and generate responses. The server keeps the model up-to-date and provides more appropriate responses by regularly incorporating updates from external sources.
[0843] As a concrete example, suppose a user types "I'm tired today" into their device. The server, upon receiving this sentence, uses an emotion engine to recognize the negative emotion and generates a response such as "You had a tough day, well done." This generated response is sent back to the user via the device in either voice or text, resulting in a natural conversation.
[0844] The generative AI model in this system plays a crucial role in flexibly adjusting responses based on the user's emotions. Examples of prompts include the following:
[0845] example:
[0846] User: I'm really tired today.
[0847] System: That sounds tough. Is there anything I can do to help?
[0848] In this way, the terminal and server can work together to provide personalized conversational services to users.
[0849] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0850] Step 1:
[0851] The user uses a device to input messages in natural language. This input includes voice or text data. The input data is received by the device and sent to the server.
[0852] Step 2:
[0853] The server uses the Google Speech-to-Text API to convert the received audio data into text. This conversion process analyzes the audio waveform data and generates corresponding text data. This text data serves as input for the next step in sentiment analysis.
[0854] Step 3:
[0855] The server processes text data using a sentiment engine based on a custom model from OpenAI's GPT-3. In this step, the user's emotions are classified from the text data as positive, negative, or neutral, and sentiment metadata is generated based on the results.
[0856] Step 4:
[0857] The server uses user profile data and the latest external information to generate an appropriate response for the user. This process utilizes a generative AI model to generate the optimal response text based on the prompt. The tone and context of the response are adjusted based on the user's sentiment metadata.
[0858] Step 5:
[0859] The server sends the generated response data to the terminal. The terminal provides feedback to the user as voice or text. For voice feedback, a speech synthesis tool converts the text data into natural-sounding speech.
[0860] Step 6:
[0861] The user receives feedback through the device, and the interaction is completed. New input or responses from the user restart the process from step 1, enabling continuous interaction.
[0862] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0863] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0864] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0865] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0866] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0867] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0868] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0869] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0870] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0871] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0872] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0873] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0874] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0875] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0876] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0877] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0878] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0879] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0880] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0881] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0882] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0883] The following is further disclosed regarding the embodiments described above.
[0884] (Claim 1)
[0885] An information processing system for enabling natural language communication with a specific virtual character,
[0886] Means for collecting and storing data on users' preferences and interests,
[0887] A means of training a model using information processing technology to imitate the personality and speech patterns of a specified character,
[0888] A means of updating the model based on the latest information obtained from external sources,
[0889] A means for analyzing user input and generating an appropriate response using a trained and updated model,
[0890] A means of generating an automated message at the appropriate time using real-world time information,
[0891] A system that includes this.
[0892] (Claim 2)
[0893] The system according to claim 1, wherein the information processing system generates an individualized response based on the user's profile.
[0894] (Claim 3)
[0895] The system according to claim 1, wherein the information processing system generates messages in conjunction with real-world schedule information.
[0896] "Example 1"
[0897] (Claim 1)
[0898] A means of collecting data from users via a terminal and sending that data to a server,
[0899] A means of updating the database based on the data received by the server and generating individual user profiles,
[0900] A means of training a model using generative AI technology to mimic the personality and conversational style of a virtual character based on a generated user profile,
[0901] A means of obtaining current news and weather from external sources and reflecting them in the character's knowledge base,
[0902] A means for analyzing messages sent by users and generating responses based on user profiles and latest information,
[0903] A means of sending time-sensitive automated messages to users using real-world time information,
[0904] A system that includes this.
[0905] (Claim 2)
[0906] The system according to claim 1, which personalizes the responses of a virtual character based on the user's profile.
[0907] (Claim 3)
[0908] The system according to claim 1, which generates an automated message based on real-world time information and sends it to a user.
[0909] "Application Example 1"
[0910] (Claim 1)
[0911] An information processing device for enabling dialogue with a specific virtual entity using natural language,
[0912] Means for collecting and storing data on users' preferences and interests,
[0913] A means for training a computational model using information processing technology to mimic the characteristics and speech patterns of a specified virtual entity,
[0914] A means of improving the computational model based on the latest data obtained from external sources,
[0915] A means for analyzing user input and generating an appropriate response using a trained and improved computational model,
[0916] A means of generating an automatic message at the appropriate time using real-world time information,
[0917] A means of providing interactive dialogue with virtual entities across different devices,
[0918] A device that includes this.
[0919] (Claim 2)
[0920] The apparatus according to claim 1, wherein the information processing device generates personalized content based on the user's profile.
[0921] (Claim 3)
[0922] The apparatus according to claim 1, wherein the information processing device generates a message in conjunction with real-world event information.
[0923] "Example 2 of combining an emotion engine"
[0924] (Claim 1)
[0925] Means for collecting and storing data on users' preferences and interests,
[0926] A means of training a model using generative AI technology to imitate the personality and speaking style of a specified character,
[0927] A means of analyzing user input using an emotion engine to identify emotional states such as positive, negative, and neutral,
[0928] A means of updating the model based on the latest information obtained from external sources and generating personalized responses,
[0929] A means of generating automatic messages according to time using real-world time information,
[0930] A system that includes this.
[0931] (Claim 2)
[0932] The system according to claim 1, which adjusts the tone and content of the response based on the emotional state analyzed by the emotion engine.
[0933] (Claim 3)
[0934] The system according to claim 1, which generates user-friendly messages in conjunction with real-world schedule information.
[0935] "Application example 2 when combining with an emotional engine"
[0936] (Claim 1)
[0937] A data processing system for enabling natural language interaction with a specific virtual character,
[0938] A device for collecting and storing information about users' preferences and interests,
[0939] A device that uses data processing technology to train a model in order to imitate the personality and speech patterns of a specified character,
[0940] A device that updates the model based on the latest information obtained from external sources,
[0941] A device that analyzes user input and generates appropriate responses using trained and updated models,
[0942] A device that automatically generates messages at appropriate times using real-world time information,
[0943] A device that uses an emotion engine to identify the user's emotions and adjust the content and expression of the response,
[0944] A system that includes this.
[0945] (Claim 2)
[0946] The system according to claim 1, wherein the information processing system is installed in a general service robot in an environment and enables dialogue that is attentive to the user's inner state.
[0947] (Claim 3)
[0948] The system according to claim 1, wherein the information processing system generates and provides schedule adjustments and reminders based on the user's behavior history and circumstances. [Explanation of symbols]
[0949] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An information processing system for enabling natural language communication with a specific virtual character, Means for collecting and storing data on users' preferences and interests, A means of training a model using information processing technology to imitate the personality and speech patterns of a specified character, A means of updating the model based on the latest information obtained from external sources, A means for analyzing user input and generating an appropriate response using a trained and updated model, A means of generating an automated message at the appropriate time using real-world time information, A system that includes this.
2. The system according to claim 1, wherein the information processing system generates an individualized response based on the user's profile.
3. The system according to claim 1, wherein the information processing system generates messages in conjunction with real-world schedule information.