System
The system facilitates realistic conversations with deceased or historical figures by generating AI models that mimic their thought patterns and speaking styles, addressing the limitations of conventional technologies and enhancing user experience.
Patent Information
- Application Number
- JP2024118095
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Conventional technologies lack the means to enable realistic conversations with deceased individuals or historical figures, limiting personal healing and academic research due to the difficulty in accurately mimicking their thought patterns and speaking styles.
A system that allows users to select a person for interaction, retrieves relevant data, analyzes it to generate an AI model mimicking the person's thought patterns and speaking style, and generates responses to user questions, using natural language processing and incremental learning to enhance accuracy.
Enables highly realistic conversations with deceased or historical figures, contributing to academic research and personal healing, and improving user experience through accurate and natural interactions.
Smart Images

Figure 2026017313000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The problem that this invention aims to solve is to fulfill people's desire to converse with deceased people or historical figures. Conventional technology lacks a means to realize such dialogue, and there are limited ways to enjoy conversations with deceased loved ones. It is also difficult to relate the thoughts and speaking styles of historical figures to contemporary issues and discuss them. This has resulted in a lack of means to contribute to academic research and personal healing. [Means for solving the problem]
[0005] The present invention solves these problems by providing a system that includes a means for a user to select a person with whom they want to converse, a means for retrieving data about the selected person from a database, a means for analyzing the retrieved data and generating an AI model to mimic the person's thought patterns and speaking style, a means for receiving a question from the user and generating a response based on the AI model, and a means for displaying the generated response to the user. This system allows users to enjoy natural conversations with deceased people and historical figures, and can also contribute to academic research and personal healing.
[0006] "User" refers to any individual or entity who wishes to interact with the system.
[0007] "Selection means" refers to the interface or method by which a user selects the person on the system with whom they wish to interact.
[0008] "Database" refers to a storage device within the system that stores and manages various information related to the person being interacted with, and the contents of that information.
[0009] "Data Retrieval Methods" refers to the processes and techniques used to search and retrieve information related to selected persons from a database.
[0010] "Analysis methods" refers to the process and techniques for analyzing a person's thought patterns and speaking style based on the data obtained.
[0011] An "AI model" refers to a machine learning algorithm that generates dialogue by imitating a person's way of thinking and speaking based on analyzed information.
[0012] "Question receiving means" refers to the interface and technology that receives a question input by a user and transmits it to the system.
[0013] "Response Generation" refers to the processes and techniques that generate appropriate answers to user questions based on AI models.
[0014] "Response display means" refers to interfaces and techniques that visually present generated responses to a user.
[0015] "Natural language processing technology" refers to computer algorithms that analyze text data in natural language and understand its grammar and meaning.
[0016] "Incremental learning" refers to a machine learning technique that continuously learns by adding new data to an existing AI model. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user, generates an AI model, and realizes the conversation.
[0039] System Overview
[0040] The user accesses the system and selects the person they want to talk to. The device sends that information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. This AI model is used to generate responses to the user's questions and displays them to the user via the device.
[0041] Program processing
[0042] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[0043] 1. A user accesses the system
[0044] A user accesses the system using a web browser or application. A login screen appears and the user enters their authentication information to log in to the system.
[0045] 2. The user selects the person they want to interact with
[0046] The user clicks the "Start a conversation" button on the main screen and selects historical figure A from the displayed list or search field.
[0047] 3. The device sends the selection information to the server
[0048] The device sends the user's selection information to the server as an HTTP request, which includes the ID of the selected person.
[0049] 4. The server retrieves the data
[0050] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[0051] 5. The server analyzes the data and generates an AI model
[0052] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[0053] 6. User enters question
[0054] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[0055] 7. The server generates a response
[0056] The server analyzes the question and uses an AI model to generate an appropriate response, such as, "The modern economy contains many of the contradictions inherent in capitalist society. The solution to this is..."
[0057] 8. The generated response is displayed to the user
[0058] The server sends the generated response to the terminal, which displays the response on the screen, allowing the user to enjoy a conversation with Great Person A.
[0059] Specific examples
[0060] Example 1: A conversation with historical figure A
[0061] A similar process is repeated if a user wants to talk with historical figure A about the "current political situation." For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..."
[0062] Example 2: Dialogue with the deceased
[0063] If a user wants to imitate a conversation with their late grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if a user mentions "how you've been taking care of your garden lately," the grandmother's AI can respond with something like, "Oh, that's wonderful. I always enjoy watching you take care of your garden."
[0064] Technical details
[0065] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. It also uses incremental learning to continuously update the AI model based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic conversations with deceased people and historical figures.
[0066] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[0067] The processing flow will be explained below.
[0068] Step 1:
[0069] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0070] Step 2:
[0071] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[0072] Step 3:
[0073] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0074] Step 4:
[0075] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[0076] Step 5:
[0077] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[0078] Step 6:
[0079] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[0080] Step 7:
[0081] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[0082] Step 8:
[0083] The server analyzes the question and generates a response based on the AI model. The server tokenizes the user's question, analyzes its meaning, and uses the AI model to generate an appropriate response. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism."
[0084] Step 9:
[0085] The server sends the generated response to the terminal, which then sends the response as an HTTP response.
[0086] Step 10:
[0087] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[0088] Step 11:
[0089] If the user wishes to enter more questions, they enter the questions again, and the same process is repeated once the new questions are entered.
[0090] Example 1
[0091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0092] Conventional communication systems have struggled to enable users to converse with deceased or historical figures in real time. Technical limitations exist, particularly in situations where accurate mimicking of the target person's thought patterns and speaking style is required. This has made it difficult to realize such systems in a wide range of applications, including education, entertainment, and personal healing.
[0093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0094] In this invention, the server includes a means for selecting a person with whom the user wants to converse, a means for acquiring information about the selected person from a database, and a means for analyzing the acquired information and generating an AI model for imitating the thought patterns and speaking style of the person, thereby enabling the user to enjoy realistic and sophisticated conversations with deceased people and historical figures.
[0095] A "user" is someone who uses the system to interact with a specific person.
[0096] A "person with whom a user wishes to converse" is a deceased person or historical figure that the user selects and with whom the user wishes to converse.
[0097] "Means for selection" is a function that provides an interface for the user to select the person with whom they want to interact from a list or search field.
[0098] "Information" refers to data such as books, papers, transcripts, letters, and audio data related to the target person.
[0099] A "database" is a digital information storage system for storing and managing information within a system.
[0100] The "means for obtaining" is a function for searching and extracting information about a selected person from a database.
[0101] An "artificial intelligence model" is a machine learning model that is generated to mimic the thought patterns and speaking style of a target person.
[0102] "Means of analysis" refers to the function of analyzing acquired information using natural language processing technology and extracting important keywords and phrases.
[0103] The "means for generating a response" is a function that uses an artificial intelligence model to generate an appropriate response based on a user's question.
[0104] The "display means" is an interface for displaying the generated response on the user's terminal screen.
[0105] "Natural language processing technology" is a technology that enables computers to understand and process human language.
[0106] "Keywords" are important words or phrases extracted during the process of analyzing information about a target person.
[0107] The present invention provides a communication system that allows users to enjoy conversations with deceased or historical figures. This system allows users to select the person they want to talk to, acquires and analyzes information about that person, generates an artificial intelligence model, and realizes the conversation.
[0108] Hardware and software used
[0109] Hardware: Servers, devices (PCs, smartphones, tablets, etc.)
[0110] Software: web browsers, applications, database management systems, natural language processing libraries (e.g., Python's NLTK, SpaCy), generative AI models (e.g., OpenAI's GPT-3, Google's BERT)
[0111] Data processing and calculation
[0112] 1. Data Acquisition: The server acquires information about the target person from a database, specifically in the form of books, papers, transcripts, letters, audio data, etc.
[0113] 2. Data Analysis: The server analyzes the acquired information using natural language processing techniques, such as Python's NLTK and SpaCy, to extract important keywords and phrases and model the target person's thought patterns and speaking style.
[0114] 3. Creating a generative AI model: Based on the analyzed data, an artificial intelligence model is generated using OpenAI's GPT-3 or Google's BERT. This model mimics the target person's speaking style and thinking patterns.
[0115] 4. User Question Analysis: The server analyzes the question entered by the user. This analysis involves using natural language processing techniques to understand the meaning and extract relevant keywords.
[0116] 5. Response Generation: Based on the analyzed question, the server uses the generated artificial intelligence model to generate an appropriate response that replicates the target person's thought patterns and speaking style.
[0117] 6. Displaying the response: The server sends the generated response to the terminal, which displays it on the user's screen, allowing the user to enjoy the interaction.
[0118] Specific examples
[0119] For example, suppose a user wants to converse with "Albert Einstein." The user enters "Albert Einstein" into the application's search field and selects it. The server collects papers, notes, and speech data related to Einstein from the database and analyzes them using natural language processing technology. Based on the analysis results, an artificial intelligence model of Einstein is generated using OpenAI's GPT-3.
[0120] When a user inputs a question such as "What do you think about the modern economy?", the server analyzes the question and uses the generated artificial intelligence model to generate a response such as "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." This response is sent to the terminal and displayed on the user's screen.
[0121] Prompt Sentence Examples
[0122] Enter: "What do you think about the modern economy?"
[0123] Response: "The modern economy contains many of the contradictions of capitalist society. The solution to these is..."
[0124] This system allows users to enjoy highly realistic and sophisticated interactions with deceased or historical figures, which is expected to have applications in a variety of fields, including education, entertainment, and personal healing.
[0125] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0126] Step 1:
[0127] A user accesses the system. The user accesses the system using a web browser or application and enters an ID and password on the login screen. The input information is sent from the terminal to the server as an HTTPS request. The server receives this information and performs authentication by checking it against a database. If authentication is successful, the server sends the main screen data to the terminal and login is complete. The input is authentication information (ID and password), and the output is the user's authentication status and the display of the main screen.
[0128] Step 2:
[0129] The user selects the person they want to talk to. The user clicks the "Start a conversation" button displayed on the main screen and selects the historical figure or deceased person they want to talk to from a search field or list. The device sends an HTTP request to the server including the selection information (person ID). The input is the selection information of the person they want to talk to, and the output is the request sent to the server.
[0130] Step 3:
[0131] The device sends the selection information to the server. The device sends an HTTP request to the server, including the ID of the selected person. The server receives this request and executes a query to retrieve information about the selected person from a database. The input is the ID of the selected person, and the output is the result of executing the database query.
[0132] Step 4:
[0133] The server retrieves the data. The server searches and retrieves information related to the selected person from its database (books, papers, transcripts, letters, audio recordings, etc.). This information is returned to the server in JSON format. The input is a database query, and the output is the information retrieved about the target person.
[0134] Step 5:
[0135] The server analyzes the data and generates an AI model. The server uses natural language processing (NLP) techniques to analyze the acquired information, using Python's NLTK and SpaCy libraries. During the analysis, important keywords and phrases are extracted, and based on this, a model is created of the target person's thought patterns and speaking style. OpenAI's GPT-3 and Google's BERT are used for generation. The input is the information to be analyzed, and the output is the generated generative AI model.
[0136] Step 6:
[0137] The user enters a question. The user enters the question into the application's input form, and the question is sent from the terminal to the server. The input is the user's question, and the output is the question sent to the server.
[0138] Step 7:
[0139] The server generates a response. The server analyzes the input question and uses a generative AI model to generate an appropriate response. It uses natural language processing technology to understand the meaning of the question and extract relevant keywords. For example, in response to the question, "What do you think about the modern economy?", it generates a response such as, "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." The input is the user's question, and the output is the generated response.
[0140] Step 8:
[0141] The generated response is displayed to the user. The server sends the generated response in JSON format to the terminal. The terminal displays the received response data on the screen, allowing the user to enjoy the interaction. The input is the generated response, and the output is the response displayed on the screen.
[0142] (Application example 1)
[0143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0144] Currently, there are technologies that allow users to enjoy conversations with historical figures and deceased people, but these conversations can only be realized on limited devices, which often limits the experience. Furthermore, it is difficult to visually display the content of the conversation, which does not improve the user experience. Furthermore, there is a lack of technology to improve the accuracy of the conversation content desired by users, making it difficult to realize realistic conversations.
[0145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0146] In this invention, the server includes: means for selecting a person the user wants to converse with, as input by the user; means for acquiring data about the selected person from a database; means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style; means for receiving the user's question and generating a response based on the AI model; means installed in the smart glasses for displaying the content of the conversation with the person selected by the user; and means for displaying the generated response to the user. This allows the user to enjoy a visual conversation in real time using the smart glasses, improving the quality of the experience. Furthermore, it is possible to accurately imitate the thought patterns and speaking style of the selected person, thereby realizing a more realistic conversation.
[0147] The "means for selecting a person to converse with input by the user" is an interface within the system that allows the user to select a specific person as a conversation partner.
[0148] The "means for obtaining data relating to a selected person from a database" is a function for obtaining information relating to a person selected by a user from a database stored within the system.
[0149] "Means for analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style" refers to technology for analyzing the data of a selected person and creating an AI model that reproduces the person's speaking style and thought patterns.
[0150] "Means for receiving a user's question and generating a response based on an AI model" refers to the function by which the system receives a question from a user and the AI model generates an appropriate response based on that question.
[0151] "A means installed in smart glasses to display the content of the conversation with a person selected by the user" is a function installed in smart glasses to visually provide the user with the response of the person with whom they want to converse.
[0152] "Means for displaying the generated response to the user" refers to an interface that allows the user to see the response generated by the AI model.
[0153] "Books, papers, transcripts, letters, and audio recordings" are sources of information related to the selected individuals from which data will be drawn and used in the analysis.
[0154] "Means of analyzing input data using natural language processing technology and extracting important keywords and phrases" is a technology that utilizes natural language processing technology to extract important information from user input and dialogue.
[0155] This invention is a system that allows a user to select a person with whom they want to converse, acquires and analyzes data about that person to generate an AI model, and realizes a dialogue with the user in real time. This system is installed in smart glasses and can visually display the dialogue content to the user. Below, a detailed description of an embodiment of this invention is provided.
[0156] System configuration
[0157] This system is configured using the following hardware and software.
[0158] Smart glasses: Devices that display interactions to the user in real time and also serve as an input / output interface.
[0159] Server: Responsible for back-end processing to acquire and analyze the data of the selected person and generate an AI model.
[0160] Database: Contains data (books, papers, transcripts, letters, audio recordings, etc.) about historical figures and deceased people.
[0161] Natural language processing technology: Technology used to analyze data and extract key keywords and phrases (e.g., GPT-2 model).
[0162] Program processing overview
[0163] 1. User inputs the person they want to talk to:
[0164] The user wearing the smart glasses selects the person they want to talk to. The user interface displays a list of people they want to talk to, allowing them to select one.
[0165] 2. Retrieving data from the database:
[0166] Data about the selected person (e.g., books, papers, transcripts, letters, audio recordings, etc.) is retrieved from a database by the server.
[0167] 3. Data analysis and AI model generation:
[0168] The server then analyzes the data using natural language processing techniques to extract key keywords and phrases, which then generates an AI model that mimics the thought patterns and speaking style of the selected person.
[0169] 4. Receiving questions and generating responses:
[0170] The user types a question through the smart glasses, which is sent to a server where an AI model analyzes it and generates an appropriate response.
[0171] 5. View the response:
[0172] The generated responses are displayed in real time on the smart glasses, enabling interaction with the user.
[0173] Specific examples
[0174] For example, consider a case where a user asks a historical scientist, "Tell me about the theory of relativity." In this case, the AI model uses the acquired data to generate a response such as, "The theory of relativity is based on the idea that time and space are relative, not absolute..." The interaction takes place in real time and is displayed through the smart glasses, improving the user experience.
[0175] Prompt Sentence Examples
[0176] User: Tell me about the theory of relativity.
[0177] AI: The theory of relativity is based on the idea that time and space are relative, not absolute...
[0178] Thus, the present invention provides a system that allows users to enjoy real-time interactions through smart glasses. The natural language processing technology used in the smart glasses can improve the quality of the interactions and the user experience.
[0179] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0180] Step 1:
[0181] The user puts on the smart glasses and selects the person they want to interact with. This selection information is displayed on the user interface, and once the user selects the target person, the selection information is retained in the smart glasses and sent to the server for the next step.
[0182] Input: User-selected person information
[0183] Output: Selection information sent to the server
[0184] Step 2:
[0185] The server retrieves data about the selected person. The server queries the database to retrieve related information such as books, papers, transcripts, letters, and audio recordings.
[0186] Input: Selection information sent from smart glasses
[0187] Output: Person-related data retrieved from the database
[0188] Step 3:
[0189] The server analyzes the acquired data using natural language processing techniques (e.g., the GPT-2 model) to extract important keywords and phrases from the data. The result is an AI model that mimics the individual's thought patterns and speaking style.
[0190] Input: Person-related data retrieved from a database
[0191] Output: AI model
[0192] Step 4:
[0193] The user inputs a question through the smart glasses, which receives the question and sends it to the server.
[0194] Input: The question entered by the user into the smart glasses
[0195] Output: The question sent to the server
[0196] Step 5:
[0197] The server uses an AI model to analyze the received question and generate an appropriate response. The AI model generates a response to the question and prepares it in text format.
[0198] Input: User-submitted question, AI model
[0199] Output: The generated response (in text format)
[0200] Step 6:
[0201] The generated response is sent to the smart glasses and displayed to the user, who can view the interaction content through the smart glasses.
[0202] Input: The generated response sent by the server
[0203] Output: Response displayed on the smart glasses
[0204] In this way, this system provides an environment in which users can interact with historical figures and deceased people in real time using smart glasses, and delivers a high-quality interactive experience by analyzing the acquired data and generating responses using an AI model.
[0205] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0206] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model, which then combines it with an emotion engine to recognize the user's emotions and realize the conversation.
[0207] System Overview
[0208] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[0209] Program processing
[0210] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[0211] 1. A user accesses the system
[0212] A user logs into the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0213] 2. The user selects the person they want to interact with
[0214] The user clicks the "Start conversation" button displayed on the main screen and selects historical figure A from the displayed list or search field.
[0215] 3. The device sends the selection information to the server
[0216] The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0217] 4. The server retrieves the data
[0218] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[0219] 5. The server analyzes the data and generates an AI model
[0220] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[0221] 6. User enters question
[0222] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[0223] 7. The server analyzes the user's emotions using an emotion engine.
[0224] The server analyzes the user's question with an emotion engine and processes the question content with emotion labels. For example, the question is tagged with emotion labels such as "interesting" and "excited."
[0225] 8. The server analyzes the question and generates a response based on the AI model
[0226] The server tokenizes the user's question, analyzes its meaning, and uses an AI model to generate an appropriate response. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism" in a calm, empathetic tone.
[0227] 9. Send the server-generated response to the device
[0228] The server sends the generated response to the terminal as an HTTP response.
[0229] 10. The device receives the response from the server and displays it to the user
[0230] The device displays the received text in a chat window on the screen for the user to view.
[0231] Specific examples
[0232] Example 1: A conversation with historical figure A
[0233] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[0234] Example 2: Dialogue with the deceased
[0235] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[0236] Technical details
[0237] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, enhancing the naturalness and intimacy of the conversation. Furthermore, by using incremental learning, the AI model is continuously updated based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures.
[0238] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[0239] The processing flow will be explained below.
[0240] Step 1:
[0241] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0242] Step 2:
[0243] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[0244] Step 3:
[0245] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0246] Step 4:
[0247] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[0248] Step 5:
[0249] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[0250] Step 6:
[0251] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[0252] Step 7:
[0253] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[0254] Step 8:
[0255] The server analyzes the user's emotions using an emotion engine. The server analyzes the user's question using the emotion engine and processes the question content along with an emotion label. For example, the question is tagged with an emotion label such as "interesting" or "excited."
[0256] Step 9:
[0257] The server analyzes the question and generates a response based on an AI model. The server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it could output something like, "The contradictions of capitalism are expanding in the modern economy," in a calm, empathetic tone.
[0258] Step 10:
[0259] The server sends the generated response to the terminal. The server sends the generated response to the terminal as an HTTP response.
[0260] Step 11:
[0261] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[0262] Step 12:
[0263] The user then enters another question, and the process repeats. The question and emotion label are sent to the server, and the AI model generates an appropriate response, which is then displayed on the device.
[0264] Specific examples
[0265] Example 1: A conversation with historical figure A
[0266] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[0267] Example 2: Dialogue with the deceased
[0268] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[0269] Example 2
[0270] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0271] Conventional communication systems have struggled to provide a natural dialogue experience for users to converse with deceased or historical figures. Furthermore, the content of the dialogue is not adjusted according to the user's emotions, resulting in a decline in dialogue quality. Furthermore, there are limitations to the accuracy of the technology used to realistically imitate a person's thought patterns and speaking style. This makes it difficult for users to obtain a sense of satisfaction from the dialogue.
[0272] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0273] In this invention, the server includes: means for selecting a person the user wants to converse with, means for retrieving data about the selected person from a database, means for analyzing the retrieved data and generating an AI model for imitating the person's thought patterns and speaking style, means for analyzing the user's emotions and assigning emotion labels, means for receiving the user's question and generating a response based on the AI model and the emotion labels, and means for displaying the generated response to the user. This allows the user to enjoy more natural and emotional conversations with deceased people and historical figures.
[0274] A "user" is a person who accesses and interacts with the system.
[0275] "People you want to talk to" refers to deceased people or historical figures with whom you want to talk through the system.
[0276] "Means for selection" refers to the interface or operation that allows a user to select a person within the system with whom they wish to interact.
[0277] "Database" refers to a magnetic disk drive or other storage system that stores information such as books, articles, transcripts, letters, email posts, audio recordings, etc. about a person with whom you wish to interact.
[0278] "Means for obtaining" refers to a process for searching and retrieving data about a selected person from a database.
[0279] "Means of analysis" refers to the process of analyzing data using natural language processing technology and extracting important keywords and phrases.
[0280] An "AI model" refers to an artificial intelligence program that is generated from acquired data and is intended to mimic the thinking patterns and speaking style of a specific person.
[0281] An "emotion engine" refers to a software module that analyzes emotions from user input and assigns emotion labels.
[0282] An "emotion label" refers to a tag that expresses an emotion, such as "interested" or "excited," in response to a user's input.
[0283] "Means for generating a response" refers to the process of analyzing the user's question and creating an appropriate response based on the AI model and emotion labels.
[0284] "Means for displaying" refers to a screen or interface for visually presenting the generated response to the user.
[0285] This invention relates to a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model. It then combines this data with an emotion engine to recognize the user's emotions and realize the conversation.
[0286] System Overview
[0287] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[0288] Program processing
[0289] When a user accesses the system and selects the person they wish to interact with, the device sends the corresponding person's ID to the server. The server retrieves data about that person from a database and analyzes it using natural language processing technology. For example, natural language processing libraries such as SpaCy and NLTK are used. As a result, important keywords and phrases are extracted and an AI model that mimics the target person's thought patterns and speaking style is generated.
[0290] When a user enters a question, the device sends it to the server. The server analyzes the question using an emotion engine and assigns an emotion label to the question using, for example, TextBlob or a sentiment analysis API. The server then tokenizes the question and generates a response based on an AI model. The tone and content of the response are adjusted based on the emotion engine's analysis results. The generated response is then sent to the device, which displays it to the user.
[0291] Specific examples
[0292] Example 1: A conversation with historical figure A
[0293] If a user wants to talk with historical figure A about the "current political situation," the user begins the conversation as follows:
[0294] Example prompt sentence:
[0295] "Ask historical figure A what he thinks about the current political situation."
[0296] The server retrieves data on great person A from the database, analyzes it, and generates an AI model. It then analyzes the user's question along with the emotion label and generates a response such as, "Compared to past historical events, the current political situation is..."
[0297] Example 2: Dialogue with the deceased
[0298] If the user wanted to mimic a conversation with their deceased grandmother, the user's prompt might look like this:
[0299] Example prompt sentence:
[0300] "Tell your late grandmother about your recent gardening."
[0301] The server generates an AI model based on the grandmother's letters and audio recordings, and generates responses such as "Oh, that's great. I always enjoyed tending to your garden" based on emotion labels.
[0302] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, improving the naturalness and intimacy of the dialogue. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures. This system has broad applications not only in academic research and personal healing, but also in education and entertainment.
[0303] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0304] Step 1:
[0305] A user accesses the system.
[0306] Specifically, a user accesses the system using a web browser or mobile app. The user enters login information (username and password) and presses the "Login" button. The input is the username and password string, and the output is the success or failure of authentication and the display of the main screen.
[0307] Step 2:
[0308] The user selects the person with whom they want to interact.
[0309] Specifically, the user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The input is the name or ID of the person they want to talk to, and the output is data including the ID of the selected person.
[0310] Step 3:
[0311] The terminal transmits the selection information to the server.
[0312] Specifically, the device generates an HTTP request including the ID of the selected person and sends it to the server. The input is JSON-formatted data including the ID of the selected person, and the output is an HTTP request to the server.
[0313] Step 4:
[0314] The server retrieves the data.
[0315] Specifically, the server searches and retrieves data about the selected person from a database. The input is the selected person's ID, and the output is data such as books, papers, transcripts, and letters retrieved from the database. The server retrieves this data from local storage or cloud storage.
[0316] Step 5:
[0317] The server analyzes the data and generates an AI model.
[0318] Specifically, the server analyzes the acquired data using natural language processing techniques (e.g., SpaCy or NLTK). Important keywords and phrases are extracted, and an AI model is generated using machine learning techniques (e.g., TensorFlow or PyTorch). The input is person data acquired from a database, and the output is an AI model that mimics the thinking patterns and speaking style of a specific person.
[0319] Step 6:
[0320] The user enters a question.
[0321] Specifically, the user enters a question into the text box on the interactive screen and presses the "Send" button. The input is the question entered by the user, and the output is the question data sent to the server.
[0322] Step 7:
[0323] The server analyzes the user's emotions using an emotion engine.
[0324] Specifically, the server analyzes the received question using an emotion engine (e.g., TextBlob or an emotion analysis API) and assigns an emotion label. The input is the user's question, and the output is the question data including the emotion label.
[0325] Step 8:
[0326] The server analyzes the question and generates a response based on an AI model.
[0327] Specifically, the server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. The input is the question data with emotion labels and the AI model, and the output is the generated response text.
[0328] Step 9:
[0329] The server generates a response and sends it to the terminal.
[0330] Specifically, the server sends the generated response in JSON format to the terminal as an HTTP response. The input is the generated response text, and the output is the HTTP response to the terminal.
[0331] Step 10:
[0332] The terminal receives the response from the server and displays it to the user.
[0333] Specifically, the device displays the received text in a chat window on the screen for the user to view. The input is the HTTP response from the server, and the output is the response text displayed in the chat window on the screen.
[0334] (Application example 2)
[0335] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0336] Conventional AI dialogue systems provide users with simple question and answer responses, but lack mechanisms for recognizing emotions and enabling natural dialogue. Furthermore, they lack an interactive experience in virtual space, resulting in low user satisfaction. The objective of this invention is to provide a dialogue system that provides a sense of realism by utilizing emotion recognition and virtual space.
[0337] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a person with whom the user wants to converse, input by the user, means for acquiring data about the selected person from a database, means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style, means including an emotion engine for recognizing and analyzing the user's emotions, means for adjusting the response content based on the analyzed emotions, and means for displaying this in a virtual space. This enables natural and realistic conversations that correspond to emotions.
[0338] The "means for selecting a person the user wishes to converse with" refers to hardware or software that provides an interface for the user to input or select a person with whom the user wishes to converse.
[0339] The "means for obtaining data on the selected person from the database" refers to a process or database communication means for accessing and obtaining data on the selected person from a database on a server or cloud.
[0340] "Means of analyzing acquired data and generating an AI model that imitates the person's thought patterns and speaking style" refers to technology that analyzes acquired data using natural language processing and machine learning techniques and generates an AI model that imitates the person.
[0341] The "means for receiving a user's question and generating a response based on an AI model" is a processing system that receives a user's question as input and generates an appropriate response using the algorithm of the AI model.
[0342] "Means for displaying the generated response to the user" refers to an interface or output device for displaying the response generated by the AI model on the user's terminal.
[0343] The "means including an emotion engine for recognizing and analyzing user emotions" refers to a software and hardware configuration for analyzing emotions from inputs and interactions from a user, and recognizing and processing the emotions.
[0344] The "means for adjusting the content of the response based on the analyzed emotions" is a technology that adaptively adjusts the tone and content of the generated response based on the user's emotions analyzed by the emotion engine.
[0345] "Means for displaying this in a virtual space" refers to the display devices and virtual space rendering technologies necessary for users to enjoy a realistic interaction in a virtual space.
[0346] System Configuration
[0347] A system for implementing this invention includes a user terminal, a server, a database, and an emotion engine. The user terminal is equipped with an interface that allows the user to select a person with whom they want to interact. The server accesses the database to obtain data about the person and generates an AI model. The server also uses the emotion engine to analyze the user's emotions and adjust the response content.
[0348] Program processing
[0349] When a user selects a person they want to talk to through their device, the device sends that information to a server. The server then retrieves data about the selected person from a database, such as books, papers, transcripts, letters, audio data, and video data. This data is analyzed using natural language processing technology to generate an AI model that mimics the person's thought patterns and speaking style. This generated AI model is then used to provide appropriate responses based on the user's questions.
[0350] When a user enters a question, it is sent to the server. The server uses an emotion engine to analyze the question's sentiment and adjusts the AI model's response based on the analysis results. After the response is generated, the server sends it to the user's device, which then displays it to the user in a virtual space.
[0351] Hardware and software used
[0352] User devices: smartphones, smart glasses, head-mounted displays
[0353] Server: Database Management System (DBMS), Web Server
[0354] Emotion engine: Emotion API (e.g. Microsoft Emotion API)
[0355] Natural language processing technology: OpenAI GPT, natural language tokenization libraries (e.g., spaCy)
[0356] Specific examples
[0357] For example, consider a scenario where a user enters a virtual cafe and chooses to interact with a great inventor. If the user asks, "What do you think about modern technology?", the server analyzes the question using its emotion engine and tags it as "interesting." Once the analysis is complete, the AI model generates a response such as, "Modern technology has advanced in ways never before imagined," which is then displayed on the user's device. The response is adjusted in tone to reflect the user's emotions.
[0358] Prompt Sentence Examples
[0359] Imagine the following dialogue:
[0360] You are a great inventor. You have had a huge impact on modern life with your inventions, and you have many inventions to your name.
[0361] User: What do you think about modern technology? [Interested]
[0362] In this way, the system can realize natural and realistic dialogue that responds to emotions, providing users with an interactive experience.
[0363] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0364] Step 1:
[0365] The user selects the person they want to talk to through their terminal. The person selection information entered by the user is sent from the terminal to the server. The ID of the person selected by the user is included as input, and the server receives it as output.
[0366] Step 2:
[0367] The server uses the received person ID to retrieve data about the selected person from a database, including books, papers, transcripts, letters, audio data, and video data. The person ID is used as input, and the target data is obtained as output.
[0368] Step 3:
[0369] The server analyzes the acquired data and generates an AI model to mimic the person's thought patterns and speaking style. It uses natural language processing technology to tokenize the data and extract important keywords and phrases. The acquired data is used as input, and the AI model is generated as output.
[0370] Step 4:
[0371] The user inputs a question through a terminal. The question is sent from the terminal to the server. The input contains the user's question, and the server receives it as output.
[0372] Step 5:
[0373] The server uses an emotion engine to perform emotion analysis based on the user's question. The analysis generates an emotion label for the question. The user's question is used as input, and the emotion label is generated as output.
[0374] Step 6:
[0375] The server uses an AI model to generate an appropriate response based on the analyzed emotion label. The response content is adjusted to take the user's emotion into account. The user's question and emotion label are used as input, and an adjusted response is generated as output.
[0376] Step 7:
[0377] The server sends the generated and adjusted response to the user terminal, which includes the generated response as input and receives it as output.
[0378] Step 8:
[0379] The terminal displays the received response to the user in the virtual space. The displayed response is reflected in the interface so that the user can see it in real time. The response sent to the terminal is used as input, and the response is displayed in the virtual space as output.
[0380] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0381] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0382] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0383] [Second embodiment]
[0384] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0385] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0386] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0387] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0388] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0389] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0390] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0391] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0392] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0393] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0394] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0395] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0396] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user, generates an AI model, and realizes the conversation.
[0397] System Overview
[0398] The user accesses the system and selects the person they want to talk to. The device sends that information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. This AI model is used to generate responses to the user's questions and displays them to the user via the device.
[0399] Program processing
[0400] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[0401] 1. A user accesses the system
[0402] A user accesses the system using a web browser or application. A login screen appears and the user enters their authentication information to log in to the system.
[0403] 2. The user selects the person they want to interact with
[0404] The user clicks the "Start a conversation" button on the main screen and selects historical figure A from the displayed list or search field.
[0405] 3. The device sends the selection information to the server
[0406] The device sends the user's selection information to the server as an HTTP request, which includes the ID of the selected person.
[0407] 4. The server retrieves the data
[0408] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[0409] 5. The server analyzes the data and generates an AI model
[0410] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[0411] 6. User enters question
[0412] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[0413] 7. The server generates a response
[0414] The server analyzes the question and uses an AI model to generate an appropriate response, such as, "The modern economy contains many of the contradictions inherent in capitalist society. The solution to this is..."
[0415] 8. The generated response is displayed to the user
[0416] The server sends the generated response to the terminal, which displays the response on the screen, allowing the user to enjoy a conversation with Great Person A.
[0417] Specific examples
[0418] Example 1: A conversation with historical figure A
[0419] A similar process is repeated if a user wants to talk with historical figure A about the "current political situation." For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..."
[0420] Example 2: Dialogue with the deceased
[0421] If a user wants to imitate a conversation with their late grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if a user mentions "how you've been taking care of your garden lately," the grandmother's AI can respond with something like, "Oh, that's wonderful. I always enjoy watching you take care of your garden."
[0422] Technical details
[0423] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. It also uses incremental learning to continuously update the AI model based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic conversations with deceased people and historical figures.
[0424] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[0425] The processing flow will be explained below.
[0426] Step 1:
[0427] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0428] Step 2:
[0429] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[0430] Step 3:
[0431] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0432] Step 4:
[0433] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[0434] Step 5:
[0435] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[0436] Step 6:
[0437] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[0438] Step 7:
[0439] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[0440] Step 8:
[0441] The server analyzes the question and generates a response based on the AI model. The server tokenizes the user's question, analyzes its meaning, and uses the AI model to generate an appropriate response. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism."
[0442] Step 9:
[0443] The server sends the generated response to the terminal, which then sends the response as an HTTP response.
[0444] Step 10:
[0445] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[0446] Step 11:
[0447] If the user wishes to enter more questions, they enter the questions again, and the same process is repeated once the new questions are entered.
[0448] Example 1
[0449] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0450] Conventional communication systems have struggled to enable users to converse with deceased or historical figures in real time. Technical limitations exist, particularly in situations where accurate mimicking of the target person's thought patterns and speaking style is required. This has made it difficult to realize such systems in a wide range of applications, including education, entertainment, and personal healing.
[0451] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0452] In this invention, the server includes a means for selecting a person with whom the user wants to converse, a means for acquiring information about the selected person from a database, and a means for analyzing the acquired information and generating an AI model for imitating the thought patterns and speaking style of the person, thereby enabling the user to enjoy realistic and sophisticated conversations with deceased people and historical figures.
[0453] A "user" is someone who uses the system to interact with a specific person.
[0454] A "person with whom a user wishes to converse" is a deceased person or historical figure that the user selects and with whom the user wishes to converse.
[0455] "Means for selection" is a function that provides an interface for the user to select the person with whom they want to interact from a list or search field.
[0456] "Information" refers to data such as books, papers, transcripts, letters, and audio data related to the target person.
[0457] A "database" is a digital information storage system for storing and managing information within a system.
[0458] The "means for obtaining" is a function for searching and extracting information about a selected person from a database.
[0459] An "artificial intelligence model" is a machine learning model that is generated to mimic the thought patterns and speaking style of a target person.
[0460] "Means of analysis" refers to the function of analyzing acquired information using natural language processing technology and extracting important keywords and phrases.
[0461] The "means for generating a response" is a function that uses an artificial intelligence model to generate an appropriate response based on a user's question.
[0462] The "display means" is an interface for displaying the generated response on the user's terminal screen.
[0463] "Natural language processing technology" is a technology that enables computers to understand and process human language.
[0464] "Keywords" are important words or phrases extracted during the process of analyzing information about a target person.
[0465] The present invention provides a communication system that allows users to enjoy conversations with deceased or historical figures. This system allows users to select the person they want to talk to, acquires and analyzes information about that person, generates an artificial intelligence model, and realizes the conversation.
[0466] Hardware and software used
[0467] Hardware: Servers, devices (PCs, smartphones, tablets, etc.)
[0468] Software: web browsers, applications, database management systems, natural language processing libraries (e.g., Python's NLTK, SpaCy), generative AI models (e.g., OpenAI's GPT-3, Google's BERT)
[0469] Data processing and calculation
[0470] 1. Data Acquisition: The server acquires information about the target person from a database, specifically in the form of books, papers, transcripts, letters, audio data, etc.
[0471] 2. Data Analysis: The server analyzes the acquired information using natural language processing techniques, such as Python's NLTK and SpaCy, to extract important keywords and phrases and model the target person's thought patterns and speaking style.
[0472] 3. Creating a generative AI model: Based on the analyzed data, an artificial intelligence model is generated using OpenAI's GPT-3 or Google's BERT. This model mimics the target person's speaking style and thinking patterns.
[0473] 4. User Question Analysis: The server analyzes the question entered by the user. This analysis involves using natural language processing techniques to understand the meaning and extract relevant keywords.
[0474] 5. Response Generation: Based on the analyzed question, the server uses the generated artificial intelligence model to generate an appropriate response that replicates the target person's thought patterns and speaking style.
[0475] 6. Displaying the response: The server sends the generated response to the terminal, which displays it on the user's screen, allowing the user to enjoy the interaction.
[0476] Specific examples
[0477] For example, suppose a user wants to converse with "Albert Einstein." The user enters "Albert Einstein" into the application's search field and selects it. The server collects papers, notes, and speech data related to Einstein from the database and analyzes them using natural language processing technology. Based on the analysis results, an artificial intelligence model of Einstein is generated using OpenAI's GPT-3.
[0478] When a user inputs a question such as "What do you think about the modern economy?", the server analyzes the question and uses the generated artificial intelligence model to generate a response such as "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." This response is sent to the terminal and displayed on the user's screen.
[0479] Prompt Sentence Examples
[0480] Enter: "What do you think about the modern economy?"
[0481] Response: "The modern economy contains many of the contradictions of capitalist society. The solution to these is..."
[0482] This system allows users to enjoy highly realistic and sophisticated interactions with deceased or historical figures, which is expected to have applications in a variety of fields, including education, entertainment, and personal healing.
[0483] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0484] Step 1:
[0485] A user accesses the system. The user accesses the system using a web browser or application and enters an ID and password on the login screen. The input information is sent from the terminal to the server as an HTTPS request. The server receives this information and performs authentication by checking it against a database. If authentication is successful, the server sends the main screen data to the terminal and login is complete. The input is authentication information (ID and password), and the output is the user's authentication status and the display of the main screen.
[0486] Step 2:
[0487] The user selects the person they want to talk to. The user clicks the "Start a conversation" button displayed on the main screen and selects the historical figure or deceased person they want to talk to from a search field or list. The device sends an HTTP request to the server including the selection information (person ID). The input is the selection information of the person they want to talk to, and the output is the request sent to the server.
[0488] Step 3:
[0489] The device sends the selection information to the server. The device sends an HTTP request to the server, including the ID of the selected person. The server receives this request and executes a query to retrieve information about the selected person from a database. The input is the ID of the selected person, and the output is the result of executing the database query.
[0490] Step 4:
[0491] The server retrieves the data. The server searches and retrieves information related to the selected person from its database (books, papers, transcripts, letters, audio recordings, etc.). This information is returned to the server in JSON format. The input is a database query, and the output is the information retrieved about the target person.
[0492] Step 5:
[0493] The server analyzes the data and generates an AI model. The server uses natural language processing (NLP) techniques to analyze the acquired information, using Python's NLTK and SpaCy libraries. During the analysis, important keywords and phrases are extracted, and based on this, a model is created of the target person's thought patterns and speaking style. OpenAI's GPT-3 and Google's BERT are used for generation. The input is the information to be analyzed, and the output is the generated generative AI model.
[0494] Step 6:
[0495] The user enters a question. The user enters the question into the application's input form, and the question is sent from the terminal to the server. The input is the user's question, and the output is the question sent to the server.
[0496] Step 7:
[0497] The server generates a response. The server analyzes the input question and uses a generative AI model to generate an appropriate response. It uses natural language processing technology to understand the meaning of the question and extract relevant keywords. For example, in response to the question, "What do you think about the modern economy?", it generates a response such as, "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." The input is the user's question, and the output is the generated response.
[0498] Step 8:
[0499] The generated response is displayed to the user. The server sends the generated response in JSON format to the terminal. The terminal displays the received response data on the screen, allowing the user to enjoy the interaction. The input is the generated response, and the output is the response displayed on the screen.
[0500] (Application example 1)
[0501] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0502] Currently, there are technologies that allow users to enjoy conversations with historical figures and deceased people, but these conversations can only be realized on limited devices, which often limits the experience. Furthermore, it is difficult to visually display the content of the conversation, which does not improve the user experience. Furthermore, there is a lack of technology to improve the accuracy of the conversation content desired by users, making it difficult to realize realistic conversations.
[0503] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0504] In this invention, the server includes: means for selecting a person the user wants to converse with, as input by the user; means for acquiring data about the selected person from a database; means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style; means for receiving the user's question and generating a response based on the AI model; means installed in the smart glasses for displaying the content of the conversation with the person selected by the user; and means for displaying the generated response to the user. This allows the user to enjoy a visual conversation in real time using the smart glasses, improving the quality of the experience. Furthermore, it is possible to accurately imitate the thought patterns and speaking style of the selected person, thereby realizing a more realistic conversation.
[0505] The "means for selecting a person to converse with input by the user" is an interface within the system that allows the user to select a specific person as a conversation partner.
[0506] The "means for obtaining data relating to a selected person from a database" is a function for obtaining information relating to a person selected by a user from a database stored within the system.
[0507] "Means for analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style" refers to technology for analyzing the data of a selected person and creating an AI model that reproduces the person's speaking style and thought patterns.
[0508] "Means for receiving a user's question and generating a response based on an AI model" refers to the function by which the system receives a question from a user and the AI model generates an appropriate response based on that question.
[0509] "A means installed in smart glasses to display the content of the conversation with a person selected by the user" is a function installed in smart glasses to visually provide the user with the response of the person with whom they want to converse.
[0510] "Means for displaying the generated response to the user" refers to an interface that allows the user to see the response generated by the AI model.
[0511] "Books, papers, transcripts, letters, and audio recordings" are sources of information related to the selected individuals from which data will be drawn and used in the analysis.
[0512] "Means of analyzing input data using natural language processing technology and extracting important keywords and phrases" is a technology that utilizes natural language processing technology to extract important information from user input and dialogue.
[0513] This invention is a system that allows a user to select a person with whom they want to converse, acquires and analyzes data about that person to generate an AI model, and realizes a dialogue with the user in real time. This system is installed in smart glasses and can visually display the dialogue content to the user. Below, a detailed description of an embodiment of this invention is provided.
[0514] System configuration
[0515] This system is configured using the following hardware and software.
[0516] Smart glasses: Devices that display interactions to the user in real time and also serve as an input / output interface.
[0517] Server: Responsible for back-end processing to acquire and analyze the data of the selected person and generate an AI model.
[0518] Database: Contains data (books, papers, transcripts, letters, audio recordings, etc.) about historical figures and deceased people.
[0519] Natural language processing technology: Technology used to analyze data and extract key keywords and phrases (e.g., GPT-2 model).
[0520] Program processing overview
[0521] 1. User inputs the person they want to talk to:
[0522] The user wearing the smart glasses selects the person they want to talk to. The user interface displays a list of people they want to talk to, allowing them to select one.
[0523] 2. Retrieving data from the database:
[0524] Data about the selected person (e.g., books, papers, transcripts, letters, audio recordings, etc.) is retrieved from a database by the server.
[0525] 3. Data analysis and AI model generation:
[0526] The server then analyzes the data using natural language processing techniques to extract key keywords and phrases, which then generates an AI model that mimics the thought patterns and speaking style of the selected person.
[0527] 4. Receiving questions and generating responses:
[0528] The user types a question through the smart glasses, which is sent to a server where an AI model analyzes it and generates an appropriate response.
[0529] 5. View the response:
[0530] The generated responses are displayed in real time on the smart glasses, enabling interaction with the user.
[0531] Specific examples
[0532] For example, consider a case where a user asks a historical scientist, "Tell me about the theory of relativity." In this case, the AI model uses the acquired data to generate a response such as, "The theory of relativity is based on the idea that time and space are relative, not absolute..." The interaction takes place in real time and is displayed through the smart glasses, improving the user experience.
[0533] Prompt Sentence Examples
[0534] User: Tell me about the theory of relativity.
[0535] AI: The theory of relativity is based on the idea that time and space are relative, not absolute...
[0536] Thus, the present invention provides a system that allows users to enjoy real-time interactions through smart glasses. The natural language processing technology used in the smart glasses can improve the quality of the interactions and the user experience.
[0537] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0538] Step 1:
[0539] The user puts on the smart glasses and selects the person they want to interact with. This selection information is displayed on the user interface, and once the user selects the target person, the selection information is retained in the smart glasses and sent to the server for the next step.
[0540] Input: User-selected person information
[0541] Output: Selection information sent to the server
[0542] Step 2:
[0543] The server retrieves data about the selected person. The server queries the database to retrieve related information such as books, papers, transcripts, letters, and audio recordings.
[0544] Input: Selection information sent from smart glasses
[0545] Output: Person-related data retrieved from the database
[0546] Step 3:
[0547] The server analyzes the acquired data using natural language processing techniques (e.g., the GPT-2 model) to extract important keywords and phrases from the data. The result is an AI model that mimics the individual's thought patterns and speaking style.
[0548] Input: Person-related data retrieved from a database
[0549] Output: AI model
[0550] Step 4:
[0551] The user inputs a question through the smart glasses, which receives the question and sends it to the server.
[0552] Input: The question entered by the user into the smart glasses
[0553] Output: The question sent to the server
[0554] Step 5:
[0555] The server uses an AI model to analyze the received question and generate an appropriate response. The AI model generates a response to the question and prepares it in text format.
[0556] Input: User-submitted question, AI model
[0557] Output: The generated response (in text format)
[0558] Step 6:
[0559] The generated response is sent to the smart glasses and displayed to the user, who can view the interaction content through the smart glasses.
[0560] Input: The generated response sent by the server
[0561] Output: Response displayed on the smart glasses
[0562] In this way, this system provides an environment in which users can interact with historical figures and deceased people in real time using smart glasses, and delivers a high-quality interactive experience by analyzing the acquired data and generating responses using an AI model.
[0563] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0564] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model, which then combines it with an emotion engine to recognize the user's emotions and realize the conversation.
[0565] System Overview
[0566] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[0567] Program processing
[0568] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[0569] 1. A user accesses the system
[0570] A user logs into the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0571] 2. The user selects the person they want to interact with
[0572] The user clicks the "Start conversation" button displayed on the main screen and selects historical figure A from the displayed list or search field.
[0573] 3. The device sends the selection information to the server
[0574] The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0575] 4. The server retrieves the data
[0576] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[0577] 5. The server analyzes the data and generates an AI model
[0578] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[0579] 6. User enters question
[0580] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[0581] 7. The server analyzes the user's emotions using an emotion engine.
[0582] The server analyzes the user's question with an emotion engine and processes the question content with emotion labels. For example, the question is tagged with emotion labels such as "interesting" and "excited."
[0583] 8. The server analyzes the question and generates a response based on the AI model
[0584] The server tokenizes the user's question, analyzes its meaning, and uses an AI model to generate an appropriate response. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism" in a calm, empathetic tone.
[0585] 9. Send the server-generated response to the device
[0586] The server sends the generated response to the terminal as an HTTP response.
[0587] 10. The device receives the response from the server and displays it to the user
[0588] The device displays the received text in a chat window on the screen for the user to view.
[0589] Specific examples
[0590] Example 1: A conversation with historical figure A
[0591] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[0592] Example 2: Dialogue with the deceased
[0593] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[0594] Technical details
[0595] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, enhancing the naturalness and intimacy of the conversation. Furthermore, by using incremental learning, the AI model is continuously updated based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures.
[0596] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[0597] The processing flow will be explained below.
[0598] Step 1:
[0599] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0600] Step 2:
[0601] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[0602] Step 3:
[0603] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0604] Step 4:
[0605] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[0606] Step 5:
[0607] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[0608] Step 6:
[0609] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[0610] Step 7:
[0611] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[0612] Step 8:
[0613] The server analyzes the user's emotions using an emotion engine. The server analyzes the user's question using the emotion engine and processes the question content along with an emotion label. For example, the question is tagged with an emotion label such as "interesting" or "excited."
[0614] Step 9:
[0615] The server analyzes the question and generates a response based on an AI model. The server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it could output something like, "The contradictions of capitalism are expanding in the modern economy," in a calm, empathetic tone.
[0616] Step 10:
[0617] The server sends the generated response to the terminal. The server sends the generated response to the terminal as an HTTP response.
[0618] Step 11:
[0619] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[0620] Step 12:
[0621] The user then enters another question, and the process repeats. The question and emotion label are sent to the server, and the AI model generates an appropriate response, which is then displayed on the device.
[0622] Specific examples
[0623] Example 1: A conversation with historical figure A
[0624] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[0625] Example 2: Dialogue with the deceased
[0626] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[0627] Example 2
[0628] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0629] Conventional communication systems have struggled to provide a natural dialogue experience for users to converse with deceased or historical figures. Furthermore, the content of the dialogue is not adjusted according to the user's emotions, resulting in a decline in dialogue quality. Furthermore, there are limitations to the accuracy of the technology used to realistically imitate a person's thought patterns and speaking style. This makes it difficult for users to obtain a sense of satisfaction from the dialogue.
[0630] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0631] In this invention, the server includes: means for selecting a person the user wants to converse with, means for retrieving data about the selected person from a database, means for analyzing the retrieved data and generating an AI model for imitating the person's thought patterns and speaking style, means for analyzing the user's emotions and assigning emotion labels, means for receiving the user's question and generating a response based on the AI model and the emotion labels, and means for displaying the generated response to the user. This allows the user to enjoy more natural and emotional conversations with deceased people and historical figures.
[0632] A "user" is a person who accesses and interacts with the system.
[0633] "People you want to talk to" refers to deceased people or historical figures with whom you want to talk through the system.
[0634] "Means for selection" refers to the interface or operation that allows a user to select a person within the system with whom they wish to interact.
[0635] "Database" refers to a magnetic disk drive or other storage system that stores information such as books, articles, transcripts, letters, email posts, audio recordings, etc. about a person with whom you wish to interact.
[0636] "Means for obtaining" refers to a process for searching and retrieving data about a selected person from a database.
[0637] "Means of analysis" refers to the process of analyzing data using natural language processing technology and extracting important keywords and phrases.
[0638] An "AI model" refers to an artificial intelligence program that is generated from acquired data and is intended to mimic the thinking patterns and speaking style of a specific person.
[0639] An "emotion engine" refers to a software module that analyzes emotions from user input and assigns emotion labels.
[0640] An "emotion label" refers to a tag that expresses an emotion, such as "interested" or "excited," in response to a user's input.
[0641] "Means for generating a response" refers to the process of analyzing the user's question and creating an appropriate response based on the AI model and emotion labels.
[0642] "Means for displaying" refers to a screen or interface for visually presenting the generated response to the user.
[0643] This invention relates to a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model. It then combines this data with an emotion engine to recognize the user's emotions and realize the conversation.
[0644] System Overview
[0645] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[0646] Program processing
[0647] When a user accesses the system and selects the person they wish to interact with, the device sends the corresponding person's ID to the server. The server retrieves data about that person from a database and analyzes it using natural language processing technology. For example, natural language processing libraries such as SpaCy and NLTK are used. As a result, important keywords and phrases are extracted and an AI model that mimics the target person's thought patterns and speaking style is generated.
[0648] When a user enters a question, the device sends it to the server. The server analyzes the question using an emotion engine and assigns an emotion label to the question using, for example, TextBlob or a sentiment analysis API. The server then tokenizes the question and generates a response based on an AI model. The tone and content of the response are adjusted based on the emotion engine's analysis results. The generated response is then sent to the device, which displays it to the user.
[0649] Specific examples
[0650] Example 1: A conversation with historical figure A
[0651] If a user wants to talk with historical figure A about the "current political situation," the user begins the conversation as follows:
[0652] Example prompt sentence:
[0653] "Ask historical figure A what he thinks about the current political situation."
[0654] The server retrieves data on great person A from the database, analyzes it, and generates an AI model. It then analyzes the user's question along with the emotion label and generates a response such as, "Compared to past historical events, the current political situation is..."
[0655] Example 2: Dialogue with the deceased
[0656] If the user wanted to mimic a conversation with their deceased grandmother, the user's prompt might look like this:
[0657] Example prompt sentence:
[0658] "Tell your late grandmother about your recent gardening."
[0659] The server generates an AI model based on the grandmother's letters and audio recordings, and generates responses such as "Oh, that's great. I always enjoyed tending to your garden" based on emotion labels.
[0660] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, improving the naturalness and intimacy of the dialogue. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures. This system has broad applications not only in academic research and personal healing, but also in education and entertainment.
[0661] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0662] Step 1:
[0663] A user accesses the system.
[0664] Specifically, a user accesses the system using a web browser or mobile app. The user enters login information (username and password) and presses the "Login" button. The input is the username and password string, and the output is the success or failure of authentication and the display of the main screen.
[0665] Step 2:
[0666] The user selects the person with whom they want to interact.
[0667] Specifically, the user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The input is the name or ID of the person they want to talk to, and the output is data including the ID of the selected person.
[0668] Step 3:
[0669] The terminal transmits the selection information to the server.
[0670] Specifically, the device generates an HTTP request including the ID of the selected person and sends it to the server. The input is JSON-formatted data including the ID of the selected person, and the output is an HTTP request to the server.
[0671] Step 4:
[0672] The server retrieves the data.
[0673] Specifically, the server searches and retrieves data about the selected person from a database. The input is the selected person's ID, and the output is data such as books, papers, transcripts, and letters retrieved from the database. The server retrieves this data from local storage or cloud storage.
[0674] Step 5:
[0675] The server analyzes the data and generates an AI model.
[0676] Specifically, the server analyzes the acquired data using natural language processing techniques (e.g., SpaCy or NLTK). Important keywords and phrases are extracted, and an AI model is generated using machine learning techniques (e.g., TensorFlow or PyTorch). The input is person data acquired from a database, and the output is an AI model that mimics the thinking patterns and speaking style of a specific person.
[0677] Step 6:
[0678] The user enters a question.
[0679] Specifically, the user enters a question into the text box on the interactive screen and presses the "Send" button. The input is the question entered by the user, and the output is the question data sent to the server.
[0680] Step 7:
[0681] The server analyzes the user's emotions using an emotion engine.
[0682] Specifically, the server analyzes the received question using an emotion engine (e.g., TextBlob or an emotion analysis API) and assigns an emotion label. The input is the user's question, and the output is the question data including the emotion label.
[0683] Step 8:
[0684] The server analyzes the question and generates a response based on an AI model.
[0685] Specifically, the server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. The input is the question data with emotion labels and the AI model, and the output is the generated response text.
[0686] Step 9:
[0687] The server generates a response and sends it to the terminal.
[0688] Specifically, the server sends the generated response in JSON format to the terminal as an HTTP response. The input is the generated response text, and the output is the HTTP response to the terminal.
[0689] Step 10:
[0690] The terminal receives the response from the server and displays it to the user.
[0691] Specifically, the device displays the received text in a chat window on the screen for the user to view. The input is the HTTP response from the server, and the output is the response text displayed in the chat window on the screen.
[0692] (Application example 2)
[0693] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0694] Conventional AI dialogue systems provide users with simple question and answer responses, but lack mechanisms for recognizing emotions and enabling natural dialogue. Furthermore, they lack an interactive experience in virtual space, resulting in low user satisfaction. The objective of this invention is to provide a dialogue system that provides a sense of realism by utilizing emotion recognition and virtual space.
[0695] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a person with whom the user wants to converse, input by the user, means for acquiring data about the selected person from a database, means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style, means including an emotion engine for recognizing and analyzing the user's emotions, means for adjusting the response content based on the analyzed emotions, and means for displaying this in a virtual space. This enables natural and realistic conversations that correspond to emotions.
[0696] The "means for selecting a person the user wishes to converse with" refers to hardware or software that provides an interface for the user to input or select a person with whom the user wishes to converse.
[0697] The "means for obtaining data on the selected person from the database" refers to a process or database communication means for accessing and obtaining data on the selected person from a database on a server or cloud.
[0698] "Means of analyzing acquired data and generating an AI model that imitates the person's thought patterns and speaking style" refers to technology that analyzes acquired data using natural language processing and machine learning techniques and generates an AI model that imitates the person.
[0699] The "means for receiving a user's question and generating a response based on an AI model" is a processing system that receives a user's question as input and generates an appropriate response using the algorithm of the AI model.
[0700] "Means for displaying the generated response to the user" refers to an interface or output device for displaying the response generated by the AI model on the user's terminal.
[0701] The "means including an emotion engine for recognizing and analyzing user emotions" refers to a software and hardware configuration for analyzing emotions from inputs and interactions from a user, and recognizing and processing the emotions.
[0702] The "means for adjusting the content of the response based on the analyzed emotions" is a technology that adaptively adjusts the tone and content of the generated response based on the user's emotions analyzed by the emotion engine.
[0703] "Means for displaying this in a virtual space" refers to the display devices and virtual space rendering technologies necessary for users to enjoy a realistic interaction in a virtual space.
[0704] System Configuration
[0705] A system for implementing this invention includes a user terminal, a server, a database, and an emotion engine. The user terminal is equipped with an interface that allows the user to select a person with whom they want to interact. The server accesses the database to obtain data about the person and generates an AI model. The server also uses the emotion engine to analyze the user's emotions and adjust the response content.
[0706] Program processing
[0707] When a user selects a person they want to talk to through their device, the device sends that information to a server. The server then retrieves data about the selected person from a database, such as books, papers, transcripts, letters, audio data, and video data. This data is analyzed using natural language processing technology to generate an AI model that mimics the person's thought patterns and speaking style. This generated AI model is then used to provide appropriate responses based on the user's questions.
[0708] When a user enters a question, it is sent to the server. The server uses an emotion engine to analyze the question's sentiment and adjusts the AI model's response based on the analysis results. After the response is generated, the server sends it to the user's device, which then displays it to the user in a virtual space.
[0709] Hardware and software used
[0710] User devices: smartphones, smart glasses, head-mounted displays
[0711] Server: Database Management System (DBMS), Web Server
[0712] Emotion engine: Emotion API (e.g. Microsoft Emotion API)
[0713] Natural language processing technology: OpenAI GPT, natural language tokenization libraries (e.g., spaCy)
[0714] Specific examples
[0715] For example, consider a scenario where a user enters a virtual cafe and chooses to interact with a great inventor. If the user asks, "What do you think about modern technology?", the server analyzes the question using its emotion engine and tags it as "interesting." Once the analysis is complete, the AI model generates a response such as, "Modern technology has advanced in ways never before imagined," which is then displayed on the user's device. The response is adjusted in tone to reflect the user's emotions.
[0716] Prompt Sentence Examples
[0717] Imagine the following dialogue:
[0718] You are a great inventor. You have had a huge impact on modern life with your inventions, and you have many inventions to your name.
[0719] User: What do you think about modern technology? [Interested]
[0720] In this way, the system can realize natural and realistic dialogue that responds to emotions, providing users with an interactive experience.
[0721] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0722] Step 1:
[0723] The user selects the person they want to talk to through their terminal. The person selection information entered by the user is sent from the terminal to the server. The ID of the person selected by the user is included as input, and the server receives it as output.
[0724] Step 2:
[0725] The server uses the received person ID to retrieve data about the selected person from a database, including books, papers, transcripts, letters, audio data, and video data. The person ID is used as input, and the target data is obtained as output.
[0726] Step 3:
[0727] The server analyzes the acquired data and generates an AI model to mimic the person's thought patterns and speaking style. It uses natural language processing technology to tokenize the data and extract important keywords and phrases. The acquired data is used as input, and the AI model is generated as output.
[0728] Step 4:
[0729] The user inputs a question through a terminal. The question is sent from the terminal to the server. The input contains the user's question, and the server receives it as output.
[0730] Step 5:
[0731] The server uses an emotion engine to perform emotion analysis based on the user's question. The analysis generates an emotion label for the question. The user's question is used as input, and the emotion label is generated as output.
[0732] Step 6:
[0733] The server uses an AI model to generate an appropriate response based on the analyzed emotion label. The response content is adjusted to take the user's emotion into account. The user's question and emotion label are used as input, and an adjusted response is generated as output.
[0734] Step 7:
[0735] The server sends the generated and adjusted response to the user terminal, which includes the generated response as input and receives it as output.
[0736] Step 8:
[0737] The terminal displays the received response to the user in the virtual space. The displayed response is reflected in the interface so that the user can see it in real time. The response sent to the terminal is used as input, and the response is displayed in the virtual space as output.
[0738] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0739] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0740] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0741] [Third embodiment]
[0742] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0743] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0744] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0745] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0746] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0747] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0748] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0749] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0750] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0751] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0752] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0753] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0754] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user, generates an AI model, and realizes the conversation.
[0755] System Overview
[0756] The user accesses the system and selects the person they want to talk to. The device sends that information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. This AI model is used to generate responses to the user's questions and displays them to the user via the device.
[0757] Program processing
[0758] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[0759] 1. A user accesses the system
[0760] A user accesses the system using a web browser or application. A login screen appears and the user enters their authentication information to log in to the system.
[0761] 2. The user selects the person they want to interact with
[0762] The user clicks the "Start a conversation" button on the main screen and selects historical figure A from the displayed list or search field.
[0763] 3. The device sends the selection information to the server
[0764] The device sends the user's selection information to the server as an HTTP request, which includes the ID of the selected person.
[0765] 4. The server retrieves the data
[0766] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[0767] 5. The server analyzes the data and generates an AI model
[0768] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[0769] 6. User enters question
[0770] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[0771] 7. The server generates a response
[0772] The server analyzes the question and uses an AI model to generate an appropriate response, such as, "The modern economy contains many of the contradictions inherent in capitalist society. The solution to this is..."
[0773] 8. The generated response is displayed to the user
[0774] The server sends the generated response to the terminal, which displays the response on the screen, allowing the user to enjoy a conversation with Great Person A.
[0775] Specific examples
[0776] Example 1: A conversation with historical figure A
[0777] A similar process is repeated if a user wants to talk with historical figure A about the "current political situation." For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..."
[0778] Example 2: Dialogue with the deceased
[0779] If a user wants to imitate a conversation with their late grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if a user mentions "how you've been taking care of your garden lately," the grandmother's AI can respond with something like, "Oh, that's wonderful. I always enjoy watching you take care of your garden."
[0780] Technical details
[0781] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. It also uses incremental learning to continuously update the AI model based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic conversations with deceased people and historical figures.
[0782] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[0783] The processing flow will be explained below.
[0784] Step 1:
[0785] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0786] Step 2:
[0787] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[0788] Step 3:
[0789] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0790] Step 4:
[0791] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[0792] Step 5:
[0793] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[0794] Step 6:
[0795] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[0796] Step 7:
[0797] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[0798] Step 8:
[0799] The server analyzes the question and generates a response based on the AI model. The server tokenizes the user's question, analyzes its meaning, and uses the AI model to generate an appropriate response. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism."
[0800] Step 9:
[0801] The server sends the generated response to the terminal, which then sends the response as an HTTP response.
[0802] Step 10:
[0803] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[0804] Step 11:
[0805] If the user wishes to enter more questions, they enter the questions again, and the same process is repeated once the new questions are entered.
[0806] Example 1
[0807] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0808] Conventional communication systems have struggled to enable users to converse with deceased or historical figures in real time. Technical limitations exist, particularly in situations where accurate mimicking of the target person's thought patterns and speaking style is required. This has made it difficult to realize such systems in a wide range of applications, including education, entertainment, and personal healing.
[0809] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0810] In this invention, the server includes a means for selecting a person with whom the user wants to converse, a means for acquiring information about the selected person from a database, and a means for analyzing the acquired information and generating an AI model for imitating the thought patterns and speaking style of the person, thereby enabling the user to enjoy realistic and sophisticated conversations with deceased people and historical figures.
[0811] A "user" is someone who uses the system to interact with a specific person.
[0812] A "person with whom a user wishes to converse" is a deceased person or historical figure that the user selects and with whom the user wishes to converse.
[0813] "Means for selection" is a function that provides an interface for the user to select the person with whom they want to interact from a list or search field.
[0814] "Information" refers to data such as books, papers, transcripts, letters, and audio data related to the target person.
[0815] A "database" is a digital information storage system for storing and managing information within a system.
[0816] The "means for obtaining" is a function for searching and extracting information about a selected person from a database.
[0817] An "artificial intelligence model" is a machine learning model that is generated to mimic the thought patterns and speaking style of a target person.
[0818] "Means of analysis" refers to the function of analyzing acquired information using natural language processing technology and extracting important keywords and phrases.
[0819] The "means for generating a response" is a function that uses an artificial intelligence model to generate an appropriate response based on a user's question.
[0820] The "display means" is an interface for displaying the generated response on the user's terminal screen.
[0821] "Natural language processing technology" is a technology that enables computers to understand and process human language.
[0822] "Keywords" are important words or phrases extracted during the process of analyzing information about a target person.
[0823] The present invention provides a communication system that allows users to enjoy conversations with deceased or historical figures. This system allows users to select the person they want to talk to, acquires and analyzes information about that person, generates an artificial intelligence model, and realizes the conversation.
[0824] Hardware and software used
[0825] Hardware: Servers, devices (PCs, smartphones, tablets, etc.)
[0826] Software: web browsers, applications, database management systems, natural language processing libraries (e.g., Python's NLTK, SpaCy), generative AI models (e.g., OpenAI's GPT-3, Google's BERT)
[0827] Data processing and calculation
[0828] 1. Data Acquisition: The server acquires information about the target person from a database, specifically in the form of books, papers, transcripts, letters, audio data, etc.
[0829] 2. Data Analysis: The server analyzes the acquired information using natural language processing techniques, such as Python's NLTK and SpaCy, to extract important keywords and phrases and model the target person's thought patterns and speaking style.
[0830] 3. Creating a generative AI model: Based on the analyzed data, an artificial intelligence model is generated using OpenAI's GPT-3 or Google's BERT. This model mimics the target person's speaking style and thinking patterns.
[0831] 4. User Question Analysis: The server analyzes the question entered by the user. This analysis involves using natural language processing techniques to understand the meaning and extract relevant keywords.
[0832] 5. Response Generation: Based on the analyzed question, the server uses the generated artificial intelligence model to generate an appropriate response that replicates the target person's thought patterns and speaking style.
[0833] 6. Displaying the response: The server sends the generated response to the terminal, which displays it on the user's screen, allowing the user to enjoy the interaction.
[0834] Specific examples
[0835] For example, suppose a user wants to converse with "Albert Einstein." The user enters "Albert Einstein" into the application's search field and selects it. The server collects papers, notes, and speech data related to Einstein from the database and analyzes them using natural language processing technology. Based on the analysis results, an artificial intelligence model of Einstein is generated using OpenAI's GPT-3.
[0836] When a user inputs a question such as "What do you think about the modern economy?", the server analyzes the question and uses the generated artificial intelligence model to generate a response such as "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." This response is sent to the terminal and displayed on the user's screen.
[0837] Prompt Sentence Examples
[0838] Enter: "What do you think about the modern economy?"
[0839] Response: "The modern economy contains many of the contradictions of capitalist society. The solution to these is..."
[0840] This system allows users to enjoy highly realistic and sophisticated interactions with deceased or historical figures, which is expected to have applications in a variety of fields, including education, entertainment, and personal healing.
[0841] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0842] Step 1:
[0843] A user accesses the system. The user accesses the system using a web browser or application and enters an ID and password on the login screen. The input information is sent from the terminal to the server as an HTTPS request. The server receives this information and performs authentication by checking it against a database. If authentication is successful, the server sends the main screen data to the terminal and login is complete. The input is authentication information (ID and password), and the output is the user's authentication status and the display of the main screen.
[0844] Step 2:
[0845] The user selects the person they want to talk to. The user clicks the "Start a conversation" button displayed on the main screen and selects the historical figure or deceased person they want to talk to from a search field or list. The device sends an HTTP request to the server including the selection information (person ID). The input is the selection information of the person they want to talk to, and the output is the request sent to the server.
[0846] Step 3:
[0847] The device sends the selection information to the server. The device sends an HTTP request to the server, including the ID of the selected person. The server receives this request and executes a query to retrieve information about the selected person from a database. The input is the ID of the selected person, and the output is the result of executing the database query.
[0848] Step 4:
[0849] The server retrieves the data. The server searches and retrieves information related to the selected person from its database (books, papers, transcripts, letters, audio recordings, etc.). This information is returned to the server in JSON format. The input is a database query, and the output is the information retrieved about the target person.
[0850] Step 5:
[0851] The server analyzes the data and generates an AI model. The server uses natural language processing (NLP) techniques to analyze the acquired information, using Python's NLTK and SpaCy libraries. During the analysis, important keywords and phrases are extracted, and based on this, a model is created of the target person's thought patterns and speaking style. OpenAI's GPT-3 and Google's BERT are used for generation. The input is the information to be analyzed, and the output is the generated generative AI model.
[0852] Step 6:
[0853] The user enters a question. The user enters the question into the application's input form, and the question is sent from the terminal to the server. The input is the user's question, and the output is the question sent to the server.
[0854] Step 7:
[0855] The server generates a response. The server analyzes the input question and uses a generative AI model to generate an appropriate response. It uses natural language processing technology to understand the meaning of the question and extract relevant keywords. For example, in response to the question, "What do you think about the modern economy?", it generates a response such as, "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." The input is the user's question, and the output is the generated response.
[0856] Step 8:
[0857] The generated response is displayed to the user. The server sends the generated response in JSON format to the terminal. The terminal displays the received response data on the screen, allowing the user to enjoy the interaction. The input is the generated response, and the output is the response displayed on the screen.
[0858] (Application example 1)
[0859] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0860] Currently, there are technologies that allow users to enjoy conversations with historical figures and deceased people, but these conversations can only be realized on limited devices, which often limits the experience. Furthermore, it is difficult to visually display the content of the conversation, which does not improve the user experience. Furthermore, there is a lack of technology to improve the accuracy of the conversation content desired by users, making it difficult to realize realistic conversations.
[0861] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0862] In this invention, the server includes: means for selecting a person the user wants to converse with, as input by the user; means for acquiring data about the selected person from a database; means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style; means for receiving the user's question and generating a response based on the AI model; means installed in the smart glasses for displaying the content of the conversation with the person selected by the user; and means for displaying the generated response to the user. This allows the user to enjoy a visual conversation in real time using the smart glasses, improving the quality of the experience. Furthermore, it is possible to accurately imitate the thought patterns and speaking style of the selected person, thereby realizing a more realistic conversation.
[0863] The "means for selecting a person to converse with input by the user" is an interface within the system that allows the user to select a specific person as a conversation partner.
[0864] The "means for obtaining data relating to a selected person from a database" is a function for obtaining information relating to a person selected by a user from a database stored within the system.
[0865] "Means for analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style" refers to technology for analyzing the data of a selected person and creating an AI model that reproduces the person's speaking style and thought patterns.
[0866] "Means for receiving a user's question and generating a response based on an AI model" refers to the function by which the system receives a question from a user and the AI model generates an appropriate response based on that question.
[0867] "A means installed in smart glasses to display the content of the conversation with a person selected by the user" is a function installed in smart glasses to visually provide the user with the response of the person with whom they want to converse.
[0868] "Means for displaying the generated response to the user" refers to an interface that allows the user to see the response generated by the AI model.
[0869] "Books, papers, transcripts, letters, and audio recordings" are sources of information related to the selected individuals from which data will be drawn and used in the analysis.
[0870] "Means of analyzing input data using natural language processing technology and extracting important keywords and phrases" is a technology that utilizes natural language processing technology to extract important information from user input and dialogue.
[0871] This invention is a system that allows a user to select a person with whom they want to converse, acquires and analyzes data about that person to generate an AI model, and realizes a dialogue with the user in real time. This system is installed in smart glasses and can visually display the dialogue content to the user. Below, a detailed description of an embodiment of this invention is provided.
[0872] System configuration
[0873] This system is configured using the following hardware and software.
[0874] Smart glasses: Devices that display interactions to the user in real time and also serve as an input / output interface.
[0875] Server: Responsible for back-end processing to acquire and analyze the data of the selected person and generate an AI model.
[0876] Database: Contains data (books, papers, transcripts, letters, audio recordings, etc.) about historical figures and deceased people.
[0877] Natural language processing technology: Technology used to analyze data and extract key keywords and phrases (e.g., GPT-2 model).
[0878] Program processing overview
[0879] 1. User inputs the person they want to talk to:
[0880] The user wearing the smart glasses selects the person they want to talk to. The user interface displays a list of people they want to talk to, allowing them to select one.
[0881] 2. Retrieving data from the database:
[0882] Data about the selected person (e.g., books, papers, transcripts, letters, audio recordings, etc.) is retrieved from a database by the server.
[0883] 3. Data analysis and AI model generation:
[0884] The server then analyzes the data using natural language processing techniques to extract key keywords and phrases, which then generates an AI model that mimics the thought patterns and speaking style of the selected person.
[0885] 4. Receiving questions and generating responses:
[0886] The user types a question through the smart glasses, which is sent to a server where an AI model analyzes it and generates an appropriate response.
[0887] 5. View the response:
[0888] The generated responses are displayed in real time on the smart glasses, enabling interaction with the user.
[0889] Specific examples
[0890] For example, consider a case where a user asks a historical scientist, "Tell me about the theory of relativity." In this case, the AI model uses the acquired data to generate a response such as, "The theory of relativity is based on the idea that time and space are relative, not absolute..." The interaction takes place in real time and is displayed through the smart glasses, improving the user experience.
[0891] Prompt Sentence Examples
[0892] User: Tell me about the theory of relativity.
[0893] AI: The theory of relativity is based on the idea that time and space are relative, not absolute...
[0894] Thus, the present invention provides a system that allows users to enjoy real-time interactions through smart glasses. The natural language processing technology used in the smart glasses can improve the quality of the interactions and the user experience.
[0895] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0896] Step 1:
[0897] The user puts on the smart glasses and selects the person they want to interact with. This selection information is displayed on the user interface, and once the user selects the target person, the selection information is retained in the smart glasses and sent to the server for the next step.
[0898] Input: User-selected person information
[0899] Output: Selection information sent to the server
[0900] Step 2:
[0901] The server retrieves data about the selected person. The server queries the database to retrieve related information such as books, papers, transcripts, letters, and audio recordings.
[0902] Input: Selection information sent from smart glasses
[0903] Output: Person-related data retrieved from the database
[0904] Step 3:
[0905] The server analyzes the acquired data using natural language processing techniques (e.g., the GPT-2 model) to extract important keywords and phrases from the data. The result is an AI model that mimics the individual's thought patterns and speaking style.
[0906] Input: Person-related data retrieved from a database
[0907] Output: AI model
[0908] Step 4:
[0909] The user inputs a question through the smart glasses, which receives the question and sends it to the server.
[0910] Input: The question entered by the user into the smart glasses
[0911] Output: The question sent to the server
[0912] Step 5:
[0913] The server uses an AI model to analyze the received question and generate an appropriate response. The AI model generates a response to the question and prepares it in text format.
[0914] Input: User-submitted question, AI model
[0915] Output: The generated response (in text format)
[0916] Step 6:
[0917] The generated response is sent to the smart glasses and displayed to the user, who can view the interaction content through the smart glasses.
[0918] Input: The generated response sent by the server
[0919] Output: Response displayed on the smart glasses
[0920] In this way, this system provides an environment in which users can interact with historical figures and deceased people in real time using smart glasses, and delivers a high-quality interactive experience by analyzing the acquired data and generating responses using an AI model.
[0921] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0922] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model, which then combines it with an emotion engine to recognize the user's emotions and realize the conversation.
[0923] System Overview
[0924] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[0925] Program processing
[0926] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[0927] 1. A user accesses the system
[0928] A user logs into the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0929] 2. The user selects the person they want to interact with
[0930] The user clicks the "Start conversation" button displayed on the main screen and selects historical figure A from the displayed list or search field.
[0931] 3. The device sends the selection information to the server
[0932] The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0933] 4. The server retrieves the data
[0934] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[0935] 5. The server analyzes the data and generates an AI model
[0936] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[0937] 6. User enters question
[0938] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[0939] 7. The server analyzes the user's emotions using an emotion engine.
[0940] The server analyzes the user's question with an emotion engine and processes the question content with emotion labels. For example, the question is tagged with emotion labels such as "interesting" and "excited."
[0941] 8. The server analyzes the question and generates a response based on the AI model
[0942] The server tokenizes the user's question, analyzes its meaning, and uses an AI model to generate an appropriate response. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism" in a calm, empathetic tone.
[0943] 9. Send the server-generated response to the device
[0944] The server sends the generated response to the terminal as an HTTP response.
[0945] 10. The device receives the response from the server and displays it to the user
[0946] The device displays the received text in a chat window on the screen for the user to view.
[0947] Specific examples
[0948] Example 1: A conversation with historical figure A
[0949] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[0950] Example 2: Dialogue with the deceased
[0951] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[0952] Technical details
[0953] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, enhancing the naturalness and intimacy of the conversation. Furthermore, by using incremental learning, the AI model is continuously updated based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures.
[0954] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[0955] The processing flow will be explained below.
[0956] Step 1:
[0957] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[0958] Step 2:
[0959] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[0960] Step 3:
[0961] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[0962] Step 4:
[0963] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[0964] Step 5:
[0965] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[0966] Step 6:
[0967] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[0968] Step 7:
[0969] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[0970] Step 8:
[0971] The server analyzes the user's emotions using an emotion engine. The server analyzes the user's question using the emotion engine and processes the question content along with an emotion label. For example, the question is tagged with an emotion label such as "interesting" or "excited."
[0972] Step 9:
[0973] The server analyzes the question and generates a response based on an AI model. The server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it could output something like, "The contradictions of capitalism are expanding in the modern economy," in a calm, empathetic tone.
[0974] Step 10:
[0975] The server sends the generated response to the terminal. The server sends the generated response to the terminal as an HTTP response.
[0976] Step 11:
[0977] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[0978] Step 12:
[0979] The user then enters another question, and the process repeats. The question and emotion label are sent to the server, and the AI model generates an appropriate response, which is then displayed on the device.
[0980] Specific examples
[0981] Example 1: A conversation with historical figure A
[0982] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[0983] Example 2: Dialogue with the deceased
[0984] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[0985] Example 2
[0986] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0987] Conventional communication systems have struggled to provide a natural dialogue experience for users to converse with deceased or historical figures. Furthermore, the content of the dialogue is not adjusted according to the user's emotions, resulting in a decline in dialogue quality. Furthermore, there are limitations to the accuracy of the technology used to realistically imitate a person's thought patterns and speaking style. This makes it difficult for users to obtain a sense of satisfaction from the dialogue.
[0988] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0989] In this invention, the server includes: means for selecting a person the user wants to converse with, means for retrieving data about the selected person from a database, means for analyzing the retrieved data and generating an AI model for imitating the person's thought patterns and speaking style, means for analyzing the user's emotions and assigning emotion labels, means for receiving the user's question and generating a response based on the AI model and the emotion labels, and means for displaying the generated response to the user. This allows the user to enjoy more natural and emotional conversations with deceased people and historical figures.
[0990] A "user" is a person who accesses and interacts with the system.
[0991] "People you want to talk to" refers to deceased people or historical figures with whom you want to talk through the system.
[0992] "Means for selection" refers to the interface or operation that allows a user to select a person within the system with whom they wish to interact.
[0993] "Database" refers to a magnetic disk drive or other storage system that stores information such as books, articles, transcripts, letters, email posts, audio recordings, etc. about a person with whom you wish to interact.
[0994] "Means for obtaining" refers to a process for searching and retrieving data about a selected person from a database.
[0995] "Means of analysis" refers to the process of analyzing data using natural language processing technology and extracting important keywords and phrases.
[0996] An "AI model" refers to an artificial intelligence program that is generated from acquired data and is intended to mimic the thinking patterns and speaking style of a specific person.
[0997] An "emotion engine" refers to a software module that analyzes emotions from user input and assigns emotion labels.
[0998] An "emotion label" refers to a tag that expresses an emotion, such as "interested" or "excited," in response to a user's input.
[0999] "Means for generating a response" refers to the process of analyzing the user's question and creating an appropriate response based on the AI model and emotion labels.
[1000] "Means for displaying" refers to a screen or interface for visually presenting the generated response to the user.
[1001] This invention relates to a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model. It then combines this data with an emotion engine to recognize the user's emotions and realize the conversation.
[1002] System Overview
[1003] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[1004] Program processing
[1005] When a user accesses the system and selects the person they wish to interact with, the device sends the corresponding person's ID to the server. The server retrieves data about that person from a database and analyzes it using natural language processing technology. For example, natural language processing libraries such as SpaCy and NLTK are used. As a result, important keywords and phrases are extracted and an AI model that mimics the target person's thought patterns and speaking style is generated.
[1006] When a user enters a question, the device sends it to the server. The server analyzes the question using an emotion engine and assigns an emotion label to the question using, for example, TextBlob or a sentiment analysis API. The server then tokenizes the question and generates a response based on an AI model. The tone and content of the response are adjusted based on the emotion engine's analysis results. The generated response is then sent to the device, which displays it to the user.
[1007] Specific examples
[1008] Example 1: A conversation with historical figure A
[1009] If a user wants to talk with historical figure A about the "current political situation," the user begins the conversation as follows:
[1010] Example prompt sentence:
[1011] "Ask historical figure A what he thinks about the current political situation."
[1012] The server retrieves data on great person A from the database, analyzes it, and generates an AI model. It then analyzes the user's question along with the emotion label and generates a response such as, "Compared to past historical events, the current political situation is..."
[1013] Example 2: Dialogue with the deceased
[1014] If the user wanted to mimic a conversation with their deceased grandmother, the user's prompt might look like this:
[1015] Example prompt sentence:
[1016] "Tell your late grandmother about your recent gardening."
[1017] The server generates an AI model based on the grandmother's letters and audio recordings, and generates responses such as "Oh, that's great. I always enjoyed tending to your garden" based on emotion labels.
[1018] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, improving the naturalness and intimacy of the dialogue. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures. This system has broad applications not only in academic research and personal healing, but also in education and entertainment.
[1019] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1020] Step 1:
[1021] A user accesses the system.
[1022] Specifically, a user accesses the system using a web browser or mobile app. The user enters login information (username and password) and presses the "Login" button. The input is the username and password string, and the output is the success or failure of authentication and the display of the main screen.
[1023] Step 2:
[1024] The user selects the person with whom they want to interact.
[1025] Specifically, the user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The input is the name or ID of the person they want to talk to, and the output is data including the ID of the selected person.
[1026] Step 3:
[1027] The terminal transmits the selection information to the server.
[1028] Specifically, the device generates an HTTP request including the ID of the selected person and sends it to the server. The input is JSON-formatted data including the ID of the selected person, and the output is an HTTP request to the server.
[1029] Step 4:
[1030] The server retrieves the data.
[1031] Specifically, the server searches and retrieves data about the selected person from a database. The input is the selected person's ID, and the output is data such as books, papers, transcripts, and letters retrieved from the database. The server retrieves this data from local storage or cloud storage.
[1032] Step 5:
[1033] The server analyzes the data and generates an AI model.
[1034] Specifically, the server analyzes the acquired data using natural language processing techniques (e.g., SpaCy or NLTK). Important keywords and phrases are extracted, and an AI model is generated using machine learning techniques (e.g., TensorFlow or PyTorch). The input is person data acquired from a database, and the output is an AI model that mimics the thinking patterns and speaking style of a specific person.
[1035] Step 6:
[1036] The user enters a question.
[1037] Specifically, the user enters a question into the text box on the interactive screen and presses the "Send" button. The input is the question entered by the user, and the output is the question data sent to the server.
[1038] Step 7:
[1039] The server analyzes the user's emotions using an emotion engine.
[1040] Specifically, the server analyzes the received question using an emotion engine (e.g., TextBlob or an emotion analysis API) and assigns an emotion label. The input is the user's question, and the output is the question data including the emotion label.
[1041] Step 8:
[1042] The server analyzes the question and generates a response based on an AI model.
[1043] Specifically, the server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. The input is the question data with emotion labels and the AI model, and the output is the generated response text.
[1044] Step 9:
[1045] The server generates a response and sends it to the terminal.
[1046] Specifically, the server sends the generated response in JSON format to the terminal as an HTTP response. The input is the generated response text, and the output is the HTTP response to the terminal.
[1047] Step 10:
[1048] The terminal receives the response from the server and displays it to the user.
[1049] Specifically, the device displays the received text in a chat window on the screen for the user to view. The input is the HTTP response from the server, and the output is the response text displayed in the chat window on the screen.
[1050] (Application example 2)
[1051] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1052] Conventional AI dialogue systems provide users with simple question and answer responses, but lack mechanisms for recognizing emotions and enabling natural dialogue. Furthermore, they lack an interactive experience in virtual space, resulting in low user satisfaction. The objective of this invention is to provide a dialogue system that provides a sense of realism by utilizing emotion recognition and virtual space.
[1053] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a person with whom the user wants to converse, input by the user, means for acquiring data about the selected person from a database, means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style, means including an emotion engine for recognizing and analyzing the user's emotions, means for adjusting the response content based on the analyzed emotions, and means for displaying this in a virtual space. This enables natural and realistic conversations that correspond to emotions.
[1054] The "means for selecting a person the user wishes to converse with" refers to hardware or software that provides an interface for the user to input or select a person with whom the user wishes to converse.
[1055] The "means for obtaining data on the selected person from the database" refers to a process or database communication means for accessing and obtaining data on the selected person from a database on a server or cloud.
[1056] "Means of analyzing acquired data and generating an AI model that imitates the person's thought patterns and speaking style" refers to technology that analyzes acquired data using natural language processing and machine learning techniques and generates an AI model that imitates the person.
[1057] The "means for receiving a user's question and generating a response based on an AI model" is a processing system that receives a user's question as input and generates an appropriate response using the algorithm of the AI model.
[1058] "Means for displaying the generated response to the user" refers to an interface or output device for displaying the response generated by the AI model on the user's terminal.
[1059] The "means including an emotion engine for recognizing and analyzing user emotions" refers to a software and hardware configuration for analyzing emotions from inputs and interactions from a user, and recognizing and processing the emotions.
[1060] The "means for adjusting the content of the response based on the analyzed emotions" is a technology that adaptively adjusts the tone and content of the generated response based on the user's emotions analyzed by the emotion engine.
[1061] "Means for displaying this in a virtual space" refers to the display devices and virtual space rendering technologies necessary for users to enjoy a realistic interaction in a virtual space.
[1062] System Configuration
[1063] A system for implementing this invention includes a user terminal, a server, a database, and an emotion engine. The user terminal is equipped with an interface that allows the user to select a person with whom they want to interact. The server accesses the database to obtain data about the person and generates an AI model. The server also uses the emotion engine to analyze the user's emotions and adjust the response content.
[1064] Program processing
[1065] When a user selects a person they want to talk to through their device, the device sends that information to a server. The server then retrieves data about the selected person from a database, such as books, papers, transcripts, letters, audio data, and video data. This data is analyzed using natural language processing technology to generate an AI model that mimics the person's thought patterns and speaking style. This generated AI model is then used to provide appropriate responses based on the user's questions.
[1066] When a user enters a question, it is sent to the server. The server uses an emotion engine to analyze the question's sentiment and adjusts the AI model's response based on the analysis results. After the response is generated, the server sends it to the user's device, which then displays it to the user in a virtual space.
[1067] Hardware and software used
[1068] User devices: smartphones, smart glasses, head-mounted displays
[1069] Server: Database Management System (DBMS), Web Server
[1070] Emotion engine: Emotion API (e.g. Microsoft Emotion API)
[1071] Natural language processing technology: OpenAI GPT, natural language tokenization libraries (e.g., spaCy)
[1072] Specific examples
[1073] For example, consider a scenario where a user enters a virtual cafe and chooses to interact with a great inventor. If the user asks, "What do you think about modern technology?", the server analyzes the question using its emotion engine and tags it as "interesting." Once the analysis is complete, the AI model generates a response such as, "Modern technology has advanced in ways never before imagined," which is then displayed on the user's device. The response is adjusted in tone to reflect the user's emotions.
[1074] Prompt Sentence Examples
[1075] Imagine the following dialogue:
[1076] You are a great inventor. You have had a huge impact on modern life with your inventions, and you have many inventions to your name.
[1077] User: What do you think about modern technology? [Interested]
[1078] In this way, the system can realize natural and realistic dialogue that responds to emotions, providing users with an interactive experience.
[1079] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1080] Step 1:
[1081] The user selects the person they want to talk to through their terminal. The person selection information entered by the user is sent from the terminal to the server. The ID of the person selected by the user is included as input, and the server receives it as output.
[1082] Step 2:
[1083] The server uses the received person ID to retrieve data about the selected person from a database, including books, papers, transcripts, letters, audio data, and video data. The person ID is used as input, and the target data is obtained as output.
[1084] Step 3:
[1085] The server analyzes the acquired data and generates an AI model to mimic the person's thought patterns and speaking style. It uses natural language processing technology to tokenize the data and extract important keywords and phrases. The acquired data is used as input, and the AI model is generated as output.
[1086] Step 4:
[1087] The user inputs a question through a terminal. The question is sent from the terminal to the server. The input contains the user's question, and the server receives it as output.
[1088] Step 5:
[1089] The server uses an emotion engine to perform emotion analysis based on the user's question. The analysis generates an emotion label for the question. The user's question is used as input, and the emotion label is generated as output.
[1090] Step 6:
[1091] The server uses an AI model to generate an appropriate response based on the analyzed emotion label. The response content is adjusted to take the user's emotion into account. The user's question and emotion label are used as input, and an adjusted response is generated as output.
[1092] Step 7:
[1093] The server sends the generated and adjusted response to the user terminal, which includes the generated response as input and receives it as output.
[1094] Step 8:
[1095] The terminal displays the received response to the user in the virtual space. The displayed response is reflected in the interface so that the user can see it in real time. The response sent to the terminal is used as input, and the response is displayed in the virtual space as output.
[1096] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1097] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1098] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1099] [Fourth embodiment]
[1100] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1101] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1102] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1103] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1104] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1105] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1106] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1107] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1108] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1109] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1110] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1111] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1112] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1113] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user, generates an AI model, and realizes the conversation.
[1114] System Overview
[1115] The user accesses the system and selects the person they want to talk to. The device sends that information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. This AI model is used to generate responses to the user's questions and displays them to the user via the device.
[1116] Program processing
[1117] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[1118] 1. A user accesses the system
[1119] A user accesses the system using a web browser or application. A login screen appears and the user enters their authentication information to log in to the system.
[1120] 2. The user selects the person they want to interact with
[1121] The user clicks the "Start a conversation" button on the main screen and selects historical figure A from the displayed list or search field.
[1122] 3. The device sends the selection information to the server
[1123] The device sends the user's selection information to the server as an HTTP request, which includes the ID of the selected person.
[1124] 4. The server retrieves the data
[1125] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[1126] 5. The server analyzes the data and generates an AI model
[1127] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[1128] 6. User enters question
[1129] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[1130] 7. The server generates a response
[1131] The server analyzes the question and uses an AI model to generate an appropriate response, such as, "The modern economy contains many of the contradictions inherent in capitalist society. The solution to this is..."
[1132] 8. The generated response is displayed to the user
[1133] The server sends the generated response to the terminal, which displays the response on the screen, allowing the user to enjoy a conversation with Great Person A.
[1134] Specific examples
[1135] Example 1: A conversation with historical figure A
[1136] A similar process is repeated if a user wants to talk with historical figure A about the "current political situation." For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..."
[1137] Example 2: Dialogue with the deceased
[1138] If a user wants to imitate a conversation with their late grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if a user mentions "how you've been taking care of your garden lately," the grandmother's AI can respond with something like, "Oh, that's wonderful. I always enjoy watching you take care of your garden."
[1139] Technical details
[1140] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. It also uses incremental learning to continuously update the AI model based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic conversations with deceased people and historical figures.
[1141] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[1142] The processing flow will be explained below.
[1143] Step 1:
[1144] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[1145] Step 2:
[1146] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[1147] Step 3:
[1148] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[1149] Step 4:
[1150] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[1151] Step 5:
[1152] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[1153] Step 6:
[1154] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[1155] Step 7:
[1156] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[1157] Step 8:
[1158] The server analyzes the question and generates a response based on the AI model. The server tokenizes the user's question, analyzes its meaning, and uses the AI model to generate an appropriate response. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism."
[1159] Step 9:
[1160] The server sends the generated response to the terminal, which then sends the response as an HTTP response.
[1161] Step 10:
[1162] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[1163] Step 11:
[1164] If the user wishes to enter more questions, they enter the questions again, and the same process is repeated once the new questions are entered.
[1165] Example 1
[1166] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1167] Conventional communication systems have struggled to enable users to converse with deceased or historical figures in real time. Technical limitations exist, particularly in situations where accurate mimicking of the target person's thought patterns and speaking style is required. This has made it difficult to realize such systems in a wide range of applications, including education, entertainment, and personal healing.
[1168] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1169] In this invention, the server includes a means for selecting a person with whom the user wants to converse, a means for acquiring information about the selected person from a database, and a means for analyzing the acquired information and generating an AI model for imitating the thought patterns and speaking style of the person, thereby enabling the user to enjoy realistic and sophisticated conversations with deceased people and historical figures.
[1170] A "user" is someone who uses the system to interact with a specific person.
[1171] A "person with whom a user wishes to converse" is a deceased person or historical figure that the user selects and with whom the user wishes to converse.
[1172] "Means for selection" is a function that provides an interface for the user to select the person with whom they want to interact from a list or search field.
[1173] "Information" refers to data such as books, papers, transcripts, letters, and audio data related to the target person.
[1174] A "database" is a digital information storage system for storing and managing information within a system.
[1175] The "means for obtaining" is a function for searching and extracting information about a selected person from a database.
[1176] An "artificial intelligence model" is a machine learning model that is generated to mimic the thought patterns and speaking style of a target person.
[1177] "Means of analysis" refers to the function of analyzing acquired information using natural language processing technology and extracting important keywords and phrases.
[1178] The "means for generating a response" is a function that uses an artificial intelligence model to generate an appropriate response based on a user's question.
[1179] The "display means" is an interface for displaying the generated response on the user's terminal screen.
[1180] "Natural language processing technology" is a technology that enables computers to understand and process human language.
[1181] "Keywords" are important words or phrases extracted during the process of analyzing information about a target person.
[1182] The present invention provides a communication system that allows users to enjoy conversations with deceased or historical figures. This system allows users to select the person they want to talk to, acquires and analyzes information about that person, generates an artificial intelligence model, and realizes the conversation.
[1183] Hardware and software used
[1184] Hardware: Servers, devices (PCs, smartphones, tablets, etc.)
[1185] Software: web browsers, applications, database management systems, natural language processing libraries (e.g., Python's NLTK, SpaCy), generative AI models (e.g., OpenAI's GPT-3, Google's BERT)
[1186] Data processing and calculation
[1187] 1. Data Acquisition: The server acquires information about the target person from a database, specifically in the form of books, papers, transcripts, letters, audio data, etc.
[1188] 2. Data Analysis: The server analyzes the acquired information using natural language processing techniques, such as Python's NLTK and SpaCy, to extract important keywords and phrases and model the target person's thought patterns and speaking style.
[1189] 3. Creating a generative AI model: Based on the analyzed data, an artificial intelligence model is generated using OpenAI's GPT-3 or Google's BERT. This model mimics the target person's speaking style and thinking patterns.
[1190] 4. User Question Analysis: The server analyzes the question entered by the user. This analysis involves using natural language processing techniques to understand the meaning and extract relevant keywords.
[1191] 5. Response Generation: Based on the analyzed question, the server uses the generated artificial intelligence model to generate an appropriate response that replicates the target person's thought patterns and speaking style.
[1192] 6. Displaying the response: The server sends the generated response to the terminal, which displays it on the user's screen, allowing the user to enjoy the interaction.
[1193] Specific examples
[1194] For example, suppose a user wants to converse with "Albert Einstein." The user enters "Albert Einstein" into the application's search field and selects it. The server collects papers, notes, and speech data related to Einstein from the database and analyzes them using natural language processing technology. Based on the analysis results, an artificial intelligence model of Einstein is generated using OpenAI's GPT-3.
[1195] When a user inputs a question such as "What do you think about the modern economy?", the server analyzes the question and uses the generated artificial intelligence model to generate a response such as "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." This response is sent to the terminal and displayed on the user's screen.
[1196] Prompt Sentence Examples
[1197] Enter: "What do you think about the modern economy?"
[1198] Response: "The modern economy contains many of the contradictions of capitalist society. The solution to these is..."
[1199] This system allows users to enjoy highly realistic and sophisticated interactions with deceased or historical figures, which is expected to have applications in a variety of fields, including education, entertainment, and personal healing.
[1200] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1201] Step 1:
[1202] A user accesses the system. The user accesses the system using a web browser or application and enters an ID and password on the login screen. The input information is sent from the terminal to the server as an HTTPS request. The server receives this information and performs authentication by checking it against a database. If authentication is successful, the server sends the main screen data to the terminal and login is complete. The input is authentication information (ID and password), and the output is the user's authentication status and the display of the main screen.
[1203] Step 2:
[1204] The user selects the person they want to talk to. The user clicks the "Start a conversation" button displayed on the main screen and selects the historical figure or deceased person they want to talk to from a search field or list. The device sends an HTTP request to the server including the selection information (person ID). The input is the selection information of the person they want to talk to, and the output is the request sent to the server.
[1205] Step 3:
[1206] The device sends the selection information to the server. The device sends an HTTP request to the server, including the ID of the selected person. The server receives this request and executes a query to retrieve information about the selected person from a database. The input is the ID of the selected person, and the output is the result of executing the database query.
[1207] Step 4:
[1208] The server retrieves the data. The server searches and retrieves information related to the selected person from its database (books, papers, transcripts, letters, audio recordings, etc.). This information is returned to the server in JSON format. The input is a database query, and the output is the information retrieved about the target person.
[1209] Step 5:
[1210] The server analyzes the data and generates an AI model. The server uses natural language processing (NLP) techniques to analyze the acquired information, using Python's NLTK and SpaCy libraries. During the analysis, important keywords and phrases are extracted, and based on this, a model is created of the target person's thought patterns and speaking style. OpenAI's GPT-3 and Google's BERT are used for generation. The input is the information to be analyzed, and the output is the generated generative AI model.
[1211] Step 6:
[1212] The user enters a question. The user enters the question into the application's input form, and the question is sent from the terminal to the server. The input is the user's question, and the output is the question sent to the server.
[1213] Step 7:
[1214] The server generates a response. The server analyzes the input question and uses a generative AI model to generate an appropriate response. It uses natural language processing technology to understand the meaning of the question and extract relevant keywords. For example, in response to the question, "What do you think about the modern economy?", it generates a response such as, "The modern economy contains many of the contradictions of capitalist society. The solution to this is..." The input is the user's question, and the output is the generated response.
[1215] Step 8:
[1216] The generated response is displayed to the user. The server sends the generated response in JSON format to the terminal. The terminal displays the received response data on the screen, allowing the user to enjoy the interaction. The input is the generated response, and the output is the response displayed on the screen.
[1217] (Application example 1)
[1218] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1219] Currently, there are technologies that allow users to enjoy conversations with historical figures and deceased people, but these conversations can only be realized on limited devices, which often limits the experience. Furthermore, it is difficult to visually display the content of the conversation, which does not improve the user experience. Furthermore, there is a lack of technology to improve the accuracy of the conversation content desired by users, making it difficult to realize realistic conversations.
[1220] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1221] In this invention, the server includes: means for selecting a person the user wants to converse with, as input by the user; means for acquiring data about the selected person from a database; means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style; means for receiving the user's question and generating a response based on the AI model; means installed in the smart glasses for displaying the content of the conversation with the person selected by the user; and means for displaying the generated response to the user. This allows the user to enjoy a visual conversation in real time using the smart glasses, improving the quality of the experience. Furthermore, it is possible to accurately imitate the thought patterns and speaking style of the selected person, thereby realizing a more realistic conversation.
[1222] The "means for selecting a person to converse with input by the user" is an interface within the system that allows the user to select a specific person as a conversation partner.
[1223] The "means for obtaining data relating to a selected person from a database" is a function for obtaining information relating to a person selected by a user from a database stored within the system.
[1224] "Means for analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style" refers to technology for analyzing the data of a selected person and creating an AI model that reproduces the person's speaking style and thought patterns.
[1225] "Means for receiving a user's question and generating a response based on an AI model" refers to the function by which the system receives a question from a user and the AI model generates an appropriate response based on that question.
[1226] "A means installed in smart glasses to display the content of the conversation with a person selected by the user" is a function installed in smart glasses to visually provide the user with the response of the person with whom they want to converse.
[1227] "Means for displaying the generated response to the user" refers to an interface that allows the user to see the response generated by the AI model.
[1228] "Books, papers, transcripts, letters, and audio recordings" are sources of information related to the selected individuals from which data will be drawn and used in the analysis.
[1229] "Means of analyzing input data using natural language processing technology and extracting important keywords and phrases" is a technology that utilizes natural language processing technology to extract important information from user input and dialogue.
[1230] This invention is a system that allows a user to select a person with whom they want to converse, acquires and analyzes data about that person to generate an AI model, and realizes a dialogue with the user in real time. This system is installed in smart glasses and can visually display the dialogue content to the user. Below, a detailed description of an embodiment of this invention is provided.
[1231] System configuration
[1232] This system is configured using the following hardware and software.
[1233] Smart glasses: Devices that display interactions to the user in real time and also serve as an input / output interface.
[1234] Server: Responsible for back-end processing to acquire and analyze the data of the selected person and generate an AI model.
[1235] Database: Contains data (books, papers, transcripts, letters, audio recordings, etc.) about historical figures and deceased people.
[1236] Natural language processing technology: Technology used to analyze data and extract key keywords and phrases (e.g., GPT-2 model).
[1237] Program processing overview
[1238] 1. User inputs the person they want to talk to:
[1239] The user wearing the smart glasses selects the person they want to talk to. The user interface displays a list of people they want to talk to, allowing them to select one.
[1240] 2. Retrieving data from the database:
[1241] Data about the selected person (e.g., books, papers, transcripts, letters, audio recordings, etc.) is retrieved from a database by the server.
[1242] 3. Data analysis and AI model generation:
[1243] The server then analyzes the data using natural language processing techniques to extract key keywords and phrases, which then generates an AI model that mimics the thought patterns and speaking style of the selected person.
[1244] 4. Receiving questions and generating responses:
[1245] The user types a question through the smart glasses, which is sent to a server where an AI model analyzes it and generates an appropriate response.
[1246] 5. View the response:
[1247] The generated responses are displayed in real time on the smart glasses, enabling interaction with the user.
[1248] Specific examples
[1249] For example, consider a case where a user asks a historical scientist, "Tell me about the theory of relativity." In this case, the AI model uses the acquired data to generate a response such as, "The theory of relativity is based on the idea that time and space are relative, not absolute..." The interaction takes place in real time and is displayed through the smart glasses, improving the user experience.
[1250] Prompt Sentence Examples
[1251] User: Tell me about the theory of relativity.
[1252] AI: The theory of relativity is based on the idea that time and space are relative, not absolute...
[1253] Thus, the present invention provides a system that allows users to enjoy real-time interactions through smart glasses. The natural language processing technology used in the smart glasses can improve the quality of the interactions and the user experience.
[1254] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1255] Step 1:
[1256] The user puts on the smart glasses and selects the person they want to interact with. This selection information is displayed on the user interface, and once the user selects the target person, the selection information is retained in the smart glasses and sent to the server for the next step.
[1257] Input: User-selected person information
[1258] Output: Selection information sent to the server
[1259] Step 2:
[1260] The server retrieves data about the selected person. The server queries the database to retrieve related information such as books, papers, transcripts, letters, and audio recordings.
[1261] Input: Selection information sent from smart glasses
[1262] Output: Person-related data retrieved from the database
[1263] Step 3:
[1264] The server analyzes the acquired data using natural language processing techniques (e.g., the GPT-2 model) to extract important keywords and phrases from the data. The result is an AI model that mimics the individual's thought patterns and speaking style.
[1265] Input: Person-related data retrieved from a database
[1266] Output: AI model
[1267] Step 4:
[1268] The user inputs a question through the smart glasses, which receives the question and sends it to the server.
[1269] Input: The question entered by the user into the smart glasses
[1270] Output: The question sent to the server
[1271] Step 5:
[1272] The server uses an AI model to analyze the received question and generate an appropriate response. The AI model generates a response to the question and prepares it in text format.
[1273] Input: User-submitted question, AI model
[1274] Output: The generated response (in text format)
[1275] Step 6:
[1276] The generated response is sent to the smart glasses and displayed to the user, who can view the interaction content through the smart glasses.
[1277] Input: The generated response sent by the server
[1278] Output: Response displayed on the smart glasses
[1279] In this way, this system provides an environment in which users can interact with historical figures and deceased people in real time using smart glasses, and delivers a high-quality interactive experience by analyzing the acquired data and generating responses using an AI model.
[1280] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1281] This invention provides a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model, which then combines it with an emotion engine to recognize the user's emotions and realize the conversation.
[1282] System Overview
[1283] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[1284] Program processing
[1285] As a specific example of processing, consider a case where a user wants to discuss modern economics with historical figure A. The processing will be explained below in natural language.
[1286] 1. A user accesses the system
[1287] A user logs into the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[1288] 2. The user selects the person they want to interact with
[1289] The user clicks the "Start conversation" button displayed on the main screen and selects historical figure A from the displayed list or search field.
[1290] 3. The device sends the selection information to the server
[1291] The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[1292] 4. The server retrieves the data
[1293] The server searches and retrieves data about the selected person from the database (books, papers, transcripts, letters, etc.), which comprehensively covers the person's thoughts and statements.
[1294] 5. The server analyzes the data and generates an AI model
[1295] The acquired data is analyzed on a server using natural language processing (NLP) technology. Important keywords and phrases are extracted, and the individual's thought patterns and speaking style are modeled. The generated AI model can mimic the words and thoughts of historical figure A.
[1296] 6. User enters question
[1297] The user uses the input form to enter the question, "What do you think about the current economy?" The input information is sent from the terminal to the server.
[1298] 7. The server analyzes the user's emotions using an emotion engine.
[1299] The server analyzes the user's question with an emotion engine and processes the question content with emotion labels. For example, the question is tagged with emotion labels such as "interesting" and "excited."
[1300] 8. The server analyzes the question and generates a response based on the AI model
[1301] The server tokenizes the user's question, analyzes its meaning, and uses an AI model to generate an appropriate response. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it might output something like, "The modern economy is seeing the growing contradictions of capitalism" in a calm, empathetic tone.
[1302] 9. Send the server-generated response to the device
[1303] The server sends the generated response to the terminal as an HTTP response.
[1304] 10. The device receives the response from the server and displays it to the user
[1305] The device displays the received text in a chat window on the screen for the user to view.
[1306] Specific examples
[1307] Example 1: A conversation with historical figure A
[1308] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[1309] Example 2: Dialogue with the deceased
[1310] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[1311] Technical details
[1312] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, enhancing the naturalness and intimacy of the conversation. Furthermore, by using incremental learning, the AI model is continuously updated based on the conversation history, achieving highly accurate conversations. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures.
[1313] The present invention has wide applicability not only to academic research and personal healing, but also to the fields of education and entertainment.
[1314] The processing flow will be explained below.
[1315] Step 1:
[1316] A user accesses the system. The user logs in to the system using a web browser or application. After the user enters their login information and is successfully authenticated, the main screen is displayed.
[1317] Step 2:
[1318] The user selects the person they want to talk to. The user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The selection results are displayed on the screen.
[1319] Step 3:
[1320] The device sends the selection information to the server. The device generates an HTTP request including the ID of the person selected by the user and sends this request to the server.
[1321] Step 4:
[1322] The server retrieves the data of the selected person from the database. The server analyzes the request received from the user and searches and retrieves the data corresponding to the selected person from the database. This data may include books, papers, transcripts, letters, etc.
[1323] Step 5:
[1324] The server analyzes the acquired data and generates an AI model that mimics the person. The server uses natural language processing (NLP) technology to analyze the acquired data and extract important keywords and phrases. The server then trains the AI model based on the analysis results and generates a model that mimics the person's thought patterns and speaking style.
[1325] Step 6:
[1326] A user enters a question into an input form, for example, "What do you think about the modern economy?", and clicks the "Submit" button.
[1327] Step 7:
[1328] The device sends the user's question to the server. The device sends the entered question to the server as an HTTP request. This request includes the user's question.
[1329] Step 8:
[1330] The server analyzes the user's emotions using an emotion engine. The server analyzes the user's question using the emotion engine and processes the question content along with an emotion label. For example, the question is tagged with an emotion label such as "interesting" or "excited."
[1331] Step 9:
[1332] The server analyzes the question and generates a response based on an AI model. The server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. For example, it could output something like, "The contradictions of capitalism are expanding in the modern economy," in a calm, empathetic tone.
[1333] Step 10:
[1334] The server sends the generated response to the terminal. The server sends the generated response to the terminal as an HTTP response.
[1335] Step 11:
[1336] The terminal receives the response from the server and displays it to the user. The terminal displays the received text in a chat window on the screen for the user to view.
[1337] Step 12:
[1338] The user then enters another question, and the process repeats. The question and emotion label are sent to the server, and the AI model generates an appropriate response, which is then displayed on the device.
[1339] Specific examples
[1340] Example 1: A conversation with historical figure A
[1341] If a user wants to talk to historical figure A about the "current political situation," the same process is repeated. For example, historical figure A can generate a response such as, "The current political situation is compared to past historical events..." Also, if the question contains "anger," the AI model of historical figure A adjusts the tone of the response to remain calm.
[1342] Example 2: Dialogue with the deceased
[1343] If a user wants to imitate a conversation with their deceased grandmother, the AI model can be generated based on her letters and recordings. Using a similar process, if the user mentions "how you've been tending to your garden lately," the grandmother's AI can respond with a tone of voice that adjusts based on the user's emotions, such as "Oh, that's wonderful. I always enjoyed tending to your garden."
[1344] Example 2
[1345] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1346] Conventional communication systems have struggled to provide a natural dialogue experience for users to converse with deceased or historical figures. Furthermore, the content of the dialogue is not adjusted according to the user's emotions, resulting in a decline in dialogue quality. Furthermore, there are limitations to the accuracy of the technology used to realistically imitate a person's thought patterns and speaking style. This makes it difficult for users to obtain a sense of satisfaction from the dialogue.
[1347] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1348] In this invention, the server includes: means for selecting a person the user wants to converse with, means for retrieving data about the selected person from a database, means for analyzing the retrieved data and generating an AI model for imitating the person's thought patterns and speaking style, means for analyzing the user's emotions and assigning emotion labels, means for receiving the user's question and generating a response based on the AI model and the emotion labels, and means for displaying the generated response to the user. This allows the user to enjoy more natural and emotional conversations with deceased people and historical figures.
[1349] A "user" is a person who accesses and interacts with the system.
[1350] "People you want to talk to" refers to deceased people or historical figures with whom you want to talk through the system.
[1351] "Means for selection" refers to the interface or operation that allows a user to select a person within the system with whom they wish to interact.
[1352] "Database" refers to a magnetic disk drive or other storage system that stores information such as books, articles, transcripts, letters, email posts, audio recordings, etc. about a person with whom you wish to interact.
[1353] "Means for obtaining" refers to a process for searching and retrieving data about a selected person from a database.
[1354] "Means of analysis" refers to the process of analyzing data using natural language processing technology and extracting important keywords and phrases.
[1355] An "AI model" refers to an artificial intelligence program that is generated from acquired data and is intended to mimic the thinking patterns and speaking style of a specific person.
[1356] An "emotion engine" refers to a software module that analyzes emotions from user input and assigns emotion labels.
[1357] An "emotion label" refers to a tag that expresses an emotion, such as "interested" or "excited," in response to a user's input.
[1358] "Means for generating a response" refers to the process of analyzing the user's question and creating an appropriate response based on the AI model and emotion labels.
[1359] "Means for displaying" refers to a screen or interface for visually presenting the generated response to the user.
[1360] This invention relates to a communication system that allows users to enjoy conversations with deceased people and historical figures. This system acquires and analyzes data on people selected by the user to generate an AI model. It then combines this data with an emotion engine to recognize the user's emotions and realize the conversation.
[1361] System Overview
[1362] The user accesses the system and selects the person they want to talk to. The device sends this information to the server, which retrieves the corresponding person's information from a database. The retrieved information is analyzed by the server, which generates an AI model that mimics the person's thought patterns and speaking style. In addition, the emotion engine analyzes the user's emotions, and the generated responses are adjusted according to the user's emotions. This combination of the AI model and emotion engine allows users to enjoy more natural and emotionally rich conversations.
[1363] Program processing
[1364] When a user accesses the system and selects the person they wish to interact with, the device sends the corresponding person's ID to the server. The server retrieves data about that person from a database and analyzes it using natural language processing technology. For example, natural language processing libraries such as SpaCy and NLTK are used. As a result, important keywords and phrases are extracted and an AI model that mimics the target person's thought patterns and speaking style is generated.
[1365] When a user enters a question, the device sends it to the server. The server analyzes the question using an emotion engine and assigns an emotion label to the question using, for example, TextBlob or a sentiment analysis API. The server then tokenizes the question and generates a response based on an AI model. The tone and content of the response are adjusted based on the emotion engine's analysis results. The generated response is then sent to the device, which displays it to the user.
[1366] Specific examples
[1367] Example 1: A conversation with historical figure A
[1368] If a user wants to talk with historical figure A about the "current political situation," the user begins the conversation as follows:
[1369] Example prompt sentence:
[1370] "Ask historical figure A what he thinks about the current political situation."
[1371] The server retrieves data on great person A from the database, analyzes it, and generates an AI model. It then analyzes the user's question along with the emotion label and generates a response such as, "Compared to past historical events, the current political situation is..."
[1372] Example 2: Dialogue with the deceased
[1373] If the user wanted to mimic a conversation with their deceased grandmother, the user's prompt might look like this:
[1374] Example prompt sentence:
[1375] "Tell your late grandmother about your recent gardening."
[1376] The server generates an AI model based on the grandmother's letters and audio recordings, and generates responses such as "Oh, that's great. I always enjoyed tending to your garden" based on emotion labels.
[1377] This system uses natural language processing technology to perform detailed analysis of the target person's data and generate an advanced AI model. Furthermore, by combining it with an emotion engine, it recognizes the user's emotions and generates responses accordingly, improving the naturalness and intimacy of the dialogue. This allows users to enjoy highly realistic and emotionally rich conversations with deceased people and historical figures. This system has broad applications not only in academic research and personal healing, but also in education and entertainment.
[1378] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1379] Step 1:
[1380] A user accesses the system.
[1381] Specifically, a user accesses the system using a web browser or mobile app. The user enters login information (username and password) and presses the "Login" button. The input is the username and password string, and the output is the success or failure of authentication and the display of the main screen.
[1382] Step 2:
[1383] The user selects the person with whom they want to interact.
[1384] Specifically, the user clicks the "Start a conversation" button on the main screen and selects the person they want to talk to from the displayed list or search field. The input is the name or ID of the person they want to talk to, and the output is data including the ID of the selected person.
[1385] Step 3:
[1386] The terminal transmits the selection information to the server.
[1387] Specifically, the device generates an HTTP request including the ID of the selected person and sends it to the server. The input is JSON-formatted data including the ID of the selected person, and the output is an HTTP request to the server.
[1388] Step 4:
[1389] The server retrieves the data.
[1390] Specifically, the server searches and retrieves data about the selected person from a database. The input is the selected person's ID, and the output is data such as books, papers, transcripts, and letters retrieved from the database. The server retrieves this data from local storage or cloud storage.
[1391] Step 5:
[1392] The server analyzes the data and generates an AI model.
[1393] Specifically, the server analyzes the acquired data using natural language processing techniques (e.g., SpaCy or NLTK). Important keywords and phrases are extracted, and an AI model is generated using machine learning techniques (e.g., TensorFlow or PyTorch). The input is person data acquired from a database, and the output is an AI model that mimics the thinking patterns and speaking style of a specific person.
[1394] Step 6:
[1395] The user enters a question.
[1396] Specifically, the user enters a question into the text box on the interactive screen and presses the "Send" button. The input is the question entered by the user, and the output is the question data sent to the server.
[1397] Step 7:
[1398] The server analyzes the user's emotions using an emotion engine.
[1399] Specifically, the server analyzes the received question using an emotion engine (e.g., TextBlob or an emotion analysis API) and assigns an emotion label. The input is the user's question, and the output is the question data including the emotion label.
[1400] Step 8:
[1401] The server analyzes the question and generates a response based on an AI model.
[1402] Specifically, the server tokenizes the user's question, analyzes its meaning, and generates an appropriate response using an AI model. The tone and content of the response are adjusted based on the analysis results of the emotion engine. The input is the question data with emotion labels and the AI model, and the output is the generated response text.
[1403] Step 9:
[1404] The server generates a response and sends it to the terminal.
[1405] Specifically, the server sends the generated response in JSON format to the terminal as an HTTP response. The input is the generated response text, and the output is the HTTP response to the terminal.
[1406] Step 10:
[1407] The terminal receives the response from the server and displays it to the user.
[1408] Specifically, the device displays the received text in a chat window on the screen for the user to view. The input is the HTTP response from the server, and the output is the response text displayed in the chat window on the screen.
[1409] (Application example 2)
[1410] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1411] Conventional AI dialogue systems provide users with simple question and answer responses, but lack mechanisms for recognizing emotions and enabling natural dialogue. Furthermore, they lack an interactive experience in virtual space, resulting in low user satisfaction. The objective of this invention is to provide a dialogue system that provides a sense of realism by utilizing emotion recognition and virtual space.
[1412] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for selecting a person with whom the user wants to converse, input by the user, means for acquiring data about the selected person from a database, means for analyzing the acquired data and generating an AI model for imitating the person's thought patterns and speaking style, means including an emotion engine for recognizing and analyzing the user's emotions, means for adjusting the response content based on the analyzed emotions, and means for displaying this in a virtual space. This enables natural and realistic conversations that correspond to emotions.
[1413] The "means for selecting a person the user wishes to converse with" refers to hardware or software that provides an interface for the user to input or select a person with whom the user wishes to converse.
[1414] The "means for obtaining data on the selected person from the database" refers to a process or database communication means for accessing and obtaining data on the selected person from a database on a server or cloud.
[1415] "Means of analyzing acquired data and generating an AI model that imitates the person's thought patterns and speaking style" refers to technology that analyzes acquired data using natural language processing and machine learning techniques and generates an AI model that imitates the person.
[1416] The "means for receiving a user's question and generating a response based on an AI model" is a processing system that receives a user's question as input and generates an appropriate response using the algorithm of the AI model.
[1417] "Means for displaying the generated response to the user" refers to an interface or output device for displaying the response generated by the AI model on the user's terminal.
[1418] The "means including an emotion engine for recognizing and analyzing user emotions" refers to a software and hardware configuration for analyzing emotions from inputs and interactions from a user, and recognizing and processing the emotions.
[1419] The "means for adjusting the content of the response based on the analyzed emotions" is a technology that adaptively adjusts the tone and content of the generated response based on the user's emotions analyzed by the emotion engine.
[1420] "Means for displaying this in a virtual space" refers to the display devices and virtual space rendering technologies necessary for users to enjoy a realistic interaction in a virtual space.
[1421] System Configuration
[1422] A system for implementing this invention includes a user terminal, a server, a database, and an emotion engine. The user terminal is equipped with an interface that allows the user to select a person with whom they want to interact. The server accesses the database to obtain data about the person and generates an AI model. The server also uses the emotion engine to analyze the user's emotions and adjust the response content.
[1423] Program processing
[1424] When a user selects a person they want to talk to through their device, the device sends that information to a server. The server then retrieves data about the selected person from a database, such as books, papers, transcripts, letters, audio data, and video data. This data is analyzed using natural language processing technology to generate an AI model that mimics the person's thought patterns and speaking style. This generated AI model is then used to provide appropriate responses based on the user's questions.
[1425] When a user enters a question, it is sent to the server. The server uses an emotion engine to analyze the question's sentiment and adjusts the AI model's response based on the analysis results. After the response is generated, the server sends it to the user's device, which then displays it to the user in a virtual space.
[1426] Hardware and software used
[1427] User devices: smartphones, smart glasses, head-mounted displays
[1428] Server: Database Management System (DBMS), Web Server
[1429] Emotion engine: Emotion API (e.g. Microsoft Emotion API)
[1430] Natural language processing technology: OpenAI GPT, natural language tokenization libraries (e.g., spaCy)
[1431] Specific examples
[1432] For example, consider a scenario where a user enters a virtual cafe and chooses to interact with a great inventor. If the user asks, "What do you think about modern technology?", the server analyzes the question using its emotion engine and tags it as "interesting." Once the analysis is complete, the AI model generates a response such as, "Modern technology has advanced in ways never before imagined," which is then displayed on the user's device. The response is adjusted in tone to reflect the user's emotions.
[1433] Prompt Sentence Examples
[1434] Imagine the following dialogue:
[1435] You are a great inventor. You have had a huge impact on modern life with your inventions, and you have many inventions to your name.
[1436] User: What do you think about modern technology? [Interested]
[1437] In this way, the system can realize natural and realistic dialogue that responds to emotions, providing users with an interactive experience.
[1438] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1439] Step 1:
[1440] The user selects the person they want to talk to through their terminal. The person selection information entered by the user is sent from the terminal to the server. The ID of the person selected by the user is included as input, and the server receives it as output.
[1441] Step 2:
[1442] The server uses the received person ID to retrieve data about the selected person from a database, including books, papers, transcripts, letters, audio data, and video data. The person ID is used as input, and the target data is obtained as output.
[1443] Step 3:
[1444] The server analyzes the acquired data and generates an AI model to mimic the person's thought patterns and speaking style. It uses natural language processing technology to tokenize the data and extract important keywords and phrases. The acquired data is used as input, and the AI model is generated as output.
[1445] Step 4:
[1446] The user inputs a question through a terminal. The question is sent from the terminal to the server. The input contains the user's question, and the server receives it as output.
[1447] Step 5:
[1448] The server uses an emotion engine to perform emotion analysis based on the user's question. The analysis generates an emotion label for the question. The user's question is used as input, and the emotion label is generated as output.
[1449] Step 6:
[1450] The server uses an AI model to generate an appropriate response based on the analyzed emotion label. The response content is adjusted to take the user's emotion into account. The user's question and emotion label are used as input, and an adjusted response is generated as output.
[1451] Step 7:
[1452] The server sends the generated and adjusted response to the user terminal, which includes the generated response as input and receives it as output.
[1453] Step 8:
[1454] The terminal displays the received response to the user in the virtual space. The displayed response is reflected in the interface so that the user can see it in real time. The response sent to the terminal is used as input, and the response is displayed in the virtual space as output.
[1455] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1456] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1457] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1458] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1459] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1460] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1461] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1462] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1463] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1464] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1465] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1466] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1467] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1468] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1469] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1470] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1471] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1472] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1473] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1474] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1475] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1476] The following is further disclosed regarding the above embodiment.
[1477] (Claim 1)
[1478] A means for selecting a person input by the user with whom the user wishes to converse;
[1479] means for retrieving data about the selected person from a database;
[1480] A means of analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style; and
[1481] means for receiving a user question and generating a response based on an AI model;
[1482] means for displaying the generated response to a user;
[1483] A system including:
[1484] (Claim 2)
[1485] The system of claim 1, wherein the database stores data relating to the selected person, such as books, papers, transcripts, letters, social media posts, and audio recordings.
[1486] (Claim 3)
[1487] The system of claim 1, further comprising means for analyzing input data using natural language processing techniques to generate an AI model and extracting important keywords and phrases.
[1488] (Claim 4)
[1489] The system of claim 1, further comprising means for generating a response using the generated AI model while taking into account past conversation history.
[1490] (Claim 5)
[1491] 10. The system of claim 1, further comprising means for incrementally updating the AI model using data input by a user during interaction.
[1492] "Example 1"
[1493] (Claim 1)
[1494] A means for selecting a person input by the user with whom the user wishes to converse;
[1495] means for retrieving information about the selected person from a database;
[1496] means for analyzing the acquired information and generating an artificial intelligence model for imitating the person's thought patterns and speaking style;
[1497] means for receiving a user question and generating a response based on an artificial intelligence model;
[1498] means for displaying the generated response to a user;
[1499] A system including:
[1500] (Claim 2)
[1501] 2. The system according to claim 1, wherein the information about the selected person includes books, papers, transcripts, letters, audio data, etc. stored in the database.
[1502] (Claim 3)
[1503] 2. The system of claim 1, further comprising means for analyzing input data using natural language processing techniques to generate an artificial intelligence model and extracting important keywords and phrases.
[1504] "Application Example 1"
[1505] (Claim 1)
[1506] A means for selecting a person input by the user with whom the user wishes to converse;
[1507] means for retrieving data about the selected person from a database;
[1508] A means of analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style; and
[1509] means for receiving a user question and generating a response based on an AI model;
[1510] means installed on the smart glasses for displaying the content of a conversation between the user and a person selected by the user;
[1511] means for displaying the generated response to a user;
[1512] A system including:
[1513] (Claim 2)
[1514] 2. The system according to claim 1, wherein the database stores books, papers, transcripts, letters, audio data, etc. as data relating to the selected person.
[1515] (Claim 3)
[1516] The system of claim 1, further comprising means for analyzing input data using natural language processing techniques to generate an AI model and extracting important keywords and phrases.
[1517] "Example 2: Combining Emotion Engines"
[1518] (Claim 1)
[1519] A means for selecting a person input by the user with whom the user wishes to converse;
[1520] means for retrieving data about the selected person from a database;
[1521] A means of analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style; and
[1522] means for assigning emotion labels, the emotion engine including:
[1523] means for receiving a user question and generating a response based on the AI model and the emotion label;
[1524] means for displaying the generated response to a user;
[1525] A system including:
[1526] (Claim 2)
[1527] 2. The system of claim 1, wherein the database stores data relating to the selected person, such as books, papers, transcripts, letters, email posts, and audio recordings.
[1528] (Claim 3)
[1529] The system of claim 1, further comprising means for analyzing input data using natural language processing techniques to generate an AI model and extracting important keywords and phrases.
[1530] "Application example 2 when combining emotion engines"
[1531] (Claim 1)
[1532] A means for selecting a person input by the user with whom the user wishes to converse;
[1533] means for retrieving data about the selected person from a database;
[1534] A means of analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style; and
[1535] means for receiving a user question and generating a response based on an AI model;
[1536] means for displaying the generated response to a user;
[1537] means including an emotion engine for recognizing and analyzing the emotions of a user;
[1538] means for adjusting response content based on the analyzed emotion;
[1539] A means for displaying this in virtual space;
[1540] A system including:
[1541] (Claim 2)
[1542] 2. The system according to claim 1, wherein the database stores books, papers, transcripts, letters, audio data, video data, etc., relating to the selected person.
[1543] (Claim 3)
[1544] The system of claim 1, further comprising means for analyzing input data using natural language processing techniques to generate an AI model and extracting important keywords and phrases. [Explanation of symbols]
[1545] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for selecting a person input by the user with whom the user wishes to converse; means for retrieving data about the selected person from a database; A means of analyzing the acquired data and generating an AI model to mimic the person's thought patterns and speaking style; and means for receiving a user question and generating a response based on an AI model; means for displaying the generated response to a user; A system including:
2. The system according to claim 1, wherein the database stores data relating to the selected person, such as books, papers, transcripts, letters, social media posts, and audio recordings.
3. The system according to claim 1, further comprising means for analyzing input data using natural language processing techniques to generate an AI model and extracting important keywords and phrases.
4. The system according to claim 1, further comprising means for generating a response using the generated AI model while taking into account past conversation history.
5. 2. The system of claim 1, further comprising means for incrementally updating the AI model using data input by a user during interaction.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A