System

The system enables interactive and immersive learning by managing historical figure data, generating responses through AI, and simulating realistic interactions, addressing the limitations of traditional educational methods.

JP2026025681APending Publication Date: 2026-02-16SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024128493
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-02
Publication Date
2026-02-16

AI Technical Summary

Technical Problem

Existing educational and training systems lack immersive and interactive methods to provide deep insights into historical figures, limiting the understanding of their thoughts and personalities, and failing to facilitate meaningful interactions among them.

Method used

A system that includes a database for managing historical figure data, an interface for user input, a communication mechanism for request analysis, a generative AI model for answer generation, and an output mechanism for response transmission, along with lip-sync and simulation capabilities to create realistic interactions.

Benefits of technology

Enhances learning experiences by allowing users to engage in realistic conversations with historical figures, providing new perspectives and deeper understanding through interactive and immersive learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026025681000001_ABST
    Figure 2026025681000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: a database means for managing information about historical figures; an interface means for accepting input from a user; a communication means for analyzing a request from the user and obtaining information about a specified historical figure from the database means; a generative AI model means for providing a generated answer; and an outputting means for transmitting the generated answer to the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] When learning history, textbooks and other books provide only limited information, making it difficult to gain a deep understanding of the thoughts and personalities of people who actually lived in that era. Furthermore, there are a lack of ways to have different historical figures interact with each other, allowing learners to gain new perspectives. Furthermore, there are limited visual and audio interaction methods to enhance the immersive learning experience. [Means for solving the problem]

[0005] The present invention solves these problems with a system including a database means for managing data on historical figures, an interface means for accepting input from a user, a communication means for analyzing a request from a user and retrieving information on the specified historical figure from the database means, a generation AI model means for providing a generated answer, and an output means for transmitting the generated answer to the user. Furthermore, by including a lip-sync means for performing lip-sync processing on a facial photograph of the historical figure based on the generated answer, and a simulation means for having different historical figures interact with each other and simulating the development of that interaction, it is possible to enhance the immersive feeling of learning and provide new perspectives.

[0006] "Database means" is a part of a system that manages information about historical figures and makes that information available when needed.

[0007] An "interface means" is an input device or screen through which a user accesses the system and inputs questions.

[0008] The "communication means" is a system component that analyzes requests from users, acquires specified information from database means, and exchanges data with other system parts.

[0009] The "generative AI model means" is an artificial intelligence algorithm that generates appropriate answers based on information obtained from a database and the user's question.

[0010] "Output means" refers to a device or function for providing the answer generated by the generation AI model means to the user.

[0011] The "lip sync means" is a processing system for synchronizing the mouth movements with the facial photograph of a historical figure based on the generated answer.

[0012] "Simulation means" is a system function that allows different historical figures to converse with each other and reproduce the development of that conversation. [Brief explanation of the drawings]

[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0021] [First embodiment]

[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0034] The present invention relates to a system for managing data on historical figures and allowing users to have conversations as if they were the historical figures. The system includes a database, an interface, a communication means, a generative AI model, an output means, a lip-sync means, and a simulation means.

[0035] System Components

[0036] server

[0037] The server is the core of the system and includes a database, communication, AI model generation, and output. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[0038] Terminal

[0039] The terminal provides an interface for the user to interact with the system: it accepts user input, transmits it to the server, and displays responses received from the server to the user.

[0040] User

[0041] Through this system, users interact with historical figures by operating a terminal, inputting questions, and reading or listening to the answers provided by the server.

[0042] Program Processing Overview

[0043] Terminal

[0044] The terminal provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying answers to the user. If a lip-sync function is included, a photo of the person's face is lip-synced and displayed.

[0045] Examples:

[0046] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[0047] server

[0048] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. The generated answer is sent to the terminal through the output means.

[0049] Examples:

[0050] The server retrieves information about "Napoleon" and "tactics" from a database, inputs it into a generative AI model, and generates the answer, "The tactics I value most are rapid advance and attacking the enemy's weak points."

[0051] Lip Sync Means

[0052] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[0053] Examples:

[0054] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[0055] Simulation Method

[0056] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[0057] Examples:

[0058] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[0059] The system allows users to interact with historical figures in real time, providing a deeper historical understanding and learning experience.

[0060] The processing flow will be explained below.

[0061] Specific processing flow of the program

[0062] Step 1:

[0063] The user enters a question into the terminal interface and clicks the send button.

[0064] Step 2:

[0065] The device formats the user's question as an API request and sends it to the server.

[0066] Step 3:

[0067] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[0068] Step 4:

[0069] The server searches and acquires information about the designated historical person from the database means.

[0070] Step 5:

[0071] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[0072] Step 6:

[0073] The server formats the generated answer and creates response data.

[0074] Step 7:

[0075] The server sends the response data to the terminal.

[0076] Step 8:

[0077] The terminal receives the response data from the server and formats it for display to the user.

[0078] Step 9:

[0079] The terminal displays the answer to the user, and if there is lip synchronization means, generates and displays the mouth movements.

[0080] Step 10:

[0081] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[0082] This process allows users to simulate real-time interactions with historical figures, enhancing the learning experience.

[0083] Example 1

[0084] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0085] In the modern education system, history classes tend to provide only one-sided information, making it difficult for students to acquire interest or concern. In particular, there is a lack of concrete ways to deepen knowledge about historical figures. There is a growing need for interactive educational tools that will engage students and trainees in learning.

[0086] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0087] In this invention, the server includes an information aggregation means for managing data on historical figures, an input means for accepting input from a user, a receiving means for analyzing a request from the user and acquiring information on the specified historical figure from the information aggregation means, a generation AI model means for providing a generated answer, and a transmitting means for transmitting the generated answer to the user. This allows the user to have an interactive conversation with the historical figure, thereby improving the learning effect.

[0088] "Information aggregation means" refers to the means for managing data on historical figures and collecting and organizing necessary information.

[0089] The "input means" is an interface that accepts input from the user, and is a means by which the user inputs questions or requests.

[0090] The "receiving means" is a means for analyzing a request from a user and obtaining information about a specified historical person from the information aggregating means.

[0091] A "generative AI model means" is a means that uses an artificial intelligence model to generate an answer based on acquired information.

[0092] The "transmission means" is a means for transmitting the generated answer to the user.

[0093] The "lip sync means" is a means for synthesizing lip movements in real time with an image of a historical figure based on the generated answer.

[0094] A "simulation method" is a method for arranging a dialogue between different historical figures and reproducing the development of that dialogue.

[0095] This invention is a system for managing data on historical figures and allowing users to have conversations as if they were those figures. The system includes an information aggregation unit, an input unit, a receiving unit, a generating AI model unit, a transmitting unit, a lip-sync unit, and a simulation unit.

[0096] server

[0097] The server is the central part of the system and is the hardware that controls various functions. The server has the following means for processing data:

[0098] Information aggregation tool: A tool for managing data about historical figures. This data is stored in a database and includes detailed historical information about the person and records of their statements.

[0099] Reception means: A means for receiving requests sent by users. The received requests are analyzed and necessary keywords and contexts are extracted.

[0100] Generative AI model means: A means for obtaining necessary data from the information aggregation means based on the analysis results obtained from the receiving means, and then using an AI model to generate an appropriate response. The generative AI model is, for example, a natural language generation model such as GPT (Generative Pre-trained Transformer).

[0101] Transmission means: A means for transmitting the generated answer to the terminal, allowing the user to receive the generated answer.

[0102] Terminal

[0103] A terminal is a piece of hardware that provides an interface for users to interact with a system. Specifically, it operates as follows:

[0104] Input means: A means that provides an interface for users to input questions. Users input their questions using this means and send them to the server by pressing the send button.

[0105] Lip-syncing: A method for synthesizing lip movements in real time onto images of historical figures based on answers sent from the server. This process provides a more natural and interactive dialogue experience.

[0106] Examples:

[0107] The user types "Tell me about your tactics" into the device interface and presses the send button. The server receives this question, analyzes it, and extracts the keywords "Napoleon" and "tactics." The server obtains related data from the information aggregation means and inputs it into the generative AI model. The generative AI model generates an answer, "The tactics I value most are rapid advancement and attacking the enemy's weak points," and sends this to the device. The device receives this answer and displays it to the user using lip-syncing means, replicating the mouth movements of an image of Napoleon.

[0108] Prompt Sentence Examples

[0109] "Tell me about your tactics."

[0110] "Generate detailed information about Napoleon's tactics."

[0111] conclusion

[0112] This system allows users to interact with historical figures, enhancing learning effectiveness, and its lip-sync function provides a realistic experience both visually and aurally.

[0113] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0114] Processing Steps

[0115] Step 1:

[0116] User enters question and submits

[0117] The user uses the device interface to input a question for the historical figure. When input is complete, the user presses the "Send" button. This operation sends the question data from the device to the server.

[0118] Input: "Tell me about your tactics."

[0119] Output: Question data (e.g., "Tell me about your tactics")

[0120] Step 2:

[0121] The server receives and analyzes the query

[0122] The server receives the question data sent from the device, analyzes the received question data, and extracts keywords and context, thereby identifying the type and subject of the requested information.

[0123] Input: Question data (e.g., "Tell me about your tactics")

[0124] Output: Analysis results (e.g. "Napoleon" "Tactics")

[0125] Step 3:

[0126] The server retrieves the information from the database

[0127] Based on the analysis results, the server searches for information about related historical figures from the information aggregation means. The information aggregation means stores a large amount of data and efficiently retrieves the required information.

[0128] Input: Analysis result (e.g. "Napoleon" "Tactics")

[0129] Output: relevant information (e.g. "Information about Napoleon's tactics")

[0130] Step 4:

[0131] The server inputs information into the generated AI model

[0132] The server inputs the relevant information retrieved from the database and the user's question into the generative AI model means, which then generates the optimal answer based on this information.

[0133] Input: Relevant information and user questions (e.g., "Information about Napoleon's tactics" and "Tell me about your tactics")

[0134] Output: Generated answer (e.g., "My most important tactics are rapid advance and attacking the enemy's weak points.")

[0135] Step 5:

[0136] The server generates a response and sends it to the device.

[0137] The generated answer is returned to the server, which then transmits it to the terminal.

[0138] Input: Generated answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0139] Output: The submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0140] Step 6:

[0141] The device will lip sync and display the answer

[0142] The device displays the answer received from the server to the user. If the device includes a lip-sync function, it performs lip-sync processing on the image of the person and reproduces natural mouth movements using voice synthesis.

[0143] Input: Submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0144] Output: Displayed and voice-synthesized answer (e.g., "My most important tactical priorities are rapid advance and striking the enemy's weak points.")

[0145] Step 7:

[0146] User observations and follow-up questions

[0147] The user observes the displayed answers, and if they have further questions, they again use the terminal interface to enter and submit new questions, and the process repeats.

[0148] Input: Follow-up question (e.g., "Can you give us an example of a specific tactic?")

[0149] Output: New question data (e.g., "Can you give us an example of a specific tactic?")

[0150] (Application example 1)

[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0152] In today's factories, managers and workers are required to have a high level of specialized knowledge and experience in order to carry out efficient work management and training. However, this knowledge depends on the experience of each individual, making it difficult to achieve consistent and efficient work management. Another problem is that training and nurturing new workers takes a lot of time and money.

[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0154] In this invention, the server includes a database for managing data on historical figures, an interface for receiving user input, a communication unit for analyzing user requests and retrieving information about the specified historical figures from the database, a generation AI model for providing generated answers, and an output unit for transmitting the generated answers to the user. This allows factory managers and workers to apply historical knowledge and strategies to improve work efficiency and quickly train new workers. Furthermore, the voice synthesis and lip-sync functions allow users to enjoy a more realistic experience.

[0155] The "database means" is a system or device that centrally manages data related to historical figures and quickly provides relevant data in response to a user's request.

[0156] "Interface means" refers to an input device or interface that allows a user to input questions or requests to the system, and has the function of accepting user input.

[0157] "Communication means" refers to a system or device with communication capabilities that analyzes requests from users, retrieves the necessary data from a database, and passes it to the generative AI model.

[0158] "Generative AI model means" refers to an artificial intelligence model that generates an appropriate answer based on information obtained from a database in response to a question from a user.

[0159] "Output means" refers to a system or device for transmitting and displaying the answer generated by the generative AI model to the user.

[0160] The "lip sync means" is a system or device that has the function of synthesizing the generated answer into voice and then performing lip sync processing on a photograph of a historical figure's face.

[0161] "Factory training and support measures" refers to the functions and equipment that support the effective management of machines and systems used within a factory and the training of workers.

[0162] This invention is a system that applies the tactics and knowledge of historical figures to factory robots, enabling factory management and work support. The system includes database means, interface means, communication means, generative AI model means, output means, lip-sync means, and factory training and support means.

[0163] System Configuration

[0164] server

[0165] The server is the core of this system and includes a database, communication, AI model generation, and output functions. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[0166] Terminal

[0167] The terminal provides an interface for the user to interact with the system, accepts user input, transmits it to the server, and displays responses received from the server. It also includes functionality for displaying a lip-synced person's face.

[0168] User

[0169] The user operates the interface means to input a question and receive a response from the system, thereby obtaining advice and knowledge from historical figures.

[0170] Hardware

[0171] Robot: An industrial machine used in a factory (e.g., an industrial robot).

[0172] Server: Runs on a cloud service (e.g. AWS EC2).

[0173] Camera and microphone: Sensor devices built into the robot.

[0174] Display: A display device that allows the user to see the content of the interaction.

[0175] software

[0176] Natural Language Processing (NLP) library: OpenAI's GPT-4.

[0177] Speech synthesis library: Google Text-to-Speech.

[0178] Lip sync library: Avatarify.

[0179] Database: MySQL.

[0180] Program Processing Overview

[0181] Processing on the terminal

[0182] The device provides an interface for users to input questions for historical figures. When a user inputs and submits a question, the data is sent to the server. Furthermore, if the device includes a lip-sync function, it has the function of displaying a lip-synced photograph of the historical figure.

[0183] Examples:

[0184] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[0185] Processing on the server

[0186] The server receives the request from the device and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from a database. The retrieved information and the user's question are input into a generative AI model to generate an appropriate answer. The generated answer is synthesized into voice, lip-synced, and sent to the device.

[0187] Examples:

[0188] The server retrieves information about a specified historical figure and tactics from a database, inputs it into a generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points."

[0189] Lip Sync Means

[0190] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them to the facial photographs of historical figures, providing a more realistic conversation experience.

[0191] Examples:

[0192] A lip sync means generates and displays in real time mouth movements based on the generated answers for the photographs of the faces of historical figures.

[0193] Prompt Sentence Examples

[0194] "As the designated historical figure, answer this question: Tell us about your tactics."

[0195] "As a designated historical figure, please assess the progress of the production line this week."

[0196] This system allows users to receive factory management and work support that applies historical knowledge, enabling efficient work management and rapid training.

[0197] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0198] Step 1:

[0199] The user enters a question into the terminal interface and clicks the send button. The entered question is saved on the terminal and then sent to the server.

[0200] Specific operation: The user enters "Tell us about your tactics" into the input field on the device and presses the send button. The input data is sent from the device to the server.

[0201] Step 2:

[0202] The server receives and analyzes a request (question) from the user. Based on the analysis results, it retrieves information about the specified historical figure from the database.

[0203] Input: User question data

[0204] Output: Database query results

[0205] Specific operation: The server searches the database for information about the "specified historical figure" and "tactic" and retrieves related information.

[0206] Step 3:

[0207] The server inputs the acquired information and the user's question into a generative AI model to generate an appropriate answer.

[0208] Input: Information about historical figures retrieved from the database, user questions

[0209] Output: The generated answer

[0210] How it works: The server inputs a prompt to a generative AI model (e.g., OpenAI GPT-4) saying, "As a specified historical figure, tell me about a tactical battle," and receives the generated answer.

[0211] Step 4:

[0212] The generated answer is synthesized into voice.

[0213] Input: Generated answer text

[0214] Output: Audio file

[0215] Specific behavior: The server uses a speech synthesis library (e.g., Google Text-to-Speech) to convert the generated answer text into an audio file.

[0216] Step 5:

[0217] Lip-sync processing is performed using the generated answer audio file.

[0218] Input: Photo of a historical figure, audio file

[0219] Output: Lip-synced video

[0220] What happens: The server uses a lip-sync library (e.g., Avatarify) to generate audio-synced mouth movements for a photograph of a historical figure, creating a lip-synced video.

[0221] Step 6:

[0222] The server sends the lip-synced video to the terminal.

[0223] Input: Lip-synced video

[0224] Output: The video displayed on the user's device

[0225] Specific operation: The server sends lip-synced video data to the device, which then displays the video to the user.

[0226] Step 7:

[0227] Users watch lip-synced videos on their devices and receive answers from historical figures.

[0228] Input: Lip-synced video

[0229] Output: User understanding and real-time interaction experience

[0230] What happens: The user watches a lip-synced video of a historical figure on the device display and confirms their answers.

[0231] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0232] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[0233] System Components

[0234] server

[0235] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[0236] Terminal

[0237] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[0238] User

[0239] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[0240] Program Processing Overview

[0241] Terminal

[0242] The device provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If a lip-sync function is included, the device displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[0243] Examples:

[0244] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[0245] server

[0246] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[0247] Examples:

[0248] The server retrieves information about "Napoleon" and "tactics" from the database, inputs it into the generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points." The emotion engine analyzes the user's interested facial expression and adds a slightly more detailed explanation to the answer.

[0249] Lip Sync Means

[0250] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[0251] Examples:

[0252] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[0253] Simulation Method

[0254] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[0255] Examples:

[0256] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[0257] Emotion Engine

[0258] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[0259] Examples:

[0260] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[0261] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[0262] The processing flow will be explained below.

[0263] Specific processing flow of the program (including the emotion engine)

[0264] Step 1:

[0265] The user enters a question into the terminal interface and clicks the send button.

[0266] Step 2:

[0267] The device formats the user's question as an API request and sends it to the server, and also captures the user's facial expressions and voice and sends them to the server as emotion data.

[0268] Step 3:

[0269] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[0270] Step 4:

[0271] The server searches and acquires information about the designated historical person from the database means.

[0272] Step 5:

[0273] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[0274] Step 6:

[0275] The emotion engine analyzes the user's emotion data transmitted from the terminal and determines the user's emotional state.

[0276] Examples:

[0277] The emotion engine analyzes the user's facial expressions and determines whether the user is interested.

[0278] Step 7:

[0279] The sentiment engine adjusts the generated answer based on the determined user sentiment, for example adding detailed explanations to the answer if the user indicates interest.

[0280] Examples:

[0281] "My most important tactic is to advance quickly and attack the enemy's weak points," one generated answer might explain, with the further detail, "This tactic has won me many battles."

[0282] Step 8:

[0283] The server formats the generated answer and creates response data.

[0284] Step 9:

[0285] The server sends the response data to the terminal.

[0286] Step 10:

[0287] The terminal receives the response data from the server and formats it for display to the user.

[0288] Step 11:

[0289] The lip-syncing means synthesizes the generated responses into voice and lip-syncs them to the portraits of historical figures, providing a more realistic dialogue experience.

[0290] Step 12:

[0291] The terminal displays the answer to the user and plays the video with lip-sync processing completed.

[0292] Step 13:

[0293] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[0294] This process allows users to simulate real-time interactions with historical figures and receive personalized, emotion-based responses to enhance their learning experience.

[0295] Example 2

[0296] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0297] Conventional dialogue systems only provide simple text-based answers when users seek information about historical figures, resulting in a lack of realism and immersion. Furthermore, they are unable to provide personalized answers based on the user's emotions, preventing an improved learning experience. Furthermore, they are unable to simulate conversations between different historical figures, limiting opportunities to learn history from new perspectives.

[0298] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0299] In this invention, the server includes a database means for managing data on historical figures, a communication means for analyzing a request from a user and obtaining information on the specified historical figure from the database means, a generative AI model means for providing a generated answer, an emotion engine for recognizing emotions by analyzing the user's facial expressions and voice, and a means for adjusting the tone and content of the generated answer based on the recognized emotion. This allows the user to interact with the historical figure in real time and receive personalized answers based on their emotions, enabling a deeper learning experience.

[0300] The "database means" is a storage or management system for managing data relating to historical figures and retrieving information as needed.

[0301] An "interface means" is an input device or software interface for accepting input from a user.

[0302] "Communication means" refers to a network device or communication protocol for transmitting user input data to a server and transmitting data from the server to the user.

[0303] The "generative AI model means" is an artificial intelligence model or natural language processing system that analyzes a user's question and generates an appropriate answer.

[0304] The "output means" is a display device or an audio output device for displaying or reproducing the generated answer to the user.

[0305] An "emotion engine" is an analysis device or software engine that analyzes a user's facial expressions and voice data to recognize the user's emotions.

[0306] "Lip sync means" refers to a processing device or software that synthesizes the generated response into voice and synchronizes the lip movements with the image of the historical figure.

[0307] A "simulation means" is a system for arranging a dialogue between different historical figures and reproducing or simulating the development of that dialogue.

[0308] A "user" is someone who operates the system and enters questions to obtain information about historical figures.

[0309] The "server" is the core of the system, and is a computer system that controls the database means, generative AI model means, emotion engine, etc., and manages and executes the overall processing.

[0310] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[0311] server

[0312] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[0313] Terminal

[0314] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[0315] User

[0316] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[0317] Program Processing Overview

[0318] server

[0319] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[0320] Terminal

[0321] The device provides an interface for users to input questions for historical figures. When the user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If the device includes a lip-sync function, it displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[0322] Examples:

[0323] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[0324] Lip Sync Means

[0325] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[0326] Examples:

[0327] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[0328] Simulation Method

[0329] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[0330] Examples:

[0331] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[0332] Emotion Engine

[0333] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[0334] Examples:

[0335] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[0336] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[0337] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0338] Step 1:

[0339] The user types a question on the device

[0340] Using the device's interface, the user enters a text question for the historical figure, such as "Napoleon, tell me about tactics." The user then clicks a submit button to complete the question.

[0341] Input: A question entered by the user as text

[0342] Output: The device sends the input data to the server

[0343] Step 2:

[0344] The device sends user input and emotion data to the server.

[0345] The device receives text data entered by the user and captures the user's facial expressions and voice using the built-in camera and microphone. The facial expression and voice data are analyzed and sent to the server as emotional data.

[0346] Input: User question, and captured facial expressions and voice

[0347] Output: Send all data to the server along with the analyzed emotion data

[0348] Step 3:

[0349] The server parses the request and retrieves information from the database

[0350] The server receives the request sent from the terminal and analyzes its contents. For example, it extracts keywords such as "Napoleon" and "tactics." Then, based on these keywords, it retrieves related information from the database means.

[0351] Input: Text data and emotion data sent from the device

[0352] Output: Historical information retrieved from the database

[0353] Step 4:

[0354] The server generates an answer using a generative AI model

[0355] The server inputs the information retrieved from the database and the user's question into a generative AI model (e.g., a natural language processing model). The generative AI model is used to generate an appropriate answer based on the context.

[0356] Input: Retrieved historical information and user question

[0357] Output: The answer generated by the generative AI model

[0358] Step 5:

[0359] The server uses an emotion engine to tailor the response

[0360] The server uses an emotion engine to tailor the generated answer based on the analyzed user's emotion data, for example adding more details and explanations to the answer if the user seems interested.

[0361] Input: Generated answers and sentiment data

[0362] Output: An answer tailored based on the sentiment

[0363] Step 6:

[0364] The server performs lip sync processing

[0365] The server uses lip-syncing to synthesize the generated responses and synchronize the mouth movements with the image of the historical figure, providing a realistic dialogue experience, as if Napoleon were actually speaking.

[0366] Input: Adjusted answer

[0367] Output: Lip-synced video and audio

[0368] Step 7:

[0369] The server sends the answer to the device

[0370] The server sends a response to the terminal indicating that the lip-sync process has been completed. The data sent includes text information, audio files, and lip-synced video.

[0371] Input: Lip-synced video and audio data

[0372] Output: Answer data sent to the device

[0373] Step 8:

[0374] The device displays the answer to the user

[0375] The device receives the answer data sent from the server and displays it to the user. The lip-synced faces of historical figures are displayed on the screen, and the answers are played back aloud.

[0376] Input: Response data sent from the server

[0377] Output: Answers shown to the user (text, audio, lip-sync video)

[0378] (Application example 2)

[0379] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0380] Conventional educational systems and content distribution services tend to provide users with a passive experience when learning history. Furthermore, one-way information provision without considering the user's emotions makes it difficult to maintain interest in learning. Furthermore, interactive systems lack the technology to provide a realistic dialogue experience, and the depictions of important historical figures are static. Therefore, there is a need for the development of a system that allows users to experience interactive dialogue with historical figures who actually have emotions while learning history.

[0381] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes database means for managing data on historical figures, interface means for accepting input from a user, means for capturing the user's facial expressions and voice and analyzing their emotions, communication means for analyzing the user's request and obtaining information on the specified historical figure from the database means, generation AI model means for adjusting the generated answer based on the user's emotions, output means for sending the generated answer to the user, and rendering means for displaying the lip-synced answer as a 3D image. This allows the user to interact with the historical figure in real time and enjoy an interactive learning experience personalized based on their emotions.

[0382] "Database means" is a device that has the function of systematically managing and storing information about historical figures.

[0383] An "interface means" is a device that accepts input from a user and provides a screen and operating means for the user to interact with the system.

[0384] The "means for analyzing emotions" is a device that has the function of capturing the user's facial expressions and voice and analyzing them to identify the user's emotional state.

[0385] The "communication means" is a device that receives a request from a user, retrieves information from a database based on the request, and passes the result to the generation AI model means.

[0386] The "generative AI model means" is a device that has the function of generating an appropriate response based on a user request and adjusting the tone and content of the response based on the user's emotional data.

[0387] The "output means" is a device for transmitting the generated answer to the user, and has the function of displaying or audibly transmitting the answer to the user in real time.

[0388] The "rendering means" is a device that renders the lip-synced answer as a three-dimensional model in real time and provides it visually to the user.

[0389] The "lip sync means" is a device that synchronizes the mouth movements of a three-dimensional model with the generated voice, thereby providing a realistic dialogue experience.

[0390] The "simulation means" is a device that has the function of having different historical figures converse with each other and simulating the development of that conversation in real time.

[0391] This invention provides a system for providing users with information about historical figures and realizing interactive dialogue. The system includes a database, an interface, an emotion analysis, a communication, a generative AI model, an output, a lip-sync, a rendering, and a simulation.

[0392] System components and functions

[0393] 1. Database Means

[0394] The server is equipped with a database that systematically manages and stores data on historical figures. This database is constructed using a relational database management system (RDBMS) such as MySQL.

[0395] 2. Interface Method

[0396] Users interact with the system through devices such as smartphones and head-mounted displays (HMDs). The devices are equipped with a user interface (UI) built using Unity, which allows users to enter input.

[0397] 3. Emotion analysis method

[0398] The device has the ability to capture the user's facial expressions and voice and analyze their emotions using APIs from Google ML Kit and Amazon Rekognition, making it possible to analyze the user's emotional state in real time.

[0399] 4. Means of communication

[0400] A request from a user is sent to the server using a RESTful API. The communication means has the function to analyze this request and retrieve the necessary information from the database.

[0401] 5. Generative AI Model Means

[0402] The generative AI model means installed on the server generates appropriate answers to questions using a natural language processing model such as OpenAI, and has the function of adjusting the tone and content of the answers based on user sentiment analysis data.

[0403] 6. Output Method

[0404] The generated answers are sent from the server to the device and provided to the user, using voice synthesis technology such as Amazon Polly.

[0405] 7. Lip Sync Method

[0406] The server performs lip-sync processing based on the generated audio, which uses the lip-sync function of MetaHuman Creator to generate mouth movements for the 3D model.

[0407] 8. Rendering Method

[0408] Using Unity's 3D rendering engine, three-dimensional models that have undergone lip-sync processing are rendered in real time and provided to users.

[0409] 9. Simulation Methods

[0410] The server has the ability to have different historical figures interact with each other and simulate the progress of the interaction in real time, allowing users to learn about history from a new perspective.

[0411] Specific examples of processing

[0412] For example, suppose a user types "Tell me about your days as Napoleon" into the device's interface and presses the send button. At this time, the device's camera captures and analyzes the user's facial expression. The emotion analysis means recognizes that the user is interested. This request is then sent to the server via the communication means, and data about Napoleon is retrieved from the database.

[0413] Next, the generating AI model means generates an appropriate answer to "Life as Napoleon" and adjusts the answer taking into account the user's emotions. The server then synthesizes the generated answer into speech and generates mouth movements for the three-dimensional model of Napoleon using the lip-sync means. Finally, the rendering means renders the three-dimensional model of Napoleon in real time and displays it on the device.

[0414] Prompt Sentence Examples

[0415] "Tell me about your days as Napoleon."

[0416] "Alexander the Great, what is your most amazing tactic?"

[0417] Using this system, users can interact with historical figures in real time and enjoy a personalized, interactive learning experience that responds to their emotions.

[0418] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0419] Step 1: The user enters a question for a historical figure through the device interface and presses the send button. At this time, the interface captures the text of the user's question and uses the camera and microphone to collect the user's facial expressions and voice in real time. The input is the user's question text, facial expressions, and voice data, and the output is a process that compiles this data into a single request.

[0420] Step 2: The device uses Google ML Kit and Amazon Rekognition API on the edge device to analyze the collected facial and voice data in real time to identify the user's emotional state. The input is the user's facial and voice data, and the output is the analyzed emotional data.

[0421] Step 3: The device sends a request containing the user's question text and the parsed emotion data to the server via a RESTful API. The input is the request containing the question text and emotion data, and the output is the request data sent to the server.

[0422] Step 4: The server receives the request and retrieves information about the specified historical figure using a database (e.g., MySQL). The input is the request data from the user (question text and emotion data), and the output is the historical information retrieved from the database.

[0423] Step 5: The server inputs the acquired historical information into a generative AI model (such as OpenAI) to generate an answer to the user's question. The generated answer is also adjusted based on the emotional data. The input is the historical information acquired from the database and the user's emotional data, and the output is an emotion-adjusted answer text.

[0424] Step 6: The adjusted answer text is converted into speech data using a speech synthesis means (e.g., Amazon Polly). During this process, the generated speech data is sent to a lip-sync means (e.g., MetaHuman Creator) to generate mouth movements corresponding to a three-dimensional model of the historical figure. The input is the emotion-adjusted answer text, and the output is speech data and lip-sync data.

[0425] Step 7: The server renders the audio data and lip-sync data in real time using a rendering engine (such as Unity's 3D rendering engine) and sends it to the device. The input is the audio data and lip-sync data, and the output is the rendered 3D model.

[0426] Step 8: The terminal displays the rendered 3D model to the user and plays the audio data. This allows the user to interact with the historical figure in real time and enjoy personalized responses based on their emotions. The input is the rendered 3D model and audio data from the server, and the output is audiovisual information provided to the user.

[0427] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0428] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0429] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0430] [Second embodiment]

[0431] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0432] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0433] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0434] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0435] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0436] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0437] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0438] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0439] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0440] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0441] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0442] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0443] The present invention relates to a system for managing data on historical figures and allowing users to have conversations as if they were the historical figures. The system includes a database, an interface, a communication means, a generative AI model, an output means, a lip-sync means, and a simulation means.

[0444] System Components

[0445] server

[0446] The server is the core of the system and includes a database, communication, AI model generation, and output. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[0447] Terminal

[0448] The terminal provides an interface for the user to interact with the system: it accepts user input, transmits it to the server, and displays responses received from the server to the user.

[0449] User

[0450] Through this system, users interact with historical figures by operating a terminal, inputting questions, and reading or listening to the answers provided by the server.

[0451] Program Processing Overview

[0452] Terminal

[0453] The terminal provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying answers to the user. If a lip-sync function is included, a photo of the person's face is lip-synced and displayed.

[0454] Examples:

[0455] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[0456] server

[0457] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. The generated answer is sent to the terminal through the output means.

[0458] Examples:

[0459] The server retrieves information about "Napoleon" and "tactics" from a database, inputs it into a generative AI model, and generates the answer, "The tactics I value most are rapid advance and attacking the enemy's weak points."

[0460] Lip Sync Means

[0461] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[0462] Examples:

[0463] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[0464] Simulation Method

[0465] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[0466] Examples:

[0467] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[0468] The system allows users to interact with historical figures in real time, providing a deeper historical understanding and learning experience.

[0469] The processing flow will be explained below.

[0470] Specific processing flow of the program

[0471] Step 1:

[0472] The user enters a question into the terminal interface and clicks the send button.

[0473] Step 2:

[0474] The device formats the user's question as an API request and sends it to the server.

[0475] Step 3:

[0476] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[0477] Step 4:

[0478] The server searches and acquires information about the designated historical person from the database means.

[0479] Step 5:

[0480] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[0481] Step 6:

[0482] The server formats the generated answer and creates response data.

[0483] Step 7:

[0484] The server sends the response data to the terminal.

[0485] Step 8:

[0486] The terminal receives the response data from the server and formats it for display to the user.

[0487] Step 9:

[0488] The terminal displays the answer to the user, and if there is lip synchronization means, generates and displays the mouth movements.

[0489] Step 10:

[0490] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[0491] This process allows users to simulate real-time interactions with historical figures, enhancing the learning experience.

[0492] Example 1

[0493] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0494] In the modern education system, history classes tend to provide only one-sided information, making it difficult for students to acquire interest or concern. In particular, there is a lack of concrete ways to deepen knowledge about historical figures. There is a growing need for interactive educational tools that will engage students and trainees in learning.

[0495] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0496] In this invention, the server includes an information aggregation means for managing data on historical figures, an input means for accepting input from a user, a receiving means for analyzing a request from the user and acquiring information on the specified historical figure from the information aggregation means, a generation AI model means for providing a generated answer, and a transmitting means for transmitting the generated answer to the user. This allows the user to have an interactive conversation with the historical figure, thereby improving the learning effect.

[0497] "Information aggregation means" refers to the means for managing data on historical figures and collecting and organizing necessary information.

[0498] The "input means" is an interface that accepts input from the user, and is a means by which the user inputs questions or requests.

[0499] The "receiving means" is a means for analyzing a request from a user and obtaining information about a specified historical person from the information aggregating means.

[0500] A "generative AI model means" is a means that uses an artificial intelligence model to generate an answer based on acquired information.

[0501] The "transmission means" is a means for transmitting the generated answer to the user.

[0502] The "lip sync means" is a means for synthesizing lip movements in real time with an image of a historical figure based on the generated answer.

[0503] A "simulation method" is a method for arranging a dialogue between different historical figures and reproducing the development of that dialogue.

[0504] This invention is a system for managing data on historical figures and allowing users to have conversations as if they were those figures. The system includes an information aggregation unit, an input unit, a receiving unit, a generating AI model unit, a transmitting unit, a lip-sync unit, and a simulation unit.

[0505] server

[0506] The server is the central part of the system and is the hardware that controls various functions. The server has the following means for processing data:

[0507] Information aggregation tool: A tool for managing data about historical figures. This data is stored in a database and includes detailed historical information about the person and records of their statements.

[0508] Reception means: A means for receiving requests sent by users. The received requests are analyzed and necessary keywords and contexts are extracted.

[0509] Generative AI model means: A means for obtaining necessary data from the information aggregation means based on the analysis results obtained from the receiving means, and then using an AI model to generate an appropriate response. The generative AI model is, for example, a natural language generation model such as GPT (Generative Pre-trained Transformer).

[0510] Transmission means: A means for transmitting the generated answer to the terminal, allowing the user to receive the generated answer.

[0511] Terminal

[0512] A terminal is a piece of hardware that provides an interface for users to interact with a system. Specifically, it operates as follows:

[0513] Input means: A means that provides an interface for users to input questions. Users input their questions using this means and send them to the server by pressing the send button.

[0514] Lip-syncing: A method for synthesizing lip movements in real time onto images of historical figures based on answers sent from the server. This process provides a more natural and interactive dialogue experience.

[0515] Examples:

[0516] The user types "Tell me about your tactics" into the device interface and presses the send button. The server receives this question, analyzes it, and extracts the keywords "Napoleon" and "tactics." The server obtains related data from the information aggregation means and inputs it into the generative AI model. The generative AI model generates an answer, "The tactics I value most are rapid advancement and attacking the enemy's weak points," and sends this to the device. The device receives this answer and displays it to the user using lip-syncing means, replicating the mouth movements of an image of Napoleon.

[0517] Prompt Sentence Examples

[0518] "Tell me about your tactics."

[0519] "Generate detailed information about Napoleon's tactics."

[0520] conclusion

[0521] This system allows users to interact with historical figures, enhancing learning effectiveness, and its lip-sync function provides a realistic experience both visually and aurally.

[0522] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0523] Processing Steps

[0524] Step 1:

[0525] User enters question and submits

[0526] The user uses the device interface to input a question for the historical figure. When input is complete, the user presses the "Send" button. This operation sends the question data from the device to the server.

[0527] Input: "Tell me about your tactics."

[0528] Output: Question data (e.g., "Tell me about your tactics")

[0529] Step 2:

[0530] The server receives and analyzes the query

[0531] The server receives the question data sent from the device, analyzes the received question data, and extracts keywords and context, thereby identifying the type and subject of the requested information.

[0532] Input: Question data (e.g., "Tell me about your tactics")

[0533] Output: Analysis results (e.g. "Napoleon" "Tactics")

[0534] Step 3:

[0535] The server retrieves the information from the database

[0536] Based on the analysis results, the server searches for information about related historical figures from the information aggregation means. The information aggregation means stores a large amount of data and efficiently retrieves the required information.

[0537] Input: Analysis result (e.g. "Napoleon" "Tactics")

[0538] Output: relevant information (e.g. "Information about Napoleon's tactics")

[0539] Step 4:

[0540] The server inputs information into the generated AI model

[0541] The server inputs the relevant information retrieved from the database and the user's question into the generative AI model means, which then generates the optimal answer based on this information.

[0542] Input: Relevant information and user questions (e.g., "Information about Napoleon's tactics" and "Tell me about your tactics")

[0543] Output: Generated answer (e.g., "My most important tactics are rapid advance and attacking the enemy's weak points.")

[0544] Step 5:

[0545] The server generates a response and sends it to the device.

[0546] The generated answer is returned to the server, which then transmits it to the terminal.

[0547] Input: Generated answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0548] Output: The submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0549] Step 6:

[0550] The device will lip sync and display the answer

[0551] The device displays the answer received from the server to the user. If the device includes a lip-sync function, it performs lip-sync processing on the image of the person and reproduces natural mouth movements using voice synthesis.

[0552] Input: Submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0553] Output: Displayed and voice-synthesized answer (e.g., "My most important tactical priorities are rapid advance and striking the enemy's weak points.")

[0554] Step 7:

[0555] User observations and follow-up questions

[0556] The user observes the displayed answers, and if they have further questions, they again use the terminal interface to enter and submit new questions, and the process repeats.

[0557] Input: Follow-up question (e.g., "Can you give us an example of a specific tactic?")

[0558] Output: New question data (e.g., "Can you give us an example of a specific tactic?")

[0559] (Application example 1)

[0560] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0561] In today's factories, managers and workers are required to have a high level of specialized knowledge and experience in order to carry out efficient work management and training. However, this knowledge depends on the experience of each individual, making it difficult to achieve consistent and efficient work management. Another problem is that training and nurturing new workers takes a lot of time and money.

[0562] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0563] In this invention, the server includes a database for managing data on historical figures, an interface for receiving user input, a communication unit for analyzing user requests and retrieving information about the specified historical figures from the database, a generation AI model for providing generated answers, and an output unit for transmitting the generated answers to the user. This allows factory managers and workers to apply historical knowledge and strategies to improve work efficiency and quickly train new workers. Furthermore, the voice synthesis and lip-sync functions allow users to enjoy a more realistic experience.

[0564] The "database means" is a system or device that centrally manages data related to historical figures and quickly provides relevant data in response to a user's request.

[0565] "Interface means" refers to an input device or interface that allows a user to input questions or requests to the system, and has the function of accepting user input.

[0566] "Communication means" refers to a system or device with communication capabilities that analyzes requests from users, retrieves the necessary data from a database, and passes it to the generative AI model.

[0567] "Generative AI model means" refers to an artificial intelligence model that generates an appropriate answer based on information obtained from a database in response to a question from a user.

[0568] "Output means" refers to a system or device for transmitting and displaying the answer generated by the generative AI model to the user.

[0569] The "lip sync means" is a system or device that has the function of synthesizing the generated answer into voice and then performing lip sync processing on a photograph of a historical figure's face.

[0570] "Factory training and support measures" refers to the functions and equipment that support the effective management of machines and systems used within a factory and the training of workers.

[0571] This invention is a system that applies the tactics and knowledge of historical figures to factory robots, enabling factory management and work support. The system includes database means, interface means, communication means, generative AI model means, output means, lip-sync means, and factory training and support means.

[0572] System Configuration

[0573] server

[0574] The server is the core of this system and includes a database, communication, AI model generation, and output functions. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[0575] Terminal

[0576] The terminal provides an interface for the user to interact with the system, accepts user input, transmits it to the server, and displays responses received from the server. It also includes functionality for displaying a lip-synced person's face.

[0577] User

[0578] The user operates the interface means to input a question and receive a response from the system, thereby obtaining advice and knowledge from historical figures.

[0579] Hardware

[0580] Robot: An industrial machine used in a factory (e.g., an industrial robot).

[0581] Server: Runs on a cloud service (e.g. AWS EC2).

[0582] Camera and microphone: Sensor devices built into the robot.

[0583] Display: A display device that allows the user to see the content of the interaction.

[0584] software

[0585] Natural Language Processing (NLP) library: OpenAI's GPT-4.

[0586] Speech synthesis library: Google Text-to-Speech.

[0587] Lip sync library: Avatarify.

[0588] Database: MySQL.

[0589] Program Processing Overview

[0590] Processing on the terminal

[0591] The device provides an interface for users to input questions for historical figures. When a user inputs and submits a question, the data is sent to the server. Furthermore, if the device includes a lip-sync function, it has the function of displaying a lip-synced photograph of the historical figure.

[0592] Examples:

[0593] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[0594] Processing on the server

[0595] The server receives the request from the device and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from a database. The retrieved information and the user's question are input into a generative AI model to generate an appropriate answer. The generated answer is synthesized into voice, lip-synced, and sent to the device.

[0596] Examples:

[0597] The server retrieves information about a specified historical figure and tactics from a database, inputs it into a generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points."

[0598] Lip Sync Means

[0599] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them to the facial photographs of historical figures, providing a more realistic conversation experience.

[0600] Examples:

[0601] A lip sync means generates and displays in real time mouth movements based on the generated answers for the photographs of the faces of historical figures.

[0602] Prompt Sentence Examples

[0603] "As the designated historical figure, answer this question: Tell us about your tactics."

[0604] "As a designated historical figure, please assess the progress of the production line this week."

[0605] This system allows users to receive factory management and work support that applies historical knowledge, enabling efficient work management and rapid training.

[0606] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0607] Step 1:

[0608] The user enters a question into the terminal interface and clicks the send button. The entered question is saved on the terminal and then sent to the server.

[0609] Specific operation: The user enters "Tell us about your tactics" into the input field on the device and presses the send button. The input data is sent from the device to the server.

[0610] Step 2:

[0611] The server receives and analyzes a request (question) from the user. Based on the analysis results, it retrieves information about the specified historical figure from the database.

[0612] Input: User question data

[0613] Output: Database query results

[0614] Specific operation: The server searches the database for information about the "specified historical figure" and "tactic" and retrieves related information.

[0615] Step 3:

[0616] The server inputs the acquired information and the user's question into a generative AI model to generate an appropriate answer.

[0617] Input: Information about historical figures retrieved from the database, user questions

[0618] Output: The generated answer

[0619] How it works: The server inputs a prompt to a generative AI model (e.g., OpenAI GPT-4) saying, "As a specified historical figure, tell me about a tactical battle," and receives the generated answer.

[0620] Step 4:

[0621] The generated answer is synthesized into voice.

[0622] Input: Generated answer text

[0623] Output: Audio file

[0624] Specific behavior: The server uses a speech synthesis library (e.g., Google Text-to-Speech) to convert the generated answer text into an audio file.

[0625] Step 5:

[0626] Lip-sync processing is performed using the generated answer audio file.

[0627] Input: Photo of a historical figure, audio file

[0628] Output: Lip-synced video

[0629] What happens: The server uses a lip-sync library (e.g., Avatarify) to generate audio-synced mouth movements for a photograph of a historical figure, creating a lip-synced video.

[0630] Step 6:

[0631] The server sends the lip-synced video to the terminal.

[0632] Input: Lip-synced video

[0633] Output: The video displayed on the user's device

[0634] Specific operation: The server sends lip-synced video data to the device, which then displays the video to the user.

[0635] Step 7:

[0636] Users watch lip-synced videos on their devices and receive answers from historical figures.

[0637] Input: Lip-synced video

[0638] Output: User understanding and real-time interaction experience

[0639] What happens: The user watches a lip-synced video of a historical figure on the device display and confirms their answers.

[0640] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0641] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[0642] System Components

[0643] server

[0644] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[0645] Terminal

[0646] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[0647] User

[0648] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[0649] Program Processing Overview

[0650] Terminal

[0651] The device provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If a lip-sync function is included, the device displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[0652] Examples:

[0653] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[0654] server

[0655] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[0656] Examples:

[0657] The server retrieves information about "Napoleon" and "tactics" from the database, inputs it into the generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points." The emotion engine analyzes the user's interested facial expression and adds a slightly more detailed explanation to the answer.

[0658] Lip Sync Means

[0659] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[0660] Examples:

[0661] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[0662] Simulation Method

[0663] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[0664] Examples:

[0665] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[0666] Emotion Engine

[0667] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[0668] Examples:

[0669] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[0670] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[0671] The processing flow will be explained below.

[0672] Specific processing flow of the program (including the emotion engine)

[0673] Step 1:

[0674] The user enters a question into the terminal interface and clicks the send button.

[0675] Step 2:

[0676] The device formats the user's question as an API request and sends it to the server, and also captures the user's facial expressions and voice and sends them to the server as emotion data.

[0677] Step 3:

[0678] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[0679] Step 4:

[0680] The server searches and acquires information about the designated historical person from the database means.

[0681] Step 5:

[0682] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[0683] Step 6:

[0684] The emotion engine analyzes the user's emotion data transmitted from the terminal and determines the user's emotional state.

[0685] Examples:

[0686] The emotion engine analyzes the user's facial expressions and determines whether the user is interested.

[0687] Step 7:

[0688] The sentiment engine adjusts the generated answer based on the determined user sentiment, for example adding detailed explanations to the answer if the user indicates interest.

[0689] Examples:

[0690] "My most important tactic is to advance quickly and attack the enemy's weak points," one generated answer might explain, with the further detail, "This tactic has won me many battles."

[0691] Step 8:

[0692] The server formats the generated answer and creates response data.

[0693] Step 9:

[0694] The server sends the response data to the terminal.

[0695] Step 10:

[0696] The terminal receives the response data from the server and formats it for display to the user.

[0697] Step 11:

[0698] The lip-syncing means synthesizes the generated responses into voice and lip-syncs them to the portraits of historical figures, providing a more realistic dialogue experience.

[0699] Step 12:

[0700] The terminal displays the answer to the user and plays the video with lip-sync processing completed.

[0701] Step 13:

[0702] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[0703] This process allows users to simulate real-time interactions with historical figures and receive personalized, emotion-based responses to enhance their learning experience.

[0704] Example 2

[0705] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0706] Conventional dialogue systems only provide simple text-based answers when users seek information about historical figures, resulting in a lack of realism and immersion. Furthermore, they are unable to provide personalized answers based on the user's emotions, preventing an improved learning experience. Furthermore, they are unable to simulate conversations between different historical figures, limiting opportunities to learn history from new perspectives.

[0707] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0708] In this invention, the server includes a database means for managing data on historical figures, a communication means for analyzing a request from a user and obtaining information on the specified historical figure from the database means, a generative AI model means for providing a generated answer, an emotion engine for recognizing emotions by analyzing the user's facial expressions and voice, and a means for adjusting the tone and content of the generated answer based on the recognized emotion. This allows the user to interact with the historical figure in real time and receive personalized answers based on their emotions, enabling a deeper learning experience.

[0709] The "database means" is a storage or management system for managing data relating to historical figures and retrieving information as needed.

[0710] An "interface means" is an input device or software interface for accepting input from a user.

[0711] "Communication means" refers to a network device or communication protocol for transmitting user input data to a server and transmitting data from the server to the user.

[0712] The "generative AI model means" is an artificial intelligence model or natural language processing system that analyzes a user's question and generates an appropriate answer.

[0713] The "output means" is a display device or an audio output device for displaying or reproducing the generated answer to the user.

[0714] An "emotion engine" is an analysis device or software engine that analyzes a user's facial expressions and voice data to recognize the user's emotions.

[0715] "Lip sync means" refers to a processing device or software that synthesizes the generated response into voice and synchronizes the lip movements with the image of the historical figure.

[0716] A "simulation means" is a system for arranging a dialogue between different historical figures and reproducing or simulating the development of that dialogue.

[0717] A "user" is someone who operates the system and enters questions to obtain information about historical figures.

[0718] The "server" is the core of the system, and is a computer system that controls the database means, generative AI model means, emotion engine, etc., and manages and executes the overall processing.

[0719] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[0720] server

[0721] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[0722] Terminal

[0723] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[0724] User

[0725] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[0726] Program Processing Overview

[0727] server

[0728] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[0729] Terminal

[0730] The device provides an interface for users to input questions for historical figures. When the user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If the device includes a lip-sync function, it displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[0731] Examples:

[0732] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[0733] Lip Sync Means

[0734] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[0735] Examples:

[0736] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[0737] Simulation Method

[0738] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[0739] Examples:

[0740] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[0741] Emotion Engine

[0742] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[0743] Examples:

[0744] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[0745] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[0746] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0747] Step 1:

[0748] The user types a question on the device

[0749] Using the device's interface, the user enters a text question for the historical figure, such as "Napoleon, tell me about tactics." The user then clicks a submit button to complete the question.

[0750] Input: A question entered by the user as text

[0751] Output: The device sends the input data to the server

[0752] Step 2:

[0753] The device sends user input and emotion data to the server.

[0754] The device receives text data entered by the user and captures the user's facial expressions and voice using the built-in camera and microphone. The facial expression and voice data are analyzed and sent to the server as emotional data.

[0755] Input: User question, and captured facial expressions and voice

[0756] Output: Send all data to the server along with the analyzed emotion data

[0757] Step 3:

[0758] The server parses the request and retrieves information from the database

[0759] The server receives the request sent from the terminal and analyzes its contents. For example, it extracts keywords such as "Napoleon" and "tactics." Then, based on these keywords, it retrieves related information from the database means.

[0760] Input: Text data and emotion data sent from the device

[0761] Output: Historical information retrieved from the database

[0762] Step 4:

[0763] The server generates an answer using a generative AI model

[0764] The server inputs the information retrieved from the database and the user's question into a generative AI model (e.g., a natural language processing model). The generative AI model is used to generate an appropriate answer based on the context.

[0765] Input: Retrieved historical information and user question

[0766] Output: The answer generated by the generative AI model

[0767] Step 5:

[0768] The server uses an emotion engine to tailor the response

[0769] The server uses an emotion engine to tailor the generated answer based on the analyzed user's emotion data, for example adding more details and explanations to the answer if the user seems interested.

[0770] Input: Generated answers and sentiment data

[0771] Output: An answer tailored based on the sentiment

[0772] Step 6:

[0773] The server performs lip sync processing

[0774] The server uses lip-syncing to synthesize the generated responses and synchronize the mouth movements with the image of the historical figure, providing a realistic dialogue experience, as if Napoleon were actually speaking.

[0775] Input: Adjusted answer

[0776] Output: Lip-synced video and audio

[0777] Step 7:

[0778] The server sends the answer to the device

[0779] The server sends a response to the terminal indicating that the lip-sync process has been completed. The data sent includes text information, audio files, and lip-synced video.

[0780] Input: Lip-synced video and audio data

[0781] Output: Answer data sent to the device

[0782] Step 8:

[0783] The device displays the answer to the user

[0784] The device receives the answer data sent from the server and displays it to the user. The lip-synced faces of historical figures are displayed on the screen, and the answers are played back aloud.

[0785] Input: Response data sent from the server

[0786] Output: Answers shown to the user (text, audio, lip-sync video)

[0787] (Application example 2)

[0788] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0789] Conventional educational systems and content distribution services tend to provide users with a passive experience when learning history. Furthermore, one-way information provision without considering the user's emotions makes it difficult to maintain interest in learning. Furthermore, interactive systems lack the technology to provide a realistic dialogue experience, and the depictions of important historical figures are static. Therefore, there is a need for the development of a system that allows users to experience interactive dialogue with historical figures who actually have emotions while learning history.

[0790] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes database means for managing data on historical figures, interface means for accepting input from a user, means for capturing the user's facial expressions and voice and analyzing their emotions, communication means for analyzing the user's request and obtaining information on the specified historical figure from the database means, generation AI model means for adjusting the generated answer based on the user's emotions, output means for sending the generated answer to the user, and rendering means for displaying the lip-synced answer as a 3D image. This allows the user to interact with the historical figure in real time and enjoy an interactive learning experience personalized based on their emotions.

[0791] "Database means" is a device that has the function of systematically managing and storing information about historical figures.

[0792] An "interface means" is a device that accepts input from a user and provides a screen and operating means for the user to interact with the system.

[0793] The "means for analyzing emotions" is a device that has the function of capturing the user's facial expressions and voice and analyzing them to identify the user's emotional state.

[0794] The "communication means" is a device that receives a request from a user, retrieves information from a database based on the request, and passes the result to the generation AI model means.

[0795] The "generative AI model means" is a device that has the function of generating an appropriate response based on a user request and adjusting the tone and content of the response based on the user's emotional data.

[0796] The "output means" is a device for transmitting the generated answer to the user, and has the function of displaying or audibly transmitting the answer to the user in real time.

[0797] The "rendering means" is a device that renders the lip-synced answer as a three-dimensional model in real time and provides it visually to the user.

[0798] The "lip sync means" is a device that synchronizes the mouth movements of a three-dimensional model with the generated voice, thereby providing a realistic dialogue experience.

[0799] The "simulation means" is a device that has the function of having different historical figures converse with each other and simulating the development of that conversation in real time.

[0800] This invention provides a system for providing users with information about historical figures and realizing interactive dialogue. The system includes a database, an interface, an emotion analysis, a communication, a generative AI model, an output, a lip-sync, a rendering, and a simulation.

[0801] System components and functions

[0802] 1. Database Means

[0803] The server is equipped with a database that systematically manages and stores data on historical figures. This database is constructed using a relational database management system (RDBMS) such as MySQL.

[0804] 2. Interface Method

[0805] Users interact with the system through devices such as smartphones and head-mounted displays (HMDs). The devices are equipped with a user interface (UI) built using Unity, which allows users to enter input.

[0806] 3. Emotion analysis method

[0807] The device has the ability to capture the user's facial expressions and voice and analyze their emotions using APIs from Google ML Kit and Amazon Rekognition, making it possible to analyze the user's emotional state in real time.

[0808] 4. Means of communication

[0809] A request from a user is sent to the server using a RESTful API. The communication means has the function to analyze this request and retrieve the necessary information from the database.

[0810] 5. Generative AI Model Means

[0811] The generative AI model means installed on the server generates appropriate answers to questions using a natural language processing model such as OpenAI, and has the function of adjusting the tone and content of the answers based on user sentiment analysis data.

[0812] 6. Output Method

[0813] The generated answers are sent from the server to the device and provided to the user, using voice synthesis technology such as Amazon Polly.

[0814] 7. Lip Sync Method

[0815] The server performs lip-sync processing based on the generated audio, which uses the lip-sync function of MetaHuman Creator to generate mouth movements for the 3D model.

[0816] 8. Rendering Method

[0817] Using Unity's 3D rendering engine, three-dimensional models that have undergone lip-sync processing are rendered in real time and provided to users.

[0818] 9. Simulation Methods

[0819] The server has the ability to have different historical figures interact with each other and simulate the progress of the interaction in real time, allowing users to learn about history from a new perspective.

[0820] Specific examples of processing

[0821] For example, suppose a user types "Tell me about your days as Napoleon" into the device's interface and presses the send button. At this time, the device's camera captures and analyzes the user's facial expression. The emotion analysis means recognizes that the user is interested. This request is then sent to the server via the communication means, and data about Napoleon is retrieved from the database.

[0822] Next, the generating AI model means generates an appropriate answer to "Life as Napoleon" and adjusts the answer taking into account the user's emotions. The server then synthesizes the generated answer into speech and generates mouth movements for the three-dimensional model of Napoleon using the lip-sync means. Finally, the rendering means renders the three-dimensional model of Napoleon in real time and displays it on the device.

[0823] Prompt Sentence Examples

[0824] "Tell me about your days as Napoleon."

[0825] "Alexander the Great, what is your most amazing tactic?"

[0826] Using this system, users can interact with historical figures in real time and enjoy a personalized, interactive learning experience that responds to their emotions.

[0827] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0828] Step 1: The user enters a question for a historical figure through the device interface and presses the send button. At this time, the interface captures the text of the user's question and uses the camera and microphone to collect the user's facial expressions and voice in real time. The input is the user's question text, facial expressions, and voice data, and the output is a process that compiles this data into a single request.

[0829] Step 2: The device uses Google ML Kit and Amazon Rekognition API on the edge device to analyze the collected facial and voice data in real time to identify the user's emotional state. The input is the user's facial and voice data, and the output is the analyzed emotional data.

[0830] Step 3: The device sends a request containing the user's question text and the parsed emotion data to the server via a RESTful API. The input is the request containing the question text and emotion data, and the output is the request data sent to the server.

[0831] Step 4: The server receives the request and retrieves information about the specified historical figure using a database (e.g., MySQL). The input is the request data from the user (question text and emotion data), and the output is the historical information retrieved from the database.

[0832] Step 5: The server inputs the acquired historical information into a generative AI model (such as OpenAI) to generate an answer to the user's question. The generated answer is also adjusted based on the emotional data. The input is the historical information acquired from the database and the user's emotional data, and the output is an emotion-adjusted answer text.

[0833] Step 6: The adjusted answer text is converted into speech data using a speech synthesis means (e.g., Amazon Polly). During this process, the generated speech data is sent to a lip-sync means (e.g., MetaHuman Creator) to generate mouth movements corresponding to a three-dimensional model of the historical figure. The input is the emotion-adjusted answer text, and the output is speech data and lip-sync data.

[0834] Step 7: The server renders the audio data and lip-sync data in real time using a rendering engine (such as Unity's 3D rendering engine) and sends it to the device. The input is the audio data and lip-sync data, and the output is the rendered 3D model.

[0835] Step 8: The terminal displays the rendered 3D model to the user and plays the audio data. This allows the user to interact with the historical figure in real time and enjoy personalized responses based on their emotions. The input is the rendered 3D model and audio data from the server, and the output is audiovisual information provided to the user.

[0836] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0837] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0838] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0839] [Third embodiment]

[0840] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0841] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0842] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0843] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0844] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0845] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0846] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0847] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0848] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0849] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0850] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0851] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0852] The present invention relates to a system for managing data on historical figures and allowing users to have conversations as if they were the historical figures. The system includes a database, an interface, a communication means, a generative AI model, an output means, a lip-sync means, and a simulation means.

[0853] System Components

[0854] server

[0855] The server is the core of the system and includes a database, communication, AI model generation, and output. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[0856] Terminal

[0857] The terminal provides an interface for the user to interact with the system: it accepts user input, transmits it to the server, and displays responses received from the server to the user.

[0858] User

[0859] Through this system, users interact with historical figures by operating a terminal, inputting questions, and reading or listening to the answers provided by the server.

[0860] Program Processing Overview

[0861] Terminal

[0862] The terminal provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying answers to the user. If a lip-sync function is included, a photo of the person's face is lip-synced and displayed.

[0863] Examples:

[0864] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[0865] server

[0866] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. The generated answer is sent to the terminal through the output means.

[0867] Examples:

[0868] The server retrieves information about "Napoleon" and "tactics" from a database, inputs it into a generative AI model, and generates the answer, "The tactics I value most are rapid advance and attacking the enemy's weak points."

[0869] Lip Sync Means

[0870] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[0871] Examples:

[0872] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[0873] Simulation Method

[0874] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[0875] Examples:

[0876] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[0877] The system allows users to interact with historical figures in real time, providing a deeper historical understanding and learning experience.

[0878] The processing flow will be explained below.

[0879] Specific processing flow of the program

[0880] Step 1:

[0881] The user enters a question into the terminal interface and clicks the send button.

[0882] Step 2:

[0883] The device formats the user's question as an API request and sends it to the server.

[0884] Step 3:

[0885] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[0886] Step 4:

[0887] The server searches and acquires information about the designated historical person from the database means.

[0888] Step 5:

[0889] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[0890] Step 6:

[0891] The server formats the generated answer and creates response data.

[0892] Step 7:

[0893] The server sends the response data to the terminal.

[0894] Step 8:

[0895] The terminal receives the response data from the server and formats it for display to the user.

[0896] Step 9:

[0897] The terminal displays the answer to the user, and if there is lip synchronization means, generates and displays the mouth movements.

[0898] Step 10:

[0899] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[0900] This process allows users to simulate real-time interactions with historical figures, enhancing the learning experience.

[0901] Example 1

[0902] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0903] In the modern education system, history classes tend to provide only one-sided information, making it difficult for students to acquire interest or concern. In particular, there is a lack of concrete ways to deepen knowledge about historical figures. There is a growing need for interactive educational tools that will engage students and trainees in learning.

[0904] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0905] In this invention, the server includes an information aggregation means for managing data on historical figures, an input means for accepting input from a user, a receiving means for analyzing a request from the user and acquiring information on the specified historical figure from the information aggregation means, a generation AI model means for providing a generated answer, and a transmitting means for transmitting the generated answer to the user. This allows the user to have an interactive conversation with the historical figure, thereby improving the learning effect.

[0906] "Information aggregation means" refers to the means for managing data on historical figures and collecting and organizing necessary information.

[0907] The "input means" is an interface that accepts input from the user, and is a means by which the user inputs questions or requests.

[0908] The "receiving means" is a means for analyzing a request from a user and obtaining information about a specified historical person from the information aggregating means.

[0909] A "generative AI model means" is a means that uses an artificial intelligence model to generate an answer based on acquired information.

[0910] The "transmission means" is a means for transmitting the generated answer to the user.

[0911] The "lip sync means" is a means for synthesizing lip movements in real time with an image of a historical figure based on the generated answer.

[0912] A "simulation method" is a method for arranging a dialogue between different historical figures and reproducing the development of that dialogue.

[0913] This invention is a system for managing data on historical figures and allowing users to have conversations as if they were those figures. The system includes an information aggregation unit, an input unit, a receiving unit, a generating AI model unit, a transmitting unit, a lip-sync unit, and a simulation unit.

[0914] server

[0915] The server is the central part of the system and is the hardware that controls various functions. The server has the following means for processing data:

[0916] Information aggregation tool: A tool for managing data about historical figures. This data is stored in a database and includes detailed historical information about the person and records of their statements.

[0917] Reception means: A means for receiving requests sent by users. The received requests are analyzed and necessary keywords and contexts are extracted.

[0918] Generative AI model means: A means for obtaining necessary data from the information aggregation means based on the analysis results obtained from the receiving means, and then using an AI model to generate an appropriate response. The generative AI model is, for example, a natural language generation model such as GPT (Generative Pre-trained Transformer).

[0919] Transmission means: A means for transmitting the generated answer to the terminal, allowing the user to receive the generated answer.

[0920] Terminal

[0921] A terminal is a piece of hardware that provides an interface for users to interact with a system. Specifically, it operates as follows:

[0922] Input means: A means that provides an interface for users to input questions. Users input their questions using this means and send them to the server by pressing the send button.

[0923] Lip-syncing: A method for synthesizing lip movements in real time onto images of historical figures based on answers sent from the server. This process provides a more natural and interactive dialogue experience.

[0924] Examples:

[0925] The user types "Tell me about your tactics" into the device interface and presses the send button. The server receives this question, analyzes it, and extracts the keywords "Napoleon" and "tactics." The server obtains related data from the information aggregation means and inputs it into the generative AI model. The generative AI model generates an answer, "The tactics I value most are rapid advancement and attacking the enemy's weak points," and sends this to the device. The device receives this answer and displays it to the user using lip-syncing means, replicating the mouth movements of an image of Napoleon.

[0926] Prompt Sentence Examples

[0927] "Tell me about your tactics."

[0928] "Generate detailed information about Napoleon's tactics."

[0929] conclusion

[0930] This system allows users to interact with historical figures, enhancing learning effectiveness, and its lip-sync function provides a realistic experience both visually and aurally.

[0931] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0932] Processing Steps

[0933] Step 1:

[0934] User enters question and submits

[0935] The user uses the device interface to input a question for the historical figure. When input is complete, the user presses the "Send" button. This operation sends the question data from the device to the server.

[0936] Input: "Tell me about your tactics."

[0937] Output: Question data (e.g., "Tell me about your tactics")

[0938] Step 2:

[0939] The server receives and analyzes the query

[0940] The server receives the question data sent from the device, analyzes the received question data, and extracts keywords and context, thereby identifying the type and subject of the requested information.

[0941] Input: Question data (e.g., "Tell me about your tactics")

[0942] Output: Analysis results (e.g. "Napoleon" "Tactics")

[0943] Step 3:

[0944] The server retrieves the information from the database

[0945] Based on the analysis results, the server searches for information about related historical figures from the information aggregation means. The information aggregation means stores a large amount of data and efficiently retrieves the required information.

[0946] Input: Analysis result (e.g. "Napoleon" "Tactics")

[0947] Output: relevant information (e.g. "Information about Napoleon's tactics")

[0948] Step 4:

[0949] The server inputs information into the generated AI model

[0950] The server inputs the relevant information retrieved from the database and the user's question into the generative AI model means, which then generates the optimal answer based on this information.

[0951] Input: Relevant information and user questions (e.g., "Information about Napoleon's tactics" and "Tell me about your tactics")

[0952] Output: Generated answer (e.g., "My most important tactics are rapid advance and attacking the enemy's weak points.")

[0953] Step 5:

[0954] The server generates a response and sends it to the device.

[0955] The generated answer is returned to the server, which then transmits it to the terminal.

[0956] Input: Generated answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0957] Output: The submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0958] Step 6:

[0959] The device will lip sync and display the answer

[0960] The device displays the answer received from the server to the user. If the device includes a lip-sync function, it performs lip-sync processing on the image of the person and reproduces natural mouth movements using voice synthesis.

[0961] Input: Submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[0962] Output: Displayed and voice-synthesized answer (e.g., "My most important tactical priorities are rapid advance and striking the enemy's weak points.")

[0963] Step 7:

[0964] User observations and follow-up questions

[0965] The user observes the displayed answers, and if they have further questions, they again use the terminal interface to enter and submit new questions, and the process repeats.

[0966] Input: Follow-up question (e.g., "Can you give us an example of a specific tactic?")

[0967] Output: New question data (e.g., "Can you give us an example of a specific tactic?")

[0968] (Application example 1)

[0969] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0970] In today's factories, managers and workers are required to have a high level of specialized knowledge and experience in order to carry out efficient work management and training. However, this knowledge depends on the experience of each individual, making it difficult to achieve consistent and efficient work management. Another problem is that training and nurturing new workers takes a lot of time and money.

[0971] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0972] In this invention, the server includes a database for managing data on historical figures, an interface for receiving user input, a communication unit for analyzing user requests and retrieving information about the specified historical figures from the database, a generation AI model for providing generated answers, and an output unit for transmitting the generated answers to the user. This allows factory managers and workers to apply historical knowledge and strategies to improve work efficiency and quickly train new workers. Furthermore, the voice synthesis and lip-sync functions allow users to enjoy a more realistic experience.

[0973] The "database means" is a system or device that centrally manages data related to historical figures and quickly provides relevant data in response to a user's request.

[0974] "Interface means" refers to an input device or interface that allows a user to input questions or requests to the system, and has the function of accepting user input.

[0975] "Communication means" refers to a system or device with communication capabilities that analyzes requests from users, retrieves the necessary data from a database, and passes it to the generative AI model.

[0976] "Generative AI model means" refers to an artificial intelligence model that generates an appropriate answer based on information obtained from a database in response to a question from a user.

[0977] "Output means" refers to a system or device for transmitting and displaying the answer generated by the generative AI model to the user.

[0978] The "lip sync means" is a system or device that has the function of synthesizing the generated answer into voice and then performing lip sync processing on a photograph of a historical figure's face.

[0979] "Factory training and support measures" refers to the functions and equipment that support the effective management of machines and systems used within a factory and the training of workers.

[0980] This invention is a system that applies the tactics and knowledge of historical figures to factory robots, enabling factory management and work support. The system includes database means, interface means, communication means, generative AI model means, output means, lip-sync means, and factory training and support means.

[0981] System Configuration

[0982] server

[0983] The server is the core of this system and includes a database, communication, AI model generation, and output functions. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[0984] Terminal

[0985] The terminal provides an interface for the user to interact with the system, accepts user input, transmits it to the server, and displays responses received from the server. It also includes functionality for displaying a lip-synced person's face.

[0986] User

[0987] The user operates the interface means to input a question and receive a response from the system, thereby obtaining advice and knowledge from historical figures.

[0988] Hardware

[0989] Robot: An industrial machine used in a factory (e.g., an industrial robot).

[0990] Server: Runs on a cloud service (e.g. AWS EC2).

[0991] Camera and microphone: Sensor devices built into the robot.

[0992] Display: A display device that allows the user to see the content of the interaction.

[0993] software

[0994] Natural Language Processing (NLP) library: OpenAI's GPT-4.

[0995] Speech synthesis library: Google Text-to-Speech.

[0996] Lip sync library: Avatarify.

[0997] Database: MySQL.

[0998] Program Processing Overview

[0999] Processing on the terminal

[1000] The device provides an interface for users to input questions for historical figures. When a user inputs and submits a question, the data is sent to the server. Furthermore, if the device includes a lip-sync function, it has the function of displaying a lip-synced photograph of the historical figure.

[1001] Examples:

[1002] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[1003] Processing on the server

[1004] The server receives the request from the device and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from a database. The retrieved information and the user's question are input into a generative AI model to generate an appropriate answer. The generated answer is synthesized into voice, lip-synced, and sent to the device.

[1005] Examples:

[1006] The server retrieves information about a specified historical figure and tactics from a database, inputs it into a generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points."

[1007] Lip Sync Means

[1008] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them to the facial photographs of historical figures, providing a more realistic conversation experience.

[1009] Examples:

[1010] A lip sync means generates and displays in real time mouth movements based on the generated answers for the photographs of the faces of historical figures.

[1011] Prompt Sentence Examples

[1012] "As the designated historical figure, answer this question: Tell us about your tactics."

[1013] "As a designated historical figure, please assess the progress of the production line this week."

[1014] This system allows users to receive factory management and work support that applies historical knowledge, enabling efficient work management and rapid training.

[1015] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1016] Step 1:

[1017] The user enters a question into the terminal interface and clicks the send button. The entered question is saved on the terminal and then sent to the server.

[1018] Specific operation: The user enters "Tell us about your tactics" into the input field on the device and presses the send button. The input data is sent from the device to the server.

[1019] Step 2:

[1020] The server receives and analyzes a request (question) from the user. Based on the analysis results, it retrieves information about the specified historical figure from the database.

[1021] Input: User question data

[1022] Output: Database query results

[1023] Specific operation: The server searches the database for information about the "specified historical figure" and "tactic" and retrieves related information.

[1024] Step 3:

[1025] The server inputs the acquired information and the user's question into a generative AI model to generate an appropriate answer.

[1026] Input: Information about historical figures retrieved from the database, user questions

[1027] Output: The generated answer

[1028] How it works: The server inputs a prompt to a generative AI model (e.g., OpenAI GPT-4) saying, "As a specified historical figure, tell me about a tactical battle," and receives the generated answer.

[1029] Step 4:

[1030] The generated answer is synthesized into voice.

[1031] Input: Generated answer text

[1032] Output: Audio file

[1033] Specific behavior: The server uses a speech synthesis library (e.g., Google Text-to-Speech) to convert the generated answer text into an audio file.

[1034] Step 5:

[1035] Lip-sync processing is performed using the generated answer audio file.

[1036] Input: Photo of a historical figure, audio file

[1037] Output: Lip-synced video

[1038] What happens: The server uses a lip-sync library (e.g., Avatarify) to generate audio-synced mouth movements for a photograph of a historical figure, creating a lip-synced video.

[1039] Step 6:

[1040] The server sends the lip-synced video to the terminal.

[1041] Input: Lip-synced video

[1042] Output: The video displayed on the user's device

[1043] Specific operation: The server sends lip-synced video data to the device, which then displays the video to the user.

[1044] Step 7:

[1045] Users watch lip-synced videos on their devices and receive answers from historical figures.

[1046] Input: Lip-synced video

[1047] Output: User understanding and real-time interaction experience

[1048] What happens: The user watches a lip-synced video of a historical figure on the device display and confirms their answers.

[1049] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1050] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[1051] System Components

[1052] server

[1053] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[1054] Terminal

[1055] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[1056] User

[1057] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[1058] Program Processing Overview

[1059] Terminal

[1060] The device provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If a lip-sync function is included, the device displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[1061] Examples:

[1062] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[1063] server

[1064] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[1065] Examples:

[1066] The server retrieves information about "Napoleon" and "tactics" from the database, inputs it into the generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points." The emotion engine analyzes the user's interested facial expression and adds a slightly more detailed explanation to the answer.

[1067] Lip Sync Means

[1068] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[1069] Examples:

[1070] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[1071] Simulation Method

[1072] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[1073] Examples:

[1074] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[1075] Emotion Engine

[1076] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[1077] Examples:

[1078] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[1079] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[1080] The processing flow will be explained below.

[1081] Specific processing flow of the program (including the emotion engine)

[1082] Step 1:

[1083] The user enters a question into the terminal interface and clicks the send button.

[1084] Step 2:

[1085] The device formats the user's question as an API request and sends it to the server, and also captures the user's facial expressions and voice and sends them to the server as emotion data.

[1086] Step 3:

[1087] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[1088] Step 4:

[1089] The server searches and acquires information about the designated historical person from the database means.

[1090] Step 5:

[1091] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[1092] Step 6:

[1093] The emotion engine analyzes the user's emotion data transmitted from the terminal and determines the user's emotional state.

[1094] Examples:

[1095] The emotion engine analyzes the user's facial expressions and determines whether the user is interested.

[1096] Step 7:

[1097] The sentiment engine adjusts the generated answer based on the determined user sentiment, for example adding detailed explanations to the answer if the user indicates interest.

[1098] Examples:

[1099] "My most important tactic is to advance quickly and attack the enemy's weak points," one generated answer might explain, with the further detail, "This tactic has won me many battles."

[1100] Step 8:

[1101] The server formats the generated answer and creates response data.

[1102] Step 9:

[1103] The server sends the response data to the terminal.

[1104] Step 10:

[1105] The terminal receives the response data from the server and formats it for display to the user.

[1106] Step 11:

[1107] The lip-syncing means synthesizes the generated responses into voice and lip-syncs them to the portraits of historical figures, providing a more realistic dialogue experience.

[1108] Step 12:

[1109] The terminal displays the answer to the user and plays the video with lip-sync processing completed.

[1110] Step 13:

[1111] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[1112] This process allows users to simulate real-time interactions with historical figures and receive personalized, emotion-based responses to enhance their learning experience.

[1113] Example 2

[1114] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1115] Conventional dialogue systems only provide simple text-based answers when users seek information about historical figures, resulting in a lack of realism and immersion. Furthermore, they are unable to provide personalized answers based on the user's emotions, preventing an improved learning experience. Furthermore, they are unable to simulate conversations between different historical figures, limiting opportunities to learn history from new perspectives.

[1116] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1117] In this invention, the server includes a database means for managing data on historical figures, a communication means for analyzing a request from a user and obtaining information on the specified historical figure from the database means, a generative AI model means for providing a generated answer, an emotion engine for recognizing emotions by analyzing the user's facial expressions and voice, and a means for adjusting the tone and content of the generated answer based on the recognized emotion. This allows the user to interact with the historical figure in real time and receive personalized answers based on their emotions, enabling a deeper learning experience.

[1118] The "database means" is a storage or management system for managing data relating to historical figures and retrieving information as needed.

[1119] An "interface means" is an input device or software interface for accepting input from a user.

[1120] "Communication means" refers to a network device or communication protocol for transmitting user input data to a server and transmitting data from the server to the user.

[1121] The "generative AI model means" is an artificial intelligence model or natural language processing system that analyzes a user's question and generates an appropriate answer.

[1122] The "output means" is a display device or an audio output device for displaying or reproducing the generated answer to the user.

[1123] An "emotion engine" is an analysis device or software engine that analyzes a user's facial expressions and voice data to recognize the user's emotions.

[1124] "Lip sync means" refers to a processing device or software that synthesizes the generated response into voice and synchronizes the lip movements with the image of the historical figure.

[1125] A "simulation means" is a system for arranging a dialogue between different historical figures and reproducing or simulating the development of that dialogue.

[1126] A "user" is someone who operates the system and enters questions to obtain information about historical figures.

[1127] The "server" is the core of the system, and is a computer system that controls the database means, generative AI model means, emotion engine, etc., and manages and executes the overall processing.

[1128] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[1129] server

[1130] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[1131] Terminal

[1132] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[1133] User

[1134] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[1135] Program Processing Overview

[1136] server

[1137] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[1138] Terminal

[1139] The device provides an interface for users to input questions for historical figures. When the user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If the device includes a lip-sync function, it displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[1140] Examples:

[1141] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[1142] Lip Sync Means

[1143] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[1144] Examples:

[1145] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[1146] Simulation Method

[1147] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[1148] Examples:

[1149] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[1150] Emotion Engine

[1151] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[1152] Examples:

[1153] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[1154] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[1155] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1156] Step 1:

[1157] The user types a question on the device

[1158] Using the device's interface, the user enters a text question for the historical figure, such as "Napoleon, tell me about tactics." The user then clicks a submit button to complete the question.

[1159] Input: A question entered by the user as text

[1160] Output: The device sends the input data to the server

[1161] Step 2:

[1162] The device sends user input and emotion data to the server.

[1163] The device receives text data entered by the user and captures the user's facial expressions and voice using the built-in camera and microphone. The facial expression and voice data are analyzed and sent to the server as emotional data.

[1164] Input: User question, and captured facial expressions and voice

[1165] Output: Send all data to the server along with the analyzed emotion data

[1166] Step 3:

[1167] The server parses the request and retrieves information from the database

[1168] The server receives the request sent from the terminal and analyzes its contents. For example, it extracts keywords such as "Napoleon" and "tactics." Then, based on these keywords, it retrieves related information from the database means.

[1169] Input: Text data and emotion data sent from the device

[1170] Output: Historical information retrieved from the database

[1171] Step 4:

[1172] The server generates an answer using a generative AI model

[1173] The server inputs the information retrieved from the database and the user's question into a generative AI model (e.g., a natural language processing model). The generative AI model is used to generate an appropriate answer based on the context.

[1174] Input: Retrieved historical information and user question

[1175] Output: The answer generated by the generative AI model

[1176] Step 5:

[1177] The server uses an emotion engine to tailor the response

[1178] The server uses an emotion engine to tailor the generated answer based on the analyzed user's emotion data, for example adding more details and explanations to the answer if the user seems interested.

[1179] Input: Generated answers and sentiment data

[1180] Output: An answer tailored based on the sentiment

[1181] Step 6:

[1182] The server performs lip sync processing

[1183] The server uses lip-syncing to synthesize the generated responses and synchronize the mouth movements with the image of the historical figure, providing a realistic dialogue experience, as if Napoleon were actually speaking.

[1184] Input: Adjusted answer

[1185] Output: Lip-synced video and audio

[1186] Step 7:

[1187] The server sends the answer to the device

[1188] The server sends a response to the terminal indicating that the lip-sync process has been completed. The data sent includes text information, audio files, and lip-synced video.

[1189] Input: Lip-synced video and audio data

[1190] Output: Answer data sent to the device

[1191] Step 8:

[1192] The device displays the answer to the user

[1193] The device receives the answer data sent from the server and displays it to the user. The lip-synced faces of historical figures are displayed on the screen, and the answers are played back aloud.

[1194] Input: Response data sent from the server

[1195] Output: Answers shown to the user (text, audio, lip-sync video)

[1196] (Application example 2)

[1197] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1198] Conventional educational systems and content distribution services tend to provide users with a passive experience when learning history. Furthermore, one-way information provision without considering the user's emotions makes it difficult to maintain interest in learning. Furthermore, interactive systems lack the technology to provide a realistic dialogue experience, and the depictions of important historical figures are static. Therefore, there is a need for the development of a system that allows users to experience interactive dialogue with historical figures who actually have emotions while learning history.

[1199] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes database means for managing data on historical figures, interface means for accepting input from a user, means for capturing the user's facial expressions and voice and analyzing their emotions, communication means for analyzing the user's request and obtaining information on the specified historical figure from the database means, generation AI model means for adjusting the generated answer based on the user's emotions, output means for sending the generated answer to the user, and rendering means for displaying the lip-synced answer as a 3D image. This allows the user to interact with the historical figure in real time and enjoy an interactive learning experience personalized based on their emotions.

[1200] "Database means" is a device that has the function of systematically managing and storing information about historical figures.

[1201] An "interface means" is a device that accepts input from a user and provides a screen and operating means for the user to interact with the system.

[1202] The "means for analyzing emotions" is a device that has the function of capturing the user's facial expressions and voice and analyzing them to identify the user's emotional state.

[1203] The "communication means" is a device that receives a request from a user, retrieves information from a database based on the request, and passes the result to the generation AI model means.

[1204] The "generative AI model means" is a device that has the function of generating an appropriate response based on a user request and adjusting the tone and content of the response based on the user's emotional data.

[1205] The "output means" is a device for transmitting the generated answer to the user, and has the function of displaying or audibly transmitting the answer to the user in real time.

[1206] The "rendering means" is a device that renders the lip-synced answer as a three-dimensional model in real time and provides it visually to the user.

[1207] The "lip sync means" is a device that synchronizes the mouth movements of a three-dimensional model with the generated voice, thereby providing a realistic dialogue experience.

[1208] The "simulation means" is a device that has the function of having different historical figures converse with each other and simulating the development of that conversation in real time.

[1209] This invention provides a system for providing users with information about historical figures and realizing interactive dialogue. The system includes a database, an interface, an emotion analysis, a communication, a generative AI model, an output, a lip-sync, a rendering, and a simulation.

[1210] System components and functions

[1211] 1. Database Means

[1212] The server is equipped with a database that systematically manages and stores data on historical figures. This database is constructed using a relational database management system (RDBMS) such as MySQL.

[1213] 2. Interface Method

[1214] Users interact with the system through devices such as smartphones and head-mounted displays (HMDs). The devices are equipped with a user interface (UI) built using Unity, which allows users to enter input.

[1215] 3. Emotion analysis method

[1216] The device has the ability to capture the user's facial expressions and voice and analyze their emotions using APIs from Google ML Kit and Amazon Rekognition, making it possible to analyze the user's emotional state in real time.

[1217] 4. Means of communication

[1218] A request from a user is sent to the server using a RESTful API. The communication means has the function to analyze this request and retrieve the necessary information from the database.

[1219] 5. Generative AI Model Means

[1220] The generative AI model means installed on the server generates appropriate answers to questions using a natural language processing model such as OpenAI, and has the function of adjusting the tone and content of the answers based on user sentiment analysis data.

[1221] 6. Output Method

[1222] The generated answers are sent from the server to the device and provided to the user, using voice synthesis technology such as Amazon Polly.

[1223] 7. Lip Sync Method

[1224] The server performs lip-sync processing based on the generated audio, which uses the lip-sync function of MetaHuman Creator to generate mouth movements for the 3D model.

[1225] 8. Rendering Method

[1226] Using Unity's 3D rendering engine, three-dimensional models that have undergone lip-sync processing are rendered in real time and provided to users.

[1227] 9. Simulation Methods

[1228] The server has the ability to have different historical figures interact with each other and simulate the progress of the interaction in real time, allowing users to learn about history from a new perspective.

[1229] Specific examples of processing

[1230] For example, suppose a user types "Tell me about your days as Napoleon" into the device's interface and presses the send button. At this time, the device's camera captures and analyzes the user's facial expression. The emotion analysis means recognizes that the user is interested. This request is then sent to the server via the communication means, and data about Napoleon is retrieved from the database.

[1231] Next, the generating AI model means generates an appropriate answer to "Life as Napoleon" and adjusts the answer taking into account the user's emotions. The server then synthesizes the generated answer into speech and generates mouth movements for the three-dimensional model of Napoleon using the lip-sync means. Finally, the rendering means renders the three-dimensional model of Napoleon in real time and displays it on the device.

[1232] Prompt Sentence Examples

[1233] "Tell me about your days as Napoleon."

[1234] "Alexander the Great, what is your most amazing tactic?"

[1235] Using this system, users can interact with historical figures in real time and enjoy a personalized, interactive learning experience that responds to their emotions.

[1236] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1237] Step 1: The user enters a question for a historical figure through the device interface and presses the send button. At this time, the interface captures the text of the user's question and uses the camera and microphone to collect the user's facial expressions and voice in real time. The input is the user's question text, facial expressions, and voice data, and the output is a process that compiles this data into a single request.

[1238] Step 2: The device uses Google ML Kit and Amazon Rekognition API on the edge device to analyze the collected facial and voice data in real time to identify the user's emotional state. The input is the user's facial and voice data, and the output is the analyzed emotional data.

[1239] Step 3: The device sends a request containing the user's question text and the parsed emotion data to the server via a RESTful API. The input is the request containing the question text and emotion data, and the output is the request data sent to the server.

[1240] Step 4: The server receives the request and retrieves information about the specified historical figure using a database (e.g., MySQL). The input is the request data from the user (question text and emotion data), and the output is the historical information retrieved from the database.

[1241] Step 5: The server inputs the acquired historical information into a generative AI model (such as OpenAI) to generate an answer to the user's question. The generated answer is also adjusted based on the emotional data. The input is the historical information acquired from the database and the user's emotional data, and the output is an emotion-adjusted answer text.

[1242] Step 6: The adjusted answer text is converted into speech data using a speech synthesis means (e.g., Amazon Polly). During this process, the generated speech data is sent to a lip-sync means (e.g., MetaHuman Creator) to generate mouth movements corresponding to a three-dimensional model of the historical figure. The input is the emotion-adjusted answer text, and the output is speech data and lip-sync data.

[1243] Step 7: The server renders the audio data and lip-sync data in real time using a rendering engine (such as Unity's 3D rendering engine) and sends it to the device. The input is the audio data and lip-sync data, and the output is the rendered 3D model.

[1244] Step 8: The terminal displays the rendered 3D model to the user and plays the audio data. This allows the user to interact with the historical figure in real time and enjoy personalized responses based on their emotions. The input is the rendered 3D model and audio data from the server, and the output is audiovisual information provided to the user.

[1245] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1246] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1247] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1248] [Fourth embodiment]

[1249] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1250] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1251] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1252] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1253] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1254] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1255] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1256] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1257] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1258] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1259] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1260] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1261] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1262] The present invention relates to a system for managing data on historical figures and allowing users to have conversations as if they were the historical figures. The system includes a database, an interface, a communication means, a generative AI model, an output means, a lip-sync means, and a simulation means.

[1263] System Components

[1264] server

[1265] The server is the core of the system and includes a database, communication, AI model generation, and output. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[1266] Terminal

[1267] The terminal provides an interface for the user to interact with the system: it accepts user input, transmits it to the server, and displays responses received from the server to the user.

[1268] User

[1269] Through this system, users interact with historical figures by operating a terminal, inputting questions, and reading or listening to the answers provided by the server.

[1270] Program Processing Overview

[1271] Terminal

[1272] The terminal provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying answers to the user. If a lip-sync function is included, a photo of the person's face is lip-synced and displayed.

[1273] Examples:

[1274] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[1275] server

[1276] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. The generated answer is sent to the terminal through the output means.

[1277] Examples:

[1278] The server retrieves information about "Napoleon" and "tactics" from a database, inputs it into a generative AI model, and generates the answer, "The tactics I value most are rapid advance and attacking the enemy's weak points."

[1279] Lip Sync Means

[1280] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[1281] Examples:

[1282] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[1283] Simulation Method

[1284] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[1285] Examples:

[1286] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[1287] The system allows users to interact with historical figures in real time, providing a deeper historical understanding and learning experience.

[1288] The processing flow will be explained below.

[1289] Specific processing flow of the program

[1290] Step 1:

[1291] The user enters a question into the terminal interface and clicks the send button.

[1292] Step 2:

[1293] The device formats the user's question as an API request and sends it to the server.

[1294] Step 3:

[1295] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[1296] Step 4:

[1297] The server searches and acquires information about the designated historical person from the database means.

[1298] Step 5:

[1299] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[1300] Step 6:

[1301] The server formats the generated answer and creates response data.

[1302] Step 7:

[1303] The server sends the response data to the terminal.

[1304] Step 8:

[1305] The terminal receives the response data from the server and formats it for display to the user.

[1306] Step 9:

[1307] The terminal displays the answer to the user, and if there is lip synchronization means, generates and displays the mouth movements.

[1308] Step 10:

[1309] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[1310] This process allows users to simulate real-time interactions with historical figures, enhancing the learning experience.

[1311] Example 1

[1312] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1313] In the modern education system, history classes tend to provide only one-sided information, making it difficult for students to acquire interest or concern. In particular, there is a lack of concrete ways to deepen knowledge about historical figures. There is a growing need for interactive educational tools that will engage students and trainees in learning.

[1314] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1315] In this invention, the server includes an information aggregation means for managing data on historical figures, an input means for accepting input from a user, a receiving means for analyzing a request from the user and acquiring information on the specified historical figure from the information aggregation means, a generation AI model means for providing a generated answer, and a transmitting means for transmitting the generated answer to the user. This allows the user to have an interactive conversation with the historical figure, thereby improving the learning effect.

[1316] "Information aggregation means" refers to the means for managing data on historical figures and collecting and organizing necessary information.

[1317] The "input means" is an interface that accepts input from the user, and is a means by which the user inputs questions or requests.

[1318] The "receiving means" is a means for analyzing a request from a user and obtaining information about a specified historical person from the information aggregating means.

[1319] A "generative AI model means" is a means that uses an artificial intelligence model to generate an answer based on acquired information.

[1320] The "transmission means" is a means for transmitting the generated answer to the user.

[1321] The "lip sync means" is a means for synthesizing lip movements in real time with an image of a historical figure based on the generated answer.

[1322] A "simulation method" is a method for arranging a dialogue between different historical figures and reproducing the development of that dialogue.

[1323] This invention is a system for managing data on historical figures and allowing users to have conversations as if they were those figures. The system includes an information aggregation unit, an input unit, a receiving unit, a generating AI model unit, a transmitting unit, a lip-sync unit, and a simulation unit.

[1324] server

[1325] The server is the central part of the system and is the hardware that controls various functions. The server has the following means for processing data:

[1326] Information aggregation tool: A tool for managing data about historical figures. This data is stored in a database and includes detailed historical information about the person and records of their statements.

[1327] Reception means: A means for receiving requests sent by users. The received requests are analyzed and necessary keywords and contexts are extracted.

[1328] Generative AI model means: A means for obtaining necessary data from the information aggregation means based on the analysis results obtained from the receiving means, and then using an AI model to generate an appropriate response. The generative AI model is, for example, a natural language generation model such as GPT (Generative Pre-trained Transformer).

[1329] Transmission means: A means for transmitting the generated answer to the terminal, allowing the user to receive the generated answer.

[1330] Terminal

[1331] A terminal is a piece of hardware that provides an interface for users to interact with a system. Specifically, it operates as follows:

[1332] Input means: A means that provides an interface for users to input questions. Users input their questions using this means and send them to the server by pressing the send button.

[1333] Lip-syncing: A method for synthesizing lip movements in real time onto images of historical figures based on answers sent from the server. This process provides a more natural and interactive dialogue experience.

[1334] Examples:

[1335] The user types "Tell me about your tactics" into the device interface and presses the send button. The server receives this question, analyzes it, and extracts the keywords "Napoleon" and "tactics." The server obtains related data from the information aggregation means and inputs it into the generative AI model. The generative AI model generates an answer, "The tactics I value most are rapid advancement and attacking the enemy's weak points," and sends this to the device. The device receives this answer and displays it to the user using lip-syncing means, replicating the mouth movements of an image of Napoleon.

[1336] Prompt Sentence Examples

[1337] "Tell me about your tactics."

[1338] "Generate detailed information about Napoleon's tactics."

[1339] conclusion

[1340] This system allows users to interact with historical figures, enhancing learning effectiveness, and its lip-sync function provides a realistic experience both visually and aurally.

[1341] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1342] Processing Steps

[1343] Step 1:

[1344] User enters question and submits

[1345] The user uses the device interface to input a question for the historical figure. When input is complete, the user presses the "Send" button. This operation sends the question data from the device to the server.

[1346] Input: "Tell me about your tactics."

[1347] Output: Question data (e.g., "Tell me about your tactics")

[1348] Step 2:

[1349] The server receives and analyzes the query

[1350] The server receives the question data sent from the device, analyzes the received question data, and extracts keywords and context, thereby identifying the type and subject of the requested information.

[1351] Input: Question data (e.g., "Tell me about your tactics")

[1352] Output: Analysis results (e.g. "Napoleon" "Tactics")

[1353] Step 3:

[1354] The server retrieves the information from the database

[1355] Based on the analysis results, the server searches for information about related historical figures from the information aggregation means. The information aggregation means stores a large amount of data and efficiently retrieves the required information.

[1356] Input: Analysis result (e.g. "Napoleon" "Tactics")

[1357] Output: relevant information (e.g. "Information about Napoleon's tactics")

[1358] Step 4:

[1359] The server inputs information into the generated AI model

[1360] The server inputs the relevant information retrieved from the database and the user's question into the generative AI model means, which then generates the optimal answer based on this information.

[1361] Input: Relevant information and user questions (e.g., "Information about Napoleon's tactics" and "Tell me about your tactics")

[1362] Output: Generated answer (e.g., "My most important tactics are rapid advance and attacking the enemy's weak points.")

[1363] Step 5:

[1364] The server generates a response and sends it to the device.

[1365] The generated answer is returned to the server, which then transmits it to the terminal.

[1366] Input: Generated answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[1367] Output: The submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[1368] Step 6:

[1369] The device will lip sync and display the answer

[1370] The device displays the answer received from the server to the user. If the device includes a lip-sync function, it performs lip-sync processing on the image of the person and reproduces natural mouth movements using voice synthesis.

[1371] Input: Submitted answer (e.g., "My most important tactic is to advance quickly and attack the enemy's weak points.")

[1372] Output: Displayed and voice-synthesized answer (e.g., "My most important tactical priorities are rapid advance and striking the enemy's weak points.")

[1373] Step 7:

[1374] User observations and follow-up questions

[1375] The user observes the displayed answers, and if they have further questions, they again use the terminal interface to enter and submit new questions, and the process repeats.

[1376] Input: Follow-up question (e.g., "Can you give us an example of a specific tactic?")

[1377] Output: New question data (e.g., "Can you give us an example of a specific tactic?")

[1378] (Application example 1)

[1379] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1380] In today's factories, managers and workers are required to have a high level of specialized knowledge and experience in order to carry out efficient work management and training. However, this knowledge depends on the experience of each individual, making it difficult to achieve consistent and efficient work management. Another problem is that training and nurturing new workers takes a lot of time and money.

[1381] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1382] In this invention, the server includes a database for managing data on historical figures, an interface for receiving user input, a communication unit for analyzing user requests and retrieving information about the specified historical figures from the database, a generation AI model for providing generated answers, and an output unit for transmitting the generated answers to the user. This allows factory managers and workers to apply historical knowledge and strategies to improve work efficiency and quickly train new workers. Furthermore, the voice synthesis and lip-sync functions allow users to enjoy a more realistic experience.

[1383] The "database means" is a system or device that centrally manages data related to historical figures and quickly provides relevant data in response to a user's request.

[1384] "Interface means" refers to an input device or interface that allows a user to input questions or requests to the system, and has the function of accepting user input.

[1385] "Communication means" refers to a system or device with communication capabilities that analyzes requests from users, retrieves the necessary data from a database, and passes it to the generative AI model.

[1386] "Generative AI model means" refers to an artificial intelligence model that generates an appropriate answer based on information obtained from a database in response to a question from a user.

[1387] "Output means" refers to a system or device for transmitting and displaying the answer generated by the generative AI model to the user.

[1388] The "lip sync means" is a system or device that has the function of synthesizing the generated answer into voice and then performing lip sync processing on a photograph of a historical figure's face.

[1389] "Factory training and support measures" refers to the functions and equipment that support the effective management of machines and systems used within a factory and the training of workers.

[1390] This invention is a system that applies the tactics and knowledge of historical figures to factory robots, enabling factory management and work support. The system includes database means, interface means, communication means, generative AI model means, output means, lip-sync means, and factory training and support means.

[1391] System Configuration

[1392] server

[1393] The server is the core of this system and includes a database, communication, AI model generation, and output functions. The server receives requests from users, analyzes them, generates appropriate responses, and sends them to the users.

[1394] Terminal

[1395] The terminal provides an interface for the user to interact with the system, accepts user input, transmits it to the server, and displays responses received from the server. It also includes functionality for displaying a lip-synced person's face.

[1396] User

[1397] The user operates the interface means to input a question and receive a response from the system, thereby obtaining advice and knowledge from historical figures.

[1398] Hardware

[1399] Robot: An industrial machine used in a factory (e.g., an industrial robot).

[1400] Server: Runs on a cloud service (e.g. AWS EC2).

[1401] Camera and microphone: Sensor devices built into the robot.

[1402] Display: A display device that allows the user to see the content of the interaction.

[1403] software

[1404] Natural Language Processing (NLP) library: OpenAI's GPT-4.

[1405] Speech synthesis library: Google Text-to-Speech.

[1406] Lip sync library: Avatarify.

[1407] Database: MySQL.

[1408] Program Processing Overview

[1409] Processing on the terminal

[1410] The device provides an interface for users to input questions for historical figures. When a user inputs and submits a question, the data is sent to the server. Furthermore, if the device includes a lip-sync function, it has the function of displaying a lip-synced photograph of the historical figure.

[1411] Examples:

[1412] The user types "Tell me about your tactics" into the device interface and clicks the send button.

[1413] Processing on the server

[1414] The server receives the request from the device and analyzes its contents. Based on the analysis results, it retrieves information about the requested historical figure from a database. The retrieved information and the user's question are input into a generative AI model to generate an appropriate answer. The generated answer is synthesized into voice, lip-synced, and sent to the device.

[1415] Examples:

[1416] The server retrieves information about a specified historical figure and tactics from a database, inputs it into a generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points."

[1417] Lip Sync Means

[1418] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them to the facial photographs of historical figures, providing a more realistic conversation experience.

[1419] Examples:

[1420] A lip sync means generates and displays in real time mouth movements based on the generated answers for the photographs of the faces of historical figures.

[1421] Prompt Sentence Examples

[1422] "As the designated historical figure, answer this question: Tell us about your tactics."

[1423] "As a designated historical figure, please assess the progress of the production line this week."

[1424] This system allows users to receive factory management and work support that applies historical knowledge, enabling efficient work management and rapid training.

[1425] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1426] Step 1:

[1427] The user enters a question into the terminal interface and clicks the send button. The entered question is saved on the terminal and then sent to the server.

[1428] Specific operation: The user enters "Tell us about your tactics" into the input field on the device and presses the send button. The input data is sent from the device to the server.

[1429] Step 2:

[1430] The server receives and analyzes a request (question) from the user. Based on the analysis results, it retrieves information about the specified historical figure from the database.

[1431] Input: User question data

[1432] Output: Database query results

[1433] Specific operation: The server searches the database for information about the "specified historical figure" and "tactic" and retrieves related information.

[1434] Step 3:

[1435] The server inputs the acquired information and the user's question into a generative AI model to generate an appropriate answer.

[1436] Input: Information about historical figures retrieved from the database, user questions

[1437] Output: The generated answer

[1438] How it works: The server inputs a prompt to a generative AI model (e.g., OpenAI GPT-4) saying, "As a specified historical figure, tell me about a tactical battle," and receives the generated answer.

[1439] Step 4:

[1440] The generated answer is synthesized into voice.

[1441] Input: Generated answer text

[1442] Output: Audio file

[1443] Specific behavior: The server uses a speech synthesis library (e.g., Google Text-to-Speech) to convert the generated answer text into an audio file.

[1444] Step 5:

[1445] Lip-sync processing is performed using the generated answer audio file.

[1446] Input: Photo of a historical figure, audio file

[1447] Output: Lip-synced video

[1448] What happens: The server uses a lip-sync library (e.g., Avatarify) to generate audio-synced mouth movements for a photograph of a historical figure, creating a lip-synced video.

[1449] Step 6:

[1450] The server sends the lip-synced video to the terminal.

[1451] Input: Lip-synced video

[1452] Output: The video displayed on the user's device

[1453] Specific operation: The server sends lip-synced video data to the device, which then displays the video to the user.

[1454] Step 7:

[1455] Users watch lip-synced videos on their devices and receive answers from historical figures.

[1456] Input: Lip-synced video

[1457] Output: User understanding and real-time interaction experience

[1458] What happens: The user watches a lip-synced video of a historical figure on the device display and confirms their answers.

[1459] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1460] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[1461] System Components

[1462] server

[1463] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[1464] Terminal

[1465] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[1466] User

[1467] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[1468] Program Processing Overview

[1469] Terminal

[1470] The device provides an interface for users to input questions for historical figures. When a user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If a lip-sync function is included, the device displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[1471] Examples:

[1472] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[1473] server

[1474] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[1475] Examples:

[1476] The server retrieves information about "Napoleon" and "tactics" from the database, inputs it into the generative AI model, and generates an answer such as, "The tactics I value most are rapid advancement and attacking the enemy's weak points." The emotion engine analyzes the user's interested facial expression and adds a slightly more detailed explanation to the answer.

[1477] Lip Sync Means

[1478] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[1479] Examples:

[1480] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[1481] Simulation Method

[1482] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[1483] Examples:

[1484] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[1485] Emotion Engine

[1486] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[1487] Examples:

[1488] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[1489] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[1490] The processing flow will be explained below.

[1491] Specific processing flow of the program (including the emotion engine)

[1492] Step 1:

[1493] The user enters a question into the terminal interface and clicks the send button.

[1494] Step 2:

[1495] The device formats the user's question as an API request and sends it to the server, and also captures the user's facial expressions and voice and sends them to the server as emotion data.

[1496] Step 3:

[1497] The server receives the API request, analyzes the request data, and extracts the question and the name of the historical figure.

[1498] Step 4:

[1499] The server searches and acquires information about the designated historical person from the database means.

[1500] Step 5:

[1501] The server inputs the acquired personal information and the user's question into a generation AI model means to generate an appropriate answer.

[1502] Step 6:

[1503] The emotion engine analyzes the user's emotion data transmitted from the terminal and determines the user's emotional state.

[1504] Examples:

[1505] The emotion engine analyzes the user's facial expressions and determines whether the user is interested.

[1506] Step 7:

[1507] The sentiment engine adjusts the generated answer based on the determined user sentiment, for example adding detailed explanations to the answer if the user indicates interest.

[1508] Examples:

[1509] "My most important tactic is to advance quickly and attack the enemy's weak points," one generated answer might explain, with the further detail, "This tactic has won me many battles."

[1510] Step 8:

[1511] The server formats the generated answer and creates response data.

[1512] Step 9:

[1513] The server sends the response data to the terminal.

[1514] Step 10:

[1515] The terminal receives the response data from the server and formats it for display to the user.

[1516] Step 11:

[1517] The lip-syncing means synthesizes the generated responses into voice and lip-syncs them to the portraits of historical figures, providing a more realistic dialogue experience.

[1518] Step 12:

[1519] The terminal displays the answer to the user and plays the video with lip-sync processing completed.

[1520] Step 13:

[1521] The user reads the displayed answers and, if necessary, enters the next question or ends the dialogue.

[1522] This process allows users to simulate real-time interactions with historical figures and receive personalized, emotion-based responses to enhance their learning experience.

[1523] Example 2

[1524] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1525] Conventional dialogue systems only provide simple text-based answers when users seek information about historical figures, resulting in a lack of realism and immersion. Furthermore, they are unable to provide personalized answers based on the user's emotions, preventing an improved learning experience. Furthermore, they are unable to simulate conversations between different historical figures, limiting opportunities to learn history from new perspectives.

[1526] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1527] In this invention, the server includes a database means for managing data on historical figures, a communication means for analyzing a request from a user and obtaining information on the specified historical figure from the database means, a generative AI model means for providing a generated answer, an emotion engine for recognizing emotions by analyzing the user's facial expressions and voice, and a means for adjusting the tone and content of the generated answer based on the recognized emotion. This allows the user to interact with the historical figure in real time and receive personalized answers based on their emotions, enabling a deeper learning experience.

[1528] The "database means" is a storage or management system for managing data relating to historical figures and retrieving information as needed.

[1529] An "interface means" is an input device or software interface for accepting input from a user.

[1530] "Communication means" refers to a network device or communication protocol for transmitting user input data to a server and transmitting data from the server to the user.

[1531] The "generative AI model means" is an artificial intelligence model or natural language processing system that analyzes a user's question and generates an appropriate answer.

[1532] The "output means" is a display device or an audio output device for displaying or reproducing the generated answer to the user.

[1533] An "emotion engine" is an analysis device or software engine that analyzes a user's facial expressions and voice data to recognize the user's emotions.

[1534] "Lip sync means" refers to a processing device or software that synthesizes the generated response into voice and synchronizes the lip movements with the image of the historical figure.

[1535] A "simulation means" is a system for arranging a dialogue between different historical figures and reproducing or simulating the development of that dialogue.

[1536] A "user" is someone who operates the system and enters questions to obtain information about historical figures.

[1537] The "server" is the core of the system, and is a computer system that controls the database means, generative AI model means, emotion engine, etc., and manages and executes the overall processing.

[1538] The present invention relates to a system that manages data on historical figures, allows a user to have a conversation as if they were that person, and is equipped with an emotion engine that recognizes the user's emotions and adjusts responses based on those emotions. The system includes a database means, an interface means, a communication means, a generative AI model means, an output means, a lip-sync means, a simulation means, and an emotion engine.

[1539] server

[1540] The server is the core of the system and includes a database, communication, generative AI model, output, and emotion engine. The server receives user requests, analyzes them, generates appropriate responses, and sends them to the user. The emotion engine also analyzes the user's emotions and adjusts the responses accordingly.

[1541] Terminal

[1542] The terminal provides an interface for the user to interact with the system. The terminal accepts user input, sends it to the server, and displays the response received from the server to the user. It also has the functionality to capture the user's facial expressions and voice to analyze their emotions.

[1543] User

[1544] Through this system, users can interact with historical figures by operating their devices, inputting questions, and reading or listening to the answers provided by the server. The system can then reflect the user's emotions, enabling more personalized interactions.

[1545] Program Processing Overview

[1546] server

[1547] The server receives a request from the terminal and analyzes its contents. Based on the analysis results, information about the requested historical figure is retrieved from the database means. The retrieved information and the user's question are input into the generation AI model means, and the AI ​​model generates an appropriate answer. Furthermore, the emotion engine analyzes the user's emotion data and adjusts the tone and content of the answer according to the user's interests. The generated answer is sent to the terminal through the output means.

[1548] Terminal

[1549] The device provides an interface for users to input questions for historical figures. When the user inputs a question and presses the send button, the data is sent to the server. It also provides a screen for displaying the answer to the user. If the device includes a lip-sync function, it displays a photo of the person's face after lip-syncing. The device also captures the user's facial expressions and voice and sends them to the server as emotion data.

[1550] Examples:

[1551] A user types "Napoleon, tell me about tactics" into the device's interface and clicks the send button. The device captures the user's facial expression and determines that the user looks interested.

[1552] Lip Sync Means

[1553] The lip-syncing means synthesizes the generated answers into voice and lip-syncs them with the facial photographs of historical figures, providing a more realistic conversation experience.

[1554] Examples:

[1555] A lip sync means generates and displays in real time mouth movements based on the answers to a photograph of Napoleon's face.

[1556] Simulation Method

[1557] The simulation means provides a function to have different historical figures interact with each other and replay the development of that interaction, allowing users to learn history from a new perspective.

[1558] Examples:

[1559] The simulation device recreates a conversation between Napoleon and Alexander the Great, with the conversation focusing on the topic of "the importance of speed in tactics."

[1560] Emotion Engine

[1561] The emotion engine analyzes the user's facial expressions and voice data sent from the device and recognizes the user's emotions. Based on the recognition results, it adjusts the tone and content of the responses generated by the generative AI model to personalize responses to the user.

[1562] Examples:

[1563] The emotion engine analyzes the user's smile and determines that the user is having fun, and based on this, adds humorous elements to the generated answers.

[1564] The system allows users to interact with historical figures in real time and receive personalized, emotion-based responses, enhancing their learning experience.

[1565] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1566] Step 1:

[1567] The user types a question on the device

[1568] Using the device's interface, the user enters a text question for the historical figure, such as "Napoleon, tell me about tactics." The user then clicks a submit button to complete the question.

[1569] Input: A question entered by the user as text

[1570] Output: The device sends the input data to the server

[1571] Step 2:

[1572] The device sends user input and emotion data to the server.

[1573] The device receives text data entered by the user and captures the user's facial expressions and voice using the built-in camera and microphone. The facial expression and voice data are analyzed and sent to the server as emotional data.

[1574] Input: User question, and captured facial expressions and voice

[1575] Output: Send all data to the server along with the analyzed emotion data

[1576] Step 3:

[1577] The server parses the request and retrieves information from the database

[1578] The server receives the request sent from the terminal and analyzes its contents. For example, it extracts keywords such as "Napoleon" and "tactics." Then, based on these keywords, it retrieves related information from the database means.

[1579] Input: Text data and emotion data sent from the device

[1580] Output: Historical information retrieved from the database

[1581] Step 4:

[1582] The server generates an answer using a generative AI model

[1583] The server inputs the information retrieved from the database and the user's question into a generative AI model (e.g., a natural language processing model). The generative AI model is used to generate an appropriate answer based on the context.

[1584] Input: Retrieved historical information and user question

[1585] Output: The answer generated by the generative AI model

[1586] Step 5:

[1587] The server uses an emotion engine to tailor the response

[1588] The server uses an emotion engine to tailor the generated answer based on the analyzed user's emotion data, for example adding more details and explanations to the answer if the user seems interested.

[1589] Input: Generated answers and sentiment data

[1590] Output: An answer tailored based on the sentiment

[1591] Step 6:

[1592] The server performs lip sync processing

[1593] The server uses lip-syncing to synthesize the generated responses and synchronize the mouth movements with the image of the historical figure, providing a realistic dialogue experience, as if Napoleon were actually speaking.

[1594] Input: Adjusted answer

[1595] Output: Lip-synced video and audio

[1596] Step 7:

[1597] The server sends the answer to the device

[1598] The server sends a response to the terminal indicating that the lip-sync process has been completed. The data sent includes text information, audio files, and lip-synced video.

[1599] Input: Lip-synced video and audio data

[1600] Output: Answer data sent to the device

[1601] Step 8:

[1602] The device displays the answer to the user

[1603] The device receives the answer data sent from the server and displays it to the user. The lip-synced faces of historical figures are displayed on the screen, and the answers are played back aloud.

[1604] Input: Response data sent from the server

[1605] Output: Answers shown to the user (text, audio, lip-sync video)

[1606] (Application example 2)

[1607] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1608] Conventional educational systems and content distribution services tend to provide users with a passive experience when learning history. Furthermore, one-way information provision without considering the user's emotions makes it difficult to maintain interest in learning. Furthermore, interactive systems lack the technology to provide a realistic dialogue experience, and the depictions of important historical figures are static. Therefore, there is a need for the development of a system that allows users to experience interactive dialogue with historical figures who actually have emotions while learning history.

[1609] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes database means for managing data on historical figures, interface means for accepting input from a user, means for capturing the user's facial expressions and voice and analyzing their emotions, communication means for analyzing the user's request and obtaining information on the specified historical figure from the database means, generation AI model means for adjusting the generated answer based on the user's emotions, output means for sending the generated answer to the user, and rendering means for displaying the lip-synced answer as a 3D image. This allows the user to interact with the historical figure in real time and enjoy an interactive learning experience personalized based on their emotions.

[1610] "Database means" is a device that has the function of systematically managing and storing information about historical figures.

[1611] An "interface means" is a device that accepts input from a user and provides a screen and operating means for the user to interact with the system.

[1612] The "means for analyzing emotions" is a device that has the function of capturing the user's facial expressions and voice and analyzing them to identify the user's emotional state.

[1613] The "communication means" is a device that receives a request from a user, retrieves information from a database based on the request, and passes the result to the generation AI model means.

[1614] The "generative AI model means" is a device that has the function of generating an appropriate response based on a user request and adjusting the tone and content of the response based on the user's emotional data.

[1615] The "output means" is a device for transmitting the generated answer to the user, and has the function of displaying or audibly transmitting the answer to the user in real time.

[1616] The "rendering means" is a device that renders the lip-synced answer as a three-dimensional model in real time and provides it visually to the user.

[1617] The "lip sync means" is a device that synchronizes the mouth movements of a three-dimensional model with the generated voice, thereby providing a realistic dialogue experience.

[1618] The "simulation means" is a device that has the function of having different historical figures converse with each other and simulating the development of that conversation in real time.

[1619] This invention provides a system for providing users with information about historical figures and realizing interactive dialogue. The system includes a database, an interface, an emotion analysis, a communication, a generative AI model, an output, a lip-sync, a rendering, and a simulation.

[1620] System components and functions

[1621] 1. Database Means

[1622] The server is equipped with a database that systematically manages and stores data on historical figures. This database is constructed using a relational database management system (RDBMS) such as MySQL.

[1623] 2. Interface Method

[1624] Users interact with the system through devices such as smartphones and head-mounted displays (HMDs). The devices are equipped with a user interface (UI) built using Unity, which allows users to enter input.

[1625] 3. Emotion analysis method

[1626] The device has the ability to capture the user's facial expressions and voice and analyze their emotions using APIs from Google ML Kit and Amazon Rekognition, making it possible to analyze the user's emotional state in real time.

[1627] 4. Means of communication

[1628] A request from a user is sent to the server using a RESTful API. The communication means has the function to analyze this request and retrieve the necessary information from the database.

[1629] 5. Generative AI Model Means

[1630] The generative AI model means installed on the server generates appropriate answers to questions using a natural language processing model such as OpenAI, and has the function of adjusting the tone and content of the answers based on user sentiment analysis data.

[1631] 6. Output Method

[1632] The generated answers are sent from the server to the device and provided to the user, using voice synthesis technology such as Amazon Polly.

[1633] 7. Lip Sync Method

[1634] The server performs lip-sync processing based on the generated audio, which uses the lip-sync function of MetaHuman Creator to generate mouth movements for the 3D model.

[1635] 8. Rendering Method

[1636] Using Unity's 3D rendering engine, three-dimensional models that have undergone lip-sync processing are rendered in real time and provided to users.

[1637] 9. Simulation Methods

[1638] The server has the ability to have different historical figures interact with each other and simulate the progress of the interaction in real time, allowing users to learn about history from a new perspective.

[1639] Specific examples of processing

[1640] For example, suppose a user types "Tell me about your days as Napoleon" into the device's interface and presses the send button. At this time, the device's camera captures and analyzes the user's facial expression. The emotion analysis means recognizes that the user is interested. This request is then sent to the server via the communication means, and data about Napoleon is retrieved from the database.

[1641] Next, the generating AI model means generates an appropriate answer to "Life as Napoleon" and adjusts the answer taking into account the user's emotions. The server then synthesizes the generated answer into speech and generates mouth movements for the three-dimensional model of Napoleon using the lip-sync means. Finally, the rendering means renders the three-dimensional model of Napoleon in real time and displays it on the device.

[1642] Prompt Sentence Examples

[1643] "Tell me about your days as Napoleon."

[1644] "Alexander the Great, what is your most amazing tactic?"

[1645] Using this system, users can interact with historical figures in real time and enjoy a personalized, interactive learning experience that responds to their emotions.

[1646] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1647] Step 1: The user enters a question for a historical figure through the device interface and presses the send button. At this time, the interface captures the text of the user's question and uses the camera and microphone to collect the user's facial expressions and voice in real time. The input is the user's question text, facial expressions, and voice data, and the output is a process that compiles this data into a single request.

[1648] Step 2: The device uses Google ML Kit and Amazon Rekognition API on the edge device to analyze the collected facial and voice data in real time to identify the user's emotional state. The input is the user's facial and voice data, and the output is the analyzed emotional data.

[1649] Step 3: The device sends a request containing the user's question text and the parsed emotion data to the server via a RESTful API. The input is the request containing the question text and emotion data, and the output is the request data sent to the server.

[1650] Step 4: The server receives the request and retrieves information about the specified historical figure using a database (e.g., MySQL). The input is the request data from the user (question text and emotion data), and the output is the historical information retrieved from the database.

[1651] Step 5: The server inputs the acquired historical information into a generative AI model (such as OpenAI) to generate an answer to the user's question. The generated answer is also adjusted based on the emotional data. The input is the historical information acquired from the database and the user's emotional data, and the output is an emotion-adjusted answer text.

[1652] Step 6: The adjusted answer text is converted into speech data using a speech synthesis means (e.g., Amazon Polly). During this process, the generated speech data is sent to a lip-sync means (e.g., MetaHuman Creator) to generate mouth movements corresponding to a three-dimensional model of the historical figure. The input is the emotion-adjusted answer text, and the output is speech data and lip-sync data.

[1653] Step 7: The server renders the audio data and lip-sync data in real time using a rendering engine (such as Unity's 3D rendering engine) and sends it to the device. The input is the audio data and lip-sync data, and the output is the rendered 3D model.

[1654] Step 8: The terminal displays the rendered 3D model to the user and plays the audio data. This allows the user to interact with the historical figure in real time and enjoy personalized responses based on their emotions. The input is the rendered 3D model and audio data from the server, and the output is audiovisual information provided to the user.

[1655] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1656] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1657] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1658] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1659] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1660] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1661] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1662] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1663] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1664] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1665] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1666] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1667] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1668] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1669] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1670] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1671] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1672] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1673] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1674] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1675] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1676] The following is further disclosed regarding the above embodiment.

[1677] (Claim 1)

[1678] a database means for managing data relating to historical figures;

[1679] an interface means for accepting input from a user;

[1680] a communication means for analyzing a request from a user and retrieving information about the specified historical person from a database means;

[1681] a generative AI model means for providing a generated answer;

[1682] output means for transmitting the generated answer to the user;

[1683] A system including:

[1684] (Claim 2)

[1685] 2. The system according to claim 1, further comprising lip-sync means for performing lip-sync processing on a photograph of a face of a historical figure based on the generated answer.

[1686] (Claim 3)

[1687] 2. The system according to claim 1, further comprising simulation means for simulating the development of a dialogue between different historical figures.

[1688] "Example 1"

[1689] (Claim 1)

[1690] an information aggregation means for managing data on historical figures;

[1691] an input means for accepting input from a user;

[1692] a receiving means for analyzing a request from a user and obtaining information about a designated historical figure from an information aggregation means;

[1693] a generative AI model means for providing a generated answer;

[1694] a transmitting means for transmitting the generated answer to the user;

[1695] A system including:

[1696] (Claim 2)

[1697] 2. The system of claim 1, further comprising lip-sync means for performing lip-sync processing on an image of the historical figure based on the generated answer.

[1698] (Claim 3)

[1699] 2. The system according to claim 1, further comprising simulation means for simulating the development of a dialogue between different historical figures.

[1700] "Application Example 1"

[1701] (Claim 1)

[1702] a database means for managing data relating to historical figures;

[1703] an interface means for accepting input from a user;

[1704] a communication means for analyzing a request from a user and retrieving information about the specified historical person from a database means;

[1705] a generative AI model means for providing a generated answer;

[1706] output means for transmitting the generated answer to the user;

[1707] lip-sync means for performing voice synthesis on the generated answer and lip-sync processing;

[1708] means including factory training and assistance means for managing industrial machinery;

[1709] A system including:

[1710] (Claim 2)

[1711] 10. The system of claim 1, including factory training and support means for managing industrial machinery.

[1712] (Claim 3)

[1713] 2. The system according to claim 1, further comprising simulation means for simulating the development of a dialogue between different historical figures.

[1714] "Example 2: Combining Emotion Engines"

[1715] (Claim 1)

[1716] a database means for managing data relating to historical figures;

[1717] an interface means for accepting input from a user;

[1718] a communication means for analyzing a request from a user and retrieving information about the specified historical person from a database means;

[1719] a generative AI model means for providing a generated answer;

[1720] output means for transmitting the generated answer to the user;

[1721] An emotion engine that analyzes the user's facial expressions and voice to recognize their emotions;

[1722] a means for adjusting the tone and content of generated responses based on perceived sentiment;

[1723] A system including:

[1724] (Claim 2)

[1725] 2. The system of claim 1, further comprising lip-sync means for performing lip-sync processing on an image of the historical figure based on the generated answer.

[1726] (Claim 3)

[1727] 2. The system according to claim 1, further comprising simulation means for simulating the development of a dialogue between different historical figures.

[1728] "Application example 2 when combining emotion engines"

[1729] (Claim 1)

[1730] a database means for managing data relating to historical figures;

[1731] an interface means for accepting input from a user;

[1732] A means of capturing the user's facial expressions and voice and analyzing their emotions;

[1733] a communication means for analyzing a request from a user and retrieving information about a designated historical figure from a database means;

[1734] a generative AI model means for adjusting the generated answer based on the user's emotion;

[1735] output means for transmitting the generated answer to the user;

[1736] a rendering means for displaying the lip-synced answer as a stereoscopic image;

[1737] A system including:

[1738] (Claim 2)

[1739] 2. The system of claim 1, further comprising lip-sync means for performing lip-sync processing on a three-dimensional model of the historical figure based on the generated answer.

[1740] (Claim 3)

[1741] 2. The system according to claim 1, further comprising simulation means for simulating the development of a dialogue between different historical figures. [Explanation of symbols]

[1742] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a database means for managing data relating to historical figures; an interface means for accepting input from a user; a communication means for analyzing a request from a user and retrieving information about the specified historical person from a database means; a generative AI model means for providing a generated answer; output means for transmitting the generated answer to the user; A system including:

2. 2. The system according to claim 1, further comprising lip synchronization means for performing lip synchronization processing on a photograph of a face of a historical person based on the generated answer.

3. 2. The system according to claim 1, further comprising simulation means for simulating the development of a dialogue between different historical figures.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A