System

The system addresses the challenge of creating a virtual personality that reflects deceased individuals by analyzing user data to generate and adapt a virtual persona, ensuring high-quality conversations and emotional support.

JP2026034159APending Publication Date: 2026-02-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024137280
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Conventional technologies struggle to generate a virtual personality that accurately reflects the deceased's behavior and characteristics, leading to poor conversation quality and difficulty in adapting to user feedback.

Method used

A system that collects user data, analyzes it to extract features, generates a virtual personality, and adapts based on user feedback, using natural language processing and machine learning to mimic the user's behavior and personality, with speech recognition and synthesis for interaction.

Benefits of technology

Enables natural interaction with virtual personalities of deceased loved ones, providing emotional support by accurately mimicking their behavior and adapting to user feedback for improved conversation quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026034159000001_ABST
    Figure 2026034159000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting recorded data of a user; means for analyzing the collected recorded data and extracting features; means for generating a virtual personality based on the extracted features; means for interacting the virtual personality with the user; and means for receiving feedback from the user during the interaction and adapting the virtual personality.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The desire to continue communicating with loved ones after their death is an emotional challenge faced by many people. However, conventional technologies have difficulty generating a virtual personality that accurately reflects the deceased's behavior and characteristics and conducting natural conversations based on that personality. In particular, many technical challenges remain regarding how to analyze collected data, reflect it in the virtual personality, and even adapt it during the conversation. [Means for solving the problem]

[0005] The present invention provides a means for collecting a user's recorded data, analyzing the data, and extracting features. It then provides a means for generating a virtual personality based on the extracted features and having the virtual personality interact with the user. It also provides a system including a means for receiving feedback from the user during the interaction and adapting the virtual personality based on the feedback. The system further includes a means for collecting the user's voice, converting the voice data into text data on a server, generating a response based on the text data, converting the response into speech using a speech synthesizer, and playing the synthesized speech to the user. It also includes a means for providing an interface for the user to provide consent and for starting data collection based on the user's consent. This makes it possible to realize a system that allows users to interact naturally with the virtual personalities of deceased loved ones or loved ones.

[0006] A "user" is an entity that uses the service and interacts with a virtual personality.

[0007] "Recorded data" refers to all data that can be analyzed, such as user social media posts, call history, and search history.

[0008] "Means for collecting" refers to a method or device for acquiring user record data and transmitting it to the server.

[0009] "Means for analyzing" refers to a method or device for extracting characteristics from collected recorded data and obtaining information for generating a virtual personality.

[0010] "Characteristics" are elements necessary to construct a virtual personality, such as the user's way of speaking, behavioral patterns, and hobbies.

[0011] A "virtual personality" is a virtual entity that imitates the characteristics of the user and is generated based on analyzed features.

[0012] "Means for creating a dialogue" refers to a method or device for realizing a conversation between a user and a virtual personality.

[0013] "Feedback" refers to opinions and requests provided by users during a dialogue.

[0014] "Adaptation means" means a method or device for updating and adjusting the virtual personality based on user feedback.

[0015] "Voice data" refers to data that represents in digital form the voice uttered by the user.

[0016] "Text data" refers to data obtained by converting voice data into character information.

[0017] The "means for generating a response" refers to a method or device for creating an appropriate response based on text data.

[0018] A "speech synthesizer" is a device or software that converts text data into speech data.

[0019] "Speech" is the process of converting generated text data into a form that can be reproduced as audio data.

[0020] An "interface" is the means by which a consumer provides consent to a service.

[0021] A "means for initiating data collection" is a method or device that initiates the process of collecting recorded data based on the user's consent. [Brief explanation of the drawings]

[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4]FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0024] First, the terms used in the following description will be explained.

[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0030] [First embodiment]

[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0043] To practice the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[0044] First, regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[0045] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[0046] During the conversation simulation phase, the user interacts with the virtual personality via their device. The device captures the user's speech as voice data and transmits it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using speech synthesis technology and played back to the user via their device.

[0047] Feedback and Adaptation: Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that conversation," they can send them to the server through the feedback function. The server analyzes this feedback and adapts it to the AI ​​model, making future conversations more natural and satisfying.

[0048] Explanation with concrete examples

[0049] For example, suppose the family of a deceased person, Mr. B, wants to use this system.

[0050] 1. Data Collection

[0051] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[0052] A dedicated application installed on the device collects this data and sends it to a server.

[0053] The server stores the received data in a database.

[0054] 2. Data analysis and model generation

[0055] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[0056] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[0057] 3. Conversation Simulation

[0058] The user (B's family member) talks to the virtual B through the terminal.

[0059] The device captures the audio and sends it to the server.

[0060] The server converts the speech to text and uses an AI model to generate an appropriate response.

[0061] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[0062] The device plays back the response of virtual person B.

[0063] 4. Feedback and Adaptation

[0064] The user provides feedback on the interaction with virtual Person B.

[0065] The server analyzes the feedback and adjusts the AI ​​model.

[0066] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[0067] The processing flow will be explained below.

[0068] Step 1:

[0069] User: Agrees to use the service and agrees to provide the required data.

[0070] Device: Install the dedicated application and configure it to allow access to social media accounts, call history, and search history.

[0071] Step 2:

[0072] Device: Collects data such as users' social media posts, call history, and search history, and sends it to a server.

[0073] Server: Receives collected data and stores it securely in a database.

[0074] Step 3:

[0075] Server: Runs long-text data analysis algorithms to analyze the collected data and extract users' language usage patterns, characteristic phrases, and behavioral patterns.

[0076] Step 4:

[0077] Server: Generates an AI model based on the extracted information. This model creates a virtual personality that mimics the user's behavior and characteristics.

[0078] Step 5:

[0079] User: Uses the device to initiate a conversation with the virtual personality.

[0080] Device: The user's voice is captured by a microphone and sent to the server in real time.

[0081] Step 6:

[0082] Server: Converts the received voice data into text data using voice recognition technology.

[0083] Server: Inputs the converted text data into the AI ​​model and generates an appropriate response.

[0084] Server: The generated response is converted into voice data using speech synthesis technology, and then matched to the voice specified by the user using a voice changer.

[0085] Step 7:

[0086] Server: Sends the synthesized voice data to the device.

[0087] Terminal: The synthesized voice data is played back to the user through a speaker.

[0088] Step 8:

[0089] Users: Provide feedback during a conversation, for example by sending requests such as "I'd like more humor" or "I'd like more detail on a particular topic" through the feedback feature.

[0090] Server: Analyzes the received feedback and updates and adapts the AI ​​model to make future conversations more natural and satisfying.

[0091] Example 1

[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0093] In modern society, there is a demand for technology that can recreate conversations with deceased loved ones and provide emotional support. In particular, conventional systems have had problems with insufficient analysis of collected data, resulting in poor conversation quality. Furthermore, it has been difficult to quickly and appropriately incorporate user feedback. This has led to issues such as reduced user satisfaction.

[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0095] In this invention, the server includes: [means for collecting user record data and securely transmitting it to the server;] [means for analyzing the collected record data and extracting features using natural language processing technology; and] [means for generating a virtual personality using machine learning based on the extracted features.] This makes it possible [to accurately mimic the user's behavior and personality, achieve high-quality dialogue, and quickly and appropriately reflect feedback].

[0096] "Means for collecting user record data and safely transmitting it to a server" refers to a mechanism for collecting data such as users' social media accounts, call history, and search history, and transferring it to a server using security technologies such as encryption.

[0097] "Means of analyzing collected recorded data and extracting features using natural language processing technology" refers to a method of analyzing data received by the server and extracting phrases, behavioral patterns, interests, etc. from large amounts of text data.

[0098] "Means for generating a virtual personality using machine learning based on extracted features" refers to the process of using a machine learning algorithm to generate a virtual person that mimics the user's behavior and personality based on the analyzed data.

[0099] The "means for implementing speech recognition and speech synthesis technology to enable the generated virtual personality to interact with the user" refers to a series of processes for converting the user's speech input into text and converting the generated virtual personality's responses back into speech to provide to the user.

[0100] "Means for receiving feedback from users during a dialogue, analyzing it, and adapting the virtual personality" refers to a mechanism that collects opinions and requests provided by users during a dialogue, and uses that data to improve and adapt the behavior and comments of the generated virtual personality.

[0101] "Means for collecting user voice and transmitting it to a server in real time" refers to a method for capturing voice data in real time using a device such as a microphone and immediately transmitting it to a server.

[0102] "Means for using speech recognition technology to convert voice data into text data at the server" refers to technology that uses a speech recognition algorithm to accurately convert voice data into text data.

[0103] "Means of generating a response using an AI model based on text data and converting that response into voice using speech synthesis technology" refers to the process of converting a text response generated by an AI model into voice data using a speech synthesis algorithm.

[0104] "Means for playing the synthesized voice to the user through the terminal" refers to a method for delivering the generated voice data to the user using the terminal's speaker or the like.

[0105] "Means for providing an interface for users to provide consent to data provision" refers to a mechanism that provides a screen or interactive procedure for users to agree to data provision.

[0106] "Means for automatically starting data collection upon receiving user consent" refers to the process by which the system automatically starts collecting the necessary data after the user's consent is obtained.

[0107] MODE FOR CARRYING OUT THE INVENTION

[0108] To implement the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[0109] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is encrypted and securely sent from the device to the server, where it is stored in a database.

[0110] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technologies such as Google® Speech-to-Text and Amazon Polly are used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[0111] During the conversation simulation phase, the user interacts with the virtual persona via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using voice synthesis technology such as Amazon Polly and played back to the user via the device.

[0112] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that story," they can send them to the server through the feedback function. The server then analyzes this feedback and adapts the AI ​​model, making future conversations more natural and satisfying.

[0113] Explanation with concrete examples

[0114] For example, suppose the family of a deceased person, Mr. B, wants to use this system:

[0115] 1. Data Collection

[0116] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[0117] A dedicated application installed on the device collects this data and sends it to a server.

[0118] The server stores the received data in a database.

[0119] 2. Data analysis and model generation

[0120] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[0121] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[0122] 3. Conversation Simulation

[0123] The user (B's family member) talks to the virtual B through the terminal.

[0124] The device captures the audio and sends it to the server.

[0125] The server converts the speech to text and uses an AI model to generate an appropriate response.

[0126] The generated response is converted into voice data using voice synthesis technology such as Amazon Polly and sent to the device.

[0127] The device plays back the response of virtual person B.

[0128] 4. Feedback and Adaptation

[0129] The user provides feedback on the interaction with virtual Person B.

[0130] The server analyzes the feedback and adjusts the AI ​​model.

[0131] Prompt Sentence Examples

[0132] "Person B often uses the phrase 'do your best'. Please generate words of encouragement in a tone that is characteristic of Person B."

[0133] "I'd like you to tell me about a movie that Mr. B liked."

[0134] "Tell me how Mr. B tells jokes."

[0135] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[0136] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0137] Step 1: Data collection

[0138] The user agrees to use the service and consents to the provision of data such as social media accounts, call history, search history, etc. User consent and permission are required as input.

[0139] A dedicated application installed on the device acquires the data that can be collected. In operation, the application periodically and automatically collects posts from social media accounts, call history, and user search history.

[0140] The device securely encrypts the collected data and sends it to the server, resulting in the encrypted data being sent to the server.

[0141] The server stores the received data in a database. In operation, the server appropriately classifies the received data and saves it in a database.

[0142] Step 2: Data analysis and model generation

[0143] The server analyzes the data stored in the database, using saved user social media posts, call history, search history, etc. as input.

[0144] The server uses natural language processing technology to extract features such as the user's phrasing and behavioral patterns. In operation, the NLP algorithm extracts the user's unique expressions and interests from the data. The extracted feature data is obtained as the output.

[0145] The server uses machine learning algorithms to generate an AI model based on the extracted features. The AI ​​model is built using tools such as Python and Tensorflow®. The output is a virtual personality model that mimics the user's behavior.

[0146] Step 3: Conversation simulation

[0147] The user initiates a dialogue with the virtual persona through a terminal, and the user's voice input is required.

[0148] The device captures the user's speech as voice data and sends it to the server in real time. The device's microphone collects the voice and uses a voice recognition API to transfer the data to the server. The output is real-time voice data sent to the server.

[0149] The server converts the voice data into text data and generates a response using an AI model. The operation uses the Google Speech-to-Text API to convert the voice into text, and then uses an AI model (e.g., GPT-3 (registered trademark)) to generate an appropriate response. The output is the generated text response.

[0150] The text response generated by the server is converted into voice data using speech synthesis technology and sent to the device. The text is converted into voice using speech synthesis technology such as Amazon Polly and returned to the device. The voice data is then sent to the device as an output.

[0151] The device plays back the virtual persona's response in real time. In operation, the device's speaker plays back the received audio data. As an output, the virtual persona's response is provided to the user through audio.

[0152] Step 4: Feedback and Adapt

[0153] The user provides feedback on the interaction with the virtual personality. The input requires the user's feedback, which may include specific requests such as "Tell more jokes."

[0154] The server analyzes the feedback provided by the user and adjusts the AI ​​model. The operation is to analyze the feedback and incorporate it as new data into the training of the AI ​​model. The output is the adjusted AI model.

[0155] By repeating this series of processes, the system adapts its interactions with the user to make them more natural and satisfying.

[0156] (Application example 1)

[0157] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0158] In modern society, there is a growing demand for content distribution services that go beyond simply viewing video content such as movies and dramas to provide a more fulfilling entertainment experience. However, conventional systems lack a means to empathize with users' feelings of loneliness or the enjoyment of watching a movie. Furthermore, interactive means for providing emotional support are limited. There is a need to solve these issues and provide users with a more enjoyable and fulfilling viewing experience.

[0159] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0160] In this invention, the server includes means for collecting recorded data of a user, means for analyzing the collected recorded data and extracting features, means for generating a virtual personality based on the extracted features, means for the virtual personality to interact with the user, means for receiving feedback from the user during the interaction and adapting the virtual personality, and means for the user and the virtual personality to interact while playing video content. This allows the user to interact with a virtual partner while watching a movie or drama, making for a more fulfilling viewing experience.

[0161] "User" means an individual or organization that uses this system.

[0162] "Recorded data" refers to activity data such as a user's social media accounts, call history, and search history.

[0163] "Means of collection" refers to methods of collecting data using applications or sensors installed on the user's device.

[0164] "Means for analysis" refers to software or algorithms used to analyze the collected recorded data and extract features.

[0165] "Characteristics" are patterns that indicate tendencies in a user's behavior, personality, interests, etc.

[0166] A "virtual personality" is an AI model that is generated based on collected characteristics and allows users to interact with it.

[0167] "Means for dialogue" refers to a function that allows the user and the virtual personality to communicate via voice or text.

[0168] "Means for receiving feedback" and "adapting" refer to methods for updating the virtual personality generation model by reflecting opinions and requests from users.

[0169] "Video content" refers to media content that can be viewed, such as movies and dramas.

[0170] "Means for a user to have a conversation with a virtual character during playback" refers to a technology that allows a user to have a conversation with a virtual character in real time while viewing video content.

[0171] To implement this invention, the following system configuration and programs are required: This system is made up of several main components, each of which plays a specific role.

[0172] System Configuration

[0173] 1. User's device

[0174] Data Collection Applications

[0175] Video playback application

[0176] Audio input and output devices (microphone, speaker)

[0177] 2. Server

[0178] Database

[0179] Data Analysis Module

[0180] AI model generation module

[0181] Speech-to-text module (Google Speech-to-Text API)

[0182] Text-to-speech module (Google Text-to-Speech API)

[0183] WebSocket Server

[0184] Processing flow

[0185] 1. Data Collection

[0186] The user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, search history, etc. The device sends this data to the server, which then stores the collected data in a database.

[0187] 2. Data analysis and AI model generation

[0188] The server retrieves the collected data from the database and extracts features using natural language processing technology (TensorFlow, NLTK, spaCy). Based on these features, a virtual personality is generated that mimics the user's behavior and personality.

[0189] 3. Playback and interaction with video content

[0190] While a user is playing a movie or TV drama using a video playback application, a virtual persona will engage in conversation related to the video content. When the user speaks, the device captures the audio data and sends it to the server via WebSocket. The server converts the audio data into text and uses an AI model to generate an appropriate response. This response is then converted back into audio data by a text-to-speech module and played back to the user via the device.

[0191] 4. Feedback and Adaptation

[0192] Users can provide feedback to the virtual personality during the interaction, for example, if they want it to tell more jokes, they can send that feedback to the server through the feedback function, which will analyze this feedback and adapt the AI ​​model.

[0193] Specific examples

[0194] For example, suppose a user is watching the movie "Castle in the Sky," and the virtual partner imitates the personality of their deceased best friend, Mr. A.

[0195] Prompt Sentence Examples

[0196] The user is watching the movie "Castle in the Sky." The virtual partner is the deceased best friend, Mr. A:

[0197] User: "You also liked this scene, right?"

[0198] Virtual Partner: "Yes, the music in this scene is particularly memorable."

[0199] An example of a prompt sentence to input to the generative AI model is as follows:

[0200] User social media data:

[0201] Posted yesterday: Looking forward to seeing "Castle in the Sky"!

[0202] Call History:

[0203] My best friend A and I often talked about movies.

[0204] Search History:

[0205] A moving scene from "Castle in the Sky"

[0206] feedback:

[0207] I'd like to hear more detailed feedback

[0208] This system allows users to not only watch video content, but also enjoy interacting with virtual characters, further enhancing the viewing experience.

[0209] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0210] Step 1:

[0211] A user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, and search history. Based on this consent, the device automatically collects this data and sends it to a database. The input is the user's social media accounts, call history, and search history. The output is the data securely sent to the server.

[0212] Step 2:

[0213] The server retrieves the collected data from the database and extracts features using natural language processing techniques (TensorFlow, NLTK, spaCy). Specifically, the data is first cleansed, and then a language model is applied to identify interest and behavioral patterns. The input is the user's recorded data retrieved from the database. The output is analyzed data containing the extracted features.

[0214] Step 3:

[0215] The server builds an AI model for generating a virtual personality based on the extracted features. Specifically, it trains a neural network model that mimics the user's behavior and personality based on the feature data. The input is the analysis data obtained in step 2. The output is an AI model that represents the virtual personality.

[0216] Step 4:

[0217] While a user plays a movie or drama using a video playback application, the virtual persona interacts with the user about the video content being viewed. Specifically, the voice input is captured and the voice data is sent to the server. The input is the user's voice data. The output is the data sent to the server.

[0218] Step 5:

[0219] The server converts the voice data into text using a speech-to-text module (Google Speech-to-Text API). The text data is then analyzed by an AI model to generate an appropriate response. The input is the user's voice data. The output is text data generated by the virtual personality.

[0220] Step 6:

[0221] The generated text data is converted into voice data using speech synthesis technology (Google Text-to-Speech API) and sent to the device. The input is text data generated by the virtual personality. The output is voice data.

[0222] Step 7:

[0223] The terminal plays the received voice data and continues the dialogue with the user. Specifically, the user listens to the virtual personality's response. The input is the voice data sent from the server. The output is the voice heard by the user.

[0224] Step 8:

[0225] During the interaction, the user provides feedback, for example a specific request such as "tell more jokes." The device collects this feedback and sends it to the server. The input is the user's feedback. The output is the feedback data sent to the server.

[0226] Step 9:

[0227] The server analyzes the feedback data and updates the AI ​​model. Specifically, it adjusts the neural network parameters based on the feedback and adapts the model so that future interactions are more natural. The input is the feedback data. The output is the updated AI model.

[0228] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0229] To implement the present invention, a system can be constructed and a program can be executed according to the following procedure, which includes collecting and analyzing recorded data, generating a virtual personality, conducting dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[0230] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and data collection is performed automatically by this application. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[0231] Regarding data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for generating an AI model. This AI model generates a virtual personality that mimics the user's behavior and characteristics.

[0232] Regarding the integration of the emotion engine, the server includes a model for recognizing the user's emotions using the emotion engine. The emotion engine analyzes emotions from the user's voice or text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the response of the virtual personality.

[0233] In the conversation simulation stage, the user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is converted into voice data using voice synthesis technology and adjusted to the voice specified by the user using a voice changer. The synthesized voice data is sent to the device, and the Reiwa voice data is played back to the user.

[0234] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "please use more humor" or "please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[0235] Explanation with concrete examples

[0236] For example, if the family of a deceased person, Mr. C, wants to use this system, the process is as follows:

[0237] 1. Data Collection

[0238] The user (Mr. C's family member) agrees to the service and consents to providing Mr. C's social media account, call history, and search history.

[0239] A dedicated application installed on the device collects this data and sends it to a server.

[0240] The server stores the received data in a database.

[0241] 2. Data analysis and model generation

[0242] The server analyzes the data and extracts Mr. C's characteristic phrases and interests.

[0243] An AI model is created that generates a virtual personality for Mr. C based on the analysis data.

[0244] 3. Emotion engine integration

[0245] The server recognizes emotions from the user's voice and text, and reflects that emotional data in the virtual personality's responses.

[0246] 4. Conversation Simulation

[0247] The user (Mr. C's family member) talks to the virtual Mr. C through the terminal.

[0248] The device captures the audio and sends it to the server.

[0249] The server converts the speech to text and uses an AI model to generate an appropriate response.

[0250] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[0251] The device plays back the virtual person C's reply.

[0252] 5. Feedback and Adaptation

[0253] The user provides feedback on the interaction with virtual person C.

[0254] The server analyzes the feedback and adjusts the AI ​​model and emotion engine.

[0255] In this way, the system based on this invention enables users to have conversations with their deceased loved ones and provide emotional support. The integration of an emotion engine enables more natural and emotionally sensitive conversations.

[0256] The processing flow will be explained below.

[0257] Step 1:

[0258] User: Agrees to use the service and consents to providing the necessary data. Completes procedures to allow the collection of data such as social media accounts, call history, and search history.

[0259] Step 2:

[0260] Device: Install the dedicated application and complete the setup. The application will automatically collect data such as user's social media posts, call history, and search history, and send it to the server.

[0261] Step 3:

[0262] Server: Receives recorded data sent from the device and stores it in a secure database.

[0263] Step 4:

[0264] Server: Analyzes the transmitted recorded data. Using natural language processing technology, it extracts the user's language usage patterns, characteristic phrases, and behavioral patterns.

[0265] Step 5:

[0266] Server: Generates an AI model based on the extracted feature information. This AI model creates a virtual personality that mimics the user's behavior and characteristics.

[0267] Step 6:

[0268] Server: Integrates an emotion engine into the virtual personality. The emotion engine analyzes emotions from the user's voice and text in real time and generates emotion data.

[0269] Step 7:

[0270] User: Uses the device to initiate a conversation with the virtual persona. The device captures the user's voice with a microphone.

[0271] Step 8:

[0272] Terminal: Sends captured audio data to the server in real time.

[0273] Step 9:

[0274] Server: The received voice data is converted into text data using voice recognition technology. The emotion engine then analyzes the user's emotions from the voice data and generates emotion data.

[0275] Step 10:

[0276] Server: The AI ​​model generates appropriate responses based on the text and emotional data. The responses are tailored to the characteristics of the virtual personality and take into account the emotional data.

[0277] Step 11:

[0278] Server: The generated response is converted into voice data using speech synthesis technology. A voice changer is used to match the tone of the virtual personality.

[0279] Step 12:

[0280] Server: The synthesized voice data is sent back to the device.

[0281] Step 13:

[0282] Terminal: Plays the received audio data to the user through the speaker.

[0283] Step 14:

[0284] User: Provides feedback to the virtual persona during the interaction. Feedback can include specific requests such as "Please tell me more" or "Talk about a different topic."

[0285] Step 15:

[0286] Terminal: Sends feedback data to the server in real time.

[0287] Step 16:

[0288] Server: Analyzes the received feedback and applies it to the AI ​​model and emotion engine. It adjusts the response and behavior patterns of the virtual personality based on the feedback.

[0289] In this way, the present invention provides a system that allows interaction with deceased loved ones and is flexible in adapting to the user's emotions and feedback.

[0290] Example 2

[0291] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0292] Current virtual dialogue systems lack sufficient recognition of user emotions, resulting in mechanical and unnatural dialogue. Furthermore, it is difficult to effectively incorporate user feedback, making it difficult to continuously improve the quality of dialogue. Furthermore, voice conversion lacks flexibility, and there is a lack of technology available to adjust the voice to suit the user's preferences. These challenges make it difficult to provide a satisfying dialogue experience for users.

[0293] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0294] In this invention, the server includes means for collecting user record data, means for analyzing the collected record data and extracting features, and means for generating a virtual personality based on the extracted features. This makes it possible to recognize the user's emotions using an emotion engine and reflect them in the dialogue, collect and analyze the user's voice in real time, and adapt the AI ​​model and emotion engine based on the feedback.

[0295] "User" refers to an individual who interacts with the System.

[0296] "Recorded data" refers to data provided by users, such as social media accounts, call history, and search history.

[0297] "Data collection" refers to the process of automatically collecting recorded data using a dedicated application based on the user's consent.

[0298] "Database" refers to a repository for storing recorded data collected by the Server.

[0299] "Data analysis" refers to the process of analyzing collected recorded data using natural language processing techniques and extracting features.

[0300] "Characteristics" refers to characteristics such as users' language usage patterns and behavioral patterns extracted through data analysis.

[0301] "Virtual personality" refers to a virtual conversation partner that mimics the user's characteristics and is generated using an AI model based on extracted characteristics.

[0302] "Dialogue" refers to communication between the user and the virtual personality.

[0303] "Feedback" refers to the opinions and requests that a user provides to a virtual personality during a conversation.

[0304] "Adaptation" refers to the process of updating and improving AI models and emotion engines based on the feedback provided.

[0305] An "emotion engine" refers to a system that analyzes and recognizes users' emotions from voice and text data in real time.

[0306] "Audio Data" means digital audio information that captures a user's speech.

[0307] "Text data" refers to character information obtained by converting voice data.

[0308] "Speech synthesis" refers to the process of generating digital speech from text data.

[0309] "Voice changer" refers to technology that adjusts the generated voice to a voice specified by the user.

[0310] "Server" refers to the computer system that collects and analyzes data, generates virtual personalities, and operates the emotion engine.

[0311] "Dedicated application" refers to software that is installed on the user's device and is used to collect and transmit recorded data.

[0312] "Consent interface" refers to the screen or function that allows users to consent to providing data.

[0313] To implement this invention, the system is constructed and the program is executed according to the following procedure: The system includes collecting and analyzing recorded data, generating a virtual personality, executing dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[0314] Data collection

[0315] The user agrees to use the service and consents to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[0316] Examples:

[0317] The user agrees to the terms of service and installs the dedicated application on their smartphone.

[0318] A dedicated application automatically collects social media accounts and call history, encrypts them, and sends them to a server.

[0319] Data Analysis and Model Generation

[0320] The server analyzes the collected data. Natural language processing technology (such as SpaCy or NLTK) is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. An AI model (such as the Transformer model or GPT-4 (registered trademark)) is generated based on the extracted features, creating a virtual personality that mimics the user's characteristics.

[0321] Emotion engine integration

[0322] The server contains a model for recognizing the user's emotions using an emotion engine (such as the Microsoft® Azure® Cognitive Services emotion analysis API). The emotion engine analyzes the user's voice and text in real time to determine their emotional state. The determined emotion data is used to adjust the virtual personality's responses.

[0323] Examples:

[0324] The server analyzes the collected social media posts and call history to extract the user's characteristic phrases and interests.

[0325] An AI model is created that generates a virtual personality based on the extracted data and an emotion engine is integrated.

[0326] Conversation Simulation

[0327] The user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and an AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology (such as Amazon Polly), and a voice changer adjusts it to the voice specified by the user. The synthesized voice data is sent to the device, where it is played back to the user.

[0328] Examples:

[0329] The user speaks to the virtual persona, "What's the weather like today?"

[0330] The server analyzes this, and the virtual personality generates a reply saying "It's sunny today," converts it into voice, and sends it to the terminal.

[0331] The voice that has been adjusted by the voice changer is played back to the user.

[0332] Feedback and Adaptation

[0333] Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "I want more humor" or "I want you to elaborate on a particular topic," they can send these to the server through the feedback function. The server analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[0334] Example prompt sentence:

[0335] "Based on C's social media account data, please extract characteristic phrases and behavioral patterns to create a virtual personality for C."

[0336] This invention allows users to have a more natural and emotionally relevant interaction experience.

[0337] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0338] Step 1:

[0339] Service consent and data provision consent

[0340] Input: The user agrees to the terms of use of the service and consents to providing data such as social media accounts, call history, and search history.

[0341] Specific behavior: The user clicks the "Agree" button on the screen where they agree to the terms of use and privacy policy. They are then asked to provide data, which they agree to by clicking the "Provide" button.

[0342] Output: Consent and authorization information is sent to the server.

[0343] Step 2:

[0344] Installing the dedicated application

[0345] Input: User information with consent.

[0346] How it works: The user installs the dedicated application on their smartphone or PC, and the application is downloaded from the official website or app store.

[0347] Output: The dedicated application is installed on the user's device.

[0348] Step 3:

[0349] Automatic data collection and transmission

[0350] Input: The user's device on which the dedicated application is installed.

[0351] How it works: The device automatically collects data from social media accounts, call history, and search history. The collected data is encrypted and securely sent to a server.

[0352] Output: The encrypted data is sent to the server.

[0353] Step 4:

[0354] Data storage

[0355] Input: Encrypted data.

[0356] Specific operation: The server stores the received data in a database. This database uses, for example, AWS (registered trademark) RDS.

[0357] Output: Data stored in a database.

[0358] Step 5:

[0359] Performing data analysis

[0360] Input: Data stored in a database.

[0361] What it does: The server analyzes the collected data using natural language processing techniques (e.g., SpaCy or NLTK) to extract the user's language usage patterns, characteristic phrases, and behavioral patterns.

[0362] Output: Extracted feature data.

[0363] Step 6:

[0364] AI model generation

[0365] Input: Extracted feature data.

[0366] Specific operation: The server generates an AI model (e.g., Transformer model or GPT-4) based on the feature data and creates a virtual personality that mimics the user's characteristics.

[0367] Output: A virtual personality is generated.

[0368] Step 7:

[0369] Preparing for emotion recognition

[0370] Input: User interaction data with virtual persona.

[0371] How it works: The server integrates the emotion engine, which uses the emotion analysis API from Microsoft Azure Cognitive Services and is configured to analyze user emotions in real time from voice and text data.

[0372] Output: A system capable of emotion recognition.

[0373] Step 8:

[0374] Utilizing Emotional Data

[0375] Input: Real-time analyzed emotion data.

[0376] Specific operation: The server reflects the emotional data recognized by the emotion engine in the virtual personality's response. For example, if the user makes a sad voice, the virtual personality will respond with a comforting response.

[0377] Output: The virtual personality's response depending on the emotion.

[0378] Step 9:

[0379] Start a real-time conversation

[0380] Input: What the user says.

[0381] Specific operation: The user speaks into the device, which captures what is said as audio data and sends it to the server in real time.

[0382] Output: The audio data is sent to the server.

[0383] Step 10:

[0384] Voice data conversion and response generation

[0385] Input: User's voice data.

[0386] How it works: The server uses the Google Cloud Speech-to-Text API to convert the voice data into text, and the AI ​​model generates an appropriate response based on the converted text.

[0387] Output: The generated response text data.

[0388] Step 11:

[0389] Apply voice synthesis and voice changer

[0390] Input: Generated response text data.

[0391] Specific operation: The server converts the generated text into voice data using speech synthesis technology (e.g., Amazon Polly), and then uses a voice changer to adjust the voice data to match the user's voice.

[0392] Output: The adjusted audio data.

[0393] Step 12:

[0394] Playing audio data

[0395] Input: The adjusted audio data.

[0396] Specific operation: The device plays back the voice data received from the server, and the appropriate response is provided in the voice of the virtual personality.

[0397] Output: The audio played to the user.

[0398] Step 13:

[0399] Providing feedback

[0400] Input: User feedback provided during interaction with the virtual persona.

[0401] Specific operation: The user enters and submits their opinions and requests regarding the virtual personality through a feedback form or the like.

[0402] Output: Feedback data sent to the server.

[0403] Step 14:

[0404] Feedback Analysis and Adaptation

[0405] Input: Feedback data sent to the server.

[0406] Specific operation: The server analyzes the feedback and updates and adapts the AI ​​model and emotion engine. Specifically, the feedback content is analyzed using natural language processing technology and reflected in the AI ​​model's response patterns.

[0407] Output: Updated and adapted AI model and emotion engine.

[0408] This will make future interactions more natural and satisfying.

[0409] (Application example 2)

[0410] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0411] Conventional food delivery systems often lack sufficient support when users select menus or customize their meals. It can be difficult to get appropriate advice or suggestions, especially when users have special requests or want to try new dishes. Furthermore, there is a need for a personalized experience based on users' emotions and preferences. To address these challenges, a system is needed that allows users to easily select and customize menus while interacting with a virtual chef.

[0412] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0413] In this invention, the server includes: [means for collecting user record data;] [means for analyzing the collected record data and extracting features;] [means for generating a virtual personality based on the extracted features;] [means for having the virtual personality interact with the user;] [means for receiving feedback from the user during the interaction and adapting the virtual personality;] [means for the user to select a menu while conversing with the virtual chef when ordering; and [means for assisting in menu selection and dish customization.] This enables users to easily select and customize menus, providing a more personalized experience.

[0414] "User" refers to the person using the invention, who is primarily responsible for selecting menus and customizing meals through interaction with the virtual chef.

[0415] "Recorded data" refers to information including data related to the user's use of social media accounts, call history, search history, etc.

[0416] "Collect" refers to the process of collecting user record data through a dedicated application and sending it to a server.

[0417] "Analyzing" refers to the act of analyzing collected recorded data and extracting specific patterns or characteristics.

[0418] "Extracting features" means identifying a user's language usage, behavioral patterns, preferences, etc. from the analyzed data.

[0419] A "virtual personality" is an interactive agent that uses an AI model generated based on collected and analyzed characteristics.

[0420] "Having a dialogue" refers to the virtual personality interacting with the user in a conversational format.

[0421] "Feedback" refers to the responses and requests provided by users to the system.

[0422] "Adapting" refers to updating and adjusting the virtual persona's responses and behavior based on feedback.

[0423] "Ordering" refers to the act of a user requesting a meal using a food delivery service.

[0424] A "virtual chef" is a virtual personality that has particular knowledge about cooking and assists users in choosing menus through dialogue.

[0425] "Menu selection" is the process by which the virtual chef makes appropriate meal suggestions to the user and provides the best options.

[0426] "Cuisine customization" refers to making changes or adjustments to a dish according to the user's wishes.

[0427] "Voice data" refers to data that is a digital recording of a user's voice input.

[0428] "Converting to text data" refers to analyzing the voice data and changing it into text-format data.

[0429] A "reply" is a response that the generative AI model generates in response to a user's statement.

[0430] A "speech synthesizer" is a device or software that converts text data into speech data.

[0431] "Play" refers to making the synthesized speech available to the user.

[0432] "Consent" is the act of a user explicitly giving permission to use the system.

[0433] "Interface" refers to the means or screen through which a user interacts with a system.

[0434] To implement the present invention, the following system is required: This system is configured by incorporating a user's recorded data collection, analysis, virtual personality generation, dialogue execution, feedback adaptation, and emotion engine.

[0435] Data collection

[0436] Users agree to use the service and provide the necessary data, such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects the data. The collected data is securely transmitted from the device to a server, which then stores it in a database.

[0437] Data Analysis and Model Generation

[0438] The server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for a generative AI model. This generative AI model generates a virtual personality that mimics the user's behavior and characteristics.

[0439] Emotion engine integration

[0440] The server includes a model for recognizing the user's emotions using an emotion engine. The emotion engine analyzes emotions from the user's voice and text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the responses of the virtual personality.

[0441] Conversation Simulation

[0442] The user interacts with the virtual chef via a device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the generative AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology, and the synthesized voice data is sent to the device and played back to the user.

[0443] Feedback and Adaptation

[0444] Users can provide feedback to the virtual chef during the conversation. For example, if they have specific requests, such as "I'd like more humor" or "Please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the generative AI model and emotion engine, making future conversations more natural and satisfying.

[0445] Specific examples

[0446] For example, consider a scenario in which a user uses a smartphone app to say, "I want spicy pasta." The system captures the speech, converts it into text data, and then uses a generative AI model to derive the optimal response. An example of a prompt in this case is as follows:

[0447] User: "I want spicy pasta"

[0448] Virtual Chef: "For spicy pasta, I recommend Peperoncino or Arrabbiata. Which would you like?"

[0449] In this way, the virtual chef can provide a menu tailored to the user's needs, allowing for personalized meal selection and customization.

[0450] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0451] Step 1:

[0452] The user agrees to use the service and consents to providing data such as social media accounts, call history, and search history. This data is automatically collected by a dedicated application installed on the device. The device securely transmits the collected data to the server. This input data is then stored in a database.

[0453] Step 2:

[0454] The server analyzes the user's recorded data stored in the database. During this analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. These analysis results (language usage patterns and behavioral patterns) are extracted and used as input data for the generative AI model.

[0455] Step 3:

[0456] The server generates a virtual personality based on the extracted characteristics. Using a generative AI model, it creates a virtual chef that mimics the user's behavior and characteristics. The virtual chef can then assist with menu selection and cooking customization.

[0457] Step 4:

[0458] The user interacts with the virtual chef via a terminal. The terminal captures the user's speech as voice data and transmits it to the server in real time. The server converts this voice data into text data and uses it as input.

[0459] Step 5:

[0460] The server generates a response from the virtual chef based on the text data. This response is generated using a generative AI model and output as a response text. The response text is converted into audio data using speech synthesis technology. This audio data is sent to the device and played back to the user.

[0461] Step 6:

[0462] During the conversation, the user can provide feedback to the virtual chef, such as a request for more humor. The device then sends this feedback data to the server, which analyzes it and adjusts the parameters of the generative AI model and emotion engine to reflect the feedback in the next conversation.

[0463] Step 7:

[0464] The virtual chef will help users select menus and customize dishes to suit their needs. For example, if a user says, "I want spicy pasta," the virtual chef will suggest, "For spicy pasta, we recommend peperoncino or arrabbiata. Which would you like?"

[0465] This process allows users to make personalized meal selections and improves their experience.

[0466] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0467] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0468] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0469] [Second embodiment]

[0470] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0471] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0472] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0473] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0474] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0475] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0476] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0477] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0478] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0479] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0480] In the smart glasses 214, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0481] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0482] To practice the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[0483] First, regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[0484] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[0485] During the conversation simulation phase, the user interacts with the virtual personality via their device. The device captures the user's speech as voice data and transmits it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using speech synthesis technology and played back to the user via their device.

[0486] Feedback and Adaptation: Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that conversation," they can send them to the server through the feedback function. The server analyzes this feedback and adapts it to the AI ​​model, making future conversations more natural and satisfying.

[0487] Explanation with concrete examples

[0488] For example, suppose the family of a deceased person, Mr. B, wants to use this system.

[0489] 1. Data Collection

[0490] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[0491] A dedicated application installed on the device collects this data and sends it to a server.

[0492] The server stores the received data in a database.

[0493] 2. Data analysis and model generation

[0494] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[0495] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[0496] 3. Conversation Simulation

[0497] The user (B's family member) talks to the virtual B through the terminal.

[0498] The device captures the audio and sends it to the server.

[0499] The server converts the speech to text and uses an AI model to generate an appropriate response.

[0500] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[0501] The device plays back the response of virtual person B.

[0502] 4. Feedback and Adaptation

[0503] The user provides feedback on the interaction with virtual Person B.

[0504] The server analyzes the feedback and adjusts the AI ​​model.

[0505] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[0506] The processing flow will be explained below.

[0507] Step 1:

[0508] User: Agrees to use the service and agrees to provide the required data.

[0509] Device: Install the dedicated application and configure it to allow access to social media accounts, call history, and search history.

[0510] Step 2:

[0511] Device: Collects data such as users' social media posts, call history, and search history, and sends it to a server.

[0512] Server: Receives collected data and stores it securely in a database.

[0513] Step 3:

[0514] Server: Runs long-text data analysis algorithms to analyze the collected data and extract users' language usage patterns, characteristic phrases, and behavioral patterns.

[0515] Step 4:

[0516] Server: Generates an AI model based on the extracted information. This model creates a virtual personality that mimics the user's behavior and characteristics.

[0517] Step 5:

[0518] User: Uses the device to initiate a conversation with the virtual personality.

[0519] Device: The user's voice is captured by a microphone and sent to the server in real time.

[0520] Step 6:

[0521] Server: Converts the received voice data into text data using voice recognition technology.

[0522] Server: Inputs the converted text data into the AI ​​model and generates an appropriate response.

[0523] Server: The generated response is converted into voice data using speech synthesis technology, and then matched to the voice specified by the user using a voice changer.

[0524] Step 7:

[0525] Server: Sends the synthesized voice data to the device.

[0526] Terminal: The synthesized voice data is played back to the user through a speaker.

[0527] Step 8:

[0528] Users: Provide feedback during a conversation, for example by sending requests such as "I'd like more humor" or "I'd like more detail on a particular topic" through the feedback feature.

[0529] Server: Analyzes the received feedback and updates and adapts the AI ​​model to make future conversations more natural and satisfying.

[0530] Example 1

[0531] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0532] In modern society, there is a demand for technology that can recreate conversations with deceased loved ones and provide emotional support. In particular, conventional systems have had problems with insufficient analysis of collected data, resulting in poor conversation quality. Furthermore, it has been difficult to quickly and appropriately incorporate user feedback. This has led to issues such as reduced user satisfaction.

[0533] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0534] In this invention, the server includes: [means for collecting user record data and securely transmitting it to the server;] [means for analyzing the collected record data and extracting features using natural language processing technology; and] [means for generating a virtual personality using machine learning based on the extracted features.] This makes it possible [to accurately mimic the user's behavior and personality, achieve high-quality dialogue, and quickly and appropriately reflect feedback].

[0535] "Means for collecting user record data and safely transmitting it to a server" refers to a mechanism for collecting data such as users' social media accounts, call history, and search history, and transferring it to a server using security technologies such as encryption.

[0536] "Means of analyzing collected recorded data and extracting features using natural language processing technology" refers to a method of analyzing data received by the server and extracting phrases, behavioral patterns, interests, etc. from large amounts of text data.

[0537] "Means for generating a virtual personality using machine learning based on extracted features" refers to the process of using a machine learning algorithm to generate a virtual person that mimics the user's behavior and personality based on the analyzed data.

[0538] The "means for implementing speech recognition and speech synthesis technology to enable the generated virtual personality to interact with the user" refers to a series of processes for converting the user's speech input into text and converting the generated virtual personality's responses back into speech to provide to the user.

[0539] "Means for receiving feedback from users during a dialogue, analyzing it, and adapting the virtual personality" refers to a mechanism that collects opinions and requests provided by users during a dialogue, and uses that data to improve and adapt the behavior and comments of the generated virtual personality.

[0540] "Means for collecting user voice and transmitting it to a server in real time" refers to a method for capturing voice data in real time using a device such as a microphone and immediately transmitting it to a server.

[0541] "Means for using speech recognition technology to convert voice data into text data at the server" refers to technology that uses a speech recognition algorithm to accurately convert voice data into text data.

[0542] "Means of generating a response using an AI model based on text data and converting that response into voice using speech synthesis technology" refers to the process of converting a text response generated by an AI model into voice data using a speech synthesis algorithm.

[0543] "Means for playing the synthesized voice to the user through the terminal" refers to a method for delivering the generated voice data to the user using the terminal's speaker or the like.

[0544] "Means for providing an interface for users to provide consent to data provision" refers to a mechanism that provides a screen or interactive procedure for users to agree to data provision.

[0545] "Means for automatically starting data collection upon receiving user consent" refers to the process by which the system automatically starts collecting the necessary data after the user's consent is obtained.

[0546] MODE FOR CARRYING OUT THE INVENTION

[0547] To implement the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[0548] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is encrypted and securely sent from the device to the server, where it is stored in a database.

[0549] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technologies such as Google Speech-to-Text and Amazon Polly are used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[0550] During the conversation simulation phase, the user interacts with the virtual persona via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using voice synthesis technology such as Amazon Polly and played back to the user via the device.

[0551] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that story," they can send them to the server through the feedback function. The server then analyzes this feedback and adapts the AI ​​model, making future conversations more natural and satisfying.

[0552] Explanation with concrete examples

[0553] For example, suppose the family of a deceased person, Mr. B, wants to use this system:

[0554] 1. Data Collection

[0555] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[0556] A dedicated application installed on the device collects this data and sends it to a server.

[0557] The server stores the received data in a database.

[0558] 2. Data analysis and model generation

[0559] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[0560] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[0561] 3. Conversation Simulation

[0562] The user (B's family member) talks to the virtual B through the terminal.

[0563] The device captures the audio and sends it to the server.

[0564] The server converts the speech to text and uses an AI model to generate an appropriate response.

[0565] The generated response is converted into voice data using voice synthesis technology such as Amazon Polly and sent to the device.

[0566] The device plays back the response of virtual person B.

[0567] 4. Feedback and Adaptation

[0568] The user provides feedback on the interaction with virtual Person B.

[0569] The server analyzes the feedback and adjusts the AI ​​model.

[0570] Prompt Sentence Examples

[0571] "Person B often uses the phrase 'do your best'. Please generate words of encouragement in a tone that is characteristic of Person B."

[0572] "I'd like you to tell me about a movie that Mr. B liked."

[0573] "Tell me how Mr. B tells jokes."

[0574] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[0575] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0576] Step 1: Data collection

[0577] The user agrees to use the service and consents to the provision of data such as social media accounts, call history, search history, etc. User consent and permission are required as input.

[0578] A dedicated application installed on the device acquires the data that can be collected. In operation, the application periodically and automatically collects posts from social media accounts, call history, and user search history.

[0579] The device securely encrypts the collected data and sends it to the server, resulting in the encrypted data being sent to the server.

[0580] The server stores the received data in a database. In operation, the server appropriately classifies the received data and saves it in a database.

[0581] Step 2: Data analysis and model generation

[0582] The server analyzes the data stored in the database, using saved user social media posts, call history, search history, etc. as input.

[0583] The server uses natural language processing technology to extract features such as the user's phrasing and behavioral patterns. In operation, the NLP algorithm extracts the user's unique expressions and interests from the data. The extracted feature data is obtained as the output.

[0584] The server uses machine learning algorithms to generate an AI model based on the extracted features. The AI ​​model is built using tools such as Python and TensorFlow. The output is a virtual personality model that mimics the user's behavior.

[0585] Step 3: Conversation simulation

[0586] The user initiates a dialogue with the virtual persona through a terminal, and the user's voice input is required.

[0587] The device captures the user's speech as voice data and sends it to the server in real time. The device's microphone collects the voice and uses a voice recognition API to transfer the data to the server. The output is real-time voice data sent to the server.

[0588] The server converts the voice data into text data and generates a response using an AI model. The operation is to convert the voice to text using the Google Speech-to-Text API, and then use an AI model (e.g., GPT-3) to generate an appropriate response. The output is the generated text response.

[0589] The text response generated by the server is converted into voice data using speech synthesis technology and sent to the device. The text is converted into voice using speech synthesis technology such as Amazon Polly and returned to the device. The voice data is then sent to the device as an output.

[0590] The device plays back the virtual persona's response in real time. In operation, the device's speaker plays back the received audio data. As an output, the virtual persona's response is provided to the user through audio.

[0591] Step 4: Feedback and Adapt

[0592] The user provides feedback on the interaction with the virtual personality. The input requires the user's feedback, which may include specific requests such as "Tell more jokes."

[0593] The server analyzes the feedback provided by the user and adjusts the AI ​​model. The operation is to analyze the feedback and incorporate it as new data into the training of the AI ​​model. The output is the adjusted AI model.

[0594] By repeating this series of processes, the system adapts its interactions with the user to make them more natural and satisfying.

[0595] (Application example 1)

[0596] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0597] In modern society, there is a growing demand for content distribution services that go beyond simply viewing video content such as movies and dramas to provide a more fulfilling entertainment experience. However, conventional systems lack a means to empathize with users' feelings of loneliness or the enjoyment of watching a movie. Furthermore, interactive means for providing emotional support are limited. There is a need to solve these issues and provide users with a more enjoyable and fulfilling viewing experience.

[0598] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0599] In this invention, the server includes means for collecting recorded data of a user, means for analyzing the collected recorded data and extracting features, means for generating a virtual personality based on the extracted features, means for the virtual personality to interact with the user, means for receiving feedback from the user during the interaction and adapting the virtual personality, and means for the user and the virtual personality to interact while playing video content. This allows the user to interact with a virtual partner while watching a movie or drama, making for a more fulfilling viewing experience.

[0600] "User" means an individual or organization that uses this system.

[0601] "Recorded data" refers to activity data such as a user's social media accounts, call history, and search history.

[0602] "Means of collection" refers to methods of collecting data using applications or sensors installed on the user's device.

[0603] "Means for analysis" refers to software or algorithms used to analyze the collected recorded data and extract features.

[0604] "Characteristics" are patterns that indicate tendencies in a user's behavior, personality, interests, etc.

[0605] A "virtual personality" is an AI model that is generated based on collected characteristics and allows users to interact with it.

[0606] "Means for dialogue" refers to a function that allows the user and the virtual personality to communicate via voice or text.

[0607] "Means for receiving feedback" and "adapting" refer to methods for updating the virtual personality generation model by reflecting opinions and requests from users.

[0608] "Video content" refers to media content that can be viewed, such as movies and dramas.

[0609] "Means for a user to have a conversation with a virtual character during playback" refers to a technology that allows a user to have a conversation with a virtual character in real time while viewing video content.

[0610] To implement this invention, the following system configuration and programs are required: This system is made up of several main components, each of which plays a specific role.

[0611] System Configuration

[0612] 1. User's device

[0613] Data Collection Applications

[0614] Video playback application

[0615] Audio input and output devices (microphone, speaker)

[0616] 2. Server

[0617] Database

[0618] Data Analysis Module

[0619] AI model generation module

[0620] Speech-to-text module (Google Speech-to-Text API)

[0621] Text-to-speech module (Google Text-to-Speech API)

[0622] WebSocket Server

[0623] Processing flow

[0624] 1. Data Collection

[0625] The user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, search history, etc. The device sends this data to the server, which then stores the collected data in a database.

[0626] 2. Data analysis and AI model generation

[0627] The server retrieves the collected data from the database and extracts features using natural language processing technology (TensorFlow, NLTK, spaCy). Based on these features, a virtual personality is generated that mimics the user's behavior and personality.

[0628] 3. Playback and interaction with video content

[0629] While a user is playing a movie or TV drama using a video playback application, a virtual persona will engage in conversation related to the video content. When the user speaks, the device captures the audio data and sends it to the server via WebSocket. The server converts the audio data into text and uses an AI model to generate an appropriate response. This response is then converted back into audio data by a text-to-speech module and played back to the user via the device.

[0630] 4. Feedback and Adaptation

[0631] Users can provide feedback to the virtual personality during the interaction, for example, if they want it to tell more jokes, they can send that feedback to the server through the feedback function, which will analyze this feedback and adapt the AI ​​model.

[0632] Specific examples

[0633] For example, suppose a user is watching the movie "Castle in the Sky," and the virtual partner imitates the personality of their deceased best friend, Mr. A.

[0634] Prompt Sentence Examples

[0635] The user is watching the movie "Castle in the Sky." The virtual partner is the deceased best friend, Mr. A:

[0636] User: "You also liked this scene, right?"

[0637] Virtual Partner: "Yes, the music in this scene is particularly memorable."

[0638] An example of a prompt sentence to input to the generative AI model is as follows:

[0639] User social media data:

[0640] Posted yesterday: Looking forward to seeing "Castle in the Sky"!

[0641] Call History:

[0642] My best friend A and I often talked about movies.

[0643] Search History:

[0644] A moving scene from "Castle in the Sky"

[0645] feedback:

[0646] I'd like to hear more detailed feedback

[0647] This system allows users to not only watch video content, but also enjoy interacting with virtual characters, further enhancing the viewing experience.

[0648] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0649] Step 1:

[0650] A user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, and search history. Based on this consent, the device automatically collects this data and sends it to a database. The input is the user's social media accounts, call history, and search history. The output is the data securely sent to the server.

[0651] Step 2:

[0652] The server retrieves the collected data from the database and extracts features using natural language processing techniques (TensorFlow, NLTK, spaCy). Specifically, the data is first cleansed, and then a language model is applied to identify interest and behavioral patterns. The input is the user's recorded data retrieved from the database. The output is analyzed data containing the extracted features.

[0653] Step 3:

[0654] The server builds an AI model for generating a virtual personality based on the extracted features. Specifically, it trains a neural network model that mimics the user's behavior and personality based on the feature data. The input is the analysis data obtained in step 2. The output is an AI model that represents the virtual personality.

[0655] Step 4:

[0656] While a user plays a movie or drama using a video playback application, the virtual persona interacts with the user about the video content being viewed. Specifically, the voice input is captured and the voice data is sent to the server. The input is the user's voice data. The output is the data sent to the server.

[0657] Step 5:

[0658] The server converts the voice data into text using a speech-to-text module (Google Speech-to-Text API). The text data is then analyzed by an AI model to generate an appropriate response. The input is the user's voice data. The output is text data generated by the virtual personality.

[0659] Step 6:

[0660] The generated text data is converted into voice data using speech synthesis technology (Google Text-to-Speech API) and sent to the device. The input is text data generated by the virtual personality. The output is voice data.

[0661] Step 7:

[0662] The terminal plays the received voice data and continues the dialogue with the user. Specifically, the user listens to the virtual personality's response. The input is the voice data sent from the server. The output is the voice heard by the user.

[0663] Step 8:

[0664] During the interaction, the user provides feedback, for example a specific request such as "tell more jokes." The device collects this feedback and sends it to the server. The input is the user's feedback. The output is the feedback data sent to the server.

[0665] Step 9:

[0666] The server analyzes the feedback data and updates the AI ​​model. Specifically, it adjusts the neural network parameters based on the feedback and adapts the model so that future interactions are more natural. The input is the feedback data. The output is the updated AI model.

[0667] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0668] To implement the present invention, a system can be constructed and a program can be executed according to the following procedure, which includes collecting and analyzing recorded data, generating a virtual personality, conducting dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[0669] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and data collection is performed automatically by this application. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[0670] Regarding data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for generating an AI model. This AI model generates a virtual personality that mimics the user's behavior and characteristics.

[0671] Regarding the integration of the emotion engine, the server includes a model for recognizing the user's emotions using the emotion engine. The emotion engine analyzes emotions from the user's voice or text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the response of the virtual personality.

[0672] In the conversation simulation stage, the user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is converted into voice data using voice synthesis technology and adjusted to the voice specified by the user using a voice changer. The synthesized voice data is sent to the device, and the Reiwa voice data is played back to the user.

[0673] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "please use more humor" or "please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[0674] Explanation with concrete examples

[0675] For example, if the family of a deceased person, Mr. C, wants to use this system, the process is as follows:

[0676] 1. Data Collection

[0677] The user (Mr. C's family member) agrees to the service and consents to providing Mr. C's social media account, call history, and search history.

[0678] A dedicated application installed on the device collects this data and sends it to a server.

[0679] The server stores the received data in a database.

[0680] 2. Data analysis and model generation

[0681] The server analyzes the data and extracts Mr. C's characteristic phrases and interests.

[0682] An AI model is created that generates a virtual personality for Mr. C based on the analysis data.

[0683] 3. Emotion engine integration

[0684] The server recognizes emotions from the user's voice and text, and reflects that emotional data in the virtual personality's responses.

[0685] 4. Conversation Simulation

[0686] The user (Mr. C's family member) talks to the virtual Mr. C through the terminal.

[0687] The device captures the audio and sends it to the server.

[0688] The server converts the speech to text and uses an AI model to generate an appropriate response.

[0689] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[0690] The device plays back the virtual person C's reply.

[0691] 5. Feedback and Adaptation

[0692] The user provides feedback on the interaction with virtual person C.

[0693] The server analyzes the feedback and adjusts the AI ​​model and emotion engine.

[0694] In this way, the system based on this invention enables users to have conversations with their deceased loved ones and provide emotional support. The integration of an emotion engine enables more natural and emotionally sensitive conversations.

[0695] The processing flow will be explained below.

[0696] Step 1:

[0697] User: Agrees to use the service and consents to providing the necessary data. Completes procedures to allow the collection of data such as social media accounts, call history, and search history.

[0698] Step 2:

[0699] Device: Install the dedicated application and complete the setup. The application will automatically collect data such as user's social media posts, call history, and search history, and send it to the server.

[0700] Step 3:

[0701] Server: Receives recorded data sent from the device and stores it in a secure database.

[0702] Step 4:

[0703] Server: Analyzes the transmitted recorded data. Using natural language processing technology, it extracts the user's language usage patterns, characteristic phrases, and behavioral patterns.

[0704] Step 5:

[0705] Server: Generates an AI model based on the extracted feature information. This AI model creates a virtual personality that mimics the user's behavior and characteristics.

[0706] Step 6:

[0707] Server: Integrates an emotion engine into the virtual personality. The emotion engine analyzes emotions from the user's voice and text in real time and generates emotion data.

[0708] Step 7:

[0709] User: Uses the device to initiate a conversation with the virtual persona. The device captures the user's voice with a microphone.

[0710] Step 8:

[0711] Terminal: Sends captured audio data to the server in real time.

[0712] Step 9:

[0713] Server: The received voice data is converted into text data using voice recognition technology. The emotion engine then analyzes the user's emotions from the voice data and generates emotion data.

[0714] Step 10:

[0715] Server: The AI ​​model generates appropriate responses based on the text and emotional data. The responses are tailored to the characteristics of the virtual personality and take into account the emotional data.

[0716] Step 11:

[0717] Server: The generated response is converted into voice data using speech synthesis technology. A voice changer is used to match the tone of the virtual personality.

[0718] Step 12:

[0719] Server: The synthesized voice data is sent back to the device.

[0720] Step 13:

[0721] Terminal: Plays the received audio data to the user through the speaker.

[0722] Step 14:

[0723] User: Provides feedback to the virtual persona during the interaction. Feedback can include specific requests such as "Please tell me more" or "Talk about a different topic."

[0724] Step 15:

[0725] Terminal: Sends feedback data to the server in real time.

[0726] Step 16:

[0727] Server: Analyzes the received feedback and applies it to the AI ​​model and emotion engine. It adjusts the response and behavior patterns of the virtual personality based on the feedback.

[0728] In this way, the present invention provides a system that allows interaction with deceased loved ones and is flexible in adapting to the user's emotions and feedback.

[0729] Example 2

[0730] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0731] Current virtual dialogue systems lack sufficient recognition of user emotions, resulting in mechanical and unnatural dialogue. Furthermore, it is difficult to effectively incorporate user feedback, making it difficult to continuously improve the quality of dialogue. Furthermore, voice conversion lacks flexibility, and there is a lack of technology available to adjust the voice to suit the user's preferences. These challenges make it difficult to provide a satisfying dialogue experience for users.

[0732] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0733] In this invention, the server includes means for collecting user record data, means for analyzing the collected record data and extracting features, and means for generating a virtual personality based on the extracted features. This makes it possible to recognize the user's emotions using an emotion engine and reflect them in the dialogue, collect and analyze the user's voice in real time, and adapt the AI ​​model and emotion engine based on the feedback.

[0734] "User" refers to an individual who interacts with the System.

[0735] "Recorded data" refers to data provided by users, such as social media accounts, call history, and search history.

[0736] "Data collection" refers to the process of automatically collecting recorded data using a dedicated application based on the user's consent.

[0737] "Database" refers to a repository for storing recorded data collected by the Server.

[0738] "Data analysis" refers to the process of analyzing collected recorded data using natural language processing techniques and extracting features.

[0739] "Characteristics" refers to characteristics such as users' language usage patterns and behavioral patterns extracted through data analysis.

[0740] "Virtual personality" refers to a virtual conversation partner that mimics the user's characteristics and is generated using an AI model based on extracted characteristics.

[0741] "Dialogue" refers to communication between the user and the virtual personality.

[0742] "Feedback" refers to the opinions and requests that a user provides to a virtual personality during a conversation.

[0743] "Adaptation" refers to the process of updating and improving AI models and emotion engines based on the feedback provided.

[0744] An "emotion engine" refers to a system that analyzes and recognizes users' emotions from voice and text data in real time.

[0745] "Audio Data" means digital audio information that captures a user's speech.

[0746] "Text data" refers to character information obtained by converting voice data.

[0747] "Speech synthesis" refers to the process of generating digital speech from text data.

[0748] "Voice changer" refers to technology that adjusts the generated voice to a voice specified by the user.

[0749] "Server" refers to the computer system that collects and analyzes data, generates virtual personalities, and operates the emotion engine.

[0750] "Dedicated application" refers to software that is installed on the user's device and is used to collect and transmit recorded data.

[0751] "Consent interface" refers to the screen or function that allows users to consent to providing data.

[0752] To implement this invention, the system is constructed and the program is executed according to the following procedure: The system includes collecting and analyzing recorded data, generating a virtual personality, executing dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[0753] Data collection

[0754] The user agrees to use the service and consents to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[0755] Examples:

[0756] The user agrees to the terms of service and installs the dedicated application on their smartphone.

[0757] A dedicated application automatically collects social media accounts and call history, encrypts them, and sends them to a server.

[0758] Data Analysis and Model Generation

[0759] The server analyzes the collected data using natural language processing technology (such as SpaCy or NLTK) to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. An AI model (such as the Transformer model or GPT-4) is generated based on the extracted features, creating a virtual personality that mimics the user's characteristics.

[0760] Emotion engine integration

[0761] The server contains a model for recognizing the user's emotions using an emotion engine (such as the Microsoft Azure Cognitive Services emotion analysis API). The emotion engine analyzes the user's voice and text in real time to determine their emotional state. The determined emotional data is used to adjust the virtual personality's responses.

[0762] Examples:

[0763] The server analyzes the collected social media posts and call history to extract the user's characteristic phrases and interests.

[0764] An AI model is created that generates a virtual personality based on the extracted data and an emotion engine is integrated.

[0765] Conversation Simulation

[0766] The user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and an AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology (such as Amazon Polly), and a voice changer adjusts it to the voice specified by the user. The synthesized voice data is sent to the device, where it is played back to the user.

[0767] Examples:

[0768] The user speaks to the virtual persona, "What's the weather like today?"

[0769] The server analyzes this, and the virtual personality generates a reply saying "It's sunny today," converts it into voice, and sends it to the terminal.

[0770] The voice that has been adjusted by the voice changer is played back to the user.

[0771] Feedback and Adaptation

[0772] Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "I want more humor" or "I want you to elaborate on a particular topic," they can send these to the server through the feedback function. The server analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[0773] Example prompt sentence:

[0774] "Based on C's social media account data, please extract characteristic phrases and behavioral patterns to create a virtual personality for C."

[0775] This invention allows users to have a more natural and emotionally relevant interaction experience.

[0776] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0777] Step 1:

[0778] Service consent and data provision consent

[0779] Input: The user agrees to the terms of use of the service and consents to providing data such as social media accounts, call history, and search history.

[0780] Specific behavior: The user clicks the "Agree" button on the screen where they agree to the terms of use and privacy policy. They are then asked to provide data, which they agree to by clicking the "Provide" button.

[0781] Output: Consent and authorization information is sent to the server.

[0782] Step 2:

[0783] Installing the dedicated application

[0784] Input: User information with consent.

[0785] How it works: The user installs the dedicated application on their smartphone or PC, and the application is downloaded from the official website or app store.

[0786] Output: The dedicated application is installed on the user's device.

[0787] Step 3:

[0788] Automatic data collection and transmission

[0789] Input: The user's device on which the dedicated application is installed.

[0790] How it works: The device automatically collects data from social media accounts, call history, and search history. The collected data is encrypted and securely sent to a server.

[0791] Output: The encrypted data is sent to the server.

[0792] Step 4:

[0793] Data storage

[0794] Input: Encrypted data.

[0795] Specific operation: The server stores the received data in a database, for example, AWS RDS.

[0796] Output: Data stored in a database.

[0797] Step 5:

[0798] Performing data analysis

[0799] Input: Data stored in a database.

[0800] What it does: The server analyzes the collected data using natural language processing techniques (e.g., SpaCy or NLTK) to extract the user's language usage patterns, characteristic phrases, and behavioral patterns.

[0801] Output: Extracted feature data.

[0802] Step 6:

[0803] AI model generation

[0804] Input: Extracted feature data.

[0805] Specific operation: The server generates an AI model (e.g., Transformer model or GPT-4) based on the feature data and creates a virtual personality that mimics the user's characteristics.

[0806] Output: A virtual personality is generated.

[0807] Step 7:

[0808] Preparing for emotion recognition

[0809] Input: User interaction data with virtual persona.

[0810] How it works: The server integrates the emotion engine, which uses the emotion analysis API from Microsoft Azure Cognitive Services and is configured to analyze user emotions in real time from voice and text data.

[0811] Output: A system capable of emotion recognition.

[0812] Step 8:

[0813] Utilizing Emotional Data

[0814] Input: Real-time analyzed emotion data.

[0815] Specific operation: The server reflects the emotional data recognized by the emotion engine in the virtual personality's response. For example, if the user makes a sad voice, the virtual personality will respond with a comforting response.

[0816] Output: The virtual personality's response depending on the emotion.

[0817] Step 9:

[0818] Start a real-time conversation

[0819] Input: What the user says.

[0820] Specific operation: The user speaks into the device, which captures what is said as audio data and sends it to the server in real time.

[0821] Output: The audio data is sent to the server.

[0822] Step 10:

[0823] Voice data conversion and response generation

[0824] Input: User's voice data.

[0825] How it works: The server uses the Google Cloud Speech-to-Text API to convert the voice data into text, and the AI ​​model generates an appropriate response based on the converted text.

[0826] Output: The generated response text data.

[0827] Step 11:

[0828] Apply voice synthesis and voice changer

[0829] Input: Generated response text data.

[0830] Specific operation: The server converts the generated text into voice data using speech synthesis technology (e.g., Amazon Polly), and then uses a voice changer to adjust the voice data to match the user's voice.

[0831] Output: The adjusted audio data.

[0832] Step 12:

[0833] Playing audio data

[0834] Input: The adjusted audio data.

[0835] Specific operation: The device plays back the voice data received from the server, and the appropriate response is provided in the voice of the virtual personality.

[0836] Output: The audio played to the user.

[0837] Step 13:

[0838] Providing feedback

[0839] Input: User feedback provided during interaction with the virtual persona.

[0840] Specific operation: The user enters and submits their opinions and requests regarding the virtual personality through a feedback form or the like.

[0841] Output: Feedback data sent to the server.

[0842] Step 14:

[0843] Feedback Analysis and Adaptation

[0844] Input: Feedback data sent to the server.

[0845] Specific operation: The server analyzes the feedback and updates and adapts the AI ​​model and emotion engine. Specifically, the feedback content is analyzed using natural language processing technology and reflected in the AI ​​model's response patterns.

[0846] Output: Updated and adapted AI model and emotion engine.

[0847] This will make future interactions more natural and satisfying.

[0848] (Application example 2)

[0849] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0850] Conventional food delivery systems often lack sufficient support when users select menus or customize their meals. It can be difficult to get appropriate advice or suggestions, especially when users have special requests or want to try new dishes. Furthermore, there is a need for a personalized experience based on users' emotions and preferences. To address these challenges, a system is needed that allows users to easily select and customize menus while interacting with a virtual chef.

[0851] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0852] In this invention, the server includes: [means for collecting user record data;] [means for analyzing the collected record data and extracting features;] [means for generating a virtual personality based on the extracted features;] [means for having the virtual personality interact with the user;] [means for receiving feedback from the user during the interaction and adapting the virtual personality;] [means for the user to select a menu while conversing with the virtual chef when ordering; and [means for assisting in menu selection and dish customization.] This enables users to easily select and customize menus, providing a more personalized experience.

[0853] "User" refers to the person using the invention, who is primarily responsible for selecting menus and customizing meals through interaction with the virtual chef.

[0854] "Recorded data" refers to information including data related to the user's use of social media accounts, call history, search history, etc.

[0855] "Collect" refers to the process of collecting user record data through a dedicated application and sending it to a server.

[0856] "Analyzing" refers to the act of analyzing collected recorded data and extracting specific patterns or characteristics.

[0857] "Extracting features" means identifying a user's language usage, behavioral patterns, preferences, etc. from the analyzed data.

[0858] A "virtual personality" is an interactive agent that uses an AI model generated based on collected and analyzed characteristics.

[0859] "Having a dialogue" refers to the virtual personality interacting with the user in a conversational format.

[0860] "Feedback" refers to the responses and requests provided by users to the system.

[0861] "Adapting" refers to updating and adjusting the virtual persona's responses and behavior based on feedback.

[0862] "Ordering" refers to the act of a user requesting a meal using a food delivery service.

[0863] A "virtual chef" is a virtual personality that has particular knowledge about cooking and assists users in choosing menus through dialogue.

[0864] "Menu selection" is the process by which the virtual chef makes appropriate meal suggestions to the user and provides the best options.

[0865] "Cuisine customization" refers to making changes or adjustments to a dish according to the user's wishes.

[0866] "Voice data" refers to data that is a digital recording of a user's voice input.

[0867] "Converting to text data" refers to analyzing the voice data and changing it into text-format data.

[0868] A "reply" is a response that the generative AI model generates in response to a user's statement.

[0869] A "speech synthesizer" is a device or software that converts text data into speech data.

[0870] "Play" refers to making the synthesized speech available to the user.

[0871] "Consent" is the act of a user explicitly giving permission to use the system.

[0872] "Interface" refers to the means or screen through which a user interacts with a system.

[0873] To implement the present invention, the following system is required: This system is configured by incorporating a user's recorded data collection, analysis, virtual personality generation, dialogue execution, feedback adaptation, and emotion engine.

[0874] Data collection

[0875] Users agree to use the service and provide the necessary data, such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects the data. The collected data is securely transmitted from the device to a server, which then stores it in a database.

[0876] Data Analysis and Model Generation

[0877] The server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for a generative AI model. This generative AI model generates a virtual personality that mimics the user's behavior and characteristics.

[0878] Emotion engine integration

[0879] The server includes a model for recognizing the user's emotions using an emotion engine. The emotion engine analyzes emotions from the user's voice and text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the responses of the virtual personality.

[0880] Conversation Simulation

[0881] The user interacts with the virtual chef via a device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the generative AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology, and the synthesized voice data is sent to the device and played back to the user.

[0882] Feedback and Adaptation

[0883] Users can provide feedback to the virtual chef during the conversation. For example, if they have specific requests, such as "I'd like more humor" or "Please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the generative AI model and emotion engine, making future conversations more natural and satisfying.

[0884] Specific examples

[0885] For example, consider a scenario in which a user uses a smartphone app to say, "I want spicy pasta." The system captures the speech, converts it into text data, and then uses a generative AI model to derive the optimal response. An example of a prompt in this case is as follows:

[0886] User: "I want spicy pasta"

[0887] Virtual Chef: "For spicy pasta, I recommend Peperoncino or Arrabbiata. Which would you like?"

[0888] In this way, the virtual chef can provide a menu tailored to the user's needs, allowing for personalized meal selection and customization.

[0889] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0890] Step 1:

[0891] The user agrees to use the service and consents to providing data such as social media accounts, call history, and search history. This data is automatically collected by a dedicated application installed on the device. The device securely transmits the collected data to the server. This input data is then stored in a database.

[0892] Step 2:

[0893] The server analyzes the user's recorded data stored in the database. During this analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. These analysis results (language usage patterns and behavioral patterns) are extracted and used as input data for the generative AI model.

[0894] Step 3:

[0895] The server generates a virtual personality based on the extracted characteristics. Using a generative AI model, it creates a virtual chef that mimics the user's behavior and characteristics. The virtual chef can then assist with menu selection and cooking customization.

[0896] Step 4:

[0897] The user interacts with the virtual chef via a terminal. The terminal captures the user's speech as voice data and transmits it to the server in real time. The server converts this voice data into text data and uses it as input.

[0898] Step 5:

[0899] The server generates a response from the virtual chef based on the text data. This response is generated using a generative AI model and output as a response text. The response text is converted into audio data using speech synthesis technology. This audio data is sent to the device and played back to the user.

[0900] Step 6:

[0901] During the conversation, the user can provide feedback to the virtual chef, such as a request for more humor. The device then sends this feedback data to the server, which analyzes it and adjusts the parameters of the generative AI model and emotion engine to reflect the feedback in the next conversation.

[0902] Step 7:

[0903] The virtual chef will help users select menus and customize dishes to suit their needs. For example, if a user says, "I want spicy pasta," the virtual chef will suggest, "For spicy pasta, we recommend peperoncino or arrabbiata. Which would you like?"

[0904] This process allows users to make personalized meal selections and improves their experience.

[0905] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0906] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0907] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0908] [Third embodiment]

[0909] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0910] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0911] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0912] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0913] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0914] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0915] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0916] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0917] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0918] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0919] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0920] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0921] To practice the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[0922] First, regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[0923] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[0924] During the conversation simulation phase, the user interacts with the virtual personality via their device. The device captures the user's speech as voice data and transmits it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using speech synthesis technology and played back to the user via their device.

[0925] Feedback and Adaptation: Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that conversation," they can send them to the server through the feedback function. The server analyzes this feedback and adapts it to the AI ​​model, making future conversations more natural and satisfying.

[0926] Explanation with concrete examples

[0927] For example, suppose the family of a deceased person, Mr. B, wants to use this system.

[0928] 1. Data Collection

[0929] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[0930] A dedicated application installed on the device collects this data and sends it to a server.

[0931] The server stores the received data in a database.

[0932] 2. Data analysis and model generation

[0933] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[0934] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[0935] 3. Conversation Simulation

[0936] The user (B's family member) talks to the virtual B through the terminal.

[0937] The device captures the audio and sends it to the server.

[0938] The server converts the speech to text and uses an AI model to generate an appropriate response.

[0939] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[0940] The device plays back the response of virtual person B.

[0941] 4. Feedback and Adaptation

[0942] The user provides feedback on the interaction with virtual Person B.

[0943] The server analyzes the feedback and adjusts the AI ​​model.

[0944] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[0945] The processing flow will be explained below.

[0946] Step 1:

[0947] User: Agrees to use the service and agrees to provide the required data.

[0948] Device: Install the dedicated application and configure it to allow access to social media accounts, call history, and search history.

[0949] Step 2:

[0950] Device: Collects data such as users' social media posts, call history, and search history, and sends it to a server.

[0951] Server: Receives collected data and stores it securely in a database.

[0952] Step 3:

[0953] Server: Runs long-text data analysis algorithms to analyze the collected data and extract users' language usage patterns, characteristic phrases, and behavioral patterns.

[0954] Step 4:

[0955] Server: Generates an AI model based on the extracted information. This model creates a virtual personality that mimics the user's behavior and characteristics.

[0956] Step 5:

[0957] User: Uses the device to initiate a conversation with the virtual personality.

[0958] Device: The user's voice is captured by a microphone and sent to the server in real time.

[0959] Step 6:

[0960] Server: Converts the received voice data into text data using voice recognition technology.

[0961] Server: Inputs the converted text data into the AI ​​model and generates an appropriate response.

[0962] Server: The generated response is converted into voice data using speech synthesis technology, and then matched to the voice specified by the user using a voice changer.

[0963] Step 7:

[0964] Server: Sends the synthesized voice data to the device.

[0965] Terminal: The synthesized voice data is played back to the user through a speaker.

[0966] Step 8:

[0967] Users: Provide feedback during a conversation, for example by sending requests such as "I'd like more humor" or "I'd like more detail on a particular topic" through the feedback feature.

[0968] Server: Analyzes the received feedback and updates and adapts the AI ​​model to make future conversations more natural and satisfying.

[0969] Example 1

[0970] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0971] In modern society, there is a demand for technology that can recreate conversations with deceased loved ones and provide emotional support. In particular, conventional systems have had problems with insufficient analysis of collected data, resulting in poor conversation quality. Furthermore, it has been difficult to quickly and appropriately incorporate user feedback. This has led to issues such as reduced user satisfaction.

[0972] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0973] In this invention, the server includes: [means for collecting user record data and securely transmitting it to the server;] [means for analyzing the collected record data and extracting features using natural language processing technology; and] [means for generating a virtual personality using machine learning based on the extracted features.] This makes it possible [to accurately mimic the user's behavior and personality, achieve high-quality dialogue, and quickly and appropriately reflect feedback].

[0974] "Means for collecting user record data and safely transmitting it to a server" refers to a mechanism for collecting data such as users' social media accounts, call history, and search history, and transferring it to a server using security technologies such as encryption.

[0975] "Means of analyzing collected recorded data and extracting features using natural language processing technology" refers to a method of analyzing data received by the server and extracting phrases, behavioral patterns, interests, etc. from large amounts of text data.

[0976] "Means for generating a virtual personality using machine learning based on extracted features" refers to the process of using a machine learning algorithm to generate a virtual person that mimics the user's behavior and personality based on the analyzed data.

[0977] The "means for implementing speech recognition and speech synthesis technology to enable the generated virtual personality to interact with the user" refers to a series of processes for converting the user's speech input into text and converting the generated virtual personality's responses back into speech to provide to the user.

[0978] "Means for receiving feedback from users during a dialogue, analyzing it, and adapting the virtual personality" refers to a mechanism that collects opinions and requests provided by users during a dialogue, and uses that data to improve and adapt the behavior and comments of the generated virtual personality.

[0979] "Means for collecting user voice and transmitting it to a server in real time" refers to a method for capturing voice data in real time using a device such as a microphone and immediately transmitting it to a server.

[0980] "Means for using speech recognition technology to convert voice data into text data at the server" refers to technology that uses a speech recognition algorithm to accurately convert voice data into text data.

[0981] "Means of generating a response using an AI model based on text data and converting that response into voice using speech synthesis technology" refers to the process of converting a text response generated by an AI model into voice data using a speech synthesis algorithm.

[0982] "Means for playing the synthesized voice to the user through the terminal" refers to a method for delivering the generated voice data to the user using the terminal's speaker or the like.

[0983] "Means for providing an interface for users to provide consent to data provision" refers to a mechanism that provides a screen or interactive procedure for users to agree to data provision.

[0984] "Means for automatically starting data collection upon receiving user consent" refers to the process by which the system automatically starts collecting the necessary data after the user's consent is obtained.

[0985] MODE FOR CARRYING OUT THE INVENTION

[0986] To implement the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[0987] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is encrypted and securely sent from the device to the server, where it is stored in a database.

[0988] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technologies such as Google Speech-to-Text and Amazon Polly are used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[0989] During the conversation simulation phase, the user interacts with the virtual persona via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using voice synthesis technology such as Amazon Polly and played back to the user via the device.

[0990] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that story," they can send them to the server through the feedback function. The server then analyzes this feedback and adapts the AI ​​model, making future conversations more natural and satisfying.

[0991] Explanation with concrete examples

[0992] For example, suppose the family of a deceased person, Mr. B, wants to use this system:

[0993] 1. Data Collection

[0994] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[0995] A dedicated application installed on the device collects this data and sends it to a server.

[0996] The server stores the received data in a database.

[0997] 2. Data analysis and model generation

[0998] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[0999] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[1000] 3. Conversation Simulation

[1001] The user (B's family member) talks to the virtual B through the terminal.

[1002] The device captures the audio and sends it to the server.

[1003] The server converts the speech to text and uses an AI model to generate an appropriate response.

[1004] The generated response is converted into voice data using voice synthesis technology such as Amazon Polly and sent to the device.

[1005] The device plays back the response of virtual person B.

[1006] 4. Feedback and Adaptation

[1007] The user provides feedback on the interaction with virtual Person B.

[1008] The server analyzes the feedback and adjusts the AI ​​model.

[1009] Prompt Sentence Examples

[1010] "Person B often uses the phrase 'do your best'. Please generate words of encouragement in a tone that is characteristic of Person B."

[1011] "I'd like you to tell me about a movie that Mr. B liked."

[1012] "Tell me how Mr. B tells jokes."

[1013] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[1014] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1015] Step 1: Data collection

[1016] The user agrees to use the service and consents to the provision of data such as social media accounts, call history, search history, etc. User consent and permission are required as input.

[1017] A dedicated application installed on the device acquires the data that can be collected. In operation, the application periodically and automatically collects posts from social media accounts, call history, and user search history.

[1018] The device securely encrypts the collected data and sends it to the server, resulting in the encrypted data being sent to the server.

[1019] The server stores the received data in a database. In operation, the server appropriately classifies the received data and saves it in a database.

[1020] Step 2: Data analysis and model generation

[1021] The server analyzes the data stored in the database, using saved user social media posts, call history, search history, etc. as input.

[1022] The server uses natural language processing technology to extract features such as the user's phrasing and behavioral patterns. In operation, the NLP algorithm extracts the user's unique expressions and interests from the data. The extracted feature data is obtained as the output.

[1023] The server uses machine learning algorithms to generate an AI model based on the extracted features. The AI ​​model is built using tools such as Python and TensorFlow. The output is a virtual personality model that mimics the user's behavior.

[1024] Step 3: Conversation simulation

[1025] The user initiates a dialogue with the virtual persona through a terminal, and the user's voice input is required.

[1026] The device captures the user's speech as voice data and sends it to the server in real time. The device's microphone collects the voice and uses a voice recognition API to transfer the data to the server. The output is real-time voice data sent to the server.

[1027] The server converts the voice data into text data and generates a response using an AI model. The operation is to convert the voice to text using the Google Speech-to-Text API, and then use an AI model (e.g., GPT-3) to generate an appropriate response. The output is the generated text response.

[1028] The text response generated by the server is converted into voice data using speech synthesis technology and sent to the device. The text is converted into voice using speech synthesis technology such as Amazon Polly and returned to the device. The voice data is then sent to the device as an output.

[1029] The device plays back the virtual persona's response in real time. In operation, the device's speaker plays back the received audio data. As an output, the virtual persona's response is provided to the user through audio.

[1030] Step 4: Feedback and Adapt

[1031] The user provides feedback on the interaction with the virtual personality. The input requires the user's feedback, which may include specific requests such as "Tell more jokes."

[1032] The server analyzes the feedback provided by the user and adjusts the AI ​​model. The operation is to analyze the feedback and incorporate it as new data into the training of the AI ​​model. The output is the adjusted AI model.

[1033] By repeating this series of processes, the system adapts its interactions with the user to make them more natural and satisfying.

[1034] (Application example 1)

[1035] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1036] In modern society, there is a growing demand for content distribution services that go beyond simply viewing video content such as movies and dramas to provide a more fulfilling entertainment experience. However, conventional systems lack a means to empathize with users' feelings of loneliness or the enjoyment of watching a movie. Furthermore, interactive means for providing emotional support are limited. There is a need to solve these issues and provide users with a more enjoyable and fulfilling viewing experience.

[1037] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1038] In this invention, the server includes means for collecting recorded data of a user, means for analyzing the collected recorded data and extracting features, means for generating a virtual personality based on the extracted features, means for the virtual personality to interact with the user, means for receiving feedback from the user during the interaction and adapting the virtual personality, and means for the user and the virtual personality to interact while playing video content. This allows the user to interact with a virtual partner while watching a movie or drama, making for a more fulfilling viewing experience.

[1039] "User" means an individual or organization that uses this system.

[1040] "Recorded data" refers to activity data such as a user's social media accounts, call history, and search history.

[1041] "Means of collection" refers to methods of collecting data using applications or sensors installed on the user's device.

[1042] "Means for analysis" refers to software or algorithms used to analyze the collected recorded data and extract features.

[1043] "Characteristics" are patterns that indicate tendencies in a user's behavior, personality, interests, etc.

[1044] A "virtual personality" is an AI model that is generated based on collected characteristics and allows users to interact with it.

[1045] "Means for dialogue" refers to a function that allows the user and the virtual personality to communicate via voice or text.

[1046] "Means for receiving feedback" and "adapting" refer to methods for updating the virtual personality generation model by reflecting opinions and requests from users.

[1047] "Video content" refers to media content that can be viewed, such as movies and dramas.

[1048] "Means for a user to have a conversation with a virtual character during playback" refers to a technology that allows a user to have a conversation with a virtual character in real time while viewing video content.

[1049] To implement this invention, the following system configuration and programs are required: This system is made up of several main components, each of which plays a specific role.

[1050] System Configuration

[1051] 1. User's device

[1052] Data Collection Applications

[1053] Video playback application

[1054] Audio input and output devices (microphone, speaker)

[1055] 2. Server

[1056] Database

[1057] Data Analysis Module

[1058] AI model generation module

[1059] Speech-to-text module (Google Speech-to-Text API)

[1060] Text-to-speech module (Google Text-to-Speech API)

[1061] WebSocket Server

[1062] Processing flow

[1063] 1. Data Collection

[1064] The user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, search history, etc. The device sends this data to the server, which then stores the collected data in a database.

[1065] 2. Data analysis and AI model generation

[1066] The server retrieves the collected data from the database and extracts features using natural language processing technology (TensorFlow, NLTK, spaCy). Based on these features, a virtual personality is generated that mimics the user's behavior and personality.

[1067] 3. Playback and interaction with video content

[1068] While a user is playing a movie or TV drama using a video playback application, a virtual persona will engage in conversation related to the video content. When the user speaks, the device captures the audio data and sends it to the server via WebSocket. The server converts the audio data into text and uses an AI model to generate an appropriate response. This response is then converted back into audio data by a text-to-speech module and played back to the user via the device.

[1069] 4. Feedback and Adaptation

[1070] Users can provide feedback to the virtual personality during the interaction, for example, if they want it to tell more jokes, they can send that feedback to the server through the feedback function, which will analyze this feedback and adapt the AI ​​model.

[1071] Specific examples

[1072] For example, suppose a user is watching the movie "Castle in the Sky," and the virtual partner imitates the personality of their deceased best friend, Mr. A.

[1073] Prompt Sentence Examples

[1074] The user is watching the movie "Castle in the Sky." The virtual partner is the deceased best friend, Mr. A:

[1075] User: "You also liked this scene, right?"

[1076] Virtual Partner: "Yes, the music in this scene is particularly memorable."

[1077] An example of a prompt sentence to input to the generative AI model is as follows:

[1078] User social media data:

[1079] Posted yesterday: Looking forward to seeing "Castle in the Sky"!

[1080] Call History:

[1081] My best friend A and I often talked about movies.

[1082] Search History:

[1083] A moving scene from "Castle in the Sky"

[1084] feedback:

[1085] I'd like to hear more detailed feedback

[1086] This system allows users to not only watch video content, but also enjoy interacting with virtual characters, further enhancing the viewing experience.

[1087] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1088] Step 1:

[1089] A user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, and search history. Based on this consent, the device automatically collects this data and sends it to a database. The input is the user's social media accounts, call history, and search history. The output is the data securely sent to the server.

[1090] Step 2:

[1091] The server retrieves the collected data from the database and extracts features using natural language processing techniques (TensorFlow, NLTK, spaCy). Specifically, the data is first cleansed, and then a language model is applied to identify interest and behavioral patterns. The input is the user's recorded data retrieved from the database. The output is analyzed data containing the extracted features.

[1092] Step 3:

[1093] The server builds an AI model for generating a virtual personality based on the extracted features. Specifically, it trains a neural network model that mimics the user's behavior and personality based on the feature data. The input is the analysis data obtained in step 2. The output is an AI model that represents the virtual personality.

[1094] Step 4:

[1095] While a user plays a movie or drama using a video playback application, the virtual persona interacts with the user about the video content being viewed. Specifically, the voice input is captured and the voice data is sent to the server. The input is the user's voice data. The output is the data sent to the server.

[1096] Step 5:

[1097] The server converts the voice data into text using a speech-to-text module (Google Speech-to-Text API). The text data is then analyzed by an AI model to generate an appropriate response. The input is the user's voice data. The output is text data generated by the virtual personality.

[1098] Step 6:

[1099] The generated text data is converted into voice data using speech synthesis technology (Google Text-to-Speech API) and sent to the device. The input is text data generated by the virtual personality. The output is voice data.

[1100] Step 7:

[1101] The terminal plays the received voice data and continues the dialogue with the user. Specifically, the user listens to the virtual personality's response. The input is the voice data sent from the server. The output is the voice heard by the user.

[1102] Step 8:

[1103] During the interaction, the user provides feedback, for example a specific request such as "tell more jokes." The device collects this feedback and sends it to the server. The input is the user's feedback. The output is the feedback data sent to the server.

[1104] Step 9:

[1105] The server analyzes the feedback data and updates the AI ​​model. Specifically, it adjusts the neural network parameters based on the feedback and adapts the model so that future interactions are more natural. The input is the feedback data. The output is the updated AI model.

[1106] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1107] To implement the present invention, a system can be constructed and a program can be executed according to the following procedure, which includes collecting and analyzing recorded data, generating a virtual personality, conducting dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[1108] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and data collection is performed automatically by this application. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[1109] Regarding data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for generating an AI model. This AI model generates a virtual personality that mimics the user's behavior and characteristics.

[1110] Regarding the integration of the emotion engine, the server includes a model for recognizing the user's emotions using the emotion engine. The emotion engine analyzes emotions from the user's voice or text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the response of the virtual personality.

[1111] In the conversation simulation stage, the user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is converted into voice data using voice synthesis technology and adjusted to the voice specified by the user using a voice changer. The synthesized voice data is sent to the device, and the Reiwa voice data is played back to the user.

[1112] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "please use more humor" or "please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[1113] Explanation with concrete examples

[1114] For example, if the family of a deceased person, Mr. C, wants to use this system, the process is as follows:

[1115] 1. Data Collection

[1116] The user (Mr. C's family member) agrees to the service and consents to providing Mr. C's social media account, call history, and search history.

[1117] A dedicated application installed on the device collects this data and sends it to a server.

[1118] The server stores the received data in a database.

[1119] 2. Data analysis and model generation

[1120] The server analyzes the data and extracts Mr. C's characteristic phrases and interests.

[1121] An AI model is created that generates a virtual personality for Mr. C based on the analysis data.

[1122] 3. Emotion engine integration

[1123] The server recognizes emotions from the user's voice and text, and reflects that emotional data in the virtual personality's responses.

[1124] 4. Conversation Simulation

[1125] The user (Mr. C's family member) talks to the virtual Mr. C through the terminal.

[1126] The device captures the audio and sends it to the server.

[1127] The server converts the speech to text and uses an AI model to generate an appropriate response.

[1128] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[1129] The device plays back the virtual person C's reply.

[1130] 5. Feedback and Adaptation

[1131] The user provides feedback on the interaction with virtual person C.

[1132] The server analyzes the feedback and adjusts the AI ​​model and emotion engine.

[1133] In this way, the system based on this invention enables users to have conversations with their deceased loved ones and provide emotional support. The integration of an emotion engine enables more natural and emotionally sensitive conversations.

[1134] The processing flow will be explained below.

[1135] Step 1:

[1136] User: Agrees to use the service and consents to providing the necessary data. Completes procedures to allow the collection of data such as social media accounts, call history, and search history.

[1137] Step 2:

[1138] Device: Install the dedicated application and complete the setup. The application will automatically collect data such as user's social media posts, call history, and search history, and send it to the server.

[1139] Step 3:

[1140] Server: Receives recorded data sent from the device and stores it in a secure database.

[1141] Step 4:

[1142] Server: Analyzes the transmitted recorded data. Using natural language processing technology, it extracts the user's language usage patterns, characteristic phrases, and behavioral patterns.

[1143] Step 5:

[1144] Server: Generates an AI model based on the extracted feature information. This AI model creates a virtual personality that mimics the user's behavior and characteristics.

[1145] Step 6:

[1146] Server: Integrates an emotion engine into the virtual personality. The emotion engine analyzes emotions from the user's voice and text in real time and generates emotion data.

[1147] Step 7:

[1148] User: Uses the device to initiate a conversation with the virtual persona. The device captures the user's voice with a microphone.

[1149] Step 8:

[1150] Terminal: Sends captured audio data to the server in real time.

[1151] Step 9:

[1152] Server: The received voice data is converted into text data using voice recognition technology. The emotion engine then analyzes the user's emotions from the voice data and generates emotion data.

[1153] Step 10:

[1154] Server: The AI ​​model generates appropriate responses based on the text and emotional data. The responses are tailored to the characteristics of the virtual personality and take into account the emotional data.

[1155] Step 11:

[1156] Server: The generated response is converted into voice data using speech synthesis technology. A voice changer is used to match the tone of the virtual personality.

[1157] Step 12:

[1158] Server: The synthesized voice data is sent back to the device.

[1159] Step 13:

[1160] Terminal: Plays the received audio data to the user through the speaker.

[1161] Step 14:

[1162] User: Provides feedback to the virtual persona during the interaction. Feedback can include specific requests such as "Please tell me more" or "Talk about a different topic."

[1163] Step 15:

[1164] Terminal: Sends feedback data to the server in real time.

[1165] Step 16:

[1166] Server: Analyzes the received feedback and applies it to the AI ​​model and emotion engine. It adjusts the response and behavior patterns of the virtual personality based on the feedback.

[1167] In this way, the present invention provides a system that allows interaction with deceased loved ones and is flexible in adapting to the user's emotions and feedback.

[1168] Example 2

[1169] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1170] Current virtual dialogue systems lack sufficient recognition of user emotions, resulting in mechanical and unnatural dialogue. Furthermore, it is difficult to effectively incorporate user feedback, making it difficult to continuously improve the quality of dialogue. Furthermore, voice conversion lacks flexibility, and there is a lack of technology available to adjust the voice to suit the user's preferences. These challenges make it difficult to provide a satisfying dialogue experience for users.

[1171] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1172] In this invention, the server includes means for collecting user record data, means for analyzing the collected record data and extracting features, and means for generating a virtual personality based on the extracted features. This makes it possible to recognize the user's emotions using an emotion engine and reflect them in the dialogue, collect and analyze the user's voice in real time, and adapt the AI ​​model and emotion engine based on the feedback.

[1173] "User" refers to an individual who interacts with the System.

[1174] "Recorded data" refers to data provided by users, such as social media accounts, call history, and search history.

[1175] "Data collection" refers to the process of automatically collecting recorded data using a dedicated application based on the user's consent.

[1176] "Database" refers to a repository for storing recorded data collected by the Server.

[1177] "Data analysis" refers to the process of analyzing collected recorded data using natural language processing techniques and extracting features.

[1178] "Characteristics" refers to characteristics such as users' language usage patterns and behavioral patterns extracted through data analysis.

[1179] "Virtual personality" refers to a virtual conversation partner that mimics the user's characteristics and is generated using an AI model based on extracted characteristics.

[1180] "Dialogue" refers to communication between the user and the virtual personality.

[1181] "Feedback" refers to the opinions and requests that a user provides to a virtual personality during a conversation.

[1182] "Adaptation" refers to the process of updating and improving AI models and emotion engines based on the feedback provided.

[1183] An "emotion engine" refers to a system that analyzes and recognizes users' emotions from voice and text data in real time.

[1184] "Audio Data" means digital audio information that captures a user's speech.

[1185] "Text data" refers to character information obtained by converting voice data.

[1186] "Speech synthesis" refers to the process of generating digital speech from text data.

[1187] "Voice changer" refers to technology that adjusts the generated voice to a voice specified by the user.

[1188] "Server" refers to the computer system that collects and analyzes data, generates virtual personalities, and operates the emotion engine.

[1189] "Dedicated application" refers to software that is installed on the user's device and is used to collect and transmit recorded data.

[1190] "Consent interface" refers to the screen or function that allows users to consent to providing data.

[1191] To implement this invention, the system is constructed and the program is executed according to the following procedure: The system includes collecting and analyzing recorded data, generating a virtual personality, executing dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[1192] Data collection

[1193] The user agrees to use the service and consents to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[1194] Examples:

[1195] The user agrees to the terms of service and installs the dedicated application on their smartphone.

[1196] A dedicated application automatically collects social media accounts and call history, encrypts them, and sends them to a server.

[1197] Data Analysis and Model Generation

[1198] The server analyzes the collected data using natural language processing technology (such as SpaCy or NLTK) to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. An AI model (such as the Transformer model or GPT-4) is generated based on the extracted features, creating a virtual personality that mimics the user's characteristics.

[1199] Emotion engine integration

[1200] The server contains a model for recognizing the user's emotions using an emotion engine (such as the Microsoft Azure Cognitive Services emotion analysis API). The emotion engine analyzes the user's voice and text in real time to determine their emotional state. The determined emotional data is used to adjust the virtual personality's responses.

[1201] Examples:

[1202] The server analyzes the collected social media posts and call history to extract the user's characteristic phrases and interests.

[1203] An AI model is created that generates a virtual personality based on the extracted data and an emotion engine is integrated.

[1204] Conversation Simulation

[1205] The user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and an AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology (such as Amazon Polly), and a voice changer adjusts it to the voice specified by the user. The synthesized voice data is sent to the device, where it is played back to the user.

[1206] Examples:

[1207] The user speaks to the virtual persona, "What's the weather like today?"

[1208] The server analyzes this, and the virtual personality generates a reply saying "It's sunny today," converts it into voice, and sends it to the terminal.

[1209] The voice that has been adjusted by the voice changer is played back to the user.

[1210] Feedback and Adaptation

[1211] Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "I want more humor" or "I want you to elaborate on a particular topic," they can send these to the server through the feedback function. The server analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[1212] Example prompt sentence:

[1213] "Based on C's social media account data, please extract characteristic phrases and behavioral patterns to create a virtual personality for C."

[1214] This invention allows users to have a more natural and emotionally relevant interaction experience.

[1215] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1216] Step 1:

[1217] Service consent and data provision consent

[1218] Input: The user agrees to the terms of use of the service and consents to providing data such as social media accounts, call history, and search history.

[1219] Specific behavior: The user clicks the "Agree" button on the screen where they agree to the terms of use and privacy policy. They are then asked to provide data, which they agree to by clicking the "Provide" button.

[1220] Output: Consent and authorization information is sent to the server.

[1221] Step 2:

[1222] Installing the dedicated application

[1223] Input: User information with consent.

[1224] How it works: The user installs the dedicated application on their smartphone or PC, and the application is downloaded from the official website or app store.

[1225] Output: The dedicated application is installed on the user's device.

[1226] Step 3:

[1227] Automatic data collection and transmission

[1228] Input: The user's device on which the dedicated application is installed.

[1229] How it works: The device automatically collects data from social media accounts, call history, and search history. The collected data is encrypted and securely sent to a server.

[1230] Output: The encrypted data is sent to the server.

[1231] Step 4:

[1232] Data storage

[1233] Input: Encrypted data.

[1234] Specific operation: The server stores the received data in a database, for example, AWS RDS.

[1235] Output: Data stored in a database.

[1236] Step 5:

[1237] Performing data analysis

[1238] Input: Data stored in a database.

[1239] What it does: The server analyzes the collected data using natural language processing techniques (e.g., SpaCy or NLTK) to extract the user's language usage patterns, characteristic phrases, and behavioral patterns.

[1240] Output: Extracted feature data.

[1241] Step 6:

[1242] AI model generation

[1243] Input: Extracted feature data.

[1244] Specific operation: The server generates an AI model (e.g., Transformer model or GPT-4) based on the feature data and creates a virtual personality that mimics the user's characteristics.

[1245] Output: A virtual personality is generated.

[1246] Step 7:

[1247] Preparing for emotion recognition

[1248] Input: User interaction data with virtual persona.

[1249] How it works: The server integrates the emotion engine, which uses the emotion analysis API from Microsoft Azure Cognitive Services and is configured to analyze user emotions in real time from voice and text data.

[1250] Output: A system capable of emotion recognition.

[1251] Step 8:

[1252] Utilizing Emotional Data

[1253] Input: Real-time analyzed emotion data.

[1254] Specific operation: The server reflects the emotional data recognized by the emotion engine in the virtual personality's response. For example, if the user makes a sad voice, the virtual personality will respond with a comforting response.

[1255] Output: The virtual personality's response depending on the emotion.

[1256] Step 9:

[1257] Start a real-time conversation

[1258] Input: What the user says.

[1259] Specific operation: The user speaks into the device, which captures what is said as audio data and sends it to the server in real time.

[1260] Output: The audio data is sent to the server.

[1261] Step 10:

[1262] Voice data conversion and response generation

[1263] Input: User's voice data.

[1264] How it works: The server uses the Google Cloud Speech-to-Text API to convert the voice data into text, and the AI ​​model generates an appropriate response based on the converted text.

[1265] Output: The generated response text data.

[1266] Step 11:

[1267] Apply voice synthesis and voice changer

[1268] Input: Generated response text data.

[1269] Specific operation: The server converts the generated text into voice data using speech synthesis technology (e.g., Amazon Polly), and then uses a voice changer to adjust the voice data to match the user's voice.

[1270] Output: The adjusted audio data.

[1271] Step 12:

[1272] Playing audio data

[1273] Input: The adjusted audio data.

[1274] Specific operation: The device plays back the voice data received from the server, and the appropriate response is provided in the voice of the virtual personality.

[1275] Output: The audio played to the user.

[1276] Step 13:

[1277] Providing feedback

[1278] Input: User feedback provided during interaction with the virtual persona.

[1279] Specific operation: The user enters and submits their opinions and requests regarding the virtual personality through a feedback form or the like.

[1280] Output: Feedback data sent to the server.

[1281] Step 14:

[1282] Feedback Analysis and Adaptation

[1283] Input: Feedback data sent to the server.

[1284] Specific operation: The server analyzes the feedback and updates and adapts the AI ​​model and emotion engine. Specifically, the feedback content is analyzed using natural language processing technology and reflected in the AI ​​model's response patterns.

[1285] Output: Updated and adapted AI model and emotion engine.

[1286] This will make future interactions more natural and satisfying.

[1287] (Application example 2)

[1288] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1289] Conventional food delivery systems often lack sufficient support when users select menus or customize their meals. It can be difficult to get appropriate advice or suggestions, especially when users have special requests or want to try new dishes. Furthermore, there is a need for a personalized experience based on users' emotions and preferences. To address these challenges, a system is needed that allows users to easily select and customize menus while interacting with a virtual chef.

[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1291] In this invention, the server includes: [means for collecting user record data;] [means for analyzing the collected record data and extracting features;] [means for generating a virtual personality based on the extracted features;] [means for having the virtual personality interact with the user;] [means for receiving feedback from the user during the interaction and adapting the virtual personality;] [means for the user to select a menu while conversing with the virtual chef when ordering; and [means for assisting in menu selection and dish customization.] This enables users to easily select and customize menus, providing a more personalized experience.

[1292] "User" refers to the person using the invention, who is primarily responsible for selecting menus and customizing meals through interaction with the virtual chef.

[1293] "Recorded data" refers to information including data related to the user's use of social media accounts, call history, search history, etc.

[1294] "Collect" refers to the process of collecting user record data through a dedicated application and sending it to a server.

[1295] "Analyzing" refers to the act of analyzing collected recorded data and extracting specific patterns or characteristics.

[1296] "Extracting features" means identifying a user's language usage, behavioral patterns, preferences, etc. from the analyzed data.

[1297] A "virtual personality" is an interactive agent that uses an AI model generated based on collected and analyzed characteristics.

[1298] "Having a dialogue" refers to the virtual personality interacting with the user in a conversational format.

[1299] "Feedback" refers to the responses and requests provided by users to the system.

[1300] "Adapting" refers to updating and adjusting the virtual persona's responses and behavior based on feedback.

[1301] "Ordering" refers to the act of a user requesting a meal using a food delivery service.

[1302] A "virtual chef" is a virtual personality that has particular knowledge about cooking and assists users in choosing menus through dialogue.

[1303] "Menu selection" is the process by which the virtual chef makes appropriate meal suggestions to the user and provides the best options.

[1304] "Cuisine customization" refers to making changes or adjustments to a dish according to the user's wishes.

[1305] "Voice data" refers to data that is a digital recording of a user's voice input.

[1306] "Converting to text data" refers to analyzing the voice data and changing it into text-format data.

[1307] A "reply" is a response that the generative AI model generates in response to a user's statement.

[1308] A "speech synthesizer" is a device or software that converts text data into speech data.

[1309] "Play" refers to making the synthesized speech available to the user.

[1310] "Consent" is the act of a user explicitly giving permission to use the system.

[1311] "Interface" refers to the means or screen through which a user interacts with a system.

[1312] To implement the present invention, the following system is required: This system is configured by incorporating a user's recorded data collection, analysis, virtual personality generation, dialogue execution, feedback adaptation, and emotion engine.

[1313] Data collection

[1314] Users agree to use the service and provide the necessary data, such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects the data. The collected data is securely transmitted from the device to a server, which then stores it in a database.

[1315] Data Analysis and Model Generation

[1316] The server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for a generative AI model. This generative AI model generates a virtual personality that mimics the user's behavior and characteristics.

[1317] Emotion engine integration

[1318] The server includes a model for recognizing the user's emotions using an emotion engine. The emotion engine analyzes emotions from the user's voice and text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the responses of the virtual personality.

[1319] Conversation Simulation

[1320] The user interacts with the virtual chef via a device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the generative AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology, and the synthesized voice data is sent to the device and played back to the user.

[1321] Feedback and Adaptation

[1322] Users can provide feedback to the virtual chef during the conversation. For example, if they have specific requests, such as "I'd like more humor" or "Please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the generative AI model and emotion engine, making future conversations more natural and satisfying.

[1323] Specific examples

[1324] For example, consider a scenario in which a user uses a smartphone app to say, "I want spicy pasta." The system captures the speech, converts it into text data, and then uses a generative AI model to derive the optimal response. An example of a prompt in this case is as follows:

[1325] User: "I want spicy pasta"

[1326] Virtual Chef: "For spicy pasta, I recommend Peperoncino or Arrabbiata. Which would you like?"

[1327] In this way, the virtual chef can provide a menu tailored to the user's needs, allowing for personalized meal selection and customization.

[1328] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1329] Step 1:

[1330] The user agrees to use the service and consents to providing data such as social media accounts, call history, and search history. This data is automatically collected by a dedicated application installed on the device. The device securely transmits the collected data to the server. This input data is then stored in a database.

[1331] Step 2:

[1332] The server analyzes the user's recorded data stored in the database. During this analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. These analysis results (language usage patterns and behavioral patterns) are extracted and used as input data for the generative AI model.

[1333] Step 3:

[1334] The server generates a virtual personality based on the extracted characteristics. Using a generative AI model, it creates a virtual chef that mimics the user's behavior and characteristics. The virtual chef can then assist with menu selection and cooking customization.

[1335] Step 4:

[1336] The user interacts with the virtual chef via a terminal. The terminal captures the user's speech as voice data and transmits it to the server in real time. The server converts this voice data into text data and uses it as input.

[1337] Step 5:

[1338] The server generates a response from the virtual chef based on the text data. This response is generated using a generative AI model and output as a response text. The response text is converted into audio data using speech synthesis technology. This audio data is sent to the device and played back to the user.

[1339] Step 6:

[1340] During the conversation, the user can provide feedback to the virtual chef, such as a request for more humor. The device then sends this feedback data to the server, which analyzes it and adjusts the parameters of the generative AI model and emotion engine to reflect the feedback in the next conversation.

[1341] Step 7:

[1342] The virtual chef will help users select menus and customize dishes to suit their needs. For example, if a user says, "I want spicy pasta," the virtual chef will suggest, "For spicy pasta, we recommend peperoncino or arrabbiata. Which would you like?"

[1343] This process allows users to make personalized meal selections and improves their experience.

[1344] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1345] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1346] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1347] [Fourth embodiment]

[1348] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1349] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1350] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1351] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1352] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1353] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1354] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1355] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1356] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1357] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1358] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1359] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1360] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1361] To practice the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[1362] First, regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[1363] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[1364] During the conversation simulation phase, the user interacts with the virtual personality via their device. The device captures the user's speech as voice data and transmits it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using speech synthesis technology and played back to the user via their device.

[1365] Feedback and Adaptation: Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that conversation," they can send them to the server through the feedback function. The server analyzes this feedback and adapts it to the AI ​​model, making future conversations more natural and satisfying.

[1366] Explanation with concrete examples

[1367] For example, suppose the family of a deceased person, Mr. B, wants to use this system.

[1368] 1. Data Collection

[1369] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[1370] A dedicated application installed on the device collects this data and sends it to a server.

[1371] The server stores the received data in a database.

[1372] 2. Data analysis and model generation

[1373] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[1374] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[1375] 3. Conversation Simulation

[1376] The user (B's family member) talks to the virtual B through the terminal.

[1377] The device captures the audio and sends it to the server.

[1378] The server converts the speech to text and uses an AI model to generate an appropriate response.

[1379] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[1380] The device plays back the response of virtual person B.

[1381] 4. Feedback and Adaptation

[1382] The user provides feedback on the interaction with virtual Person B.

[1383] The server analyzes the feedback and adjusts the AI ​​model.

[1384] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[1385] The processing flow will be explained below.

[1386] Step 1:

[1387] User: Agrees to use the service and agrees to provide the required data.

[1388] Device: Install the dedicated application and configure it to allow access to social media accounts, call history, and search history.

[1389] Step 2:

[1390] Device: Collects data such as users' social media posts, call history, and search history, and sends it to a server.

[1391] Server: Receives collected data and stores it securely in a database.

[1392] Step 3:

[1393] Server: Runs long-text data analysis algorithms to analyze the collected data and extract users' language usage patterns, characteristic phrases, and behavioral patterns.

[1394] Step 4:

[1395] Server: Generates an AI model based on the extracted information. This model creates a virtual personality that mimics the user's behavior and characteristics.

[1396] Step 5:

[1397] User: Uses the device to initiate a conversation with the virtual personality.

[1398] Device: The user's voice is captured by a microphone and sent to the server in real time.

[1399] Step 6:

[1400] Server: Converts the received voice data into text data using voice recognition technology.

[1401] Server: Inputs the converted text data into the AI ​​model and generates an appropriate response.

[1402] Server: The generated response is converted into voice data using speech synthesis technology, and then matched to the voice specified by the user using a voice changer.

[1403] Step 7:

[1404] Server: Sends the synthesized voice data to the device.

[1405] Terminal: The synthesized voice data is played back to the user through a speaker.

[1406] Step 8:

[1407] Users: Provide feedback during a conversation, for example by sending requests such as "I'd like more humor" or "I'd like more detail on a particular topic" through the feedback feature.

[1408] Server: Analyzes the received feedback and updates and adapts the AI ​​model to make future conversations more natural and satisfying.

[1409] Example 1

[1410] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1411] In modern society, there is a demand for technology that can recreate conversations with deceased loved ones and provide emotional support. In particular, conventional systems have had problems with insufficient analysis of collected data, resulting in poor conversation quality. Furthermore, it has been difficult to quickly and appropriately incorporate user feedback. This has led to issues such as reduced user satisfaction.

[1412] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1413] In this invention, the server includes: [means for collecting user record data and securely transmitting it to the server;] [means for analyzing the collected record data and extracting features using natural language processing technology; and] [means for generating a virtual personality using machine learning based on the extracted features.] This makes it possible [to accurately mimic the user's behavior and personality, achieve high-quality dialogue, and quickly and appropriately reflect feedback].

[1414] "Means for collecting user record data and safely transmitting it to a server" refers to a mechanism for collecting data such as users' social media accounts, call history, and search history, and transferring it to a server using security technologies such as encryption.

[1415] "Means of analyzing collected recorded data and extracting features using natural language processing technology" refers to a method of analyzing data received by the server and extracting phrases, behavioral patterns, interests, etc. from large amounts of text data.

[1416] "Means for generating a virtual personality using machine learning based on extracted features" refers to the process of using a machine learning algorithm to generate a virtual person that mimics the user's behavior and personality based on the analyzed data.

[1417] The "means for implementing speech recognition and speech synthesis technology to enable the generated virtual personality to interact with the user" refers to a series of processes for converting the user's speech input into text and converting the generated virtual personality's responses back into speech to provide to the user.

[1418] "Means for receiving feedback from users during a dialogue, analyzing it, and adapting the virtual personality" refers to a mechanism that collects opinions and requests provided by users during a dialogue, and uses that data to improve and adapt the behavior and comments of the generated virtual personality.

[1419] "Means for collecting user voice and transmitting it to a server in real time" refers to a method for capturing voice data in real time using a device such as a microphone and immediately transmitting it to a server.

[1420] "Means for using speech recognition technology to convert voice data into text data at the server" refers to technology that uses a speech recognition algorithm to accurately convert voice data into text data.

[1421] "Means of generating a response using an AI model based on text data and converting that response into voice using speech synthesis technology" refers to the process of converting a text response generated by an AI model into voice data using a speech synthesis algorithm.

[1422] "Means for playing the synthesized voice to the user through the terminal" refers to a method for delivering the generated voice data to the user using the terminal's speaker or the like.

[1423] "Means for providing an interface for users to provide consent to data provision" refers to a mechanism that provides a screen or interactive procedure for users to agree to data provision.

[1424] "Means for automatically starting data collection upon receiving user consent" refers to the process by which the system automatically starts collecting the necessary data after the user's consent is obtained.

[1425] MODE FOR CARRYING OUT THE INVENTION

[1426] To implement the present invention, the following procedure can be followed to build a system and run a program that collects and analyzes recorded data, generates a virtual personality, executes dialogue, and adapts feedback.

[1427] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, which automatically collects data. The collected data is encrypted and securely sent from the device to the server, where it is stored in a database.

[1428] Next, for data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technologies such as Google Speech-to-Text and Amazon Polly are used to extract features such as the user's phrasing and behavioral patterns. The extracted features become input data for generating an AI model. This AI model is used to generate a virtual personality that mimics the user's behavior and personality.

[1429] During the conversation simulation phase, the user interacts with the virtual persona via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is then converted into voice data using voice synthesis technology such as Amazon Polly and played back to the user via the device.

[1430] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "Tell more jokes" or "Remember that story," they can send them to the server through the feedback function. The server then analyzes this feedback and adapts the AI ​​model, making future conversations more natural and satisfying.

[1431] Explanation with concrete examples

[1432] For example, suppose the family of a deceased person, Mr. B, wants to use this system:

[1433] 1. Data Collection

[1434] The user (B's family member) agrees to the service and consents to providing B's social media account, call history, and search history.

[1435] A dedicated application installed on the device collects this data and sends it to a server.

[1436] The server stores the received data in a database.

[1437] 2. Data analysis and model generation

[1438] The server analyzes the data and extracts Mr. B's characteristic phrases and interests.

[1439] An AI model is created that generates a virtual personality for Mr. B based on the analysis data.

[1440] 3. Conversation Simulation

[1441] The user (B's family member) talks to the virtual B through the terminal.

[1442] The device captures the audio and sends it to the server.

[1443] The server converts the speech to text and uses an AI model to generate an appropriate response.

[1444] The generated response is converted into voice data using voice synthesis technology such as Amazon Polly and sent to the device.

[1445] The device plays back the response of virtual person B.

[1446] 4. Feedback and Adaptation

[1447] The user provides feedback on the interaction with virtual Person B.

[1448] The server analyzes the feedback and adjusts the AI ​​model.

[1449] Prompt Sentence Examples

[1450] "Person B often uses the phrase 'do your best'. Please generate words of encouragement in a tone that is characteristic of Person B."

[1451] "I'd like you to tell me about a movie that Mr. B liked."

[1452] "Tell me how Mr. B tells jokes."

[1453] In this way, the system according to the present invention allows users to interact with their deceased loved ones and provides emotional support to them.

[1454] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1455] Step 1: Data collection

[1456] The user agrees to use the service and consents to the provision of data such as social media accounts, call history, search history, etc. User consent and permission are required as input.

[1457] A dedicated application installed on the device acquires the data that can be collected. In operation, the application periodically and automatically collects posts from social media accounts, call history, and user search history.

[1458] The device securely encrypts the collected data and sends it to the server, resulting in the encrypted data being sent to the server.

[1459] The server stores the received data in a database. In operation, the server appropriately classifies the received data and saves it in a database.

[1460] Step 2: Data analysis and model generation

[1461] The server analyzes the data stored in the database, using saved user social media posts, call history, search history, etc. as input.

[1462] The server uses natural language processing technology to extract features such as the user's phrasing and behavioral patterns. In operation, the NLP algorithm extracts the user's unique expressions and interests from the data. The extracted feature data is obtained as the output.

[1463] The server uses machine learning algorithms to generate an AI model based on the extracted features. The AI ​​model is built using tools such as Python and TensorFlow. The output is a virtual personality model that mimics the user's behavior.

[1464] Step 3: Conversation simulation

[1465] The user initiates a dialogue with the virtual persona through a terminal, and the user's voice input is required.

[1466] The device captures the user's speech as voice data and sends it to the server in real time. The device's microphone collects the voice and uses a voice recognition API to transfer the data to the server. The output is real-time voice data sent to the server.

[1467] The server converts the voice data into text data and generates a response using an AI model. The operation is to convert the voice to text using the Google Speech-to-Text API, and then use an AI model (e.g., GPT-3) to generate an appropriate response. The output is the generated text response.

[1468] The text response generated by the server is converted into voice data using speech synthesis technology and sent to the device. The text is converted into voice using speech synthesis technology such as Amazon Polly and returned to the device. The voice data is then sent to the device as an output.

[1469] The device plays back the virtual persona's response in real time. In operation, the device's speaker plays back the received audio data. As an output, the virtual persona's response is provided to the user through audio.

[1470] Step 4: Feedback and Adapt

[1471] The user provides feedback on the interaction with the virtual personality. The input requires the user's feedback, which may include specific requests such as "Tell more jokes."

[1472] The server analyzes the feedback provided by the user and adjusts the AI ​​model. The operation is to analyze the feedback and incorporate it as new data into the training of the AI ​​model. The output is the adjusted AI model.

[1473] By repeating this series of processes, the system adapts its interactions with the user to make them more natural and satisfying.

[1474] (Application example 1)

[1475] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1476] In modern society, there is a growing demand for content distribution services that go beyond simply viewing video content such as movies and dramas to provide a more fulfilling entertainment experience. However, conventional systems lack a means to empathize with users' feelings of loneliness or the enjoyment of watching a movie. Furthermore, interactive means for providing emotional support are limited. There is a need to solve these issues and provide users with a more enjoyable and fulfilling viewing experience.

[1477] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1478] In this invention, the server includes means for collecting recorded data of a user, means for analyzing the collected recorded data and extracting features, means for generating a virtual personality based on the extracted features, means for the virtual personality to interact with the user, means for receiving feedback from the user during the interaction and adapting the virtual personality, and means for the user and the virtual personality to interact while playing video content. This allows the user to interact with a virtual partner while watching a movie or drama, making for a more fulfilling viewing experience.

[1479] "User" means an individual or organization that uses this system.

[1480] "Recorded data" refers to activity data such as a user's social media accounts, call history, and search history.

[1481] "Means of collection" refers to methods of collecting data using applications or sensors installed on the user's device.

[1482] "Means for analysis" refers to software or algorithms used to analyze the collected recorded data and extract features.

[1483] "Characteristics" are patterns that indicate tendencies in a user's behavior, personality, interests, etc.

[1484] A "virtual personality" is an AI model that is generated based on collected characteristics and allows users to interact with it.

[1485] "Means for dialogue" refers to a function that allows the user and the virtual personality to communicate via voice or text.

[1486] "Means for receiving feedback" and "adapting" refer to methods for updating the virtual personality generation model by reflecting opinions and requests from users.

[1487] "Video content" refers to media content that can be viewed, such as movies and dramas.

[1488] "Means for a user to have a conversation with a virtual character during playback" refers to a technology that allows a user to have a conversation with a virtual character in real time while viewing video content.

[1489] To implement this invention, the following system configuration and programs are required: This system is made up of several main components, each of which plays a specific role.

[1490] System Configuration

[1491] 1. User's device

[1492] Data Collection Applications

[1493] Video playback application

[1494] Audio input and output devices (microphone, speaker)

[1495] 2. Server

[1496] Database

[1497] Data Analysis Module

[1498] AI model generation module

[1499] Speech-to-text module (Google Speech-to-Text API)

[1500] Text-to-speech module (Google Text-to-Speech API)

[1501] WebSocket Server

[1502] Processing flow

[1503] 1. Data Collection

[1504] The user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, search history, etc. The device sends this data to the server, which then stores the collected data in a database.

[1505] 2. Data analysis and AI model generation

[1506] The server retrieves the collected data from the database and extracts features using natural language processing technology (TensorFlow, NLTK, spaCy). Based on these features, a virtual personality is generated that mimics the user's behavior and personality.

[1507] 3. Playback and interaction with video content

[1508] While a user is playing a movie or TV drama using a video playback application, a virtual persona will engage in conversation related to the video content. When the user speaks, the device captures the audio data and sends it to the server via WebSocket. The server converts the audio data into text and uses an AI model to generate an appropriate response. This response is then converted back into audio data by a text-to-speech module and played back to the user via the device.

[1509] 4. Feedback and Adaptation

[1510] Users can provide feedback to the virtual personality during the interaction, for example, if they want it to tell more jokes, they can send that feedback to the server through the feedback function, which will analyze this feedback and adapt the AI ​​model.

[1511] Specific examples

[1512] For example, suppose a user is watching the movie "Castle in the Sky," and the virtual partner imitates the personality of their deceased best friend, Mr. A.

[1513] Prompt Sentence Examples

[1514] The user is watching the movie "Castle in the Sky." The virtual partner is the deceased best friend, Mr. A:

[1515] User: "You also liked this scene, right?"

[1516] Virtual Partner: "Yes, the music in this scene is particularly memorable."

[1517] An example of a prompt sentence to input to the generative AI model is as follows:

[1518] User social media data:

[1519] Posted yesterday: Looking forward to seeing "Castle in the Sky"!

[1520] Call History:

[1521] My best friend A and I often talked about movies.

[1522] Search History:

[1523] A moving scene from "Castle in the Sky"

[1524] feedback:

[1525] I'd like to hear more detailed feedback

[1526] This system allows users to not only watch video content, but also enjoy interacting with virtual characters, further enhancing the viewing experience.

[1527] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1528] Step 1:

[1529] A user installs a data collection application on their device and agrees to the collection of data such as social media accounts, call history, and search history. Based on this consent, the device automatically collects this data and sends it to a database. The input is the user's social media accounts, call history, and search history. The output is the data securely sent to the server.

[1530] Step 2:

[1531] The server retrieves the collected data from the database and extracts features using natural language processing techniques (TensorFlow, NLTK, spaCy). Specifically, the data is first cleansed, and then a language model is applied to identify interest and behavioral patterns. The input is the user's recorded data retrieved from the database. The output is analyzed data containing the extracted features.

[1532] Step 3:

[1533] The server builds an AI model for generating a virtual personality based on the extracted features. Specifically, it trains a neural network model that mimics the user's behavior and personality based on the feature data. The input is the analysis data obtained in step 2. The output is an AI model that represents the virtual personality.

[1534] Step 4:

[1535] While a user plays a movie or drama using a video playback application, the virtual persona interacts with the user about the video content being viewed. Specifically, the voice input is captured and the voice data is sent to the server. The input is the user's voice data. The output is the data sent to the server.

[1536] Step 5:

[1537] The server converts the voice data into text using a speech-to-text module (Google Speech-to-Text API). The text data is then analyzed by an AI model to generate an appropriate response. The input is the user's voice data. The output is text data generated by the virtual personality.

[1538] Step 6:

[1539] The generated text data is converted into voice data using speech synthesis technology (Google Text-to-Speech API) and sent to the device. The input is text data generated by the virtual personality. The output is voice data.

[1540] Step 7:

[1541] The terminal plays the received voice data and continues the dialogue with the user. Specifically, the user listens to the virtual personality's response. The input is the voice data sent from the server. The output is the voice heard by the user.

[1542] Step 8:

[1543] During the interaction, the user provides feedback, for example a specific request such as "tell more jokes." The device collects this feedback and sends it to the server. The input is the user's feedback. The output is the feedback data sent to the server.

[1544] Step 9:

[1545] The server analyzes the feedback data and updates the AI ​​model. Specifically, it adjusts the neural network parameters based on the feedback and adapts the model so that future interactions are more natural. The input is the feedback data. The output is the updated AI model.

[1546] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1547] To implement the present invention, a system can be constructed and a program can be executed according to the following procedure, which includes collecting and analyzing recorded data, generating a virtual personality, conducting dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[1548] Regarding data collection, users agree to use the service and consent to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and data collection is performed automatically by this application. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[1549] Regarding data analysis and model generation, the server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for generating an AI model. This AI model generates a virtual personality that mimics the user's behavior and characteristics.

[1550] Regarding the integration of the emotion engine, the server includes a model for recognizing the user's emotions using the emotion engine. The emotion engine analyzes emotions from the user's voice or text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the response of the virtual personality.

[1551] In the conversation simulation stage, the user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the AI ​​model generates a response based on that text data. This response is converted into voice data using voice synthesis technology and adjusted to the voice specified by the user using a voice changer. The synthesized voice data is sent to the device, and the Reiwa voice data is played back to the user.

[1552] Regarding feedback and adaptation, users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "please use more humor" or "please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[1553] Explanation with concrete examples

[1554] For example, if the family of a deceased person, Mr. C, wants to use this system, the process is as follows:

[1555] 1. Data Collection

[1556] The user (Mr. C's family member) agrees to the service and consents to providing Mr. C's social media account, call history, and search history.

[1557] A dedicated application installed on the device collects this data and sends it to a server.

[1558] The server stores the received data in a database.

[1559] 2. Data analysis and model generation

[1560] The server analyzes the data and extracts Mr. C's characteristic phrases and interests.

[1561] An AI model is created that generates a virtual personality for Mr. C based on the analysis data.

[1562] 3. Emotion engine integration

[1563] The server recognizes emotions from the user's voice and text, and reflects that emotional data in the virtual personality's responses.

[1564] 4. Conversation Simulation

[1565] The user (Mr. C's family member) talks to the virtual Mr. C through the terminal.

[1566] The device captures the audio and sends it to the server.

[1567] The server converts the speech to text and uses an AI model to generate an appropriate response.

[1568] The generated response is converted into voice data using voice synthesis technology and sent to the terminal.

[1569] The device plays back the virtual person C's reply.

[1570] 5. Feedback and Adaptation

[1571] The user provides feedback on the interaction with virtual person C.

[1572] The server analyzes the feedback and adjusts the AI ​​model and emotion engine.

[1573] In this way, the system based on this invention enables users to have conversations with their deceased loved ones and provide emotional support. The integration of an emotion engine enables more natural and emotionally sensitive conversations.

[1574] The processing flow will be explained below.

[1575] Step 1:

[1576] User: Agrees to use the service and consents to providing the necessary data. Completes procedures to allow the collection of data such as social media accounts, call history, and search history.

[1577] Step 2:

[1578] Device: Install the dedicated application and complete the setup. The application will automatically collect data such as user's social media posts, call history, and search history, and send it to the server.

[1579] Step 3:

[1580] Server: Receives recorded data sent from the device and stores it in a secure database.

[1581] Step 4:

[1582] Server: Analyzes the transmitted recorded data. Using natural language processing technology, it extracts the user's language usage patterns, characteristic phrases, and behavioral patterns.

[1583] Step 5:

[1584] Server: Generates an AI model based on the extracted feature information. This AI model creates a virtual personality that mimics the user's behavior and characteristics.

[1585] Step 6:

[1586] Server: Integrates an emotion engine into the virtual personality. The emotion engine analyzes emotions from the user's voice and text in real time and generates emotion data.

[1587] Step 7:

[1588] User: Uses the device to initiate a conversation with the virtual persona. The device captures the user's voice with a microphone.

[1589] Step 8:

[1590] Terminal: Sends captured audio data to the server in real time.

[1591] Step 9:

[1592] Server: The received voice data is converted into text data using voice recognition technology. The emotion engine then analyzes the user's emotions from the voice data and generates emotion data.

[1593] Step 10:

[1594] Server: The AI ​​model generates appropriate responses based on the text and emotional data. The responses are tailored to the characteristics of the virtual personality and take into account the emotional data.

[1595] Step 11:

[1596] Server: The generated response is converted into voice data using speech synthesis technology. A voice changer is used to match the tone of the virtual personality.

[1597] Step 12:

[1598] Server: The synthesized voice data is sent back to the device.

[1599] Step 13:

[1600] Terminal: Plays the received audio data to the user through the speaker.

[1601] Step 14:

[1602] User: Provides feedback to the virtual persona during the interaction. Feedback can include specific requests such as "Please tell me more" or "Talk about a different topic."

[1603] Step 15:

[1604] Terminal: Sends feedback data to the server in real time.

[1605] Step 16:

[1606] Server: Analyzes the received feedback and applies it to the AI ​​model and emotion engine. It adjusts the response and behavior patterns of the virtual personality based on the feedback.

[1607] In this way, the present invention provides a system that allows interaction with deceased loved ones and is flexible in adapting to the user's emotions and feedback.

[1608] Example 2

[1609] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1610] Current virtual dialogue systems lack sufficient recognition of user emotions, resulting in mechanical and unnatural dialogue. Furthermore, it is difficult to effectively incorporate user feedback, making it difficult to continuously improve the quality of dialogue. Furthermore, voice conversion lacks flexibility, and there is a lack of technology available to adjust the voice to suit the user's preferences. These challenges make it difficult to provide a satisfying dialogue experience for users.

[1611] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1612] In this invention, the server includes means for collecting user record data, means for analyzing the collected record data and extracting features, and means for generating a virtual personality based on the extracted features. This makes it possible to recognize the user's emotions using an emotion engine and reflect them in the dialogue, collect and analyze the user's voice in real time, and adapt the AI ​​model and emotion engine based on the feedback.

[1613] "User" refers to an individual who interacts with the System.

[1614] "Recorded data" refers to data provided by users, such as social media accounts, call history, and search history.

[1615] "Data collection" refers to the process of automatically collecting recorded data using a dedicated application based on the user's consent.

[1616] "Database" refers to a repository for storing recorded data collected by the Server.

[1617] "Data analysis" refers to the process of analyzing collected recorded data using natural language processing techniques and extracting features.

[1618] "Characteristics" refers to characteristics such as users' language usage patterns and behavioral patterns extracted through data analysis.

[1619] "Virtual personality" refers to a virtual conversation partner that mimics the user's characteristics and is generated using an AI model based on extracted characteristics.

[1620] "Dialogue" refers to communication between the user and the virtual personality.

[1621] "Feedback" refers to the opinions and requests that a user provides to a virtual personality during a conversation.

[1622] "Adaptation" refers to the process of updating and improving AI models and emotion engines based on the feedback provided.

[1623] An "emotion engine" refers to a system that analyzes and recognizes users' emotions from voice and text data in real time.

[1624] "Audio Data" means digital audio information that captures a user's speech.

[1625] "Text data" refers to character information obtained by converting voice data.

[1626] "Speech synthesis" refers to the process of generating digital speech from text data.

[1627] "Voice changer" refers to technology that adjusts the generated voice to a voice specified by the user.

[1628] "Server" refers to the computer system that collects and analyzes data, generates virtual personalities, and operates the emotion engine.

[1629] "Dedicated application" refers to software that is installed on the user's device and is used to collect and transmit recorded data.

[1630] "Consent interface" refers to the screen or function that allows users to consent to providing data.

[1631] To implement this invention, the system is constructed and the program is executed according to the following procedure: The system includes collecting and analyzing recorded data, generating a virtual personality, executing dialogue, adapting feedback, and recognizing the user's emotions by incorporating an emotion engine.

[1632] Data collection

[1633] The user agrees to use the service and consents to providing necessary data such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects data. The collected data is securely transmitted from the device to a server, where it is stored in a database.

[1634] Examples:

[1635] The user agrees to the terms of service and installs the dedicated application on their smartphone.

[1636] A dedicated application automatically collects social media accounts and call history, encrypts them, and sends them to a server.

[1637] Data Analysis and Model Generation

[1638] The server analyzes the collected data using natural language processing technology (such as SpaCy or NLTK) to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. An AI model (such as the Transformer model or GPT-4) is generated based on the extracted features, creating a virtual personality that mimics the user's characteristics.

[1639] Emotion engine integration

[1640] The server contains a model for recognizing the user's emotions using an emotion engine (such as the Microsoft Azure Cognitive Services emotion analysis API). The emotion engine analyzes the user's voice and text in real time to determine their emotional state. The determined emotional data is used to adjust the virtual personality's responses.

[1641] Examples:

[1642] The server analyzes the collected social media posts and call history to extract the user's characteristic phrases and interests.

[1643] An AI model is created that generates a virtual personality based on the extracted data and an emotion engine is integrated.

[1644] Conversation Simulation

[1645] The user interacts with the virtual personality via the device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and an AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology (such as Amazon Polly), and a voice changer adjusts it to the voice specified by the user. The synthesized voice data is sent to the device, where it is played back to the user.

[1646] Examples:

[1647] The user speaks to the virtual persona, "What's the weather like today?"

[1648] The server analyzes this, and the virtual personality generates a reply saying "It's sunny today," converts it into voice, and sends it to the terminal.

[1649] The voice that has been adjusted by the voice changer is played back to the user.

[1650] Feedback and Adaptation

[1651] Users can provide feedback to the virtual personality during a conversation. For example, if they have specific requests, such as "I want more humor" or "I want you to elaborate on a particular topic," they can send these to the server through the feedback function. The server analyzes this feedback and updates and adapts the AI ​​model and emotion engine, making future conversations more natural and satisfying.

[1652] Example prompt sentence:

[1653] "Based on C's social media account data, please extract characteristic phrases and behavioral patterns to create a virtual personality for C."

[1654] This invention allows users to have a more natural and emotionally relevant interaction experience.

[1655] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1656] Step 1:

[1657] Service consent and data provision consent

[1658] Input: The user agrees to the terms of use of the service and consents to providing data such as social media accounts, call history, and search history.

[1659] Specific behavior: The user clicks the "Agree" button on the screen where they agree to the terms of use and privacy policy. They are then asked to provide data, which they agree to by clicking the "Provide" button.

[1660] Output: Consent and authorization information is sent to the server.

[1661] Step 2:

[1662] Installing the dedicated application

[1663] Input: User information with consent.

[1664] How it works: The user installs the dedicated application on their smartphone or PC, and the application is downloaded from the official website or app store.

[1665] Output: The dedicated application is installed on the user's device.

[1666] Step 3:

[1667] Automatic data collection and transmission

[1668] Input: The user's device on which the dedicated application is installed.

[1669] How it works: The device automatically collects data from social media accounts, call history, and search history. The collected data is encrypted and securely sent to a server.

[1670] Output: The encrypted data is sent to the server.

[1671] Step 4:

[1672] Data storage

[1673] Input: Encrypted data.

[1674] Specific operation: The server stores the received data in a database, for example, AWS RDS.

[1675] Output: Data stored in a database.

[1676] Step 5:

[1677] Performing data analysis

[1678] Input: Data stored in a database.

[1679] What it does: The server analyzes the collected data using natural language processing techniques (e.g., SpaCy or NLTK) to extract the user's language usage patterns, characteristic phrases, and behavioral patterns.

[1680] Output: Extracted feature data.

[1681] Step 6:

[1682] AI model generation

[1683] Input: Extracted feature data.

[1684] Specific operation: The server generates an AI model (e.g., Transformer model or GPT-4) based on the feature data and creates a virtual personality that mimics the user's characteristics.

[1685] Output: A virtual personality is generated.

[1686] Step 7:

[1687] Preparing for emotion recognition

[1688] Input: User interaction data with virtual persona.

[1689] How it works: The server integrates the emotion engine, which uses the emotion analysis API from Microsoft Azure Cognitive Services and is configured to analyze user emotions in real time from voice and text data.

[1690] Output: A system capable of emotion recognition.

[1691] Step 8:

[1692] Utilizing Emotional Data

[1693] Input: Real-time analyzed emotion data.

[1694] Specific operation: The server reflects the emotional data recognized by the emotion engine in the virtual personality's response. For example, if the user makes a sad voice, the virtual personality will respond with a comforting response.

[1695] Output: The virtual personality's response depending on the emotion.

[1696] Step 9:

[1697] Start a real-time conversation

[1698] Input: What the user says.

[1699] Specific operation: The user speaks into the device, which captures what is said as audio data and sends it to the server in real time.

[1700] Output: The audio data is sent to the server.

[1701] Step 10:

[1702] Voice data conversion and response generation

[1703] Input: User's voice data.

[1704] How it works: The server uses the Google Cloud Speech-to-Text API to convert the voice data into text, and the AI ​​model generates an appropriate response based on the converted text.

[1705] Output: The generated response text data.

[1706] Step 11:

[1707] Apply voice synthesis and voice changer

[1708] Input: Generated response text data.

[1709] Specific operation: The server converts the generated text into voice data using speech synthesis technology (e.g., Amazon Polly), and then uses a voice changer to adjust the voice data to match the user's voice.

[1710] Output: The adjusted audio data.

[1711] Step 12:

[1712] Playing audio data

[1713] Input: The adjusted audio data.

[1714] Specific operation: The device plays back the voice data received from the server, and the appropriate response is provided in the voice of the virtual personality.

[1715] Output: The audio played to the user.

[1716] Step 13:

[1717] Providing feedback

[1718] Input: User feedback provided during interaction with the virtual persona.

[1719] Specific operation: The user enters and submits their opinions and requests regarding the virtual personality through a feedback form or the like.

[1720] Output: Feedback data sent to the server.

[1721] Step 14:

[1722] Feedback Analysis and Adaptation

[1723] Input: Feedback data sent to the server.

[1724] Specific operation: The server analyzes the feedback and updates and adapts the AI ​​model and emotion engine. Specifically, the feedback content is analyzed using natural language processing technology and reflected in the AI ​​model's response patterns.

[1725] Output: Updated and adapted AI model and emotion engine.

[1726] This will make future interactions more natural and satisfying.

[1727] (Application example 2)

[1728] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1729] Conventional food delivery systems often lack sufficient support when users select menus or customize their meals. It can be difficult to get appropriate advice or suggestions, especially when users have special requests or want to try new dishes. Furthermore, there is a need for a personalized experience based on users' emotions and preferences. To address these challenges, a system is needed that allows users to easily select and customize menus while interacting with a virtual chef.

[1730] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1731] In this invention, the server includes: [means for collecting user record data;] [means for analyzing the collected record data and extracting features;] [means for generating a virtual personality based on the extracted features;] [means for having the virtual personality interact with the user;] [means for receiving feedback from the user during the interaction and adapting the virtual personality;] [means for the user to select a menu while conversing with the virtual chef when ordering; and [means for assisting in menu selection and dish customization.] This enables users to easily select and customize menus, providing a more personalized experience.

[1732] "User" refers to the person using the invention, who is primarily responsible for selecting menus and customizing meals through interaction with the virtual chef.

[1733] "Recorded data" refers to information including data related to the user's use of social media accounts, call history, search history, etc.

[1734] "Collect" refers to the process of collecting user record data through a dedicated application and sending it to a server.

[1735] "Analyzing" refers to the act of analyzing collected recorded data and extracting specific patterns or characteristics.

[1736] "Extracting features" means identifying a user's language usage, behavioral patterns, preferences, etc. from the analyzed data.

[1737] A "virtual personality" is an interactive agent that uses an AI model generated based on collected and analyzed characteristics.

[1738] "Having a dialogue" refers to the virtual personality interacting with the user in a conversational format.

[1739] "Feedback" refers to the responses and requests provided by users to the system.

[1740] "Adapting" refers to updating and adjusting the virtual persona's responses and behavior based on feedback.

[1741] "Ordering" refers to the act of a user requesting a meal using a food delivery service.

[1742] A "virtual chef" is a virtual personality that has particular knowledge about cooking and assists users in choosing menus through dialogue.

[1743] "Menu selection" is the process by which the virtual chef makes appropriate meal suggestions to the user and provides the best options.

[1744] "Cuisine customization" refers to making changes or adjustments to a dish according to the user's wishes.

[1745] "Voice data" refers to data that is a digital recording of a user's voice input.

[1746] "Converting to text data" refers to analyzing the voice data and changing it into text-format data.

[1747] A "reply" is a response that the generative AI model generates in response to a user's statement.

[1748] A "speech synthesizer" is a device or software that converts text data into speech data.

[1749] "Play" refers to making the synthesized speech available to the user.

[1750] "Consent" is the act of a user explicitly giving permission to use the system.

[1751] "Interface" refers to the means or screen through which a user interacts with a system.

[1752] To implement the present invention, the following system is required: This system is configured by incorporating a user's recorded data collection, analysis, virtual personality generation, dialogue execution, feedback adaptation, and emotion engine.

[1753] Data collection

[1754] Users agree to use the service and provide the necessary data, such as social media accounts, call history, and search history. A dedicated application is installed on the user's device, and this application automatically collects the data. The collected data is securely transmitted from the device to a server, which then stores it in a database.

[1755] Data Analysis and Model Generation

[1756] The server analyzes the collected data. During the analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. The extracted features become input data for a generative AI model. This generative AI model generates a virtual personality that mimics the user's behavior and characteristics.

[1757] Emotion engine integration

[1758] The server includes a model for recognizing the user's emotions using an emotion engine. The emotion engine analyzes emotions from the user's voice and text in real time and determines their emotional state. The emotion data recognized by the emotion engine is used to adjust the responses of the virtual personality.

[1759] Conversation Simulation

[1760] The user interacts with the virtual chef via a device. The device captures the user's speech as voice data and sends it to the server in real time. The server converts the voice data into text data, and the generative AI model generates a response based on that text data. This response is converted into voice data using speech synthesis technology, and the synthesized voice data is sent to the device and played back to the user.

[1761] Feedback and Adaptation

[1762] Users can provide feedback to the virtual chef during the conversation. For example, if they have specific requests, such as "I'd like more humor" or "Please elaborate on a particular topic," they can send these to the server through the feedback function. The server then analyzes this feedback and updates and adapts the generative AI model and emotion engine, making future conversations more natural and satisfying.

[1763] Specific examples

[1764] For example, consider a scenario in which a user uses a smartphone app to say, "I want spicy pasta." The system captures the speech, converts it into text data, and then uses a generative AI model to derive the optimal response. An example of a prompt in this case is as follows:

[1765] User: "I want spicy pasta"

[1766] Virtual Chef: "For spicy pasta, I recommend Peperoncino or Arrabbiata. Which would you like?"

[1767] In this way, the virtual chef can provide a menu tailored to the user's needs, allowing for personalized meal selection and customization.

[1768] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1769] Step 1:

[1770] The user agrees to use the service and consents to providing data such as social media accounts, call history, and search history. This data is automatically collected by a dedicated application installed on the device. The device securely transmits the collected data to the server. This input data is then stored in a database.

[1771] Step 2:

[1772] The server analyzes the user's recorded data stored in the database. During this analysis process, natural language processing technology is used to extract the user's language usage patterns, characteristic phrases, and behavioral patterns. These analysis results (language usage patterns and behavioral patterns) are extracted and used as input data for the generative AI model.

[1773] Step 3:

[1774] The server generates a virtual personality based on the extracted characteristics. Using a generative AI model, it creates a virtual chef that mimics the user's behavior and characteristics. The virtual chef can then assist with menu selection and cooking customization.

[1775] Step 4:

[1776] The user interacts with the virtual chef via a terminal. The terminal captures the user's speech as voice data and transmits it to the server in real time. The server converts this voice data into text data and uses it as input.

[1777] Step 5:

[1778] The server generates a response from the virtual chef based on the text data. This response is generated using a generative AI model and output as a response text. The response text is converted into audio data using speech synthesis technology. This audio data is sent to the device and played back to the user.

[1779] Step 6:

[1780] During the conversation, the user can provide feedback to the virtual chef, such as a request for more humor. The device then sends this feedback data to the server, which analyzes it and adjusts the parameters of the generative AI model and emotion engine to reflect the feedback in the next conversation.

[1781] Step 7:

[1782] The virtual chef will help users select menus and customize dishes to suit their needs. For example, if a user says, "I want spicy pasta," the virtual chef will suggest, "For spicy pasta, we recommend peperoncino or arrabbiata. Which would you like?"

[1783] This process allows users to make personalized meal selections and improves their experience.

[1784] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1785] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1786] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1787] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1788] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1789] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1790] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1791] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1792] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1793] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1794] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1795] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1796] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1797] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1798] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1799] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1800] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1801] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1802] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1803] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1804] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1805] The following is further disclosed regarding the above embodiment.

[1806] (Claim 1)

[1807] [Means for collecting user record data;

[1808] [Means for analyzing and extracting features from the collected record data;

[1809] [Means for generating a virtual personality based on the extracted features;

[1810] [Means for allowing the virtual personality to interact with the user;

[1811] [Means for receiving feedback from the user during the interaction and adapting the virtual personality;

[1812] A system including:

[1813] (Claim 2)

[1814] [Means for collecting the user's voice and transmitting it to a server;

[1815] [Means for converting voice data into text data on the server;

[1816] [Means for generating a response based on text data and converting it into voice using a speech synthesizer;

[1817] [means for playing the synthesized speech to a user; and

[1818] The system of claim 1 further comprising:

[1819] (Claim 3)

[1820] [means for providing an interface for users to provide consent; and

[1821] [Means to initiate data collection based on user consent; and

[1822] The system of claim 1 further comprising:

[1823] "Example 1"

[1824] (Claim 1)

[1825] [Means for collecting user record data and transmitting it securely to a server;

[1826] [Means for analyzing the collected record data and extracting features using natural language processing techniques;

[1827] [Means for generating a virtual personality using machine learning based on the extracted features;

[1828] [means for executing speech recognition and speech synthesis technology to allow the generated virtual personality to interact with the user;

[1829] [Means for receiving feedback from the user during the interaction, analyzing it, and adapting the virtual personality;

[1830] A system including:

[1831] (Claim 2)

[1832] [Means for collecting user voice and transmitting it to a server in real time;

[1833] [means for using speech recognition technology to convert speech data into text data at a server;

[1834] [Means of generating a response using an AI model based on text data and converting that response into voice using speech synthesis technology;

[1835] [Means for playing the synthesized speech to a user through a terminal;

[1836] The system of claim 1 further comprising:

[1837] (Claim 3)

[1838] [Means for providing an interface for users to provide consent to data provision; and

[1839] [Means to automatically initiate data collection with user consent; and

[1840] The system of claim 1 further comprising:

[1841] "Application Example 1"

[1842] (Claim 1)

[1843] [Means for collecting user record data;

[1844] [Means for analyzing and extracting features from the collected record data;

[1845] [Means for generating a virtual personality based on the extracted features;

[1846] [Means for allowing the virtual personality to interact with the user;

[1847] [Means for receiving feedback from the user during the interaction and adapting the virtual personality;

[1848] [Means for the user to have a dialogue with the virtual personality while the video content is being played;

[1849] A system including:

[1850] (Claim 2)

[1851] [Means for collecting the user's voice and transmitting it to a server;

[1852] [Means for converting voice data into text data on the server;

[1853] [Means for generating a response based on text data and converting it into voice using a speech synthesizer;

[1854] [means for playing the synthesized speech to a user; and

[1855] The system of claim 1 further comprising:

[1856] (Claim 3)

[1857] [means for providing an interface for users to provide consent; and

[1858] [Means to initiate data collection based on user consent; and

[1859] The system of claim 1 further comprising:

[1860] "Example 2: Combining Emotion Engines"

[1861] (Claim 1)

[1862] [Means for collecting user record data;

[1863] [Means for analyzing and extracting features from the collected record data;

[1864] [Means for generating a virtual personality based on the extracted features;

[1865] [Means for allowing the virtual personality to interact with the user;

[1866] [Means for receiving feedback from the user during the interaction and adapting the virtual personality;

[1867] [Means for recognizing a user's emotions using an emotion engine;

[1868] A system including:

[1869] (Claim 2)

[1870] [Means for collecting the user's voice and transmitting it to a server;

[1871] [Means for converting voice data into text data on the server;

[1872] [Means for generating a response based on text data and converting it into voice using a speech synthesizer;

[1873] [means for playing the synthesized speech to a user; and

[1874] [Means for adjusting the generated voice using a voice changer to a voice designated by a user;

[1875] The system of claim 1 further comprising:

[1876] (Claim 3)

[1877] [means for providing an interface for users to provide consent; and

[1878] [Means to initiate data collection based on user consent; and

[1879] [Means of automatically collecting data by installing a dedicated application on the device, and

[1880] The system of claim 1 further comprising:

[1881] "Application example 2 when combining emotion engines"

[1882] (Claim 1)

[1883] [Means for collecting user record data;

[1884] [Means for analyzing and extracting features from the collected record data;

[1885] [Means for generating a virtual personality based on the extracted features;

[1886] [Means for allowing the virtual personality to interact with the user;

[1887] [Means for receiving feedback from the user during the interaction and adapting the virtual personality;

[1888] [A way for users to choose a menu while talking to a virtual chef when ordering,

[1889] [Methods to assist with menu selection and food customization,

[1890] A system including:

[1891] (Claim 2)

[1892] [Means for collecting the user's voice and transmitting it to a server;

[1893] [Means for converting voice data into text data on the server;

[1894] [Means for generating a response based on text data and converting it into voice using a speech synthesizer;

[1895] [means for playing the synthesized speech to a user; and

[1896] The system of claim 1 further comprising:

[1897] (Claim 3)

[1898] [means for providing an interface for users to provide consent; and

[1899] [Means to initiate data collection based on user consent; and

[1900] The system of claim 1 further comprising: [Explanation of symbols]

[1901] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for collecting user record data; means for analyzing the collected record data and extracting features; A means for generating a virtual personality based on the extracted features; a means for allowing the virtual persona to interact with the user; means for receiving feedback from the user during the interaction and adapting the virtual persona; A system including:

2. A means for collecting a user's voice and transmitting it to a server; A means for converting voice data into text data on a server; A means for generating a response based on the text data and converting the response into voice using a voice synthesizer; means for playing the synthesized speech to a user; The system of claim 1 further comprising:

3. a means for providing an interface for a user to provide consent; A means for initiating data collection based on user consent; The system of claim 1 further comprising:

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A