System
The system enables real-time interactive conversations with fictional characters using a server, generative AI, and voice generation AI, addressing the emotional needs of fictiosexuals by providing a satisfying and immersive experience.
Patent Information
- Application Number
- JP2024125392
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-13
AI Technical Summary
Fictiosexuals experience loneliness and unfulfilled emotions due to the inability to interact with fictional characters in real life, lacking effective means to satisfy their romantic feelings for these characters.
A system that allows users to converse with fictional characters in real-time using a server with a character database, generative AI for response text generation, and voice generation AI to reproduce the character's voice, enabling interactive conversations through a terminal device.
The system provides a realistic and immersive interaction experience, alleviating feelings of loneliness and dissatisfaction by allowing users to engage with multiple characters, enhancing user satisfaction.
Smart Images

Figure 2026023457000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] There are people called "fictiosexuals" who have romantic feelings for fictional characters. These people have strong feelings for fictional characters with whom they cannot interact in real life. However, unlike real people, they have no opportunity to directly converse or interact with these characters, so they have limited means of satisfying these feelings. As a result, fictiosexuals often suffer from feelings of loneliness and unfulfilled emotions. Therefore, there is a need to enable real-life interactions with fictional characters and provide a means for them to satisfy their feelings. [Means for solving the problem]
[0005] The present invention provides a system including: a means for retrieving fictional character setting materials from a database; a means for generating a character's response text based on a user's input text using a generation AI; a means for generating an audio file of the generated response text in the character's unique voice using a voice generation AI; and a means for the character and the user to have an interactive conversation on a terminal. This system allows the user to converse with a selected character in real time and enjoy interacting with that character. This can alleviate the feelings of loneliness and dissatisfaction that fictiosexuals experience. Furthermore, the voice generation AI reproduces the character's tone and voice quality based on the character setting materials, providing a more realistic experience. Furthermore, by providing a means for the user to select from multiple characters on the terminal, the user can enjoy interacting with a variety of characters.
[0006] "Setting materials" refers to a collection of information that includes details such as a character's personality, tone of voice, background information, and appearance.
[0007] "Database" refers to a system for efficiently storing, retrieving, and managing specific data.
[0008] "Generative AI" refers to an algorithm or system that uses artificial intelligence techniques to perform natural language processing and generate appropriate text for a given input.
[0009] "Voice generation AI" refers to the technology that converts text data into voice data, specifically the artificial intelligence technology that reproduces a character's voice quality and tone of voice.
[0010] "Terminal" refers to a device such as a computer, smartphone, or tablet that is directly operated by a user.
[0011] "User" refers to a person who uses the system to converse and interact with characters.
[0012] "Interactive" refers to a form of information exchange between the user and the system.
[0013] "Character" refers to a fictional person or entity that appears in a work of fiction. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting data from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. Specific embodiments of the system are described below.
[0036] Overall system overview
[0037] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, and a voice generation AI, while the terminal is equipped with a user interface, real-time chat function, and voice output function. Users can start conversations with characters through their terminal and enjoy the results.
[0038] Overview of program processing
[0039] The system program operates as follows.
[0040] 1. Character Selection
[0041] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[0042] 2. Start a conversation
[0043] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[0044] 3. Speech generation
[0045] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0046] 4. Keep the conversation going
[0047] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[0048] Specific examples
[0049] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[0050] When a user types "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as "I went to a cafe with my friends today! It was fun." The server then uses speech generation AI to convert this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[0051] If the user then continues with, "That's nice. What did you have at the cafe?" the conversation continues using the same process.
[0052] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexual people.
[0053] The processing flow will be explained below.
[0054] Step 1:
[0055] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[0056] Step 2:
[0057] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[0058] Step 3:
[0059] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[0060] Step 4:
[0061] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[0062] Step 5:
[0063] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[0064] Step 6:
[0065] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[0066] Step 7:
[0067] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[0068] Step 8:
[0069] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[0070] Step 9:
[0071] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[0072] Step 10:
[0073] If the user wants to continue the conversation, they can enter a new message and send it to the server using the same steps. The server then uses the language generation and speech generation AI to generate a new reply and sends it to the device. This process is repeated, creating an interactive conversation.
[0074] Example 1
[0075] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0076] There is a demand for a system that allows users to enjoy real-time interactive conversations with fictional characters. Conventional systems have made it difficult for users to obtain a satisfying experience because the characters' responses are monotonous and the voice generation is unnatural. Furthermore, the ability to select from multiple characters is limited, making it impossible to meet the diverse needs of users.
[0077] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0078] In this invention, the server includes means for retrieving setting materials for a fictional character selected by a user from a database, means for generating a response text for the fictional character based on a text input by the user using a generation AI, and means for generating an audio file using a voice generation AI from the generated response text in a voice unique to the fictional character, thereby enabling the user to enjoy natural and immersive interactive conversations with the fictional character.
[0079] "Users" are people who use the system to converse with fictional characters.
[0080] "Fictional characters" refers to fictional characters and animals that appear in creative works such as animation, manga, novels, and games.
[0081] "Setting materials" refers to information that describes a fictional character in detail, such as their personality, age, occupation, hobbies, etc.
[0082] A "database" is a collection of information that stores and manages setting materials for fictional characters and can be retrieved as needed.
[0083] "Generative AI" refers to artificial intelligence technology that generates response text from fictional characters based on user input text.
[0084] "Response text" refers to the reply content of a fictional character created by the generation AI in response to user input.
[0085] "Voice generation AI" is an artificial intelligence technology that converts response text created by the generation AI into an audio file that reproduces the tone and voice quality of a fictional character.
[0086] "Audio file" refers to digital data generated by voice generation AI to play back responses in the voices of fictional characters.
[0087] A "terminal" is a device or software that a user operates to interact with fictional characters.
[0088] "Interactive conversation" means that the user and the fictional character have two-way communication in real time via the system.
[0089] The present invention provides a system for enabling users to have real-time interactive conversations with fictional characters, which system is comprised of a server, a terminal, and a user.
[0090] Overall system configuration
[0091] The server includes the following components:
[0092] Database: Accumulates and manages setting materials for fictional characters.
[0093] Generative AI: Generates response text from a fictional character based on the user's input text.
[0094] Voice generation AI: Converts generated response text into an audio file with a voice unique to a fictional character.
[0095] The terminal is a device operated by the user and has the following functions:
[0096] User Interface: An interface that allows the user to select a character, input messages, and send them.
[0097] Real-time chat function: A function that communicates with the server and sends and receives messages in real time.
[0098] Audio output function: A function to play audio files received from the server.
[0099] The user operates the terminal to enjoy conversations with fictional characters.
[0100] System Operation Overview
[0101] First, the user operates the device's interface and selects the fictional character they want to talk to. The device then sends the character's ID to the server, which then retrieves the character's profile information from a database.
[0102] Next, the user enters a message to start the conversation and presses the send button. The device then sends the entered message to the server, which uses a generation AI to generate a reply text based on the received message and character settings.
[0103] Based on the generated response text, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0104] The user can enter a new message and continue the conversation in the same process. For each message sent from the terminal, the server generates an appropriate reply and outputs it as voice.
[0105] Specific examples
[0106] For example, if a user wants to have a conversation with a specific fictional character, the user selects the character on the device interface and sends the selection information to the server. The server retrieves character information from the database and starts a conversation based on that information.
[0107] When a user types "What did you do today?", the device sends this message to the server. The server uses a generation AI to generate a response such as "I went to a cafe with a friend today! It was fun." The server then uses a voice generation AI to convert this response into an audio file in the character's voice and sends it to the device. The device plays this audio file, allowing the user to feel as if they are having a conversation with the character.
[0108] If the user then types, "That's nice. What did you have at the cafe?" the conversation continues using a similar process.
[0109] Example prompt sentence:
[0110] "Based on the character's background material, generate a response to the following message: 'What did you do today?'"
[0111] "Say in the voice of a specific fictional character, 'I went to a cafe with my friends today. It was fun.'"
[0112] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice.
[0113] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0114] Step 1: Choose your character
[0115] Input: The user selects a character in the device interface.
[0116] Operation:
[0117] The terminal acquires the ID of the character selected by the user.
[0118] The ID of the selected character is sent to the server.
[0119] Output: The character's ID is sent to the server.
[0120] Step 2: Obtaining the configuration data
[0121] Input: The server receives the character's ID.
[0122] Operation:
[0123] The server searches the database for the character's settings information.
[0124] The setting materials include information such as the character's personality, age, occupation, hobbies, etc.
[0125] Obtain the setting materials.
[0126] Output: The configuration data is available on the server.
[0127] Step 3: Start a conversation
[0128] Input: The user types a message into the terminal.
[0129] Operation:
[0130] The user inputs text into the message input field on the terminal and presses the send button.
[0131] The terminal transmits the input text data to the server.
[0132] Output: The user's input text is sent to the server.
[0133] Step 4: Generate response text
[0134] Input: The server receives the user's input text and character profile information.
[0135] Operation:
[0136] The server passes prompts to the generation AI based on the user's message and character settings.
[0137] The generative AI generates response text that matches the character's tone and personality.
[0138] Output: The generated response text.
[0139] Step 5: Generate the audio file
[0140] Input: The server receives the generated response text.
[0141] Operation:
[0142] The server inputs the generated response text into the speech generation AI.
[0143] The voice generation AI generates audio files that reproduce the character's voice quality and speaking style.
[0144] Output: The generated audio file.
[0145] Step 6: Send and play the audio file
[0146] Input: The server has the generated audio file.
[0147] Operation:
[0148] The server transmits the generated audio file to the terminal.
[0149] The terminal receives the audio file and plays it to the user using its audio output capabilities.
[0150] Output: The reply is played to the user as audio.
[0151] Step 7: Keep the conversation going
[0152] Input: The user types a new message.
[0153] Operation:
[0154] The user types another question or comment into the terminal and presses the send button.
[0155] The entered text data is sent to the server again.
[0156] The server repeats the same process, using the generation AI and voice generation AI to generate new response text and audio files.
[0157] The terminal plays the received audio file.
[0158] Output: The conversation continues, generating new messages and replies.
[0159] (Application example 1)
[0160] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0161] In traditional brick-and-mortar stores, users lacked an effective and enjoyable way to obtain product information. Furthermore, there was no technology available to provide a highly entertaining shopping experience through interactions with fictional characters. This made it difficult to improve the user experience.
[0162] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0163] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for generating a response text for the character based on a user's input text using a generation AI, means for generating an audio file in the character's unique voice using a voice generation AI from the generated response text, means for the character and the user to have an interactive conversation on a terminal, and means for the character to provide product information and services as a virtual assistant using a smart device in a physical store. This allows users to obtain product information while conversing with the fictional character in real time in the physical store, enabling them to enjoy a highly entertaining shopping experience.
[0164] A "fictional character" is a person or creature that appears in a fictional world, and whose setting is depicted in detail through media such as stories and games.
[0165] A "database" is an electronic storage device or system that systematically organizes information or data so that it can be efficiently accessed and managed.
[0166] "Generative AI" is an artificial intelligence technology that uses natural language processing and data analysis techniques to generate text and data based on user input.
[0167] "Voice generation AI" is a technology that converts text data into voice data, and is an artificial intelligence technology that can reproduce specific voice qualities and speaking styles.
[0168] A "terminal" is an electronic device that can be directly operated by a user, such as a computer, tablet, or smartphone.
[0169] "Interactive conversational means" refers to methods or functions that allow users and systems to communicate with each other in real time.
[0170] A "brick and mortar store" is a store with a physical sales floor that exists in the real world, as opposed to an online store.
[0171] A "smart device" is an electronic device with advanced functionality that can connect to the Internet and run various applications, including smartphones, smart glasses, and head-mounted displays.
[0172] A "virtual assistant" is a digital assistant that uses artificial intelligence technology to answer users' questions and provide services.
[0173] "Product Guide" is a means of providing information about products in a store and helping users consider purchasing them.
[0174] "User experience" refers to the experience and satisfaction that users feel when using a product or service.
[0175] The embodiment of the present invention is to build a system that is basically composed of a server, a terminal, and a user. An overview of the entire system and each of its main steps will be explained below.
[0176] Overall system overview
[0177] The system retrieves fictional character settings from a database, uses generative and voice-generating AI to generate responses for the characters, and outputs them as voice. The system is designed to improve the shopping experience in brick-and-mortar stores.
[0178] Hardware and software used
[0179] Hardware: Servers, smart glasses (e.g., Google Glass), head-mounted displays (e.g., Microsoft HoloLens)
[0180] Software: Language generation AI (e.g. GPT-4), speech generation AI (Google Text-to-Speech, Amazon Polly), user interface (developed with Unity or Unreal Engine)
[0181] System operation example
[0182] 1. User authentication and character selection
[0183] The user puts on the smart device and starts the application. After logging in, the user selects the character they want to talk to. For example, the user selects the character "Elizabeth." This selection information is sent to the server.
[0184] Example prompt sentence:
[0185] Character: Elizabeth
[0186] User Question: What is the product on this counter?
[0187] Preferred response language: Japanese
[0188] Information to include in the response: product type, features, special offers, etc.
[0189] As Elizabeth: "This counter is lined with the latest technology gadgets. Our most popular item is the smart wristband. It has health tracking features and is currently on sale at a special price."
[0190] 2. Obtaining character setting data
[0191] The server retrieves the character's profile information from a database, including the character's speaking style, personality, and voice quality.
[0192] 3. Conversation initiation and response generation
[0193] When a user speaks to the character, the message is sent to the server, which uses language generation AI to generate an appropriate response to the user's input. For example, if a user asks, "What products are on this counter?", the server will generate a response such as, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular and have health management functions."
[0194] 4. Speech generation and output
[0195] Based on the generated response text, the AI creates an audio file that reproduces the character's voice quality and tone. This audio file is then sent to the user's smart device and played back, making the user feel as if they are having a real-time conversation with a fictional character.
[0196] Specific examples
[0197] For example, imagine a user puts on smart glasses in a physical store, selects a character named "Elizabeth," and asks, "What product is on this counter?" This information is sent to the server, which generates a response based on the character's configuration data. Next, the voice generation AI converts the response into an audio file in the character's voice quality, which is played on the user's smart glasses. As a result, users can enjoy a highly entertaining shopping experience in which they receive product information from a fictional character.
[0198] This system allows users to interact with fictional characters in real time in physical stores while obtaining product information, providing a special shopping experience.
[0199] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0200] Step 1:
[0201] The user puts on the smart device and starts the application. The device receives the user's authentication information as input and sends it to the server. The server compares the authentication information with a database and authenticates the user. After successful authentication, the user selects the character they want to talk to from the interface on the device.
[0202] Step 2:
[0203] The terminal sends the ID of the character selected by the user as input to the server. The server retrieves the character's profile data from a database and sends information such as the character's speaking style, personality, and voice quality as output to the terminal. This information is used for subsequent processing.
[0204] Step 3:
[0205] A conversation begins when a user speaks to the device. For example, if the user types, "What is the product on this counter?", the device sends the message to the server. The server uses the received message as input and requests processing by a generative AI model (language generation AI).
[0206] Step 4:
[0207] The server calls the generative AI model and generates a response text for the character based on the user's question. The generative AI model performs data calculations based on the input user message and character setting materials, and outputs the response text, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular."
[0208] Step 5:
[0209] The generated response text is passed as input to the voice generation AI, which performs data calculations to generate an audio file that reproduces the character's voice quality and tone. Specifically, the voice generation AI generates an audio signal from the response text and the character's voice quality data, and outputs it in file format to the server.
[0210] Step 6:
[0211] The server sends the generated audio file to the device, which receives it and plays it back to the user, making the user feel as if they are having a real-time conversation with a fictional character.
[0212] Step 7:
[0213] The user can continue to ask questions or make comments, for example, "That's interesting. Do you have any other product recommendations?", and the process will repeat. The user's message will be sent to the server, which will then pass it through the generative AI model and speech generation AI to generate an appropriate response, which will then be sent to the device.
[0214] In this way, an interactive conversation between the user and the virtual assistant character is realized, providing a highly entertaining shopping experience.
[0215] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0216] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting materials from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. In addition, by combining an emotion engine, the system recognizes the user's emotions and provides more appropriate responses based on those emotions. An embodiment of the present invention is described below.
[0217] Overall system overview
[0218] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, a voice generation AI, and an emotion engine, while the terminal is equipped with a user interface, real-time chat functionality, and voice output functionality. Users can start conversations with characters via their terminal and enjoy the results.
[0219] Overview of program processing
[0220] The system program operates as follows.
[0221] 1. Character Selection
[0222] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[0223] 2. Start a conversation
[0224] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[0225] 3. Emotional Recognition
[0226] The user's input message is also input to the emotion engine, which analyzes the user's emotions. The analyzed emotional information is reflected in the response generation, ensuring that the character responds appropriately.
[0227] 4. Speech Generation
[0228] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0229] 5. Keep the conversation going
[0230] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[0231] Specific examples
[0232] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[0233] When a user types, "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as, "I went to a cafe with a friend today! It was fun." At the same time, the emotion engine analyzes the emotion in the user's message, and if the message sounds happy, for example, it generates a response based on that emotion. Next, the voice generation AI converts this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[0234] If the user then continues, "That's nice. What did you have at the cafe?", the conversation continues in the same process. The emotion engine analyzes emotions again, providing a more realistic conversational experience.
[0235] This system allows users to enjoy a realistic conversational experience with the fictional characters they choose, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexuals. At the same time, the emotional engine makes the conversation more in tune with the user's emotions.
[0236] The processing flow will be explained below.
[0237] Step 1:
[0238] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[0239] Step 2:
[0240] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[0241] Step 3:
[0242] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[0243] Step 4:
[0244] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[0245] Step 5:
[0246] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[0247] Step 6:
[0248] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[0249] Step 7:
[0250] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[0251] Step 8:
[0252] At the same time, the server inputs the user's input message into the emotion engine to analyze the user's emotion. The emotion engine analyzes the content and context of the message and recognizes what emotion the user is feeling (e.g., happy, sad, surprised, etc.).
[0253] Step 9:
[0254] The server then feeds the analysis results from the emotion engine into the language generation AI, optimizing responses based on the user's emotions, allowing the character to respond in line with the user's emotions.
[0255] Step 10:
[0256] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[0257] Step 11:
[0258] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[0259] Step 12:
[0260] If the user wants to continue the conversation, they enter a new message and send it to the server using the same steps. The server then uses the language generation AI and emotion engine to generate a new reply, and the voice generation AI to generate an audio file, which is then sent to the device. This process is repeated, creating an interactive conversation.
[0261] Example 2
[0262] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0263] In conventional dialogue systems with fictional characters, the characters' responses are fixed and lack emotional change, making it difficult for users to have a natural conversational experience. Furthermore, it is difficult to reproduce the character's unique voice and tone of voice, making it difficult to say that interactions with the character are realistic. This makes it difficult to alleviate the feelings of loneliness and dissatisfaction felt by fictiosexuals in particular. Therefore, there is a need for a system that can analyze emotions in real time based on user input, respond in line with those emotions, and reproduce the character's unique voice quality and tone of voice.
[0264] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0265] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on a user's input text, means for analyzing the user's emotions using an emotion engine and reflecting the analysis in response generation, means for using a voice generation AI to generate an audio file in the character's unique voice from the generated response text, and means for the character and the user to have an interactive conversation on a terminal. This allows the user to enjoy a realistic and natural conversational experience with the character, which is expected to have an effect of alleviating the loneliness and unfulfilled feelings that fictiosexual people may have in particular.
[0266] A "fictional character" is a fictional character or character that appears in creative works such as stories, animations, and games.
[0267] "Setting materials" are data that describe in detail the personality, background, tone of voice, behavior patterns, etc. of a fictional character.
[0268] A "database" is a system for efficiently storing, searching, and managing multiple data.
[0269] "Generative AI" refers to algorithms or software that generate text using artificial intelligence techniques, such as natural language generation (NLG).
[0270] "User input text" is a written message or question that a user enters into the system.
[0271] "Response text" is the response text that the generation AI outputs in response to the user's input text.
[0272] An "emotion engine" is a technology or system for analyzing emotions from a user's input text and generating a response based on those emotions.
[0273] "Speech generation AI" is an artificial intelligence technology for converting generated text into voice data, for example, using text-to-speech (TTS) technology.
[0274] An "audio file" is a file in which audio data is stored in digital format.
[0275] A "terminal" is a device that a user uses to access and operate the system, including, for example, a smartphone or a personal computer.
[0276] "Interactive conversation" refers to a two-way form of communication in which the user and the system send and receive messages to each other and generate responses in real time.
[0277] This invention relates to a system that allows users to enjoy real-time interactive conversations with fictional characters. This system is composed of three elements: a server, a terminal, and a user, and is realized by combining these elements. Below, we will explain the details of each element and the operation of the system.
[0278] Server Features
[0279] The server is equipped with a database that stores setting materials for fictional characters, a generation AI that generates text based on user input, an emotion engine that analyzes user emotions, and a voice generation AI that converts the generated text into voice.
[0280] Database: Stores character information such as character personality, background, speech patterns, and behavior patterns. For example, an SQL database or NoSQL database is used.
[0281] Generative AI: For example, using OpenAI's GPT-3, it generates natural-sounding response text based on user input.
[0282] Emotion engine: Analyzes emotions from the user's input text and reflects the analysis results in the generative AI's responses. Commercial sentiment analysis APIs and proprietary sentiment analysis algorithms are used.
[0283] Voice generation AI: For example, using Google's WaveNet, the generated text is converted into an audio file with a character-specific voice quality.
[0284] Device Features
[0285] The terminals are equipped with a user interface, real-time chat functionality, and voice output functionality, and can be devices such as smartphones or PCs.
[0286] User interface: It provides a character selection screen and a chat window, making it easy for users to operate.
[0287] Real-time chat function: A function that receives user input, sends it to the server, and receives a response from the server. Web sockets and HTTP communication are used.
[0288] Audio output function: Plays received audio files and provides auditory information to the user. Built-in speakers or headphones are used.
[0289] User Actions
[0290] The user accesses the device's interface, selects the character they want to talk to, and enters a message. The server processes the message, generates a response from the character, and sends it to the device. The device then plays back the received response as audio, allowing the user to enjoy a continuous conversation with the character.
[0291] Specific examples
[0292] For example, if a user wants to talk to "Character A," they select "Character A" on the device's interface. Next, the user types, "What did you do today?" and the device sends this message to the server. The server uses a generative AI to generate a response such as, "I went to a cafe with my friends today," while the emotion engine simultaneously analyzes the emotions from the user's message. The result of this analysis is reflected in the generation of a response with an appropriate emotion.
[0293] Next, the server uses a voice generation AI to convert this response into an audio file in the voice of "Character A" and send it to the device. The device plays this audio file, and the user feels as if they are having a conversation with "Character A." Below is an example of a prompt sentence for the generation AI model.
[0294] "User: What did you do today?
[0295] Character: I went to a cafe with a friend today.
[0296] In this way, users can enjoy a realistic and natural interaction experience with fictional characters.
[0297] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0298] Step 1:
[0299] Character Selection
[0300] The user selects the character they want to talk to on the device interface. For example, the user selects "Character A." The input is the user's selection information, and the device sends this information (character ID) to the server. Specific data processing involves sending an HTTP request (for example, "GET / character?id=characterA") that includes the character ID to the server. The output is the character ID sent to the server.
[0301] Step 2:
[0302] Obtaining setting materials
[0303] The server retrieves character settings from the database based on the character ID it receives. The input is the character ID, and the server executes a database query based on this (e.g., "SELECT FROM CharacterSettings WHERE ID='characterA'"). Specific data processing involves returning the settings from the database. The output is the character's settings (information such as personality, background, and tone of voice).
[0304] Step 3:
[0305] User message input
[0306] The user types a message into the terminal's conversation window and presses the send button. For example, they type "What did you do today?". The input is the user's message text, which the terminal sends to the server in JSON format (e.g., "POST / receive_message"). The output is the user's message text sent to the server.
[0307] Step 4:
[0308] Receiving messages and generating replies
[0309] Based on the user's message and character setting information received by the server, a generative AI is used to generate a character's response text. The input is the user's message and character setting information, which are sent to the AI model in the form of a prompt sentence (e.g., "User: What did you do today?\nCharacter A: "). As a specific data calculation, the generative AI model generates the response text. The output is the generated response text (e.g., "I went to a cafe with my friends today.").
[0310] Step 5:
[0311] Emotional analysis and regulation
[0312] The server uses an emotion engine to analyze the user's input message and reflects the results in generating a reply. The input is the user's message text, and the emotion engine analyzes the emotion (e.g., "emotion.analyze('What did you do today?')"). As a specific data calculation, the analysis result is added to the prompt of the generation AI and re-generated (e.g., "User: What did you do today? (Emotion: Happy)\nCharacter A: "). The output is a reply text adjusted based on the emotion.
[0313] Step 6:
[0314] Generate an audio file
[0315] The server inputs the generated response text into the voice generation AI, which generates an audio file in the character's unique voice. The input is the adjusted response text, and the voice generation AI converts this text into audio data (e.g., "voice.generate('I went to a cafe with my friends today.', 'characterA')"). The specific data processing involves generating an audio file. The output is the generated audio file.
[0316] Step 7:
[0317] Sending and playing audio files
[0318] The server sends the generated audio file to the device, and the device plays the received audio file to the user. The input is the generated audio file, which the server sends as an HTTP response (e.g., "POST / play_voice"). The device uses its built-in media player to play the received audio file. The specific operation is that the audio file is played. The output is the audio played to the user.
[0319] Step 8:
[0320] Keeping the conversation going
[0321] The user again inputs a new message and continues the conversation with the character using the same process. The input is the message text entered by the user again. The output is the sending of a new message and the same series of processes being carried out again.
[0322] (Application example 2)
[0323] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0324] In conventional virtual stores, users have to spend a lot of time searching for products and it is difficult to receive appropriate support when selecting products. Furthermore, users cannot receive appropriate advice when selecting products based on emotional aspects, which limits the user experience. Furthermore, there is a lack of ways to make the shopping experience more personalized and provide users with emotional satisfaction.
[0325] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0326] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on text input by the user, means for selecting a character in a virtual store, means for the selected virtual character to support the user's shopping experience, means for analyzing the user's emotions using an emotion engine and reflecting the emotion information in the character's response, means for using a voice generation AI to generate the generated response text into an audio file in a voice unique to the character, and means for the character and the user to have an interactive conversation on a terminal. This enables the user to have an interactive conversation with the fictional character in the virtual store and receive product suggestions based on their emotions.
[0327] A "fictional character" is a fictional person, animal, or being that appears in an imaginary story or entertainment work.
[0328] "Setting materials" refers to data that includes detailed information about a fictional character's personality, background, voice quality, behavioral characteristics, etc.
[0329] "Generative AI" refers to artificial intelligence technology that generates natural language based on user input.
[0330] "Response text" refers to the text that the generative AI creates based on the user's input.
[0331] A "virtual store" is a virtual shopping environment that exists on the Internet.
[0332] An "emotion engine" is an algorithm that analyzes emotional information from user input and makes decisions based on that information.
[0333] "Voice generation AI" is an artificial intelligence technology that converts text data into voice data that resembles a human voice.
[0334] "Audio File" means audio data stored in digital format.
[0335] "Terminal" refers to a device used by a user, such as a computer, smartphone, or head-mounted display.
[0336] "Interactive conversation" means that the user and the system communicate in two directions in real time.
[0337] "Means to support the shopping experience" refers to a method in which a virtual character suggests products based on the user's purchasing motivation and emotions and supports the purchase.
[0338] MODE FOR CARRYING OUT THE INVENTION
[0339] The present invention provides a system that allows users to interact with fictional characters in a virtual store and receive product recommendations based on their emotions. This system is composed of multiple modules and can be specifically implemented as follows:
[0340] System configuration
[0341] Hardware
[0342] 1. Device: The device used by the user, such as a smartphone or head-mounted display
[0343] 2. Server: A cloud-based server containing the database and AI model
[0344] software
[0345] 1. Generation AI: OpenAI GPT-4
[0346] 2. Speech generation AI: Google Cloud Text-to-Speech, Amazon Polly
[0347] 3. Emotion Engine: IBM Watson Tone Analyzer
[0348] 4. Database: MySQL, MongoDB
[0349] Process Overview
[0350] 1. Character Selection
[0351] The user selects a fictional character to act as a shopping assistant in the virtual store. The selected character's ID is sent to the server, which retrieves the character's profile information from a database.
[0352] 2. Start a conversation
[0353] Users can ask questions or ask for advice by text input or voice, and the device sends this message to the server. The server uses a generative AI to generate a response text for the character based on the received message and character settings.
[0354] 3. Emotional Recognition
[0355] The user's message is analyzed by the emotion engine and emotional information is added, so that the character's response will be tailored to the user's emotions.
[0356] 4. Speech Generation
[0357] Based on the generated response text, the server uses a voice generation AI to generate an audio file in the character's unique voice, which is then sent to the device and played back to the user.
[0358] 5. Product proposal
[0359] Based on the user's questions and emotions, the character will suggest related products. For example, if a user inputs "I've been feeling stressed lately and want something to help me relax," the emotion engine will recognize the stress and suggest relaxation goods.
[0360] 6. Keep the conversation going
[0361] Users can continue the conversation by entering new messages, which allows for new product suggestions and responses to inquiries, continuing the interactive dialogue.
[0362] Examples and prompts
[0363] Example 1: A conversation when a tired user is looking for something to relax
[0364] (User): "I've been feeling really tired lately and I need something to help me relax."
[0365] (Character): "That's tough. How about a nice-smelling aroma diffuser? It's perfect for relaxing."
[0366] Example 2: A conversation between a user looking for a gift and asking for advice
[0367] (User): "I'm looking for a birthday gift for my friend. Any recommendations?"
[0368] (Character): "Depending on your friend's preferences, how about some popular fashion items or accessories?"
[0369] Example prompts for generative AI models
[0370] Example prompt 1:
[0371] User text: "I've been feeling really tired lately and I need something to help me relax."
[0372] Character information: {Personality: "Kind", Product knowledge: "Abundant"}
[0373] Produces: A response suggesting products to help a tired user relax.
[0374] Example prompt 2:
[0375] User texts: "I'm looking for a birthday gift for my friend. Any recommendations?"
[0376] Character information: {Personality: "Considerate", Product knowledge: "Abundant"}
[0377] Produces: A response suggesting products that would make good birthday gifts.
[0378] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0379] System program processing flow
[0380] Step 1: Character Selection
[0381] The user selects a fictional character to act as a shopping assistant in a virtual store. The selected character's ID is sent from the device to the server. The input is the selected character's ID, and the output is the character's profile information retrieved from the database. The server retrieves the character's profile information from the database based on this ID and sends it to the device.
[0382] Step 2: Start a conversation
[0383] Users ask questions or ask questions about shopping via text input or voice. This input is sent from the device to the server. The input is text data of the user's question or inquiry, and the output is the character's reply text. The server uses a generation AI to generate the character's reply text based on the received message and character setting materials, and sends it to the device.
[0384] Step 3: Recognize emotions
[0385] The server sends the user's input text to the emotion engine, which analyzes the user's emotions. The input is the user's text data, and the output is the analyzed emotional information. The emotion engine generates this emotional information and reflects it in the response text generated by the generation AI.
[0386] Step 4: Speech generation
[0387] The server passes the generated response text to the voice generation AI, which generates an audio file in the character's unique voice. The input is the response text, and the output is an audio file generated in the character's voice. The voice generation AI converts the text data into audio data, and sends the audio file from the server to the device.
[0388] Step 5: Product proposal
[0389] Based on the user's question and emotional information, a character on the server suggests related products. The input is the user's emotional information and question, and the output is text data of product suggestions. By combining an emotion engine and generative AI, products are suggested based on the user's emotions, and the suggestions are sent to the device.
[0390] Step 6: Keep the conversation going
[0391] The user inputs a new message, which is then sent from the device to the server. Again, a new response and product suggestions are generated through the same process, continuing the interactive dialogue. The input is a new user message, and the output is a new response text, audio file, and product suggestions.
[0392] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0393] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0394] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0395] [Second embodiment]
[0396] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0397] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0398] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0399] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0400] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0401] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0402] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0403] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0404] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0405] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0406] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0407] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0408] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting data from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. Specific embodiments of the system are described below.
[0409] Overall system overview
[0410] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, and a voice generation AI, while the terminal is equipped with a user interface, real-time chat function, and voice output function. Users can start conversations with characters through their terminal and enjoy the results.
[0411] Overview of program processing
[0412] The system program operates as follows.
[0413] 1. Character Selection
[0414] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[0415] 2. Start a conversation
[0416] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[0417] 3. Speech generation
[0418] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0419] 4. Keep the conversation going
[0420] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[0421] Specific examples
[0422] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[0423] When a user types "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as "I went to a cafe with my friends today! It was fun." The server then uses speech generation AI to convert this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[0424] If the user then continues with, "That's nice. What did you have at the cafe?" the conversation continues using the same process.
[0425] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexual people.
[0426] The processing flow will be explained below.
[0427] Step 1:
[0428] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[0429] Step 2:
[0430] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[0431] Step 3:
[0432] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[0433] Step 4:
[0434] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[0435] Step 5:
[0436] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[0437] Step 6:
[0438] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[0439] Step 7:
[0440] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[0441] Step 8:
[0442] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[0443] Step 9:
[0444] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[0445] Step 10:
[0446] If the user wants to continue the conversation, they can enter a new message and send it to the server using the same steps. The server then uses the language generation and speech generation AI to generate a new reply and sends it to the device. This process is repeated, creating an interactive conversation.
[0447] Example 1
[0448] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0449] There is a demand for a system that allows users to enjoy real-time interactive conversations with fictional characters. Conventional systems have made it difficult for users to obtain a satisfying experience because the characters' responses are monotonous and the voice generation is unnatural. Furthermore, the ability to select from multiple characters is limited, making it impossible to meet the diverse needs of users.
[0450] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0451] In this invention, the server includes means for retrieving setting materials for a fictional character selected by a user from a database, means for generating a response text for the fictional character based on a text input by the user using a generation AI, and means for generating an audio file using a voice generation AI from the generated response text in a voice unique to the fictional character, thereby enabling the user to enjoy natural and immersive interactive conversations with the fictional character.
[0452] "Users" are people who use the system to converse with fictional characters.
[0453] "Fictional characters" refers to fictional characters and animals that appear in creative works such as animation, manga, novels, and games.
[0454] "Setting materials" refers to information that describes a fictional character in detail, such as their personality, age, occupation, hobbies, etc.
[0455] A "database" is a collection of information that stores and manages setting materials for fictional characters and can be retrieved as needed.
[0456] "Generative AI" refers to artificial intelligence technology that generates response text from fictional characters based on user input text.
[0457] "Response text" refers to the reply content of a fictional character created by the generation AI in response to user input.
[0458] "Voice generation AI" is an artificial intelligence technology that converts response text created by the generation AI into an audio file that reproduces the tone and voice quality of a fictional character.
[0459] "Audio file" refers to digital data generated by voice generation AI to play back responses in the voices of fictional characters.
[0460] A "terminal" is a device or software that a user operates to interact with fictional characters.
[0461] "Interactive conversation" means that the user and the fictional character have two-way communication in real time via the system.
[0462] The present invention provides a system for enabling users to have real-time interactive conversations with fictional characters, which system is comprised of a server, a terminal, and a user.
[0463] Overall system configuration
[0464] The server includes the following components:
[0465] Database: Accumulates and manages setting materials for fictional characters.
[0466] Generative AI: Generates response text from a fictional character based on the user's input text.
[0467] Voice generation AI: Converts generated response text into an audio file with a voice unique to a fictional character.
[0468] The terminal is a device operated by the user and has the following functions:
[0469] User Interface: An interface that allows the user to select a character, input messages, and send them.
[0470] Real-time chat function: A function that communicates with the server and sends and receives messages in real time.
[0471] Audio output function: A function to play audio files received from the server.
[0472] The user operates the terminal to enjoy conversations with fictional characters.
[0473] System Operation Overview
[0474] First, the user operates the device's interface and selects the fictional character they want to talk to. The device then sends the character's ID to the server, which then retrieves the character's profile information from a database.
[0475] Next, the user enters a message to start the conversation and presses the send button. The device then sends the entered message to the server, which uses a generation AI to generate a reply text based on the received message and character settings.
[0476] Based on the generated response text, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0477] The user can enter a new message and continue the conversation in the same process. For each message sent from the terminal, the server generates an appropriate reply and outputs it as voice.
[0478] Specific examples
[0479] For example, if a user wants to have a conversation with a specific fictional character, the user selects the character on the device interface and sends the selection information to the server. The server retrieves character information from the database and starts a conversation based on that information.
[0480] When a user types "What did you do today?", the device sends this message to the server. The server uses a generation AI to generate a response such as "I went to a cafe with a friend today! It was fun." The server then uses a voice generation AI to convert this response into an audio file in the character's voice and sends it to the device. The device plays this audio file, allowing the user to feel as if they are having a conversation with the character.
[0481] If the user then types, "That's nice. What did you have at the cafe?" the conversation continues using a similar process.
[0482] Example prompt sentence:
[0483] "Based on the character's background material, generate a response to the following message: 'What did you do today?'"
[0484] "Say in the voice of a specific fictional character, 'I went to a cafe with my friends today. It was fun.'"
[0485] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice.
[0486] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0487] Step 1: Choose your character
[0488] Input: The user selects a character in the device interface.
[0489] Operation:
[0490] The terminal acquires the ID of the character selected by the user.
[0491] The ID of the selected character is sent to the server.
[0492] Output: The character's ID is sent to the server.
[0493] Step 2: Obtaining the configuration data
[0494] Input: The server receives the character's ID.
[0495] Operation:
[0496] The server searches the database for the character's settings information.
[0497] The setting materials include information such as the character's personality, age, occupation, hobbies, etc.
[0498] Obtain the setting materials.
[0499] Output: The configuration data is available on the server.
[0500] Step 3: Start a conversation
[0501] Input: The user types a message into the terminal.
[0502] Operation:
[0503] The user inputs text into the message input field on the terminal and presses the send button.
[0504] The terminal transmits the input text data to the server.
[0505] Output: The user's input text is sent to the server.
[0506] Step 4: Generate response text
[0507] Input: The server receives the user's input text and character profile information.
[0508] Operation:
[0509] The server passes prompts to the generation AI based on the user's message and character settings.
[0510] The generative AI generates response text that matches the character's tone and personality.
[0511] Output: The generated response text.
[0512] Step 5: Generate the audio file
[0513] Input: The server receives the generated response text.
[0514] Operation:
[0515] The server inputs the generated response text into the speech generation AI.
[0516] The voice generation AI generates audio files that reproduce the character's voice quality and speaking style.
[0517] Output: The generated audio file.
[0518] Step 6: Send and play the audio file
[0519] Input: The server has the generated audio file.
[0520] Operation:
[0521] The server transmits the generated audio file to the terminal.
[0522] The terminal receives the audio file and plays it to the user using its audio output capabilities.
[0523] Output: The reply is played to the user as audio.
[0524] Step 7: Keep the conversation going
[0525] Input: The user types a new message.
[0526] Operation:
[0527] The user types another question or comment into the terminal and presses the send button.
[0528] The entered text data is sent to the server again.
[0529] The server repeats the same process, using the generation AI and voice generation AI to generate new response text and audio files.
[0530] The terminal plays the received audio file.
[0531] Output: The conversation continues, generating new messages and replies.
[0532] (Application example 1)
[0533] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0534] In traditional brick-and-mortar stores, users lacked an effective and enjoyable way to obtain product information. Furthermore, there was no technology available to provide a highly entertaining shopping experience through interactions with fictional characters. This made it difficult to improve the user experience.
[0535] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0536] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for generating a response text for the character based on a user's input text using a generation AI, means for generating an audio file in the character's unique voice using a voice generation AI from the generated response text, means for the character and the user to have an interactive conversation on a terminal, and means for the character to provide product information and services as a virtual assistant using a smart device in a physical store. This allows users to obtain product information while conversing with the fictional character in real time in the physical store, enabling them to enjoy a highly entertaining shopping experience.
[0537] A "fictional character" is a person or creature that appears in a fictional world, and whose setting is depicted in detail through media such as stories and games.
[0538] A "database" is an electronic storage device or system that systematically organizes information or data so that it can be efficiently accessed and managed.
[0539] "Generative AI" is an artificial intelligence technology that uses natural language processing and data analysis techniques to generate text and data based on user input.
[0540] "Voice generation AI" is a technology that converts text data into voice data, and is an artificial intelligence technology that can reproduce specific voice qualities and speaking styles.
[0541] A "terminal" is an electronic device that can be directly operated by a user, such as a computer, tablet, or smartphone.
[0542] "Interactive conversational means" refers to methods or functions that allow users and systems to communicate with each other in real time.
[0543] A "brick and mortar store" is a store with a physical sales floor that exists in the real world, as opposed to an online store.
[0544] A "smart device" is an electronic device with advanced functionality that can connect to the Internet and run various applications, including smartphones, smart glasses, and head-mounted displays.
[0545] A "virtual assistant" is a digital assistant that uses artificial intelligence technology to answer users' questions and provide services.
[0546] "Product Guide" is a means of providing information about products in a store and helping users consider purchasing them.
[0547] "User experience" refers to the experience and satisfaction that users feel when using a product or service.
[0548] The embodiment of the present invention is to build a system that is basically composed of a server, a terminal, and a user. An overview of the entire system and each of its main steps will be explained below.
[0549] Overall system overview
[0550] The system retrieves fictional character settings from a database, uses generative and voice-generating AI to generate responses for the characters, and outputs them as voice. The system is designed to improve the shopping experience in brick-and-mortar stores.
[0551] Hardware and software used
[0552] Hardware: Servers, smart glasses (e.g., Google Glass), head-mounted displays (e.g., Microsoft HoloLens)
[0553] Software: Language generation AI (e.g. GPT-4), speech generation AI (Google Text-to-Speech, Amazon Polly), user interface (developed with Unity or Unreal Engine)
[0554] System operation example
[0555] 1. User authentication and character selection
[0556] The user puts on the smart device and starts the application. After logging in, the user selects the character they want to talk to. For example, the user selects the character "Elizabeth." This selection information is sent to the server.
[0557] Example prompt sentence:
[0558] Character: Elizabeth
[0559] User Question: What is the product on this counter?
[0560] Preferred response language: Japanese
[0561] Information to include in the response: product type, features, special offers, etc.
[0562] As Elizabeth: "This counter is lined with the latest technology gadgets. Our most popular item is the smart wristband. It has health tracking features and is currently on sale at a special price."
[0563] 2. Obtaining character setting data
[0564] The server retrieves the character's profile information from a database, including the character's speaking style, personality, and voice quality.
[0565] 3. Conversation initiation and response generation
[0566] When a user speaks to the character, the message is sent to the server, which uses language generation AI to generate an appropriate response to the user's input. For example, if a user asks, "What products are on this counter?", the server will generate a response such as, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular and have health management functions."
[0567] 4. Speech generation and output
[0568] Based on the generated response text, the AI creates an audio file that reproduces the character's voice quality and tone. This audio file is then sent to the user's smart device and played back, making the user feel as if they are having a real-time conversation with a fictional character.
[0569] Specific examples
[0570] For example, imagine a user puts on smart glasses in a physical store, selects a character named "Elizabeth," and asks, "What product is on this counter?" This information is sent to the server, which generates a response based on the character's configuration data. Next, the voice generation AI converts the response into an audio file in the character's voice quality, which is played on the user's smart glasses. As a result, users can enjoy a highly entertaining shopping experience in which they receive product information from a fictional character.
[0571] This system allows users to interact with fictional characters in real time in physical stores while obtaining product information, providing a special shopping experience.
[0572] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0573] Step 1:
[0574] The user puts on the smart device and starts the application. The device receives the user's authentication information as input and sends it to the server. The server compares the authentication information with a database and authenticates the user. After successful authentication, the user selects the character they want to talk to from the interface on the device.
[0575] Step 2:
[0576] The terminal sends the ID of the character selected by the user as input to the server. The server retrieves the character's profile data from a database and sends information such as the character's speaking style, personality, and voice quality as output to the terminal. This information is used for subsequent processing.
[0577] Step 3:
[0578] A conversation begins when a user speaks to the device. For example, if the user types, "What is the product on this counter?", the device sends the message to the server. The server uses the received message as input and requests processing by a generative AI model (language generation AI).
[0579] Step 4:
[0580] The server calls the generative AI model and generates a response text for the character based on the user's question. The generative AI model performs data calculations based on the input user message and character setting materials, and outputs the response text, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular."
[0581] Step 5:
[0582] The generated response text is passed as input to the voice generation AI, which performs data calculations to generate an audio file that reproduces the character's voice quality and tone. Specifically, the voice generation AI generates an audio signal from the response text and the character's voice quality data, and outputs it in file format to the server.
[0583] Step 6:
[0584] The server sends the generated audio file to the device, which receives it and plays it back to the user, making the user feel as if they are having a real-time conversation with a fictional character.
[0585] Step 7:
[0586] The user can continue to ask questions or make comments, for example, "That's interesting. Do you have any other product recommendations?", and the process will repeat. The user's message will be sent to the server, which will then pass it through the generative AI model and speech generation AI to generate an appropriate response, which will then be sent to the device.
[0587] In this way, an interactive conversation between the user and the virtual assistant character is realized, providing a highly entertaining shopping experience.
[0588] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0589] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting materials from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. In addition, by combining an emotion engine, the system recognizes the user's emotions and provides more appropriate responses based on those emotions. An embodiment of the present invention is described below.
[0590] Overall system overview
[0591] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, a voice generation AI, and an emotion engine, while the terminal is equipped with a user interface, real-time chat functionality, and voice output functionality. Users can start conversations with characters via their terminal and enjoy the results.
[0592] Overview of program processing
[0593] The system program operates as follows.
[0594] 1. Character Selection
[0595] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[0596] 2. Start a conversation
[0597] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[0598] 3. Emotional Recognition
[0599] The user's input message is also input to the emotion engine, which analyzes the user's emotions. The analyzed emotional information is reflected in the response generation, ensuring that the character responds appropriately.
[0600] 4. Speech Generation
[0601] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0602] 5. Keep the conversation going
[0603] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[0604] Specific examples
[0605] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[0606] When a user types, "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as, "I went to a cafe with a friend today! It was fun." At the same time, the emotion engine analyzes the emotion in the user's message, and if the message sounds happy, for example, it generates a response based on that emotion. Next, the voice generation AI converts this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[0607] If the user then continues, "That's nice. What did you have at the cafe?", the conversation continues in the same process. The emotion engine analyzes emotions again, providing a more realistic conversational experience.
[0608] This system allows users to enjoy a realistic conversational experience with the fictional characters they choose, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexuals. At the same time, the emotional engine makes the conversation more in tune with the user's emotions.
[0609] The processing flow will be explained below.
[0610] Step 1:
[0611] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[0612] Step 2:
[0613] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[0614] Step 3:
[0615] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[0616] Step 4:
[0617] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[0618] Step 5:
[0619] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[0620] Step 6:
[0621] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[0622] Step 7:
[0623] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[0624] Step 8:
[0625] At the same time, the server inputs the user's input message into the emotion engine to analyze the user's emotion. The emotion engine analyzes the content and context of the message and recognizes what emotion the user is feeling (e.g., happy, sad, surprised, etc.).
[0626] Step 9:
[0627] The server then feeds the analysis results from the emotion engine into the language generation AI, optimizing responses based on the user's emotions, allowing the character to respond in line with the user's emotions.
[0628] Step 10:
[0629] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[0630] Step 11:
[0631] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[0632] Step 12:
[0633] If the user wants to continue the conversation, they enter a new message and send it to the server using the same steps. The server then uses the language generation AI and emotion engine to generate a new reply, and the voice generation AI to generate an audio file, which is then sent to the device. This process is repeated, creating an interactive conversation.
[0634] Example 2
[0635] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0636] In conventional dialogue systems with fictional characters, the characters' responses are fixed and lack emotional change, making it difficult for users to have a natural conversational experience. Furthermore, it is difficult to reproduce the character's unique voice and tone of voice, making it difficult to say that interactions with the character are realistic. This makes it difficult to alleviate the feelings of loneliness and dissatisfaction felt by fictiosexuals in particular. Therefore, there is a need for a system that can analyze emotions in real time based on user input, respond in line with those emotions, and reproduce the character's unique voice quality and tone of voice.
[0637] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0638] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on a user's input text, means for analyzing the user's emotions using an emotion engine and reflecting the analysis in response generation, means for using a voice generation AI to generate an audio file in the character's unique voice from the generated response text, and means for the character and the user to have an interactive conversation on a terminal. This allows the user to enjoy a realistic and natural conversational experience with the character, which is expected to have an effect of alleviating the loneliness and unfulfilled feelings that fictiosexual people may have in particular.
[0639] A "fictional character" is a fictional character or character that appears in creative works such as stories, animations, and games.
[0640] "Setting materials" are data that describe in detail the personality, background, tone of voice, behavior patterns, etc. of a fictional character.
[0641] A "database" is a system for efficiently storing, searching, and managing multiple data.
[0642] "Generative AI" refers to algorithms or software that generate text using artificial intelligence techniques, such as natural language generation (NLG).
[0643] "User input text" is a written message or question that a user enters into the system.
[0644] "Response text" is the response text that the generation AI outputs in response to the user's input text.
[0645] An "emotion engine" is a technology or system for analyzing emotions from a user's input text and generating a response based on those emotions.
[0646] "Speech generation AI" is an artificial intelligence technology for converting generated text into voice data, for example, using text-to-speech (TTS) technology.
[0647] An "audio file" is a file in which audio data is stored in digital format.
[0648] A "terminal" is a device that a user uses to access and operate the system, including, for example, a smartphone or a personal computer.
[0649] "Interactive conversation" refers to a two-way form of communication in which the user and the system send and receive messages to each other and generate responses in real time.
[0650] This invention relates to a system that allows users to enjoy real-time interactive conversations with fictional characters. This system is composed of three elements: a server, a terminal, and a user, and is realized by combining these elements. Below, we will explain the details of each element and the operation of the system.
[0651] Server Features
[0652] The server is equipped with a database that stores setting materials for fictional characters, a generation AI that generates text based on user input, an emotion engine that analyzes user emotions, and a voice generation AI that converts the generated text into voice.
[0653] Database: Stores character information such as character personality, background, speech patterns, and behavior patterns. For example, an SQL database or NoSQL database is used.
[0654] Generative AI: For example, using OpenAI's GPT-3, it generates natural-sounding response text based on user input.
[0655] Emotion engine: Analyzes emotions from the user's input text and reflects the analysis results in the generative AI's responses. Commercial sentiment analysis APIs and proprietary sentiment analysis algorithms are used.
[0656] Voice generation AI: For example, using Google's WaveNet, the generated text is converted into an audio file with a character-specific voice quality.
[0657] Device Features
[0658] The terminals are equipped with a user interface, real-time chat functionality, and voice output functionality, and can be devices such as smartphones or PCs.
[0659] User interface: It provides a character selection screen and a chat window, making it easy for users to operate.
[0660] Real-time chat function: A function that receives user input, sends it to the server, and receives a response from the server. Web sockets and HTTP communication are used.
[0661] Audio output function: Plays received audio files and provides auditory information to the user. Built-in speakers or headphones are used.
[0662] User Actions
[0663] The user accesses the device's interface, selects the character they want to talk to, and enters a message. The server processes the message, generates a response from the character, and sends it to the device. The device then plays back the received response as audio, allowing the user to enjoy a continuous conversation with the character.
[0664] Specific examples
[0665] For example, if a user wants to talk to "Character A," they select "Character A" on the device's interface. Next, the user types, "What did you do today?" and the device sends this message to the server. The server uses a generative AI to generate a response such as, "I went to a cafe with my friends today," while the emotion engine simultaneously analyzes the emotions from the user's message. The result of this analysis is reflected in the generation of a response with an appropriate emotion.
[0666] Next, the server uses a voice generation AI to convert this response into an audio file in the voice of "Character A" and send it to the device. The device plays this audio file, and the user feels as if they are having a conversation with "Character A." Below is an example of a prompt sentence for the generation AI model.
[0667] "User: What did you do today?
[0668] Character: I went to a cafe with a friend today.
[0669] In this way, users can enjoy a realistic and natural interaction experience with fictional characters.
[0670] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0671] Step 1:
[0672] Character Selection
[0673] The user selects the character they want to talk to on the device interface. For example, the user selects "Character A." The input is the user's selection information, and the device sends this information (character ID) to the server. Specific data processing involves sending an HTTP request (for example, "GET / character?id=characterA") that includes the character ID to the server. The output is the character ID sent to the server.
[0674] Step 2:
[0675] Obtaining setting materials
[0676] The server retrieves character settings from the database based on the character ID it receives. The input is the character ID, and the server executes a database query based on this (e.g., "SELECT FROM CharacterSettings WHERE ID='characterA'"). Specific data processing involves returning the settings from the database. The output is the character's settings (information such as personality, background, and tone of voice).
[0677] Step 3:
[0678] User message input
[0679] The user types a message into the terminal's conversation window and presses the send button. For example, they type "What did you do today?". The input is the user's message text, which the terminal sends to the server in JSON format (e.g., "POST / receive_message"). The output is the user's message text sent to the server.
[0680] Step 4:
[0681] Receiving messages and generating replies
[0682] Based on the user's message and character setting information received by the server, a generative AI is used to generate a character's response text. The input is the user's message and character setting information, which are sent to the AI model in the form of a prompt sentence (e.g., "User: What did you do today?\nCharacter A: "). As a specific data calculation, the generative AI model generates the response text. The output is the generated response text (e.g., "I went to a cafe with my friends today.").
[0683] Step 5:
[0684] Emotional analysis and regulation
[0685] The server uses an emotion engine to analyze the user's input message and reflects the results in generating a reply. The input is the user's message text, and the emotion engine analyzes the emotion (e.g., "emotion.analyze('What did you do today?')"). As a specific data calculation, the analysis result is added to the prompt of the generation AI and re-generated (e.g., "User: What did you do today? (Emotion: Happy)\nCharacter A: "). The output is a reply text adjusted based on the emotion.
[0686] Step 6:
[0687] Generate an audio file
[0688] The server inputs the generated response text into the voice generation AI, which generates an audio file in the character's unique voice. The input is the adjusted response text, and the voice generation AI converts this text into audio data (e.g., "voice.generate('I went to a cafe with my friends today.', 'characterA')"). The specific data processing involves generating an audio file. The output is the generated audio file.
[0689] Step 7:
[0690] Sending and playing audio files
[0691] The server sends the generated audio file to the device, and the device plays the received audio file to the user. The input is the generated audio file, which the server sends as an HTTP response (e.g., "POST / play_voice"). The device uses its built-in media player to play the received audio file. The specific operation is that the audio file is played. The output is the audio played to the user.
[0692] Step 8:
[0693] Keeping the conversation going
[0694] The user again inputs a new message and continues the conversation with the character using the same process. The input is the message text entered by the user again. The output is the sending of a new message and the same series of processes being carried out again.
[0695] (Application example 2)
[0696] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0697] In conventional virtual stores, users have to spend a lot of time searching for products and it is difficult to receive appropriate support when selecting products. Furthermore, users cannot receive appropriate advice when selecting products based on emotional aspects, which limits the user experience. Furthermore, there is a lack of ways to make the shopping experience more personalized and provide users with emotional satisfaction.
[0698] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0699] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on text input by the user, means for selecting a character in a virtual store, means for the selected virtual character to support the user's shopping experience, means for analyzing the user's emotions using an emotion engine and reflecting the emotion information in the character's response, means for using a voice generation AI to generate the generated response text into an audio file in a voice unique to the character, and means for the character and the user to have an interactive conversation on a terminal. This enables the user to have an interactive conversation with the fictional character in the virtual store and receive product suggestions based on their emotions.
[0700] A "fictional character" is a fictional person, animal, or being that appears in an imaginary story or entertainment work.
[0701] "Setting materials" refers to data that includes detailed information about a fictional character's personality, background, voice quality, behavioral characteristics, etc.
[0702] "Generative AI" refers to artificial intelligence technology that generates natural language based on user input.
[0703] "Response text" refers to the text that the generative AI creates based on the user's input.
[0704] A "virtual store" is a virtual shopping environment that exists on the Internet.
[0705] An "emotion engine" is an algorithm that analyzes emotional information from user input and makes decisions based on that information.
[0706] "Voice generation AI" is an artificial intelligence technology that converts text data into voice data that resembles a human voice.
[0707] "Audio File" means audio data stored in digital format.
[0708] "Terminal" refers to a device used by a user, such as a computer, smartphone, or head-mounted display.
[0709] "Interactive conversation" means that the user and the system communicate in two directions in real time.
[0710] "Means to support the shopping experience" refers to a method in which a virtual character suggests products based on the user's purchasing motivation and emotions and supports the purchase.
[0711] MODE FOR CARRYING OUT THE INVENTION
[0712] The present invention provides a system that allows users to interact with fictional characters in a virtual store and receive product recommendations based on their emotions. This system is composed of multiple modules and can be specifically implemented as follows:
[0713] System configuration
[0714] Hardware
[0715] 1. Device: The device used by the user, such as a smartphone or head-mounted display
[0716] 2. Server: A cloud-based server containing the database and AI model
[0717] software
[0718] 1. Generation AI: OpenAI GPT-4
[0719] 2. Speech generation AI: Google Cloud Text-to-Speech, Amazon Polly
[0720] 3. Emotion Engine: IBM Watson Tone Analyzer
[0721] 4. Database: MySQL, MongoDB
[0722] Process Overview
[0723] 1. Character Selection
[0724] The user selects a fictional character to act as a shopping assistant in the virtual store. The selected character's ID is sent to the server, which retrieves the character's profile information from a database.
[0725] 2. Start a conversation
[0726] Users can ask questions or ask for advice by text input or voice, and the device sends this message to the server. The server uses a generative AI to generate a response text for the character based on the received message and character settings.
[0727] 3. Emotional Recognition
[0728] The user's message is analyzed by the emotion engine and emotional information is added, so that the character's response will be tailored to the user's emotions.
[0729] 4. Speech Generation
[0730] Based on the generated response text, the server uses a voice generation AI to generate an audio file in the character's unique voice, which is then sent to the device and played back to the user.
[0731] 5. Product proposal
[0732] Based on the user's questions and emotions, the character will suggest related products. For example, if a user inputs "I've been feeling stressed lately and want something to help me relax," the emotion engine will recognize the stress and suggest relaxation goods.
[0733] 6. Keep the conversation going
[0734] Users can continue the conversation by entering new messages, which allows for new product suggestions and responses to inquiries, continuing the interactive dialogue.
[0735] Examples and prompts
[0736] Example 1: A conversation when a tired user is looking for something to relax
[0737] (User): "I've been feeling really tired lately and I need something to help me relax."
[0738] (Character): "That's tough. How about a nice-smelling aroma diffuser? It's perfect for relaxing."
[0739] Example 2: A conversation between a user looking for a gift and asking for advice
[0740] (User): "I'm looking for a birthday gift for my friend. Any recommendations?"
[0741] (Character): "Depending on your friend's preferences, how about some popular fashion items or accessories?"
[0742] Example prompts for generative AI models
[0743] Example prompt 1:
[0744] User text: "I've been feeling really tired lately and I need something to help me relax."
[0745] Character information: {Personality: "Kind", Product knowledge: "Abundant"}
[0746] Produces: A response suggesting products to help a tired user relax.
[0747] Example prompt 2:
[0748] User texts: "I'm looking for a birthday gift for my friend. Any recommendations?"
[0749] Character information: {Personality: "Considerate", Product knowledge: "Abundant"}
[0750] Produces: A response suggesting products that would make good birthday gifts.
[0751] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0752] System program processing flow
[0753] Step 1: Character Selection
[0754] The user selects a fictional character to act as a shopping assistant in a virtual store. The selected character's ID is sent from the device to the server. The input is the selected character's ID, and the output is the character's profile information retrieved from the database. The server retrieves the character's profile information from the database based on this ID and sends it to the device.
[0755] Step 2: Start a conversation
[0756] Users ask questions or ask questions about shopping via text input or voice. This input is sent from the device to the server. The input is text data of the user's question or inquiry, and the output is the character's reply text. The server uses a generation AI to generate the character's reply text based on the received message and character setting materials, and sends it to the device.
[0757] Step 3: Recognize emotions
[0758] The server sends the user's input text to the emotion engine, which analyzes the user's emotions. The input is the user's text data, and the output is the analyzed emotional information. The emotion engine generates this emotional information and reflects it in the response text generated by the generation AI.
[0759] Step 4: Speech generation
[0760] The server passes the generated response text to the voice generation AI, which generates an audio file in the character's unique voice. The input is the response text, and the output is an audio file generated in the character's voice. The voice generation AI converts the text data into audio data, and sends the audio file from the server to the device.
[0761] Step 5: Product proposal
[0762] Based on the user's question and emotional information, a character on the server suggests related products. The input is the user's emotional information and question, and the output is text data of product suggestions. By combining an emotion engine and generative AI, products are suggested based on the user's emotions, and the suggestions are sent to the device.
[0763] Step 6: Keep the conversation going
[0764] The user inputs a new message, which is then sent from the device to the server. Again, a new response and product suggestions are generated through the same process, continuing the interactive dialogue. The input is a new user message, and the output is a new response text, audio file, and product suggestions.
[0765] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0766] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0767] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0768] [Third embodiment]
[0769] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0770] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0771] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0772] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0773] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0774] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0775] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0776] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0777] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0778] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0779] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0780] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0781] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting data from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. Specific embodiments of the system are described below.
[0782] Overall system overview
[0783] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, and a voice generation AI, while the terminal is equipped with a user interface, real-time chat function, and voice output function. Users can start conversations with characters through their terminal and enjoy the results.
[0784] Overview of program processing
[0785] The system program operates as follows.
[0786] 1. Character Selection
[0787] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[0788] 2. Start a conversation
[0789] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[0790] 3. Speech generation
[0791] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0792] 4. Keep the conversation going
[0793] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[0794] Specific examples
[0795] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[0796] When a user types "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as "I went to a cafe with my friends today! It was fun." The server then uses speech generation AI to convert this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[0797] If the user then continues with, "That's nice. What did you have at the cafe?" the conversation continues using the same process.
[0798] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexual people.
[0799] The processing flow will be explained below.
[0800] Step 1:
[0801] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[0802] Step 2:
[0803] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[0804] Step 3:
[0805] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[0806] Step 4:
[0807] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[0808] Step 5:
[0809] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[0810] Step 6:
[0811] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[0812] Step 7:
[0813] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[0814] Step 8:
[0815] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[0816] Step 9:
[0817] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[0818] Step 10:
[0819] If the user wants to continue the conversation, they can enter a new message and send it to the server using the same steps. The server then uses the language generation and speech generation AI to generate a new reply and sends it to the device. This process is repeated, creating an interactive conversation.
[0820] Example 1
[0821] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0822] There is a demand for a system that allows users to enjoy real-time interactive conversations with fictional characters. Conventional systems have made it difficult for users to obtain a satisfying experience because the characters' responses are monotonous and the voice generation is unnatural. Furthermore, the ability to select from multiple characters is limited, making it impossible to meet the diverse needs of users.
[0823] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0824] In this invention, the server includes means for retrieving setting materials for a fictional character selected by a user from a database, means for generating a response text for the fictional character based on a text input by the user using a generation AI, and means for generating an audio file using a voice generation AI from the generated response text in a voice unique to the fictional character, thereby enabling the user to enjoy natural and immersive interactive conversations with the fictional character.
[0825] "Users" are people who use the system to converse with fictional characters.
[0826] "Fictional characters" refers to fictional characters and animals that appear in creative works such as animation, manga, novels, and games.
[0827] "Setting materials" refers to information that describes a fictional character in detail, such as their personality, age, occupation, hobbies, etc.
[0828] A "database" is a collection of information that stores and manages setting materials for fictional characters and can be retrieved as needed.
[0829] "Generative AI" refers to artificial intelligence technology that generates response text from fictional characters based on user input text.
[0830] "Response text" refers to the reply content of a fictional character created by the generation AI in response to user input.
[0831] "Voice generation AI" is an artificial intelligence technology that converts response text created by the generation AI into an audio file that reproduces the tone and voice quality of a fictional character.
[0832] "Audio file" refers to digital data generated by voice generation AI to play back responses in the voices of fictional characters.
[0833] A "terminal" is a device or software that a user operates to interact with fictional characters.
[0834] "Interactive conversation" means that the user and the fictional character have two-way communication in real time via the system.
[0835] The present invention provides a system for enabling users to have real-time interactive conversations with fictional characters, which system is comprised of a server, a terminal, and a user.
[0836] Overall system configuration
[0837] The server includes the following components:
[0838] Database: Accumulates and manages setting materials for fictional characters.
[0839] Generative AI: Generates response text from a fictional character based on the user's input text.
[0840] Voice generation AI: Converts generated response text into an audio file with a voice unique to a fictional character.
[0841] The terminal is a device operated by the user and has the following functions:
[0842] User Interface: An interface that allows the user to select a character, input messages, and send them.
[0843] Real-time chat function: A function that communicates with the server and sends and receives messages in real time.
[0844] Audio output function: A function to play audio files received from the server.
[0845] The user operates the terminal to enjoy conversations with fictional characters.
[0846] System Operation Overview
[0847] First, the user operates the device's interface and selects the fictional character they want to talk to. The device then sends the character's ID to the server, which then retrieves the character's profile information from a database.
[0848] Next, the user enters a message to start the conversation and presses the send button. The device then sends the entered message to the server, which uses a generation AI to generate a reply text based on the received message and character settings.
[0849] Based on the generated response text, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0850] The user can enter a new message and continue the conversation in the same process. For each message sent from the terminal, the server generates an appropriate reply and outputs it as voice.
[0851] Specific examples
[0852] For example, if a user wants to have a conversation with a specific fictional character, the user selects the character on the device interface and sends the selection information to the server. The server retrieves character information from the database and starts a conversation based on that information.
[0853] When a user types "What did you do today?", the device sends this message to the server. The server uses a generation AI to generate a response such as "I went to a cafe with a friend today! It was fun." The server then uses a voice generation AI to convert this response into an audio file in the character's voice and sends it to the device. The device plays this audio file, allowing the user to feel as if they are having a conversation with the character.
[0854] If the user then types, "That's nice. What did you have at the cafe?" the conversation continues using a similar process.
[0855] Example prompt sentence:
[0856] "Based on the character's background material, generate a response to the following message: 'What did you do today?'"
[0857] "Say in the voice of a specific fictional character, 'I went to a cafe with my friends today. It was fun.'"
[0858] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice.
[0859] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0860] Step 1: Choose your character
[0861] Input: The user selects a character in the device interface.
[0862] Operation:
[0863] The terminal acquires the ID of the character selected by the user.
[0864] The ID of the selected character is sent to the server.
[0865] Output: The character's ID is sent to the server.
[0866] Step 2: Obtaining the configuration data
[0867] Input: The server receives the character's ID.
[0868] Operation:
[0869] The server searches the database for the character's settings information.
[0870] The setting materials include information such as the character's personality, age, occupation, hobbies, etc.
[0871] Obtain the setting materials.
[0872] Output: The configuration data is available on the server.
[0873] Step 3: Start a conversation
[0874] Input: The user types a message into the terminal.
[0875] Operation:
[0876] The user inputs text into the message input field on the terminal and presses the send button.
[0877] The terminal transmits the input text data to the server.
[0878] Output: The user's input text is sent to the server.
[0879] Step 4: Generate response text
[0880] Input: The server receives the user's input text and character profile information.
[0881] Operation:
[0882] The server passes prompts to the generation AI based on the user's message and character settings.
[0883] The generative AI generates response text that matches the character's tone and personality.
[0884] Output: The generated response text.
[0885] Step 5: Generate the audio file
[0886] Input: The server receives the generated response text.
[0887] Operation:
[0888] The server inputs the generated response text into the speech generation AI.
[0889] The voice generation AI generates audio files that reproduce the character's voice quality and speaking style.
[0890] Output: The generated audio file.
[0891] Step 6: Send and play the audio file
[0892] Input: The server has the generated audio file.
[0893] Operation:
[0894] The server transmits the generated audio file to the terminal.
[0895] The terminal receives the audio file and plays it to the user using its audio output capabilities.
[0896] Output: The reply is played to the user as audio.
[0897] Step 7: Keep the conversation going
[0898] Input: The user types a new message.
[0899] Operation:
[0900] The user types another question or comment into the terminal and presses the send button.
[0901] The entered text data is sent to the server again.
[0902] The server repeats the same process, using the generation AI and voice generation AI to generate new response text and audio files.
[0903] The terminal plays the received audio file.
[0904] Output: The conversation continues, generating new messages and replies.
[0905] (Application example 1)
[0906] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0907] In traditional brick-and-mortar stores, users lacked an effective and enjoyable way to obtain product information. Furthermore, there was no technology available to provide a highly entertaining shopping experience through interactions with fictional characters. This made it difficult to improve the user experience.
[0908] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0909] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for generating a response text for the character based on a user's input text using a generation AI, means for generating an audio file in the character's unique voice using a voice generation AI from the generated response text, means for the character and the user to have an interactive conversation on a terminal, and means for the character to provide product information and services as a virtual assistant using a smart device in a physical store. This allows users to obtain product information while conversing with the fictional character in real time in the physical store, enabling them to enjoy a highly entertaining shopping experience.
[0910] A "fictional character" is a person or creature that appears in a fictional world, and whose setting is depicted in detail through media such as stories and games.
[0911] A "database" is an electronic storage device or system that systematically organizes information or data so that it can be efficiently accessed and managed.
[0912] "Generative AI" is an artificial intelligence technology that uses natural language processing and data analysis techniques to generate text and data based on user input.
[0913] "Voice generation AI" is a technology that converts text data into voice data, and is an artificial intelligence technology that can reproduce specific voice qualities and speaking styles.
[0914] A "terminal" is an electronic device that can be directly operated by a user, such as a computer, tablet, or smartphone.
[0915] "Interactive conversational means" refers to methods or functions that allow users and systems to communicate with each other in real time.
[0916] A "brick and mortar store" is a store with a physical sales floor that exists in the real world, as opposed to an online store.
[0917] A "smart device" is an electronic device with advanced functionality that can connect to the Internet and run various applications, including smartphones, smart glasses, and head-mounted displays.
[0918] A "virtual assistant" is a digital assistant that uses artificial intelligence technology to answer users' questions and provide services.
[0919] "Product Guide" is a means of providing information about products in a store and helping users consider purchasing them.
[0920] "User experience" refers to the experience and satisfaction that users feel when using a product or service.
[0921] The embodiment of the present invention is to build a system that is basically composed of a server, a terminal, and a user. An overview of the entire system and each of its main steps will be explained below.
[0922] Overall system overview
[0923] The system retrieves fictional character settings from a database, uses generative and voice-generating AI to generate responses for the characters, and outputs them as voice. The system is designed to improve the shopping experience in brick-and-mortar stores.
[0924] Hardware and software used
[0925] Hardware: Servers, smart glasses (e.g., Google Glass), head-mounted displays (e.g., Microsoft HoloLens)
[0926] Software: Language generation AI (e.g. GPT-4), speech generation AI (Google Text-to-Speech, Amazon Polly), user interface (developed with Unity or Unreal Engine)
[0927] System operation example
[0928] 1. User authentication and character selection
[0929] The user puts on the smart device and starts the application. After logging in, the user selects the character they want to talk to. For example, the user selects the character "Elizabeth." This selection information is sent to the server.
[0930] Example prompt sentence:
[0931] Character: Elizabeth
[0932] User Question: What is the product on this counter?
[0933] Preferred response language: Japanese
[0934] Information to include in the response: product type, features, special offers, etc.
[0935] As Elizabeth: "This counter is lined with the latest technology gadgets. Our most popular item is the smart wristband. It has health tracking features and is currently on sale at a special price."
[0936] 2. Obtaining character setting data
[0937] The server retrieves the character's profile information from a database, including the character's speaking style, personality, and voice quality.
[0938] 3. Conversation initiation and response generation
[0939] When a user speaks to the character, the message is sent to the server, which uses language generation AI to generate an appropriate response to the user's input. For example, if a user asks, "What products are on this counter?", the server will generate a response such as, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular and have health management functions."
[0940] 4. Speech generation and output
[0941] Based on the generated response text, the AI creates an audio file that reproduces the character's voice quality and tone. This audio file is then sent to the user's smart device and played back, making the user feel as if they are having a real-time conversation with a fictional character.
[0942] Specific examples
[0943] For example, imagine a user puts on smart glasses in a physical store, selects a character named "Elizabeth," and asks, "What product is on this counter?" This information is sent to the server, which generates a response based on the character's configuration data. Next, the voice generation AI converts the response into an audio file in the character's voice quality, which is played on the user's smart glasses. As a result, users can enjoy a highly entertaining shopping experience in which they receive product information from a fictional character.
[0944] This system allows users to interact with fictional characters in real time in physical stores while obtaining product information, providing a special shopping experience.
[0945] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0946] Step 1:
[0947] The user puts on the smart device and starts the application. The device receives the user's authentication information as input and sends it to the server. The server compares the authentication information with a database and authenticates the user. After successful authentication, the user selects the character they want to talk to from the interface on the device.
[0948] Step 2:
[0949] The terminal sends the ID of the character selected by the user as input to the server. The server retrieves the character's profile data from a database and sends information such as the character's speaking style, personality, and voice quality as output to the terminal. This information is used for subsequent processing.
[0950] Step 3:
[0951] A conversation begins when a user speaks to the device. For example, if the user types, "What is the product on this counter?", the device sends the message to the server. The server uses the received message as input and requests processing by a generative AI model (language generation AI).
[0952] Step 4:
[0953] The server calls the generative AI model and generates a response text for the character based on the user's question. The generative AI model performs data calculations based on the input user message and character setting materials, and outputs the response text, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular."
[0954] Step 5:
[0955] The generated response text is passed as input to the voice generation AI, which performs data calculations to generate an audio file that reproduces the character's voice quality and tone. Specifically, the voice generation AI generates an audio signal from the response text and the character's voice quality data, and outputs it in file format to the server.
[0956] Step 6:
[0957] The server sends the generated audio file to the device, which receives it and plays it back to the user, making the user feel as if they are having a real-time conversation with a fictional character.
[0958] Step 7:
[0959] The user can continue to ask questions or make comments, for example, "That's interesting. Do you have any other product recommendations?", and the process will repeat. The user's message will be sent to the server, which will then pass it through the generative AI model and speech generation AI to generate an appropriate response, which will then be sent to the device.
[0960] In this way, an interactive conversation between the user and the virtual assistant character is realized, providing a highly entertaining shopping experience.
[0961] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0962] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting materials from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. In addition, by combining an emotion engine, the system recognizes the user's emotions and provides more appropriate responses based on those emotions. An embodiment of the present invention is described below.
[0963] Overall system overview
[0964] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, a voice generation AI, and an emotion engine, while the terminal is equipped with a user interface, real-time chat functionality, and voice output functionality. Users can start conversations with characters via their terminal and enjoy the results.
[0965] Overview of program processing
[0966] The system program operates as follows.
[0967] 1. Character Selection
[0968] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[0969] 2. Start a conversation
[0970] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[0971] 3. Emotional Recognition
[0972] The user's input message is also input to the emotion engine, which analyzes the user's emotions. The analyzed emotional information is reflected in the response generation, ensuring that the character responds appropriately.
[0973] 4. Speech Generation
[0974] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[0975] 5. Keep the conversation going
[0976] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[0977] Specific examples
[0978] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[0979] When a user types, "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as, "I went to a cafe with a friend today! It was fun." At the same time, the emotion engine analyzes the emotion in the user's message, and if the message sounds happy, for example, it generates a response based on that emotion. Next, the voice generation AI converts this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[0980] If the user then continues, "That's nice. What did you have at the cafe?", the conversation continues in the same process. The emotion engine analyzes emotions again, providing a more realistic conversational experience.
[0981] This system allows users to enjoy a realistic conversational experience with the fictional characters they choose, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexuals. At the same time, the emotional engine makes the conversation more in tune with the user's emotions.
[0982] The processing flow will be explained below.
[0983] Step 1:
[0984] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[0985] Step 2:
[0986] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[0987] Step 3:
[0988] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[0989] Step 4:
[0990] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[0991] Step 5:
[0992] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[0993] Step 6:
[0994] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[0995] Step 7:
[0996] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[0997] Step 8:
[0998] At the same time, the server inputs the user's input message into the emotion engine to analyze the user's emotion. The emotion engine analyzes the content and context of the message and recognizes what emotion the user is feeling (e.g., happy, sad, surprised, etc.).
[0999] Step 9:
[1000] The server then feeds the analysis results from the emotion engine into the language generation AI, optimizing responses based on the user's emotions, allowing the character to respond in line with the user's emotions.
[1001] Step 10:
[1002] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[1003] Step 11:
[1004] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[1005] Step 12:
[1006] If the user wants to continue the conversation, they enter a new message and send it to the server using the same steps. The server then uses the language generation AI and emotion engine to generate a new reply, and the voice generation AI to generate an audio file, which is then sent to the device. This process is repeated, creating an interactive conversation.
[1007] Example 2
[1008] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1009] In conventional dialogue systems with fictional characters, the characters' responses are fixed and lack emotional change, making it difficult for users to have a natural conversational experience. Furthermore, it is difficult to reproduce the character's unique voice and tone of voice, making it difficult to say that interactions with the character are realistic. This makes it difficult to alleviate the feelings of loneliness and dissatisfaction felt by fictiosexuals in particular. Therefore, there is a need for a system that can analyze emotions in real time based on user input, respond in line with those emotions, and reproduce the character's unique voice quality and tone of voice.
[1010] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1011] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on a user's input text, means for analyzing the user's emotions using an emotion engine and reflecting the analysis in response generation, means for using a voice generation AI to generate an audio file in the character's unique voice from the generated response text, and means for the character and the user to have an interactive conversation on a terminal. This allows the user to enjoy a realistic and natural conversational experience with the character, which is expected to have an effect of alleviating the loneliness and unfulfilled feelings that fictiosexual people may have in particular.
[1012] A "fictional character" is a fictional character or character that appears in creative works such as stories, animations, and games.
[1013] "Setting materials" are data that describe in detail the personality, background, tone of voice, behavior patterns, etc. of a fictional character.
[1014] A "database" is a system for efficiently storing, searching, and managing multiple data.
[1015] "Generative AI" refers to algorithms or software that generate text using artificial intelligence techniques, such as natural language generation (NLG).
[1016] "User input text" is a written message or question that a user enters into the system.
[1017] "Response text" is the response text that the generation AI outputs in response to the user's input text.
[1018] An "emotion engine" is a technology or system for analyzing emotions from a user's input text and generating a response based on those emotions.
[1019] "Speech generation AI" is an artificial intelligence technology for converting generated text into voice data, for example, using text-to-speech (TTS) technology.
[1020] An "audio file" is a file in which audio data is stored in digital format.
[1021] A "terminal" is a device that a user uses to access and operate the system, including, for example, a smartphone or a personal computer.
[1022] "Interactive conversation" refers to a two-way form of communication in which the user and the system send and receive messages to each other and generate responses in real time.
[1023] This invention relates to a system that allows users to enjoy real-time interactive conversations with fictional characters. This system is composed of three elements: a server, a terminal, and a user, and is realized by combining these elements. Below, we will explain the details of each element and the operation of the system.
[1024] Server Features
[1025] The server is equipped with a database that stores setting materials for fictional characters, a generation AI that generates text based on user input, an emotion engine that analyzes user emotions, and a voice generation AI that converts the generated text into voice.
[1026] Database: Stores character information such as character personality, background, speech patterns, and behavior patterns. For example, an SQL database or NoSQL database is used.
[1027] Generative AI: For example, using OpenAI's GPT-3, it generates natural-sounding response text based on user input.
[1028] Emotion engine: Analyzes emotions from the user's input text and reflects the analysis results in the generative AI's responses. Commercial sentiment analysis APIs and proprietary sentiment analysis algorithms are used.
[1029] Voice generation AI: For example, using Google's WaveNet, the generated text is converted into an audio file with a character-specific voice quality.
[1030] Device Features
[1031] The terminals are equipped with a user interface, real-time chat functionality, and voice output functionality, and can be devices such as smartphones or PCs.
[1032] User interface: It provides a character selection screen and a chat window, making it easy for users to operate.
[1033] Real-time chat function: A function that receives user input, sends it to the server, and receives a response from the server. Web sockets and HTTP communication are used.
[1034] Audio output function: Plays received audio files and provides auditory information to the user. Built-in speakers or headphones are used.
[1035] User Actions
[1036] The user accesses the device's interface, selects the character they want to talk to, and enters a message. The server processes the message, generates a response from the character, and sends it to the device. The device then plays back the received response as audio, allowing the user to enjoy a continuous conversation with the character.
[1037] Specific examples
[1038] For example, if a user wants to talk to "Character A," they select "Character A" on the device's interface. Next, the user types, "What did you do today?" and the device sends this message to the server. The server uses a generative AI to generate a response such as, "I went to a cafe with my friends today," while the emotion engine simultaneously analyzes the emotions from the user's message. The result of this analysis is reflected in the generation of a response with an appropriate emotion.
[1039] Next, the server uses a voice generation AI to convert this response into an audio file in the voice of "Character A" and send it to the device. The device plays this audio file, and the user feels as if they are having a conversation with "Character A." Below is an example of a prompt sentence for the generation AI model.
[1040] "User: What did you do today?
[1041] Character: I went to a cafe with a friend today.
[1042] In this way, users can enjoy a realistic and natural interaction experience with fictional characters.
[1043] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1044] Step 1:
[1045] Character Selection
[1046] The user selects the character they want to talk to on the device interface. For example, the user selects "Character A." The input is the user's selection information, and the device sends this information (character ID) to the server. Specific data processing involves sending an HTTP request (for example, "GET / character?id=characterA") that includes the character ID to the server. The output is the character ID sent to the server.
[1047] Step 2:
[1048] Obtaining setting materials
[1049] The server retrieves character settings from the database based on the character ID it receives. The input is the character ID, and the server executes a database query based on this (e.g., "SELECT FROM CharacterSettings WHERE ID='characterA'"). Specific data processing involves returning the settings from the database. The output is the character's settings (information such as personality, background, and tone of voice).
[1050] Step 3:
[1051] User message input
[1052] The user types a message into the terminal's conversation window and presses the send button. For example, they type "What did you do today?". The input is the user's message text, which the terminal sends to the server in JSON format (e.g., "POST / receive_message"). The output is the user's message text sent to the server.
[1053] Step 4:
[1054] Receiving messages and generating replies
[1055] Based on the user's message and character setting information received by the server, a generative AI is used to generate a character's response text. The input is the user's message and character setting information, which are sent to the AI model in the form of a prompt sentence (e.g., "User: What did you do today?\nCharacter A: "). As a specific data calculation, the generative AI model generates the response text. The output is the generated response text (e.g., "I went to a cafe with my friends today.").
[1056] Step 5:
[1057] Emotional analysis and regulation
[1058] The server uses an emotion engine to analyze the user's input message and reflects the results in generating a reply. The input is the user's message text, and the emotion engine analyzes the emotion (e.g., "emotion.analyze('What did you do today?')"). As a specific data calculation, the analysis result is added to the prompt of the generation AI and re-generated (e.g., "User: What did you do today? (Emotion: Happy)\nCharacter A: "). The output is a reply text adjusted based on the emotion.
[1059] Step 6:
[1060] Generate an audio file
[1061] The server inputs the generated response text into the voice generation AI, which generates an audio file in the character's unique voice. The input is the adjusted response text, and the voice generation AI converts this text into audio data (e.g., "voice.generate('I went to a cafe with my friends today.', 'characterA')"). The specific data processing involves generating an audio file. The output is the generated audio file.
[1062] Step 7:
[1063] Sending and playing audio files
[1064] The server sends the generated audio file to the device, and the device plays the received audio file to the user. The input is the generated audio file, which the server sends as an HTTP response (e.g., "POST / play_voice"). The device uses its built-in media player to play the received audio file. The specific operation is that the audio file is played. The output is the audio played to the user.
[1065] Step 8:
[1066] Keeping the conversation going
[1067] The user again inputs a new message and continues the conversation with the character using the same process. The input is the message text entered by the user again. The output is the sending of a new message and the same series of processes being carried out again.
[1068] (Application example 2)
[1069] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1070] In conventional virtual stores, users have to spend a lot of time searching for products and it is difficult to receive appropriate support when selecting products. Furthermore, users cannot receive appropriate advice when selecting products based on emotional aspects, which limits the user experience. Furthermore, there is a lack of ways to make the shopping experience more personalized and provide users with emotional satisfaction.
[1071] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1072] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on text input by the user, means for selecting a character in a virtual store, means for the selected virtual character to support the user's shopping experience, means for analyzing the user's emotions using an emotion engine and reflecting the emotion information in the character's response, means for using a voice generation AI to generate the generated response text into an audio file in a voice unique to the character, and means for the character and the user to have an interactive conversation on a terminal. This enables the user to have an interactive conversation with the fictional character in the virtual store and receive product suggestions based on their emotions.
[1073] A "fictional character" is a fictional person, animal, or being that appears in an imaginary story or entertainment work.
[1074] "Setting materials" refers to data that includes detailed information about a fictional character's personality, background, voice quality, behavioral characteristics, etc.
[1075] "Generative AI" refers to artificial intelligence technology that generates natural language based on user input.
[1076] "Response text" refers to the text that the generative AI creates based on the user's input.
[1077] A "virtual store" is a virtual shopping environment that exists on the Internet.
[1078] An "emotion engine" is an algorithm that analyzes emotional information from user input and makes decisions based on that information.
[1079] "Voice generation AI" is an artificial intelligence technology that converts text data into voice data that resembles a human voice.
[1080] "Audio File" means audio data stored in digital format.
[1081] "Terminal" refers to a device used by a user, such as a computer, smartphone, or head-mounted display.
[1082] "Interactive conversation" means that the user and the system communicate in two directions in real time.
[1083] "Means to support the shopping experience" refers to a method in which a virtual character suggests products based on the user's purchasing motivation and emotions and supports the purchase.
[1084] MODE FOR CARRYING OUT THE INVENTION
[1085] The present invention provides a system that allows users to interact with fictional characters in a virtual store and receive product recommendations based on their emotions. This system is composed of multiple modules and can be specifically implemented as follows:
[1086] System configuration
[1087] Hardware
[1088] 1. Device: The device used by the user, such as a smartphone or head-mounted display
[1089] 2. Server: A cloud-based server containing the database and AI model
[1090] software
[1091] 1. Generation AI: OpenAI GPT-4
[1092] 2. Speech generation AI: Google Cloud Text-to-Speech, Amazon Polly
[1093] 3. Emotion Engine: IBM Watson Tone Analyzer
[1094] 4. Database: MySQL, MongoDB
[1095] Process Overview
[1096] 1. Character Selection
[1097] The user selects a fictional character to act as a shopping assistant in the virtual store. The selected character's ID is sent to the server, which retrieves the character's profile information from a database.
[1098] 2. Start a conversation
[1099] Users can ask questions or ask for advice by text input or voice, and the device sends this message to the server. The server uses a generative AI to generate a response text for the character based on the received message and character settings.
[1100] 3. Emotional Recognition
[1101] The user's message is analyzed by the emotion engine and emotional information is added, so that the character's response will be tailored to the user's emotions.
[1102] 4. Speech Generation
[1103] Based on the generated response text, the server uses a voice generation AI to generate an audio file in the character's unique voice, which is then sent to the device and played back to the user.
[1104] 5. Product proposal
[1105] Based on the user's questions and emotions, the character will suggest related products. For example, if a user inputs "I've been feeling stressed lately and want something to help me relax," the emotion engine will recognize the stress and suggest relaxation goods.
[1106] 6. Keep the conversation going
[1107] Users can continue the conversation by entering new messages, which allows for new product suggestions and responses to inquiries, continuing the interactive dialogue.
[1108] Examples and prompts
[1109] Example 1: A conversation when a tired user is looking for something to relax
[1110] (User): "I've been feeling really tired lately and I need something to help me relax."
[1111] (Character): "That's tough. How about a nice-smelling aroma diffuser? It's perfect for relaxing."
[1112] Example 2: A conversation between a user looking for a gift and asking for advice
[1113] (User): "I'm looking for a birthday gift for my friend. Any recommendations?"
[1114] (Character): "Depending on your friend's preferences, how about some popular fashion items or accessories?"
[1115] Example prompts for generative AI models
[1116] Example prompt 1:
[1117] User text: "I've been feeling really tired lately and I need something to help me relax."
[1118] Character information: {Personality: "Kind", Product knowledge: "Abundant"}
[1119] Produces: A response suggesting products to help a tired user relax.
[1120] Example prompt 2:
[1121] User texts: "I'm looking for a birthday gift for my friend. Any recommendations?"
[1122] Character information: {Personality: "Considerate", Product knowledge: "Abundant"}
[1123] Produces: A response suggesting products that would make good birthday gifts.
[1124] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1125] System program processing flow
[1126] Step 1: Character Selection
[1127] The user selects a fictional character to act as a shopping assistant in a virtual store. The selected character's ID is sent from the device to the server. The input is the selected character's ID, and the output is the character's profile information retrieved from the database. The server retrieves the character's profile information from the database based on this ID and sends it to the device.
[1128] Step 2: Start a conversation
[1129] Users ask questions or ask questions about shopping via text input or voice. This input is sent from the device to the server. The input is text data of the user's question or inquiry, and the output is the character's reply text. The server uses a generation AI to generate the character's reply text based on the received message and character setting materials, and sends it to the device.
[1130] Step 3: Recognize emotions
[1131] The server sends the user's input text to the emotion engine, which analyzes the user's emotions. The input is the user's text data, and the output is the analyzed emotional information. The emotion engine generates this emotional information and reflects it in the response text generated by the generation AI.
[1132] Step 4: Speech generation
[1133] The server passes the generated response text to the voice generation AI, which generates an audio file in the character's unique voice. The input is the response text, and the output is an audio file generated in the character's voice. The voice generation AI converts the text data into audio data, and sends the audio file from the server to the device.
[1134] Step 5: Product proposal
[1135] Based on the user's question and emotional information, a character on the server suggests related products. The input is the user's emotional information and question, and the output is text data of product suggestions. By combining an emotion engine and generative AI, products are suggested based on the user's emotions, and the suggestions are sent to the device.
[1136] Step 6: Keep the conversation going
[1137] The user inputs a new message, which is then sent from the device to the server. Again, a new response and product suggestions are generated through the same process, continuing the interactive dialogue. The input is a new user message, and the output is a new response text, audio file, and product suggestions.
[1138] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1139] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1140] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1141] [Fourth embodiment]
[1142] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1143] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1144] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1145] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1146] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1147] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1148] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1149] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1150] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1151] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1152] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1153] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1154] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1155] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting data from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. Specific embodiments of the system are described below.
[1156] Overall system overview
[1157] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, and a voice generation AI, while the terminal is equipped with a user interface, real-time chat function, and voice output function. Users can start conversations with characters through their terminal and enjoy the results.
[1158] Overview of program processing
[1159] The system program operates as follows.
[1160] 1. Character Selection
[1161] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[1162] 2. Start a conversation
[1163] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[1164] 3. Speech generation
[1165] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[1166] 4. Keep the conversation going
[1167] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[1168] Specific examples
[1169] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[1170] When a user types "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as "I went to a cafe with my friends today! It was fun." The server then uses speech generation AI to convert this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[1171] If the user then continues with, "That's nice. What did you have at the cafe?" the conversation continues using the same process.
[1172] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexual people.
[1173] The processing flow will be explained below.
[1174] Step 1:
[1175] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[1176] Step 2:
[1177] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[1178] Step 3:
[1179] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[1180] Step 4:
[1181] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[1182] Step 5:
[1183] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[1184] Step 6:
[1185] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[1186] Step 7:
[1187] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[1188] Step 8:
[1189] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[1190] Step 9:
[1191] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[1192] Step 10:
[1193] If the user wants to continue the conversation, they can enter a new message and send it to the server using the same steps. The server then uses the language generation and speech generation AI to generate a new reply and sends it to the device. This process is repeated, creating an interactive conversation.
[1194] Example 1
[1195] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1196] There is a demand for a system that allows users to enjoy real-time interactive conversations with fictional characters. Conventional systems have made it difficult for users to obtain a satisfying experience because the characters' responses are monotonous and the voice generation is unnatural. Furthermore, the ability to select from multiple characters is limited, making it impossible to meet the diverse needs of users.
[1197] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1198] In this invention, the server includes means for retrieving setting materials for a fictional character selected by a user from a database, means for generating a response text for the fictional character based on a text input by the user using a generation AI, and means for generating an audio file using a voice generation AI from the generated response text in a voice unique to the fictional character, thereby enabling the user to enjoy natural and immersive interactive conversations with the fictional character.
[1199] "Users" are people who use the system to converse with fictional characters.
[1200] "Fictional characters" refers to fictional characters and animals that appear in creative works such as animation, manga, novels, and games.
[1201] "Setting materials" refers to information that describes a fictional character in detail, such as their personality, age, occupation, hobbies, etc.
[1202] A "database" is a collection of information that stores and manages setting materials for fictional characters and can be retrieved as needed.
[1203] "Generative AI" refers to artificial intelligence technology that generates response text from fictional characters based on user input text.
[1204] "Response text" refers to the reply content of a fictional character created by the generation AI in response to user input.
[1205] "Voice generation AI" is an artificial intelligence technology that converts response text created by the generation AI into an audio file that reproduces the tone and voice quality of a fictional character.
[1206] "Audio file" refers to digital data generated by voice generation AI to play back responses in the voices of fictional characters.
[1207] A "terminal" is a device or software that a user operates to interact with fictional characters.
[1208] "Interactive conversation" means that the user and the fictional character have two-way communication in real time via the system.
[1209] The present invention provides a system for enabling users to have real-time interactive conversations with fictional characters, which system is comprised of a server, a terminal, and a user.
[1210] Overall system configuration
[1211] The server includes the following components:
[1212] Database: Accumulates and manages setting materials for fictional characters.
[1213] Generative AI: Generates response text from a fictional character based on the user's input text.
[1214] Voice generation AI: Converts generated response text into an audio file with a voice unique to a fictional character.
[1215] The terminal is a device operated by the user and has the following functions:
[1216] User Interface: An interface that allows the user to select a character, input messages, and send them.
[1217] Real-time chat function: A function that communicates with the server and sends and receives messages in real time.
[1218] Audio output function: A function to play audio files received from the server.
[1219] The user operates the terminal to enjoy conversations with fictional characters.
[1220] System Operation Overview
[1221] First, the user operates the device's interface and selects the fictional character they want to talk to. The device then sends the character's ID to the server, which then retrieves the character's profile information from a database.
[1222] Next, the user enters a message to start the conversation and presses the send button. The device then sends the entered message to the server, which uses a generation AI to generate a reply text based on the received message and character settings.
[1223] Based on the generated response text, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[1224] The user can enter a new message and continue the conversation in the same process. For each message sent from the terminal, the server generates an appropriate reply and outputs it as voice.
[1225] Specific examples
[1226] For example, if a user wants to have a conversation with a specific fictional character, the user selects the character on the device interface and sends the selection information to the server. The server retrieves character information from the database and starts a conversation based on that information.
[1227] When a user types "What did you do today?", the device sends this message to the server. The server uses a generation AI to generate a response such as "I went to a cafe with a friend today! It was fun." The server then uses a voice generation AI to convert this response into an audio file in the character's voice and sends it to the device. The device plays this audio file, allowing the user to feel as if they are having a conversation with the character.
[1228] If the user then types, "That's nice. What did you have at the cafe?" the conversation continues using a similar process.
[1229] Example prompt sentence:
[1230] "Based on the character's background material, generate a response to the following message: 'What did you do today?'"
[1231] "Say in the voice of a specific fictional character, 'I went to a cafe with my friends today. It was fun.'"
[1232] This system allows users to enjoy a realistic interaction experience with a fictional character of their choice.
[1233] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1234] Step 1: Choose your character
[1235] Input: The user selects a character in the device interface.
[1236] Operation:
[1237] The terminal acquires the ID of the character selected by the user.
[1238] The ID of the selected character is sent to the server.
[1239] Output: The character's ID is sent to the server.
[1240] Step 2: Obtaining the configuration data
[1241] Input: The server receives the character's ID.
[1242] Operation:
[1243] The server searches the database for the character's settings information.
[1244] The setting materials include information such as the character's personality, age, occupation, hobbies, etc.
[1245] Obtain the setting materials.
[1246] Output: The configuration data is available on the server.
[1247] Step 3: Start a conversation
[1248] Input: The user types a message into the terminal.
[1249] Operation:
[1250] The user inputs text into the message input field on the terminal and presses the send button.
[1251] The terminal transmits the input text data to the server.
[1252] Output: The user's input text is sent to the server.
[1253] Step 4: Generate response text
[1254] Input: The server receives the user's input text and character profile information.
[1255] Operation:
[1256] The server passes prompts to the generation AI based on the user's message and character settings.
[1257] The generative AI generates response text that matches the character's tone and personality.
[1258] Output: The generated response text.
[1259] Step 5: Generate the audio file
[1260] Input: The server receives the generated response text.
[1261] Operation:
[1262] The server inputs the generated response text into the speech generation AI.
[1263] The voice generation AI generates audio files that reproduce the character's voice quality and speaking style.
[1264] Output: The generated audio file.
[1265] Step 6: Send and play the audio file
[1266] Input: The server has the generated audio file.
[1267] Operation:
[1268] The server transmits the generated audio file to the terminal.
[1269] The terminal receives the audio file and plays it to the user using its audio output capabilities.
[1270] Output: The reply is played to the user as audio.
[1271] Step 7: Keep the conversation going
[1272] Input: The user types a new message.
[1273] Operation:
[1274] The user types another question or comment into the terminal and presses the send button.
[1275] The entered text data is sent to the server again.
[1276] The server repeats the same process, using the generation AI and voice generation AI to generate new response text and audio files.
[1277] The terminal plays the received audio file.
[1278] Output: The conversation continues, generating new messages and replies.
[1279] (Application example 1)
[1280] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1281] In traditional brick-and-mortar stores, users lacked an effective and enjoyable way to obtain product information. Furthermore, there was no technology available to provide a highly entertaining shopping experience through interactions with fictional characters. This made it difficult to improve the user experience.
[1282] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1283] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for generating a response text for the character based on a user's input text using a generation AI, means for generating an audio file in the character's unique voice using a voice generation AI from the generated response text, means for the character and the user to have an interactive conversation on a terminal, and means for the character to provide product information and services as a virtual assistant using a smart device in a physical store. This allows users to obtain product information while conversing with the fictional character in real time in the physical store, enabling them to enjoy a highly entertaining shopping experience.
[1284] A "fictional character" is a person or creature that appears in a fictional world, and whose setting is depicted in detail through media such as stories and games.
[1285] A "database" is an electronic storage device or system that systematically organizes information or data so that it can be efficiently accessed and managed.
[1286] "Generative AI" is an artificial intelligence technology that uses natural language processing and data analysis techniques to generate text and data based on user input.
[1287] "Voice generation AI" is a technology that converts text data into voice data, and is an artificial intelligence technology that can reproduce specific voice qualities and speaking styles.
[1288] A "terminal" is an electronic device that can be directly operated by a user, such as a computer, tablet, or smartphone.
[1289] "Interactive conversational means" refers to methods or functions that allow users and systems to communicate with each other in real time.
[1290] A "brick and mortar store" is a store with a physical sales floor that exists in the real world, as opposed to an online store.
[1291] A "smart device" is an electronic device with advanced functionality that can connect to the Internet and run various applications, including smartphones, smart glasses, and head-mounted displays.
[1292] A "virtual assistant" is a digital assistant that uses artificial intelligence technology to answer users' questions and provide services.
[1293] "Product Guide" is a means of providing information about products in a store and helping users consider purchasing them.
[1294] "User experience" refers to the experience and satisfaction that users feel when using a product or service.
[1295] The embodiment of the present invention is to build a system that is basically composed of a server, a terminal, and a user. An overview of the entire system and each of its main steps will be explained below.
[1296] Overall system overview
[1297] The system retrieves fictional character settings from a database, uses generative and voice-generating AI to generate responses for the characters, and outputs them as voice. The system is designed to improve the shopping experience in brick-and-mortar stores.
[1298] Hardware and software used
[1299] Hardware: Servers, smart glasses (e.g., Google Glass), head-mounted displays (e.g., Microsoft HoloLens)
[1300] Software: Language generation AI (e.g. GPT-4), speech generation AI (Google Text-to-Speech, Amazon Polly), user interface (developed with Unity or Unreal Engine)
[1301] System operation example
[1302] 1. User authentication and character selection
[1303] The user puts on the smart device and starts the application. After logging in, the user selects the character they want to talk to. For example, the user selects the character "Elizabeth." This selection information is sent to the server.
[1304] Example prompt sentence:
[1305] Character: Elizabeth
[1306] User Question: What is the product on this counter?
[1307] Preferred response language: Japanese
[1308] Information to include in the response: product type, features, special offers, etc.
[1309] As Elizabeth: "This counter is lined with the latest technology gadgets. Our most popular item is the smart wristband. It has health tracking features and is currently on sale at a special price."
[1310] 2. Obtaining character setting data
[1311] The server retrieves the character's profile information from a database, including the character's speaking style, personality, and voice quality.
[1312] 3. Conversation initiation and response generation
[1313] When a user speaks to the character, the message is sent to the server, which uses language generation AI to generate an appropriate response to the user's input. For example, if a user asks, "What products are on this counter?", the server will generate a response such as, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular and have health management functions."
[1314] 4. Speech generation and output
[1315] Based on the generated response text, the AI creates an audio file that reproduces the character's voice quality and tone. This audio file is then sent to the user's smart device and played back, making the user feel as if they are having a real-time conversation with a fictional character.
[1316] Specific examples
[1317] For example, imagine a user puts on smart glasses in a physical store, selects a character named "Elizabeth," and asks, "What product is on this counter?" This information is sent to the server, which generates a response based on the character's configuration data. Next, the voice generation AI converts the response into an audio file in the character's voice quality, which is played on the user's smart glasses. As a result, users can enjoy a highly entertaining shopping experience in which they receive product information from a fictional character.
[1318] This system allows users to interact with fictional characters in real time in physical stores while obtaining product information, providing a special shopping experience.
[1319] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1320] Step 1:
[1321] The user puts on the smart device and starts the application. The device receives the user's authentication information as input and sends it to the server. The server compares the authentication information with a database and authenticates the user. After successful authentication, the user selects the character they want to talk to from the interface on the device.
[1322] Step 2:
[1323] The terminal sends the ID of the character selected by the user as input to the server. The server retrieves the character's profile data from a database and sends information such as the character's speaking style, personality, and voice quality as output to the terminal. This information is used for subsequent processing.
[1324] Step 3:
[1325] A conversation begins when a user speaks to the device. For example, if the user types, "What is the product on this counter?", the device sends the message to the server. The server uses the received message as input and requests processing by a generative AI model (language generation AI).
[1326] Step 4:
[1327] The server calls the generative AI model and generates a response text for the character based on the user's question. The generative AI model performs data calculations based on the input user message and character setting materials, and outputs the response text, "This counter is lined with the latest technology gadgets. Smart wristbands are particularly popular."
[1328] Step 5:
[1329] The generated response text is passed as input to the voice generation AI, which performs data calculations to generate an audio file that reproduces the character's voice quality and tone. Specifically, the voice generation AI generates an audio signal from the response text and the character's voice quality data, and outputs it in file format to the server.
[1330] Step 6:
[1331] The server sends the generated audio file to the device, which receives it and plays it back to the user, making the user feel as if they are having a real-time conversation with a fictional character.
[1332] Step 7:
[1333] The user can continue to ask questions or make comments, for example, "That's interesting. Do you have any other product recommendations?", and the process will repeat. The user's message will be sent to the server, which will then pass it through the generative AI model and speech generation AI to generate an appropriate response, which will then be sent to the device.
[1334] In this way, an interactive conversation between the user and the virtual assistant character is realized, providing a highly entertaining shopping experience.
[1335] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1336] The present invention relates to a system that allows users to converse and interact with fictional characters in real time. This system retrieves the character's setting materials from a database, generates the character's responses using a generation AI and a voice generation AI, and outputs them as voice. In addition, by combining an emotion engine, the system recognizes the user's emotions and provides more appropriate responses based on those emotions. An embodiment of the present invention is described below.
[1337] Overall system overview
[1338] The system consists of a server, a terminal, and a user. The server has a character database, a language generation AI, a voice generation AI, and an emotion engine, while the terminal is equipped with a user interface, real-time chat functionality, and voice output functionality. Users can start conversations with characters via their terminal and enjoy the results.
[1339] Overview of program processing
[1340] The system program operates as follows.
[1341] 1. Character Selection
[1342] The user first selects the character they want to talk to from the device interface, and the device sends the character's ID to the server, which then retrieves the character's profile information from a database based on that ID.
[1343] 2. Start a conversation
[1344] The user enters a message into the device and presses the send button. The device then sends the message to the server. The server uses language generation AI to generate a response from the character based on the received message and character settings.
[1345] 3. Emotional Recognition
[1346] The user's input message is also input to the emotion engine, which analyzes the user's emotions. The analyzed emotional information is reflected in the response generation, ensuring that the character responds appropriately.
[1347] 4. Speech Generation
[1348] Based on the generated responses, the server uses voice generation AI to generate an audio file that reproduces the character's voice quality and tone, which is then sent to the device and played back to the user.
[1349] 5. Keep the conversation going
[1350] The user can enter a new message and continue the conversation using the same process. For each message sent from the device, the server generates an appropriate reply and outputs it as voice.
[1351] Specific examples
[1352] For example, suppose a user wants to talk to an anime character called "Haruka Sakurai." In this case, the user selects "Haruka Sakurai" on the device interface and sends the selection information to the server. The server retrieves information about "Haruka Sakurai" from the database and starts a conversation based on that information.
[1353] When a user types, "Haruka, what did you do today?", the device sends this message to the server. The server uses language generation AI to generate a response such as, "I went to a cafe with a friend today! It was fun." At the same time, the emotion engine analyzes the emotion in the user's message, and if the message sounds happy, for example, it generates a response based on that emotion. Next, the voice generation AI converts this response into an audio file in the voice of "Haruka Sakurai" and sends it to the device. The device plays this audio file, making the user feel as if they are having a conversation with "Haruka Sakurai."
[1354] If the user then continues, "That's nice. What did you have at the cafe?", the conversation continues in the same process. The emotion engine analyzes emotions again, providing a more realistic conversational experience.
[1355] This system allows users to enjoy a realistic conversational experience with the fictional characters they choose, helping to alleviate feelings of loneliness and dissatisfaction, especially for fictiosexuals. At the same time, the emotional engine makes the conversation more in tune with the user's emotions.
[1356] The processing flow will be explained below.
[1357] Step 1:
[1358] The user opens the device interface and is presented with a menu from which they can select the character they wish to talk to. The user selects a character from the menu and enters their selection into the device.
[1359] Step 2:
[1360] The device obtains the ID of the selected character and generates an API request to send it to the server. The API request contains the ID of the selected character.
[1361] Step 3:
[1362] The server parses the received API request, checks the character's ID, and then sends a query to retrieve the character's profile information from the database.
[1363] Step 4:
[1364] The database responds to the server's queries and returns the character's character information, such as personality, speech pattern, and profile, to the server. The server receives this information and processes it into the character's character information.
[1365] Step 5:
[1366] The server generates a profile that reflects the character's characteristics based on the acquired setting information and sends that information back to the device. The device receives the response from the server and displays the character information on the user interface.
[1367] Step 6:
[1368] The user enters a message (e.g., "What did you do today?") into the chat input field on the device and presses the send button. The device takes this message and generates an API request to send it to the server.
[1369] Step 7:
[1370] The server receives the user's input message and inputs it, along with the character's background information, into the language generation AI, which generates an appropriate response (e.g., "I went to a cafe with my friends today") based on the character's personality and tone of voice.
[1371] Step 8:
[1372] At the same time, the server inputs the user's input message into the emotion engine to analyze the user's emotion. The emotion engine analyzes the content and context of the message and recognizes what emotion the user is feeling (e.g., happy, sad, surprised, etc.).
[1373] Step 9:
[1374] The server then feeds the analysis results from the emotion engine into the language generation AI, optimizing responses based on the user's emotions, allowing the character to respond in line with the user's emotions.
[1375] Step 10:
[1376] The server receives the generated response text and inputs it into the voice generation AI, which then generates an audio file that reproduces the character's voice quality and tone.
[1377] Step 11:
[1378] The generated voice file is sent from the server to the device, which receives it and plays it back, allowing the user to hear the character's response.
[1379] Step 12:
[1380] If the user wants to continue the conversation, they enter a new message and send it to the server using the same steps. The server then uses the language generation AI and emotion engine to generate a new reply, and the voice generation AI to generate an audio file, which is then sent to the device. This process is repeated, creating an interactive conversation.
[1381] Example 2
[1382] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1383] In conventional dialogue systems with fictional characters, the characters' responses are fixed and lack emotional change, making it difficult for users to have a natural conversational experience. Furthermore, it is difficult to reproduce the character's unique voice and tone of voice, making it difficult to say that interactions with the character are realistic. This makes it difficult to alleviate the feelings of loneliness and dissatisfaction felt by fictiosexuals in particular. Therefore, there is a need for a system that can analyze emotions in real time based on user input, respond in line with those emotions, and reproduce the character's unique voice quality and tone of voice.
[1384] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1385] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on a user's input text, means for analyzing the user's emotions using an emotion engine and reflecting the analysis in response generation, means for using a voice generation AI to generate an audio file in the character's unique voice from the generated response text, and means for the character and the user to have an interactive conversation on a terminal. This allows the user to enjoy a realistic and natural conversational experience with the character, which is expected to have an effect of alleviating the loneliness and unfulfilled feelings that fictiosexual people may have in particular.
[1386] A "fictional character" is a fictional character or character that appears in creative works such as stories, animations, and games.
[1387] "Setting materials" are data that describe in detail the personality, background, tone of voice, behavior patterns, etc. of a fictional character.
[1388] A "database" is a system for efficiently storing, searching, and managing multiple data.
[1389] "Generative AI" refers to algorithms or software that generate text using artificial intelligence techniques, such as natural language generation (NLG).
[1390] "User input text" is a written message or question that a user enters into the system.
[1391] "Response text" is the response text that the generation AI outputs in response to the user's input text.
[1392] An "emotion engine" is a technology or system for analyzing emotions from a user's input text and generating a response based on those emotions.
[1393] "Speech generation AI" is an artificial intelligence technology for converting generated text into voice data, for example, using text-to-speech (TTS) technology.
[1394] An "audio file" is a file in which audio data is stored in digital format.
[1395] A "terminal" is a device that a user uses to access and operate the system, including, for example, a smartphone or a personal computer.
[1396] "Interactive conversation" refers to a two-way form of communication in which the user and the system send and receive messages to each other and generate responses in real time.
[1397] This invention relates to a system that allows users to enjoy real-time interactive conversations with fictional characters. This system is composed of three elements: a server, a terminal, and a user, and is realized by combining these elements. Below, we will explain the details of each element and the operation of the system.
[1398] Server Features
[1399] The server is equipped with a database that stores setting materials for fictional characters, a generation AI that generates text based on user input, an emotion engine that analyzes user emotions, and a voice generation AI that converts the generated text into voice.
[1400] Database: Stores character information such as character personality, background, speech patterns, and behavior patterns. For example, an SQL database or NoSQL database is used.
[1401] Generative AI: For example, using OpenAI's GPT-3, it generates natural-sounding response text based on user input.
[1402] Emotion engine: Analyzes emotions from the user's input text and reflects the analysis results in the generative AI's responses. Commercial sentiment analysis APIs and proprietary sentiment analysis algorithms are used.
[1403] Voice generation AI: For example, using Google's WaveNet, the generated text is converted into an audio file with a character-specific voice quality.
[1404] Device Features
[1405] The terminals are equipped with a user interface, real-time chat functionality, and voice output functionality, and can be devices such as smartphones or PCs.
[1406] User interface: It provides a character selection screen and a chat window, making it easy for users to operate.
[1407] Real-time chat function: A function that receives user input, sends it to the server, and receives a response from the server. Web sockets and HTTP communication are used.
[1408] Audio output function: Plays received audio files and provides auditory information to the user. Built-in speakers or headphones are used.
[1409] User Actions
[1410] The user accesses the device's interface, selects the character they want to talk to, and enters a message. The server processes the message, generates a response from the character, and sends it to the device. The device then plays back the received response as audio, allowing the user to enjoy a continuous conversation with the character.
[1411] Specific examples
[1412] For example, if a user wants to talk to "Character A," they select "Character A" on the device's interface. Next, the user types, "What did you do today?" and the device sends this message to the server. The server uses a generative AI to generate a response such as, "I went to a cafe with my friends today," while the emotion engine simultaneously analyzes the emotions from the user's message. The result of this analysis is reflected in the generation of a response with an appropriate emotion.
[1413] Next, the server uses a voice generation AI to convert this response into an audio file in the voice of "Character A" and send it to the device. The device plays this audio file, and the user feels as if they are having a conversation with "Character A." Below is an example of a prompt sentence for the generation AI model.
[1414] "User: What did you do today?
[1415] Character: I went to a cafe with a friend today.
[1416] In this way, users can enjoy a realistic and natural interaction experience with fictional characters.
[1417] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1418] Step 1:
[1419] Character Selection
[1420] The user selects the character they want to talk to on the device interface. For example, the user selects "Character A." The input is the user's selection information, and the device sends this information (character ID) to the server. Specific data processing involves sending an HTTP request (for example, "GET / character?id=characterA") that includes the character ID to the server. The output is the character ID sent to the server.
[1421] Step 2:
[1422] Obtaining setting materials
[1423] The server retrieves character settings from the database based on the character ID it receives. The input is the character ID, and the server executes a database query based on this (e.g., "SELECT FROM CharacterSettings WHERE ID='characterA'"). Specific data processing involves returning the settings from the database. The output is the character's settings (information such as personality, background, and tone of voice).
[1424] Step 3:
[1425] User message input
[1426] The user types a message into the terminal's conversation window and presses the send button. For example, they type "What did you do today?". The input is the user's message text, which the terminal sends to the server in JSON format (e.g., "POST / receive_message"). The output is the user's message text sent to the server.
[1427] Step 4:
[1428] Receiving messages and generating replies
[1429] Based on the user's message and character setting information received by the server, a generative AI is used to generate a character's response text. The input is the user's message and character setting information, which are sent to the AI model in the form of a prompt sentence (e.g., "User: What did you do today?\nCharacter A: "). As a specific data calculation, the generative AI model generates the response text. The output is the generated response text (e.g., "I went to a cafe with my friends today.").
[1430] Step 5:
[1431] Emotional analysis and regulation
[1432] The server uses an emotion engine to analyze the user's input message and reflects the results in generating a reply. The input is the user's message text, and the emotion engine analyzes the emotion (e.g., "emotion.analyze('What did you do today?')"). As a specific data calculation, the analysis result is added to the prompt of the generation AI and re-generated (e.g., "User: What did you do today? (Emotion: Happy)\nCharacter A: "). The output is a reply text adjusted based on the emotion.
[1433] Step 6:
[1434] Generate an audio file
[1435] The server inputs the generated response text into the voice generation AI, which generates an audio file in the character's unique voice. The input is the adjusted response text, and the voice generation AI converts this text into audio data (e.g., "voice.generate('I went to a cafe with my friends today.', 'characterA')"). The specific data processing involves generating an audio file. The output is the generated audio file.
[1436] Step 7:
[1437] Sending and playing audio files
[1438] The server sends the generated audio file to the device, and the device plays the received audio file to the user. The input is the generated audio file, which the server sends as an HTTP response (e.g., "POST / play_voice"). The device uses its built-in media player to play the received audio file. The specific operation is that the audio file is played. The output is the audio played to the user.
[1439] Step 8:
[1440] Keeping the conversation going
[1441] The user again inputs a new message and continues the conversation with the character using the same process. The input is the message text entered by the user again. The output is the sending of a new message and the same series of processes being carried out again.
[1442] (Application example 2)
[1443] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1444] In conventional virtual stores, users have to spend a lot of time searching for products and it is difficult to receive appropriate support when selecting products. Furthermore, users cannot receive appropriate advice when selecting products based on emotional aspects, which limits the user experience. Furthermore, there is a lack of ways to make the shopping experience more personalized and provide users with emotional satisfaction.
[1445] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1446] In this invention, the server includes means for retrieving setting materials for a fictional character from a database, means for using a generation AI to generate a response text for the character based on text input by the user, means for selecting a character in a virtual store, means for the selected virtual character to support the user's shopping experience, means for analyzing the user's emotions using an emotion engine and reflecting the emotion information in the character's response, means for using a voice generation AI to generate the generated response text into an audio file in a voice unique to the character, and means for the character and the user to have an interactive conversation on a terminal. This enables the user to have an interactive conversation with the fictional character in the virtual store and receive product suggestions based on their emotions.
[1447] A "fictional character" is a fictional person, animal, or being that appears in an imaginary story or entertainment work.
[1448] "Setting materials" refers to data that includes detailed information about a fictional character's personality, background, voice quality, behavioral characteristics, etc.
[1449] "Generative AI" refers to artificial intelligence technology that generates natural language based on user input.
[1450] "Response text" refers to the text that the generative AI creates based on the user's input.
[1451] A "virtual store" is a virtual shopping environment that exists on the Internet.
[1452] An "emotion engine" is an algorithm that analyzes emotional information from user input and makes decisions based on that information.
[1453] "Voice generation AI" is an artificial intelligence technology that converts text data into voice data that resembles a human voice.
[1454] "Audio File" means audio data stored in digital format.
[1455] "Terminal" refers to a device used by a user, such as a computer, smartphone, or head-mounted display.
[1456] "Interactive conversation" means that the user and the system communicate in two directions in real time.
[1457] "Means to support the shopping experience" refers to a method in which a virtual character suggests products based on the user's purchasing motivation and emotions and supports the purchase.
[1458] MODE FOR CARRYING OUT THE INVENTION
[1459] The present invention provides a system that allows users to interact with fictional characters in a virtual store and receive product recommendations based on their emotions. This system is composed of multiple modules and can be specifically implemented as follows:
[1460] System configuration
[1461] Hardware
[1462] 1. Device: The device used by the user, such as a smartphone or head-mounted display
[1463] 2. Server: A cloud-based server containing the database and AI model
[1464] software
[1465] 1. Generation AI: OpenAI GPT-4
[1466] 2. Speech generation AI: Google Cloud Text-to-Speech, Amazon Polly
[1467] 3. Emotion Engine: IBM Watson Tone Analyzer
[1468] 4. Database: MySQL, MongoDB
[1469] Process Overview
[1470] 1. Character Selection
[1471] The user selects a fictional character to act as a shopping assistant in the virtual store. The selected character's ID is sent to the server, which retrieves the character's profile information from a database.
[1472] 2. Start a conversation
[1473] Users can ask questions or ask for advice by text input or voice, and the device sends this message to the server. The server uses a generative AI to generate a response text for the character based on the received message and character settings.
[1474] 3. Emotional Recognition
[1475] The user's message is analyzed by the emotion engine and emotional information is added, so that the character's response will be tailored to the user's emotions.
[1476] 4. Speech Generation
[1477] Based on the generated response text, the server uses a voice generation AI to generate an audio file in the character's unique voice, which is then sent to the device and played back to the user.
[1478] 5. Product proposal
[1479] Based on the user's questions and emotions, the character will suggest related products. For example, if a user inputs "I've been feeling stressed lately and want something to help me relax," the emotion engine will recognize the stress and suggest relaxation goods.
[1480] 6. Keep the conversation going
[1481] Users can continue the conversation by entering new messages, which allows for new product suggestions and responses to inquiries, continuing the interactive dialogue.
[1482] Examples and prompts
[1483] Example 1: A conversation when a tired user is looking for something to relax
[1484] (User): "I've been feeling really tired lately and I need something to help me relax."
[1485] (Character): "That's tough. How about a nice-smelling aroma diffuser? It's perfect for relaxing."
[1486] Example 2: A conversation between a user looking for a gift and asking for advice
[1487] (User): "I'm looking for a birthday gift for my friend. Any recommendations?"
[1488] (Character): "Depending on your friend's preferences, how about some popular fashion items or accessories?"
[1489] Example prompts for generative AI models
[1490] Example prompt 1:
[1491] User text: "I've been feeling really tired lately and I need something to help me relax."
[1492] Character information: {Personality: "Kind", Product knowledge: "Abundant"}
[1493] Produces: A response suggesting products to help a tired user relax.
[1494] Example prompt 2:
[1495] User texts: "I'm looking for a birthday gift for my friend. Any recommendations?"
[1496] Character information: {Personality: "Considerate", Product knowledge: "Abundant"}
[1497] Produces: A response suggesting products that would make good birthday gifts.
[1498] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1499] System program processing flow
[1500] Step 1: Character Selection
[1501] The user selects a fictional character to act as a shopping assistant in a virtual store. The selected character's ID is sent from the device to the server. The input is the selected character's ID, and the output is the character's profile information retrieved from the database. The server retrieves the character's profile information from the database based on this ID and sends it to the device.
[1502] Step 2: Start a conversation
[1503] Users ask questions or ask questions about shopping via text input or voice. This input is sent from the device to the server. The input is text data of the user's question or inquiry, and the output is the character's reply text. The server uses a generation AI to generate the character's reply text based on the received message and character setting materials, and sends it to the device.
[1504] Step 3: Recognize emotions
[1505] The server sends the user's input text to the emotion engine, which analyzes the user's emotions. The input is the user's text data, and the output is the analyzed emotional information. The emotion engine generates this emotional information and reflects it in the response text generated by the generation AI.
[1506] Step 4: Speech generation
[1507] The server passes the generated response text to the voice generation AI, which generates an audio file in the character's unique voice. The input is the response text, and the output is an audio file generated in the character's voice. The voice generation AI converts the text data into audio data, and sends the audio file from the server to the device.
[1508] Step 5: Product proposal
[1509] Based on the user's question and emotional information, a character on the server suggests related products. The input is the user's emotional information and question, and the output is text data of product suggestions. By combining an emotion engine and generative AI, products are suggested based on the user's emotions, and the suggestions are sent to the device.
[1510] Step 6: Keep the conversation going
[1511] The user inputs a new message, which is then sent from the device to the server. Again, a new response and product suggestions are generated through the same process, continuing the interactive dialogue. The input is a new user message, and the output is a new response text, audio file, and product suggestions.
[1512] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1513] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1514] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1515] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1516] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1517] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1518] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1519] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1520] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1521] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1522] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1523] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1524] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1525] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1526] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1527] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1528] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1529] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1530] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1531] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1532] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1533] The following is further disclosed regarding the above embodiment.
[1534] (Claim 1)
[1535] A means for retrieving fictional character setting materials from a database;
[1536] A means for generating a response text for a character based on a user's input text using a generation AI;
[1537] A means for generating a voice file from the response text generated using voice generation AI in a voice unique to the character;
[1538] A means for the character and the user to have an interactive conversation on the terminal;
[1539] A system including:
[1540] (Claim 2)
[1541] The system of claim 1, wherein the voice generation AI converts the response text generated based on the character's setting materials into an audio file that reproduces the character's tone and voice quality.
[1542] (Claim 3)
[1543] 2. The system according to claim 1, wherein the terminal is provided with a means for the user to select from a plurality of characters, and responses and voices are generated according to the selected character.
[1544] "Example 1"
[1545] (Claim 1)
[1546] A means for retrieving setting materials for a fictional character selected by a user from a database;
[1547] a means for generating a response text of a fictional character based on the user's input text using a generation AI;
[1548] A means for generating the response text generated using a voice generation AI into an audio file in a voice unique to a fictional character;
[1549] means for allowing a user to have an interactive conversation with a fictional character at the terminal;
[1550] A system including:
[1551] (Claim 2)
[1552] The system of claim 1, wherein the speech generation AI converts response text generated based on setting materials for a fictional character into an audio file that reproduces the tone and voice quality of the fictional character.
[1553] (Claim 3)
[1554] 2. The system according to claim 1, wherein the terminal comprises a means for allowing a user to select from a plurality of fictional characters, and the system generates responses and voices according to the selected fictional character.
[1555] "Application Example 1"
[1556] (Claim 1)
[1557] A means for retrieving fictional character setting materials from a database;
[1558] A means for generating a response text for a character based on a user's input text using a generation AI;
[1559] A means for generating a voice file from the response text generated using voice generation AI in a voice unique to the character;
[1560] A means for the character and the user to have an interactive conversation on the terminal;
[1561] A method in which characters act as virtual assistants in real stores using smart devices to provide product information and services, and
[1562] A system including:
[1563] (Claim 2)
[1564] The system of claim 1, wherein the voice generation AI converts the response text generated based on the character's setting materials into an audio file that reproduces the character's tone and voice quality.
[1565] (Claim 3)
[1566] 2. The system according to claim 1, wherein the terminal is provided with a means for the user to select from a plurality of characters, and responses and voices are generated according to the selected character.
[1567] "Example 2: Combining Emotion Engines"
[1568] (Claim 1)
[1569] A means for retrieving fictional character setting materials from a database;
[1570] A means for generating a response text for a character based on a user's input text using a generation AI;
[1571] A means for analyzing the user's emotions using an emotion engine and reflecting the emotions in response generation;
[1572] A means for generating a voice file from the response text generated using voice generation AI in a voice unique to the character;
[1573] A means for the character and the user to have an interactive conversation on the terminal;
[1574] A system including:
[1575] (Claim 2)
[1576] The system of claim 1, wherein the voice generation AI converts the response text generated based on the character's setting materials into an audio file that reproduces the character's tone and voice quality.
[1577] (Claim 3)
[1578] 2. The system according to claim 1, wherein the terminal is provided with a means for the user to select from a plurality of characters, and responses and voices are generated according to the selected character.
[1579] "Application example 2 when combining emotion engines"
[1580] (Claim 1)
[1581] A means for retrieving fictional character setting materials from a database;
[1582] A means for generating a response text for a character based on a user's input text using a generation AI;
[1583] A means for selecting a character within the virtual store;
[1584] a means for the selected virtual character to assist the user in their shopping experience;
[1585] A means of analyzing the user's emotions using an emotion engine and reflecting the emotional information in the character's responses;
[1586] A means for generating a voice file from the response text generated using voice generation AI in a voice unique to the character;
[1587] A means for the character and the user to have an interactive conversation on the terminal;
[1588] A system including:
[1589] (Claim 2)
[1590] The system of claim 1, wherein the voice generation AI converts the response text generated based on the character's setting materials into an audio file that reproduces the character's tone and voice quality.
[1591] (Claim 3)
[1592] The system of claim 1 is provided with a terminal that allows the user to select from multiple characters in a virtual store, and generates responses and voices according to the selected character, as well as suggests products that reflect the user's emotions. [Explanation of symbols]
[1593] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for retrieving fictional character setting materials from a database; A means for generating a response text for a character based on a user's input text using a generation AI; A means for generating a voice file from the response text generated using voice generation AI in a voice unique to the character; A means for the character and the user to have an interactive conversation on the terminal; A system including:
2. The system according to claim 1, wherein the voice generation AI converts the response text generated based on the character's setting materials into a voice file that reproduces the character's tone and voice quality.
3. 2. The system according to claim 1, wherein the terminal is provided with a means for allowing the user to select from a plurality of characters, and responses and voices are generated in accordance with the selected character.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A