system
The system addresses limitations in character interactions by using a generative model and real-time user movement capture to provide personalized and realistic dialogues with favorite characters in virtual spaces.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Existing technologies limit interactions with favorite characters or idols to pre-recorded voices and stereotyped responses, lacking personalization and realism in real-time virtual spaces.
A system utilizing a generative model for character voice generation, speech recognition, artificial intelligence for response generation, speech synthesis, and real-time user movement capture and reflection in virtual spaces to enable personalized and realistic interactions.
Users can enjoy real-time, personalized dialogues and interactive experiences with their favorite characters or idols in a virtual space, enhancing the sense of realism and engagement.
Smart Images

Figure 2026062247000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Modern fans strongly desire to interact and communicate with their favorite characters or idols, but it is difficult to achieve this in the real world. In conventional technologies, interactions with characters or idols are limited to pre-recorded voices or stereotyped responses, and it has not been possible to provide a personalized experience in real time for individual users. Also, there are limitations in the actions and interactions of avatars in virtual spaces, and there is a problem of lacking a sense of reality.
Means for Solving the Problems
[0005] This invention provides a real-time and personalized dialogue experience through a system that includes means for providing a generative model for generating the voice of a character selected by the user, speech recognition means for converting the user's voice into text, artificial intelligence means for generating a response based on the generative model, speech synthesis means for converting the generated response into the character's voice, and means for playing the voice response back to the user. Furthermore, by adding means for placing user and character avatars in a virtual space, capturing user movement information in real time, and reflecting it on the avatar in the virtual space, interactive interaction between the user and the character / idol can be realized, enhancing the sense of realism. As a result, users can enjoy dialogue with their favorite characters or idols while obtaining a realistic experience even in a virtual space.
[0006] A "generative model" is a trained algorithm or dataset used to generate the voice of a character selected by the user.
[0007] "Voice recognition means" refers to technologies and devices that convert a user's voice into a digital signal and then convert that signal into text data.
[0008] "Artificial intelligence tools" refer to algorithms and software that process input text data and generate appropriate responses.
[0009] "Speech synthesis means" refers to a technology or device for outputting generated response text as speech in the voice of a specific character.
[0010] "Playback means" refers to output devices such as speakers or earphones that deliver the sound generated by the speech synthesis means to the user.
[0011] A "virtual space" is a three-dimensional digital environment created through computers and networks that users can experience interactively.
[0012] An "avatar" is a digital entity that operates in a virtual space as a representation of a user or character.
[0013] "Action information" refers to data that indicates the user's physical actions, such as location, movement, and gestures.
[0014] "Real-time" refers to a situation where user input and actions are reflected almost instantly with virtually no delay. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12]It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.
Mode for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the language used in the following description will be explained.
[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space.
[0037] User registration and login
[0038] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which receives it and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[0039] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[0040] Choosing your favorite character or idol
[0041] When the user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records the selection information in the database and loads the associated voice model and animation data.
[0042] Starting a conversation with AI
[0043] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the voice data is sent from the device to the server. The server uses a speech recognition service to convert the voice data into text and sends the text to a generative AI. The generative AI generates an appropriate response text based on the input text, which is then returned to the server. The server then uses a speech synthesis engine to convert the response text into speech, generating the voice in the character's voice. This voice data is sent to the device, which then plays the voice for the user.
[0044] For example, if the user says "Hello, Naruto!", the AI will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[0045] Providing a sense of realism in the metaverse
[0046] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. This allows the user to enjoy a realistic interactive experience in the virtual space.
[0047] The following describes the processing flow.
[0048] Step 1:
[0049] Users access the system's registration page from their smartphones or PCs.
[0050] Step 2:
[0051] The device displays a registration form to the user.
[0052] Step 3:
[0053] The user enters their name, email address, and password and submits the form.
[0054] Step 4:
[0055] The terminal sends the input data to the server.
[0056] Step 5:
[0057] The server receives the transmitted data and checks for formatting and duplicates.
[0058] Step 6:
[0059] After the server verifies the data, it saves it to the database.
[0060] Step 7:
[0061] The server sends a registration completion message to the device.
[0062] Step 8:
[0063] The device displays a registration completion message to the user.
[0064] Step 9:
[0065] The user accesses the login page and enters their email address and password.
[0066] Step 10:
[0067] The terminal sends the input data to the server.
[0068] Step 11:
[0069] The server compares the information with the registered data in the database.
[0070] Step 12:
[0071] The server matches the user information and creates a login session.
[0072] Step 13:
[0073] The server sends the login session to the terminal.
[0074] Step 14:
[0075] The device displays the user's dashboard.
[0076] Step 15:
[0077] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[0078] Step 16:
[0079] The terminal sends a request to the server.
[0080] Step 17:
[0081] The server retrieves a list of available characters / idols from the database.
[0082] Step 18:
[0083] The server sends the list to the terminal.
[0084] Step 19:
[0085] The device displays the list to the user.
[0086] Step 20:
[0087] The user selects their desired character or idol.
[0088] Step 21:
[0089] The device sends the selection information to the server.
[0090] Step 22:
[0091] The server records the selection information in a database and loads the associated voice model and animation data.
[0092] Step 23:
[0093] The user clicks the "Start Conversation" button.
[0094] Step 24:
[0095] The device turns on the microphone and captures the user's voice.
[0096] Step 25:
[0097] The device sends the captured audio data to the server.
[0098] Step 26:
[0099] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[0100] Step 27:
[0101] The server sends text data to the AI generating the response and requests that it generate a response.
[0102] Step 28:
[0103] The AI generates appropriate response text and sends it to the server.
[0104] Step 29:
[0105] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[0106] Step 30:
[0107] The server sends the generated audio file to the terminal.
[0108] Step 31:
[0109] The device plays the generated audio for the user.
[0110] Step 32:
[0111] The user selects the "Metaverse Mode" option.
[0112] Step 33:
[0113] The terminal sends a request to the server.
[0114] Step 34:
[0115] The server connects to the metaverse platform and loads information from the virtual space.
[0116] Step 35:
[0117] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[0118] Step 36:
[0119] The user puts on a VR device and accesses a virtual space.
[0120] Step 37:
[0121] The device captures the user's movements and sends them to the server in real time.
[0122] Step 38:
[0123] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[0124] Step 39:
[0125] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[0126] Step 40:
[0127] The device provides an interactive experience between the user and their favorite character or idol.
[0128] (Example 1)
[0129] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0130] In modern virtual space systems and interactive character dialogue systems, it is difficult for users to engage in natural conversations and actions with characters within the virtual space. Furthermore, there is a lack of easy registration and login methods for users accessing the system for the first time, which leads to decreased user satisfaction. This invention aims to solve these problems and provide a system that allows users to interact with characters in real time and experience natural actions within the virtual space.
[0131] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0132] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for playing the voice response back to the user; registration means for the user to input registration information into the system; means for storing the registration information in a database; login means for the user to log in to the system; and means for matching the login information with information in the database. This allows the user to easily register and log in to the system, and subsequently engage in natural conversations with characters and interactive experiences in a virtual space in real time.
[0133] A "generative model" refers to a data model used to generate the voice of a character selected by the user.
[0134] "Voice recognition means" refers to technologies and devices for converting a user's voice into text data.
[0135] "Artificial intelligence means" refers to algorithms and programs for generating appropriate responses based on generative models.
[0136] "Speech synthesis means" refers to the technology that converts generated text data into the voice of a character.
[0137] "Means of playing voice responses to the user" refers to devices and technologies for playing the generated character's voice to the user.
[0138] "Registration method" refers to a series of processes and interfaces for a user to input their information into a system and for that information to be processed.
[0139] A "database" refers to a system used to store and manage user registration information and selection information.
[0140] "Login method" refers to the process or technology used to authenticate users when they access a system.
[0141] "Verification means" refers to technology used to compare login information entered by a user with registered information in a database and confirm that they match.
[0142] A "virtual space" refers to a digital environment where user and character avatars exist and which provides an interactive experience.
[0143] An "avatar" refers to a digital character that represents a user or character within a virtual space.
[0144] "Motion information" refers to information that detects the user's physical movements and captures them as digital data.
[0145] "Means of capturing" refers to devices and technologies for acquiring user behavior information in real time.
[0146] "Means of reflection" refers to technologies that reflect acquired user behavior information onto the movements of an avatar in a virtual space.
[0147] As a specific embodiment of this invention, a system is provided that allows users to interact in real time with characters or idols of their choice and enjoy an interactive experience in a virtual space.
[0148] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form for the user to enter their name, email address, and password. Once the user enters the required information and submits it, the device sends this data to the server, which receives the data and stores it in its database. Upon completion of registration, the device displays a registration completion message to the user.
[0149] Next, the user accesses the login page and enters their email address and password. The device sends this data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard. This process allows users to easily access the system.
[0150] When a user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays this list to the user, who then selects their desired character or idol. The device sends this selection information to the server, which records the selection information in the database and loads the associated voice model and animation data. This prepares the user to interact with their chosen character.
[0151] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello," the voice data is sent from the device to the server. The server uses a speech recognition service (e.g., Google® Speech-to-Text) to convert the voice data into text. The server then sends the text to a generative AI model (e.g., OpenAI® GPT-3®) to generate an appropriate response text. The generated text is returned to the server, which uses a speech synthesis engine (e.g., Amazon Polly) to generate the character's voice. This generated voice is sent to the device and played back to the user. This process allows the user to experience a natural conversation with the character.
[0152] For example, if a user says "Hello, Naruto!", the generative AI model will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[0153] Furthermore, when a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on the VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character avatar to move naturally in response to the user's movements. Through this process, the user can enjoy a realistic interactive experience in the virtual space.
[0154] An example of a prompt message is, "I want to talk to Naruto for a bit. Turn on the microphone and speak to him." Based on this prompt, users can access the system from their smartphones or PCs and enjoy interacting with characters and experiencing virtual spaces.
[0155] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0156] Step 1:
[0157] Access the user's system registration page.
[0158] Input: The user enters the URL into their browser and accesses it.
[0159] Output: The registration page is displayed.
[0160] Specific operation: When a user enters the system's URL into their browser, the browser sends a request to the server, which generates a registration page and sends it back to the user's device.
[0161] Step 2:
[0162] The device displays a registration form to the user.
[0163] Input: Registration page data received from the server.
[0164] Output: The registration form is displayed on the screen.
[0165] Specific operation: The device uses the HTML and CSS included in the registration page to render the form on the user's screen.
[0166] Step 3:
[0167] The user enters the required information into the registration form and clicks the submit button.
[0168] Input: The action of a user entering their name, email address, and password and clicking the submit button.
[0169] Output: The input data is sent to the server.
[0170] Specific operation: The user enters, for example, "Taro Yamada", "example@mail.com", and "password123" into the form and clicks the "Submit" button. The input data is sent to the server as an HTTP POST request.
[0171] Step 4:
[0172] The server receives the input data and saves it to the database.
[0173] Input: User registration data received from the terminal.
[0174] Output: The registered data is saved in the database.
[0175] Specific operation: The server analyzes the received data and executes an INSERT query to save the data to the MySQL® database.
[0176] Step 5:
[0177] The device displays a registration completion message to the user.
[0178] Input: Registration completion response from the server.
[0179] Output: A registration complete message is displayed on the screen.
[0180] Specific actions: The server notifies the terminal that saving to the database is complete and sends a response containing that information. The terminal displays the message "Registration complete!" on its screen.
[0181] Step 6:
[0182] The user accesses the login page and enters their email address and password.
[0183] Input: Access the login page URL, enter user information (email address, password).
[0184] Output: Your email address and password will be sent to the server.
[0185] Specific operation: The user enters and accesses the login page URL, enters "example@mail.com" and "password123" in the login form, and clicks the "Login" button. The data is sent via an HTTP POST request.
[0186] Step 7:
[0187] The server verifies the login information against the registered information in the database.
[0188] Input: Login data from the device.
[0189] Output: Login authentication result.
[0190] Specific operation: The server searches the "users" table in the database for a record with a matching email address and verifies the password. If a match is found, a login session ID is generated.
[0191] Step 8:
[0192] The server creates a login session and sends it to the terminal.
[0193] Input: Result indicating successful login.
[0194] Output: Session ID.
[0195] Specific operation: The server generates a session ID and sends it to the terminal.
[0196] Step 9:
[0197] The device displays the user's dashboard.
[0198] Input: Session ID response from the server.
[0199] Output: User's dashboard screen.
[0200] Specific action: The device receives the session ID, generates a dashboard page, and displays it on the screen. This page will include a message such as "Hello, Taro Yamada!".
[0201] Step 10:
[0202] The user clicks the "Select your favorite character / idol" option.
[0203] Input: User click operation.
[0204] Output: The request is sent to the server.
[0205] Specific operation: When the user clicks the "Select Favorite Character / Idol" button, the device sends an HTTP GET request to the server.
[0206] Step 11:
[0207] The server retrieves a list of character idols from the database and sends it to the terminal.
[0208] Input: Request from the terminal.
[0209] Output: Character / Idol List.
[0210] Specific operation: The server retrieves all records from the "characters" table and sends them to the terminal.
[0211] Step 12:
[0212] The device displays a list of characters / idols to the user.
[0213] Input: List data from the server.
[0214] Output: A list of character idols is displayed on the screen.
[0215] Specific operation: Based on the received list data, the terminal displays a list of characters / idols to the user.
[0216] Step 13:
[0217] The user selects their desired character or idol.
[0218] Input: User selection operation.
[0219] Output: The selected character / idol information is sent to the server.
[0220] Specific operation: The user clicks "Naruto" from the list, and the selection information is sent to the server via an HTTP POST request.
[0221] Step 14:
[0222] The server records the selection information in a database and loads the associated voice model and animation data.
[0223] Input: Character selection information from the device.
[0224] Output: Recording to a database, loading of voice models and animation data.
[0225] Specific operation: The server records the selection information in the "users" table and processes the loading of voice models and animation data related to "Naruto".
[0226] Step 15:
[0227] The user clicks the "Start Conversation" button.
[0228] Input: User click operation.
[0229] Output: The microphone turns on.
[0230] Specific operation: When the user clicks the "Start Conversation" button, the device turns on the microphone and prepares to capture audio.
[0231] Step 16:
[0232] The device captures the user's voice and sends it to the server.
[0233] Input: User's voice.
[0234] Output: Audio data is sent to the server.
[0235] Specific operation: The device records the user's voice saying "Hello, Naruto!" and sends that data to the server via an HTTP POST request.
[0236] Step 17:
[0237] The server uses a speech recognition service to convert the speech data into text.
[0238] Input: Audio data from the device.
[0239] Output: Text data.
[0240] Specific operation: The server uses the Google Speech-to-Text API to convert the audio data into the text "Hello, Naruto!".
[0241] Step 18:
[0242] The server sends text to the AI model and generates an appropriate response text.
[0243] Input: Text data converted by speech recognition.
[0244] Output: Response text data.
[0245] Specific operation: The server sends text data to OpenAI's GPT-3 and generates the response, "Hey! What's up?"
[0246] Step 19:
[0247] The server converts the response text into speech using a speech synthesis engine.
[0248] Input: Response text from a generating AI model.
[0249] Output: Audio data.
[0250] Specific operation: The server uses Amazon Polly to generate audio data that says, "Hey! What's up?"
[0251] Step 20:
[0252] The server sends the generated audio data to the terminal, and the terminal plays it for the user.
[0253] Input: Audio data from the server.
[0254] Output: The audio played to the user.
[0255] Specific action: The device plays the received Naruto voice, and the user hears the voice say, "Hey! What's up?"
[0256] Step 21:
[0257] The user selects the "Metaverse Mode" option.
[0258] Input: User click operation.
[0259] Output: The request is sent to the server.
[0260] Specific action: The user clicks the "Metaverse Mode" button on the dashboard, and the device sends an HTTP GET request to the server.
[0261] Step 22:
[0262] The server connects to the metaverse platform and loads information from the virtual space.
[0263] Input: Request from the terminal.
[0264] Output: Data in a virtual space.
[0265] Specific operation: The server uses the API of the metaverse platform (such as VRChat) to obtain the necessary virtual space data.
[0266] Step 23:
[0267] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[0268] Input: Data from the virtual space, user and character information.
[0269] Output: An avatar placed in a virtual space.
[0270] Specific operation: The server places the user's avatar and Naruto's avatar within VRChat.
[0271] Step 24:
[0272] The user puts on a VR device and accesses a virtual space.
[0273] Input: User action.
[0274] Output: Log in to the virtual space.
[0275] Specific operation: The user wears the Oculus Quest, launches the VRChat app, and logs in to the virtual space.
[0276] Step 25:
[0277] The terminal captures the user's movements and sends them to the server in real time.
[0278] Input: Operation information of the VR device.
[0279] Output: Real-time operation data.
[0280] Specific operation: The terminal (VR device) captures the user's movements with sensors and sends them to the server in real time.
[0281] Step 26:
[0282] The server issues an instruction to move the character's avatar naturally according to the user's movement information.
[0283] Input: Real-time operation data from the terminal.
[0284] Output: The avatar of Naruto in the virtual space moves.
[0285] Specific operation: Based on the received operation data, the server instructs the appropriate actions for Naruto's avatar in VRChat and moves it naturally in real time according to the user's movements.
[0286] Through the above series of steps, the user can register and log in to the system, interact with characters and idles, and further enjoy real-time interactive experiences in the virtual space.
[0287] (Application Example 1)
[0288] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0289] Conventional voice interaction systems have limited functionality for users to interact with characters or idols, and providing a truly interactive experience in real time presented many challenges. Furthermore, character-like behavior and real-time reflection of actions in virtual space were difficult, preventing users from enjoying a consistent entertainment experience. Moreover, no system existed that integrated the entire process from speech recognition to generational AI and speech synthesis to achieve real-time responses within a virtual space.
[0290] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0291] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for displaying the voice response in a virtual space; means for controlling the character's actions in the virtual space; and means for playing the voice response back to the user. This enables the user to interact with characters and idols in real time in a virtual space and enjoy an interactive and consistent experience.
[0292] A "user" is a person who uses this system to interact with characters and idols in a virtual space.
[0293] "Character voice" refers to voice data corresponding to a specific character, and is used to reproduce the voice of the character selected by the user.
[0294] A "generative model" is a machine learning model designed to generate the voice of a character selected by the user.
[0295] "Speech recognition means" is a general term for hardware and software used to convert a user's voice into text data.
[0296] "Artificial intelligence means" refers to machine learning algorithms that generate appropriate responses from input text based on a generative model.
[0297] "Speech synthesis means" is a general term for hardware and software used to convert generated text responses into character voices.
[0298] A "virtual space" is a three-dimensional computer graphics space where users and characters engage in dialogue and interaction.
[0299] "Means for displaying voice responses" refers to the collective term for hardware and software used to visually represent voice responses generated within a virtual space.
[0300] "Means for controlling character-like behavior" refers to a general term for algorithms and systems used to control how characters in a virtual space behave naturally.
[0301] "Means for playing back voice responses" refers to an audio output device and associated software that allows the user to hear the character's voice responses.
[0302] One embodiment for carrying out this invention will be described. This system enables users to interact with characters and idols in a virtual space in real time and enjoy an interactive experience.
[0303] Hardware and software to be used
[0304] The main hardware used to implement this invention includes the following:
[0305] Terminal: A device (smartphone, PC, head-mounted display) for the user to access the system
[0306] Server: The center for performing speech recognition, generative AI, speech synthesis, and control of the virtual space
[0307] Speech input device: A microphone for capturing the user's voice
[0308] Speech output device: A speaker for playing the character's voice response
[0309] The main software includes the following:
[0310] Speech recognition software: Google's speech recognition API
[0311] Generative AI model: OpenAI GPT-3 model
[0312] Speech synthesis software: pyttsx3 library
[0313] Virtual space management software: pyqtmetaverse library
[0314] Data processing and data calculation
[0315] 1. User's voice input
[0316] The user speaks to the terminal. The speech input device captures the user's voice and sends it to the terminal.
[0317] 2. Speech recognition means
[0318] The terminal passes the captured voice data to the speech recognition software to convert it into text data.
[0319] 3. Artificial intelligence tools
[0320] The server passes text data to the generative AI model, which then generates appropriate response text. This process creates prompt text for the generative AI.
[0321] Example prompt: "Generate a character's response to the question, 'What is the best product in this store?'"
[0322] 4. Speech synthesis means
[0323] The server passes the generated text response to speech synthesis software, which converts it into a character's voice.
[0324] 5. Display in virtual space
[0325] The virtual space management software displays the generated voice responses to the user in the virtual space both visually and audibly. It also controls the character's movements to allow the user to interact naturally.
[0326] 6. Playback of voice response
[0327] The device plays character voices to the user, providing an interactive dialogue experience.
[0328] Specific example
[0329] Suppose a user asks, "What is the best item in this store?" The server captures this audio and converts it into text, "What is the best item in this store?", through speech recognition software. Based on the prompt, "Generate a character response to 'What is the best item in this store?'", the generative AI model generates the appropriate response text, "Then it's the automatic ramen cooker! It's convenient and you can enjoy delicious ramen anytime." Speech synthesis software converts this response text into the character's voice. Finally, virtual space management software displays this voice within the virtual space, and the terminal plays the voice for the user.
[0330] This allows users to enjoy natural, real-time interactions with characters and idols in a virtual space.
[0331] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0332] Step 1:
[0333] The user speaks into a voice input device (microphone). The terminal captures this audio and prepares the data.
[0334] Input: User's voice
[0335] Output: Audio data
[0336] Specific operation: When a user says, "What is the best product you recommend in this store?", the microphone captures the audio.
[0337] Step 2:
[0338] The device passes the captured audio data to speech recognition software (Google's speech recognition API) and converts it into text data.
[0339] Input: Audio data
[0340] Output: Text data
[0341] Specific operation: The device sends the captured audio data to Google's speech recognition API, which converts it into the text, "What is the best product in this store?"
[0342] Step 3:
[0343] The server receives text data and passes it to a generative AI model (OpenAI GPT-3 model) to generate appropriate response text. During this process, it generates a prompt and sends it to the AI.
[0344] Input: Text data
[0345] Output: Response text
[0346] Specific operation: Based on the text "What is the best recommended item in this store?", the server generates a prompt message "Generate a character response to 'What is the best recommended item in this store?'". The generating AI model receives this and generates the response text "Then it's the automatic ramen cooker. It's convenient, and you can enjoy delicious ramen anytime."
[0347] Step 4:
[0348] The server passes the generated response text to speech synthesis software (pyttsx3) and converts it into a character's voice.
[0349] Input: Response text
[0350] Output: Audio data
[0351] Specific operation: The server passes the text "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime." to pyttsx3, which then converts it into character voice.
[0352] Step 5:
[0353] The server uses virtual space management software (pyqtmetaverse) to display the generated audio data within the virtual space and control the character's movements.
[0354] Input: Voice data, response text
[0355] Output: Character movements and sound display in the virtual space
[0356] Specific operation: The server passes the generated audio data and response text to the pyqtmetaverse, which then controls the movement of the character in the virtual space and displays it to the user visually and audibly.
[0357] Step 6:
[0358] The device plays the generated character's voice to the user through its speaker.
[0359] Input: Audio data
[0360] Output: Audio to be played
[0361] Specific operation: The device plays a voice message from a character through its speaker saying, "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime," to the user.
[0362] Through the above processing steps, users can enjoy real-time, interactive conversations with characters and idols in a virtual space.
[0363] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0364] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space. It also has a function to recognize the user's emotions and adjust its response accordingly.
[0365] User registration and login
[0366] The user first accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which then receives the data and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[0367] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[0368] Choosing your favorite character or idol
[0369] When the user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records this information in the database and loads the associated voice model and animation data.
[0370] Initiating conversations with AI and emotion recognition
[0371] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", that voice data is sent from the device to the server. The server receives the voice data and uses a speech recognition service to convert the voice to text. At the same time, an emotion engine analyzes the voice and the user's facial expressions to recognize the user's emotions (e.g., joy, sadness, surprise, etc.).
[0372] The server sends text data and recognized emotion information to the generating AI and requests it to generate an appropriate response. The generating AI generates an appropriate response text based on the input text and emotion information and returns it to the server. The server processes the response text through a speech synthesis engine and generates audio in the character's voice. This audio data is sent to the terminal, which then plays the generated audio for the user.
[0373] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generating AI will produce the response "Hi! You seem to be doing well!" which will then be played back to the user in Naruto's voice.
[0374] Providing a sense of realism in the metaverse
[0375] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. Conversation with the generating AI also takes place in parallel, and voice and corresponding animations are reproduced in the virtual space.
[0376] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space. This allows users to not only have a realistic interactive experience in the virtual space but also receive responses that are appropriate to their emotions.
[0377] This invention allows users to enjoy interacting with their favorite characters or idols while gaining a realistic experience even in a virtual space. Emotion recognition improves the quality of the interaction, enabling deeper engagement.
[0378] The following describes the processing flow.
[0379] Step 1:
[0380] Users access the system's registration page from their smartphones or PCs.
[0381] Step 2:
[0382] The device displays a registration form to the user.
[0383] Step 3:
[0384] The user enters their name, email address, and password and submits the form.
[0385] Step 4:
[0386] The terminal sends the input data to the server.
[0387] Step 5:
[0388] The server receives the transmitted data and checks for formatting and duplicates.
[0389] Step 6:
[0390] After the server verifies the data, it saves it to the database.
[0391] Step 7:
[0392] The server sends a registration completion message to the device.
[0393] Step 8:
[0394] The device displays a registration completion message to the user.
[0395] Step 9:
[0396] The user accesses the login page and enters their email address and password.
[0397] Step 10:
[0398] The terminal sends the input data to the server.
[0399] Step 11:
[0400] The server compares the information with the registered data in the database.
[0401] Step 12:
[0402] The server matches the user information and creates a login session.
[0403] Step 13:
[0404] The server sends the login session to the terminal.
[0405] Step 14:
[0406] The device displays the user's dashboard.
[0407] Step 15:
[0408] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[0409] Step 16:
[0410] The terminal sends a request to the server.
[0411] Step 17:
[0412] The server retrieves a list of available characters / idols from the database.
[0413] Step 18:
[0414] The server sends the list to the terminal.
[0415] Step 19:
[0416] The device displays the list to the user.
[0417] Step 20:
[0418] The user selects their desired character or idol.
[0419] Step 21:
[0420] The device sends the selection information to the server.
[0421] Step 22:
[0422] The server records the selection information in a database and loads the associated voice model and animation data.
[0423] Step 23:
[0424] The user clicks the "Start Conversation" button.
[0425] Step 24:
[0426] The device turns on the microphone and captures the user's voice and facial expressions.
[0427] Step 25:
[0428] The device sends the captured audio and facial expression data to the server.
[0429] Step 26:
[0430] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[0431] Step 27:
[0432] The server uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotions.
[0433] Step 28:
[0434] The server sends text data and recognized emotional information to the AI and requests it to generate a response.
[0435] Step 29:
[0436] The generation AI generates appropriate response text based on the input text and sentiment information, and sends it to the server.
[0437] Step 30:
[0438] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[0439] Step 31:
[0440] The server sends the generated audio file to the terminal.
[0441] Step 32:
[0442] The device plays generated audio and visual information for the user.
[0443] Step 33:
[0444] The user selects the "Metaverse Mode" option.
[0445] Step 34:
[0446] The terminal sends a request to the server.
[0447] Step 35:
[0448] The server connects to the metaverse platform and loads information from the virtual space.
[0449] Step 36:
[0450] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[0451] Step 37:
[0452] The user puts on a VR device and accesses a virtual space.
[0453] Step 38:
[0454] The device captures the user's movements and sends them to the server in real time.
[0455] Step 39:
[0456] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[0457] Step 40:
[0458] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space.
[0459] Step 41:
[0460] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[0461] Step 42:
[0462] The device provides an interactive experience between the user and their favorite character or idol.
[0463] (Example 2)
[0464] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0465] Conventional voice dialogue systems simply provide monotonous responses to user input, making it difficult to recognize emotions and engage in interactive conversations. Furthermore, even in virtual reality interactions, real-time control based on user actions is insufficient, failing to provide a realistic experience. This resulted in a decline in the quality of the user experience.
[0466] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0467] In this invention, the server includes means for providing a generative model for generating the voice of a virtual character selected by the user, speech recognition means for converting the user's voice into text, and artificial intelligence means for generating a response based on the text and emotional information. This makes it possible to recognize the user's emotions and adjust the response accordingly.
[0468] A "user" is an end-user who uses the system to interact with others and have experiences in a virtual space.
[0469] A "virtual character" is a virtual character or idol that users can choose to use in conversations and interactive experiences.
[0470] A "generative model" is an artificial intelligence model used to generate the voice and behavior of a virtual character selected by the user.
[0471] "Speech recognition means" refers to technologies and devices for converting a user's speech into text.
[0472] "Artificial intelligence means" refers to technologies and devices that generate appropriate responses based on text data and emotional information.
[0473] "Voice synthesis means" refers to technologies and devices for converting generated responses into the voice of a virtual character.
[0474] "Emotion recognition means" refers to technologies and devices that recognize emotions from a user's voice and facial expressions and adjust responses based on that information.
[0475] "Means for reproducing voice responses" refers to technologies or devices for reproducing generated voice responses to the user.
[0476] An "avatar" is a graphical representation displayed in a virtual space as a digital representation of a user or virtual character.
[0477] A "virtual space" is a computer-generated three-dimensional space that users can access through VR devices or similar means.
[0478] "Motion information" refers to data about the user's body movements and is used to control the avatar in the virtual space.
[0479] "Real-time" refers to a sense of time in which processing and responses occur almost instantly, meaning that delays to user operations and inputs are kept to a minimum.
[0480] Embodiments of this invention will now be described in detail. This system aims to allow users to interact with their favorite virtual characters or idols in real time and enjoy an interactive experience in a virtual space. It also has the function of recognizing the user's emotions and adjusting its response accordingly.
[0481] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who enters necessary information such as their name, email address, and password, and sends it from the device to the server. The server stores the received data in a database and notifies the user via the device that registration is complete. Next, the user accesses the login page and enters the registered email address and password. The device sends the entered data to the server, which verifies it against the information in the database. If authentication is successful, the server generates a login session and sends it to the device, displaying the user's dashboard.
[0482] When a user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of characters and idols from the database and sends it to the device. The device displays this list to the user, and when the user selects their desired character or idol, that information is sent to the server and recorded in the database.
[0483] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the audio data is sent from the device to the server. The server converts this audio data to text using the Google Cloud Speech-to-Text API and analyzes the text and emotion information with an emotion engine (e.g., Microsoft® Azure® Emotion API). If the analysis result is positive, the server sends the text data and emotion information to a generating AI model (e.g., OpenAI GPT-3) to generate a response such as "Hi! You seem to be doing well!". The generated response text is then synthesized into a character's voice using Amazon Polly and sent to the device for playback.
[0484] Furthermore, when the user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to a Unity-based metaverse platform and loads information for the virtual space. The server places the user's avatar and character / idol avatars in the virtual space, and the user accesses the virtual space by wearing a VR device (e.g., Oculus Rift). The device captures the user's movements in real time and sends them to the server. The server controls the movements of the virtual avatar through the Unity engine and simultaneously provides a more natural and realistic interactive experience through voice interaction via a generated AI model. In addition, an emotion engine recognizes the user's emotional changes in real time and reflects that information in the movements and facial expressions of the virtual avatar.
[0485] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generative AI model will generate the response "Hi! You seem to be doing well!" This response is synthesized into Naruto's voice and played back to the user through the speaker.
[0486] Examples of prompts for operating this system include:
[0487] Prompt to convert audio data to text:
[0488] Please convert the audio data 'Hello, Naruto!' to text.
[0489] Response generation prompt:
[0490] "Based on the text 'Hello, Naruto!' and the emotion 'Joyful,' please generate an appropriate response."
[0491] Text-to-speech prompt:
[0492] "Please convert the text 'Hey! You seem to be doing well!' into speech in Naruto's voice."
[0493] Through these processes, users can gain real-time interaction and an interactive experience in a virtual space.
[0494] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0495] Program processing flow
[0496] Step 1:
[0497] Users access the system's registration page from their smartphones or PCs.
[0498] Input: Name, email address, password
[0499] The terminal receives these inputs, displays them on the registration form, and waits for user input.
[0500] Output: Input data
[0501] The user enters their name, email address, and password, and this data is sent to the device.
[0502] Step 2:
[0503] The terminal sends the input data to the server.
[0504] Input: User registration information (name, email address, password)
[0505] The server receives the input data and saves it to the database.
[0506] Output: Registration complete message
[0507] The device displays "Registration complete!" to the user.
[0508] Step 3:
[0509] The user accesses the login page and enters their email address and password.
[0510] Enter: Email address, password
[0511] The terminal sends the input data to the server.
[0512] Output: Authentication result
[0513] The server compares the information with the registration details in the database, and if a match is found, it generates a login session.
[0514] A login session is sent to the device, and the device displays the user's dashboard.
[0515] Step 4:
[0516] The user clicks the "Select Favorite Character / Idol" button in the main menu.
[0517] Input: Click Event
[0518] The terminal sends a request to the server.
[0519] Output: List of Characters / Idols
[0520] The server retrieves a list of characters and idols from the database and sends it to the terminal.
[0521] The device displays the list to the user.
[0522] Step 5:
[0523] The user selects their desired character or idol.
[0524] Input: Selected character or idol
[0525] The device sends the selection information to the server.
[0526] Output: Database update results
[0527] The server records that information in the database.
[0528] Step 6:
[0529] The user clicks the "Start Conversation" button.
[0530] Input: Click Event
[0531] The device turns on the microphone and captures the user's voice.
[0532] The user says, "Hello, [Character Name]!"
[0533] Output: Audio data
[0534] The audio data is then sent from the terminal to the server.
[0535] Step 7:
[0536] The server sends the voice data to the speech recognition service.
[0537] Input: Audio data
[0538] The speech recognition service converts speech into text.
[0539] Output: Text data
[0540] The server receives this text data.
[0541] Step 8:
[0542] The server recognizes emotions from voice and user facial expression data through an emotion engine.
[0543] Input: Voice data, user facial expression data
[0544] Output: Emotional information
[0545] The server uses text data and recognized emotion information to send to a generative AI model.
[0546] Step 9:
[0547] The generative AI model generates appropriate responses based on text data and sentiment information.
[0548] Input: Text data, sentiment information
[0549] Output: Response text
[0550] The server receives the response text.
[0551] Step 10:
[0552] The server processes the response text into a speech synthesis engine.
[0553] Input: Response text
[0554] The speech synthesis engine generates speech using the voice of a virtual character.
[0555] Output: Generated audio data
[0556] The server sends the audio data to the terminal.
[0557] Step 11:
[0558] The device plays the generated audio to the user.
[0559] Input: Audio data
[0560] The device plays audio through its speaker, and the user listens to the response.
[0561] Output: Played audio
[0562] Step 12:
[0563] The user selects the "Metaverse Mode" option.
[0564] Input: Click Event
[0565] The terminal sends a request to the server.
[0566] The server connects to the metaverse platform and loads the virtual space.
[0567] Output: Display data for the virtual space
[0568] Step 13:
[0569] The server places user avatars and character / idol avatars in the virtual space.
[0570] Input: None (automatic processing)
[0571] Users wear VR devices to access a virtual space.
[0572] Output: Initial placement of user and character avatars
[0573] Step 14:
[0574] The device captures the user's movements and sends them to the server.
[0575] Input: User activity information
[0576] The server controls the avatar in real time via the Unity engine.
[0577] Output: Avatar movement data
[0578] This allows users to enjoy interactive experiences in a virtual space.
[0579] The above is the specific processing flow of this system.
[0580] (Application Example 2)
[0581] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0582] Conventional virtual space systems have limited interaction between users and virtual characters or idols, lacking interaction based on emotions and actions. Therefore, users find it difficult to have a realistic conversational experience, resulting in a less realistic virtual experience overall. Furthermore, the lack of interactive dialogue systems utilizing emotion recognition is another challenge.
[0583] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for providing a generation model for generating the voice of a virtual character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating an appropriate response based on the generation model; speech synthesis means for converting the generated response into the voice of a virtual character; emotion recognition means for analyzing the user's emotions; means for adjusting the response based on the emotion recognition results; and means for playing the generated voice response to the user. This makes it possible for the user to enjoy a realistic dialogue experience with their favorite virtual character or idol that responds to their emotions.
[0584] A "user" is someone who uses the system to enjoy interacting with virtual characters or idols.
[0585] A "virtual space" is a three-dimensional virtual environment created by a computer program.
[0586] A "virtual character" is a character or idol selected by a user within a virtual space, whose voice and actions are generated based on AI.
[0587] A "generative model" refers to the algorithms and datasets used to generate voices and actions for virtual characters.
[0588] "Speech recognition means" refers to a technical device or program for generating text from a user's speech.
[0589] An "artificial intelligence system" is a system that uses a pre-trained model to generate an appropriate response to user input.
[0590] "Speech synthesis means" refers to a device or program that converts text data into the voice of a virtual character and generates speech data.
[0591] "Emotion recognition means" refers to technologies and programs that analyze and recognize emotions from a user's voice, facial expressions, etc.
[0592] "Means for adjusting responses" refer to technologies or programs for appropriately modifying responses generated based on the user's emotion recognition results.
[0593] "Means for reproducing voice responses" refers to devices or programs that output generated voice data to the user.
[0594] The embodiments for carrying out this invention are described in detail below.
[0595] First, the system consists of the following main components:
[0596] 1. A generative model for generating voices for virtual characters based on user selections.
[0597] 2. Speech recognition means
[0598] 3. Artificial intelligence means for generating appropriate responses
[0599] 4. Speech synthesis means for converting the generated response into speech.
[0600] 5. Emotion recognition means for analyzing user emotions
[0601] 6. Means for adjusting responses based on emotion recognition results
[0602] 7. Means for playing voice responses to the user
[0603] Specifically, when a user accesses the system and enters their email address and password on the login page, the server compares this information with existing data in the database. If the login is successful, the user's dashboard is displayed. The user then selects their preferred virtual character from the dashboard, and the server loads the character's voice model and animation data.
[0604] Hardware and software configuration
[0605] The hardware and software used in this invention are as follows:
[0606] Hardware: Smartphone, tablet, or VR headset, microphone, speaker
[0607] software:
[0608] Google Cloud Speech-to-Text: Used as a speech recognition method.
[0609] Azure Text Analytics: Used as a means of emotion recognition
[0610] OpenAI GPT-3: Used as an artificial intelligence means to generate appropriate responses.
[0611] Amazon Polly: Used as a speech synthesis method.
[0612] Data adjustment and calculation flow
[0613] 1. Speech recognition:
[0614] The audio data spoken by the user through the microphone is sent to the server and converted into text using Google Cloud Speech-to-Text.
[0615] 2. Emotion recognition:
[0616] The data, converted to text, is analyzed by Azure Text Analytics to recognize the user's sentiment.
[0617] 3. Response generation:
[0618] The server inputs the recognized emotion and text data into OpenAI GPT-3 and generates an appropriate response text.
[0619] 4. Speech synthesis:
[0620] The generated response text is converted into a virtual character's voice by Amazon Polly.
[0621] 5. Playback of voice response:
[0622] Audio data generated from the server is sent to the user's device, and the device plays the audio.
[0623] Specific example
[0624] For example, a user says "Hello, Naruto!" into the microphone. This audio is sent to the server. Google Cloud Speech-to-Text converts the audio into text "Hello, Naruto!". Next, Azure Text Analytics recognizes the user's emotion from this text as "joy". The server sends this recognized emotion and text to OpenAI GPT-3 to generate an appropriate response, "Hi! You seem to be doing well!". This response is converted into speech by Amazon Polly and played back on the user's device in the voice of the virtual character Naruto.
[0625] Example of a prompt
[0626] Joyful user: Hello, Naruto! Virtual character:
[0627] In this way, the system of the present invention enables users to enjoy a realistic dialogue experience that responds to their emotions.
[0628] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0629] Step 1:
[0630] The server receives the email address and password entered by the user and compares them with existing information in the database. If the entered email address and password match, the server creates a login session and displays the user's dashboard.
[0631] Step 2:
[0632] The user selects a virtual character from the dashboard. The device sends the selection information to the server, which loads the corresponding virtual character's voice model and animation data from the database and presents it to the user.
[0633] Step 3:
[0634] When the user presses the "Start Conversation" button, the device turns on the microphone and captures the user's voice data. This voice data is sent to the server. The server uses Google Cloud Speech-to-Text to convert this voice data into text.
[0635] Step 4:
[0636] The server sends the converted text data to Azure Text Analytics for user sentiment analysis. The sentiment analysis results are stored along with the text data.
[0637] Step 5:
[0638] The server sends text data and emotion recognition results to OpenAI GPT-3 to generate an appropriate response. Specifically, natural-sounding dialogue is generated based on the prompt text and the recognized emotion information.
[0639] Step 6:
[0640] The generated response text is converted into audio data by Amazon Polly. The server then sends the generated audio data to the device.
[0641] Step 7:
[0642] The user's device plays the received audio data. This provides the user with the experience of a virtual character responding to them verbally.
[0643] Step 8:
[0644] User movement information is also captured in real time and sent from the terminal to the server. The server controls the movement of the avatar in the virtual space and adjusts the avatar to respond to the user's movements.
[0645] This allows users to enjoy realistic, emotion-driven conversations and action-based interactions within the virtual space.
[0646] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0647] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0648] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0649] [Second Embodiment]
[0650] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0651] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0652] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0653] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0654] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0655] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0656] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0657] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0658] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0659] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0660] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0661] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0662] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space.
[0663] User registration and login
[0664] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which receives it and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[0665] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[0666] Choosing your favorite character or idol
[0667] When the user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records the selection information in the database and loads the associated voice model and animation data.
[0668] Starting a conversation with AI
[0669] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the voice data is sent from the device to the server. The server uses a speech recognition service to convert the voice data into text and sends the text to a generative AI. The generative AI generates an appropriate response text based on the input text, which is then returned to the server. The server then uses a speech synthesis engine to convert the response text into speech, generating the voice in the character's voice. This voice data is sent to the device, which then plays the voice for the user.
[0670] For example, if the user says "Hello, Naruto!", the AI will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[0671] Providing a sense of realism in the metaverse
[0672] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. This allows the user to enjoy a realistic interactive experience in the virtual space.
[0673] The following describes the processing flow.
[0674] Step 1:
[0675] Users access the system's registration page from their smartphones or PCs.
[0676] Step 2:
[0677] The device displays a registration form to the user.
[0678] Step 3:
[0679] The user enters their name, email address, and password and submits the form.
[0680] Step 4:
[0681] The terminal sends the input data to the server.
[0682] Step 5:
[0683] The server receives the transmitted data and checks for formatting and duplicates.
[0684] Step 6:
[0685] After the server verifies the data, it saves it to the database.
[0686] Step 7:
[0687] The server sends a registration completion message to the device.
[0688] Step 8:
[0689] The device displays a registration completion message to the user.
[0690] Step 9:
[0691] The user accesses the login page and enters their email address and password.
[0692] Step 10:
[0693] The terminal sends the input data to the server.
[0694] Step 11:
[0695] The server compares the information with the registered data in the database.
[0696] Step 12:
[0697] The server matches the user information and creates a login session.
[0698] Step 13:
[0699] The server sends the login session to the terminal.
[0700] Step 14:
[0701] The device displays the user's dashboard.
[0702] Step 15:
[0703] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[0704] Step 16:
[0705] The terminal sends a request to the server.
[0706] Step 17:
[0707] The server retrieves a list of available characters / idols from the database.
[0708] Step 18:
[0709] The server sends the list to the terminal.
[0710] Step 19:
[0711] The device displays the list to the user.
[0712] Step 20:
[0713] The user selects their desired character or idol.
[0714] Step 21:
[0715] The device sends the selection information to the server.
[0716] Step 22:
[0717] The server records the selection information in a database and loads the associated voice model and animation data.
[0718] Step 23:
[0719] The user clicks the "Start Conversation" button.
[0720] Step 24:
[0721] The device turns on the microphone and captures the user's voice.
[0722] Step 25:
[0723] The device sends the captured audio data to the server.
[0724] Step 26:
[0725] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[0726] Step 27:
[0727] The server sends text data to the AI generating the response and requests that it generate a response.
[0728] Step 28:
[0729] The AI generates appropriate response text and sends it to the server.
[0730] Step 29:
[0731] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[0732] Step 30:
[0733] The server sends the generated audio file to the terminal.
[0734] Step 31:
[0735] The device plays the generated audio for the user.
[0736] Step 32:
[0737] The user selects the "Metaverse Mode" option.
[0738] Step 33:
[0739] The terminal sends a request to the server.
[0740] Step 34:
[0741] The server connects to the metaverse platform and loads information from the virtual space.
[0742] Step 35:
[0743] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[0744] Step 36:
[0745] The user puts on a VR device and accesses a virtual space.
[0746] Step 37:
[0747] The device captures the user's movements and sends them to the server in real time.
[0748] Step 38:
[0749] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[0750] Step 39:
[0751] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[0752] Step 40:
[0753] The device provides an interactive experience between the user and their favorite character or idol.
[0754] (Example 1)
[0755] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0756] In modern virtual space systems and interactive character dialogue systems, it is difficult for users to engage in natural conversations and actions with characters within the virtual space. Furthermore, there is a lack of easy registration and login methods for users accessing the system for the first time, which leads to decreased user satisfaction. This invention aims to solve these problems and provide a system that allows users to interact with characters in real time and experience natural actions within the virtual space.
[0757] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0758] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for playing the voice response back to the user; registration means for the user to input registration information into the system; means for storing the registration information in a database; login means for the user to log in to the system; and means for matching the login information with information in the database. This allows the user to easily register and log in to the system, and subsequently engage in natural conversations with characters and interactive experiences in a virtual space in real time.
[0759] A "generative model" refers to a data model used to generate the voice of a character selected by the user.
[0760] "Voice recognition means" refers to technologies and devices for converting a user's voice into text data.
[0761] "Artificial intelligence means" refers to algorithms and programs for generating appropriate responses based on generative models.
[0762] "Speech synthesis means" refers to the technology that converts generated text data into the voice of a character.
[0763] "Means of playing voice responses to the user" refers to devices and technologies for playing the generated character's voice to the user.
[0764] "Registration method" refers to a series of processes and interfaces for a user to input their information into a system and for that information to be processed.
[0765] A "database" refers to a system used to store and manage user registration information and selection information.
[0766] "Login method" refers to the process or technology used to authenticate users when they access a system.
[0767] "Verification means" refers to technology used to compare login information entered by a user with registered information in a database and confirm that they match.
[0768] A "virtual space" refers to a digital environment where user and character avatars exist and which provides an interactive experience.
[0769] An "avatar" refers to a digital character that represents a user or character within a virtual space.
[0770] "Motion information" refers to information that detects the user's physical movements and captures them as digital data.
[0771] "Means of capturing" refers to devices and technologies for acquiring user behavior information in real time.
[0772] "Means of reflection" refers to technologies that reflect acquired user behavior information onto the movements of an avatar in a virtual space.
[0773] As a specific embodiment of this invention, a system is provided that allows users to interact in real time with characters or idols of their choice and enjoy an interactive experience in a virtual space.
[0774] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form for the user to enter their name, email address, and password. Once the user enters the required information and submits it, the device sends this data to the server, which receives the data and stores it in its database. Upon completion of registration, the device displays a registration completion message to the user.
[0775] Next, the user accesses the login page and enters their email address and password. The device sends this data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard. This process allows users to easily access the system.
[0776] When a user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays this list to the user, who then selects their desired character or idol. The device sends this selection information to the server, which records the selection information in the database and loads the associated voice model and animation data. This prepares the user to interact with their chosen character.
[0777] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello," the voice data is sent from the device to the server. The server uses a speech recognition service (e.g., Google Speech-to-Text) to convert the voice data into text. The server then sends the text to a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response text. The generated text is returned to the server, which uses a speech synthesis engine (e.g., Amazon Polly) to generate the character's voice. This generated voice is sent to the device and played back to the user. This process allows the user to experience a natural conversation with the character.
[0778] For example, if a user says "Hello, Naruto!", the generative AI model will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[0779] Furthermore, when a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on the VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character avatar to move naturally in response to the user's movements. Through this process, the user can enjoy a realistic interactive experience in the virtual space.
[0780] An example of a prompt message is, "I want to talk to Naruto for a bit. Turn on the microphone and speak to him." Based on this prompt, users can access the system from their smartphones or PCs and enjoy interacting with characters and experiencing virtual spaces.
[0781] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0782] Step 1:
[0783] Access the user's system registration page.
[0784] Input: The user enters the URL into their browser and accesses it.
[0785] Output: The registration page is displayed.
[0786] Specific operation: When a user enters the system's URL into their browser, the browser sends a request to the server, which generates a registration page and sends it back to the user's device.
[0787] Step 2:
[0788] The device displays a registration form to the user.
[0789] Input: Registration page data received from the server.
[0790] Output: The registration form is displayed on the screen.
[0791] Specific operation: The device uses the HTML and CSS included in the registration page to render the form on the user's screen.
[0792] Step 3:
[0793] The user enters the required information into the registration form and clicks the submit button.
[0794] Input: The action of a user entering their name, email address, and password and clicking the submit button.
[0795] Output: The input data is sent to the server.
[0796] Specific operation: The user enters, for example, "Taro Yamada", "example@mail.com", and "password123" into the form and clicks the "Submit" button. The input data is sent to the server as an HTTP POST request.
[0797] Step 4:
[0798] The server receives the input data and saves it to the database.
[0799] Input: User registration data received from the terminal.
[0800] Output: The registered data is saved in the database.
[0801] Specific operation: The server parses the received data and executes an INSERT query to save the data to the MySQL database.
[0802] Step 5:
[0803] The device displays a registration completion message to the user.
[0804] Input: Registration completion response from the server.
[0805] Output: A registration complete message is displayed on the screen.
[0806] Specific actions: The server notifies the terminal that saving to the database is complete and sends a response containing that information. The terminal displays the message "Registration complete!" on its screen.
[0807] Step 6:
[0808] The user accesses the login page and enters their email address and password.
[0809] Input: Access the login page URL, enter user information (email address, password).
[0810] Output: Your email address and password will be sent to the server.
[0811] Specific operation: The user enters and accesses the login page URL, enters "example@mail.com" and "password123" in the login form, and clicks the "Login" button. The data is sent via an HTTP POST request.
[0812] Step 7:
[0813] The server verifies the login information against the registered information in the database.
[0814] Input: Login data from the device.
[0815] Output: Login authentication result.
[0816] Specific operation: The server searches the "users" table in the database for a record with a matching email address and verifies the password. If a match is found, a login session ID is generated.
[0817] Step 8:
[0818] The server creates a login session and sends it to the terminal.
[0819] Input: Result indicating successful login.
[0820] Output: Session ID.
[0821] Specific operation: The server generates a session ID and sends it to the terminal.
[0822] Step 9:
[0823] The device displays the user's dashboard.
[0824] Input: Session ID response from the server.
[0825] Output: User's dashboard screen.
[0826] Specific action: The device receives the session ID, generates a dashboard page, and displays it on the screen. This page will include a message such as "Hello, Taro Yamada!".
[0827] Step 10:
[0828] The user clicks the "Select your favorite character / idol" option.
[0829] Input: User click operation.
[0830] Output: The request is sent to the server.
[0831] Specific operation: When the user clicks the "Select Favorite Character / Idol" button, the device sends an HTTP GET request to the server.
[0832] Step 11:
[0833] The server retrieves a list of character idols from the database and sends it to the terminal.
[0834] Input: Request from the terminal.
[0835] Output: Character / Idol List.
[0836] Specific operation: The server retrieves all records from the "characters" table and sends them to the terminal.
[0837] Step 12:
[0838] The device displays a list of characters / idols to the user.
[0839] Input: List data from the server.
[0840] Output: A list of character idols is displayed on the screen.
[0841] Specific operation: Based on the received list data, the terminal displays a list of characters / idols to the user.
[0842] Step 13:
[0843] The user selects their desired character or idol.
[0844] Input: User selection operation.
[0845] Output: The selected character / idol information is sent to the server.
[0846] Specific operation: The user clicks "Naruto" from the list, and the selection information is sent to the server via an HTTP POST request.
[0847] Step 14:
[0848] The server records the selection information in a database and loads the associated voice model and animation data.
[0849] Input: Character selection information from the device.
[0850] Output: Recording to a database, loading of voice models and animation data.
[0851] Specific operation: The server records the selection information in the "users" table and processes the loading of voice models and animation data related to "Naruto".
[0852] Step 15:
[0853] The user clicks the "Start Conversation" button.
[0854] Input: User click operation.
[0855] Output: The microphone turns on.
[0856] Specific operation: When the user clicks the "Start Conversation" button, the device turns on the microphone and prepares to capture audio.
[0857] Step 16:
[0858] The device captures the user's voice and sends it to the server.
[0859] Input: User's voice.
[0860] Output: Audio data is sent to the server.
[0861] Specific operation: The device records the user's voice saying "Hello, Naruto!" and sends that data to the server via an HTTP POST request.
[0862] Step 17:
[0863] The server uses a speech recognition service to convert the speech data into text.
[0864] Input: Audio data from the device.
[0865] Output: Text data.
[0866] Specific operation: The server uses the Google Speech-to-Text API to convert the audio data into the text "Hello, Naruto!".
[0867] Step 18:
[0868] The server sends text to the AI model and generates an appropriate response text.
[0869] Input: Text data converted by speech recognition.
[0870] Output: Response text data.
[0871] Specific operation: The server sends text data to OpenAI's GPT-3 and generates the response, "Hey! What's up?"
[0872] Step 19:
[0873] The server converts the response text into speech using a speech synthesis engine.
[0874] Input: Response text from a generating AI model.
[0875] Output: Audio data.
[0876] Specific operation: The server uses Amazon Polly to generate audio data that says, "Hey! What's up?"
[0877] Step 20:
[0878] The server sends the generated audio data to the terminal, and the terminal plays it for the user.
[0879] Input: Audio data from the server.
[0880] Output: The audio played to the user.
[0881] Specific action: The device plays the received Naruto voice, and the user hears the voice say, "Hey! What's up?"
[0882] Step 21:
[0883] The user selects the "Metaverse Mode" option.
[0884] Input: User click operation.
[0885] Output: The request is sent to the server.
[0886] Specific action: The user clicks the "Metaverse Mode" button on the dashboard, and the device sends an HTTP GET request to the server.
[0887] Step 22:
[0888] The server connects to the metaverse platform and loads information from the virtual space.
[0889] Input: Request from the terminal.
[0890] Output: Data in a virtual space.
[0891] Specific operation: The server uses the API of the metaverse platform (such as VRChat) to obtain the necessary virtual space data.
[0892] Step 23:
[0893] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[0894] Input: Data from the virtual space, user and character information.
[0895] Output: An avatar placed in a virtual space.
[0896] Specific operation: The server places the user's avatar and Naruto's avatar within VRChat.
[0897] Step 24:
[0898] The user puts on a VR device and accesses a virtual space.
[0899] Input: User action.
[0900] Output: Logged into the virtual space.
[0901] Specific actions: The user puts on an Oculus Quest, launches the VRChat app, and logs into the virtual space.
[0902] Step 25:
[0903] The device captures the user's movements and sends them to the server in real time.
[0904] Input: VR device operation information.
[0905] Output: Real-time operation data.
[0906] Specific operation: The terminal (VR device) captures the user's movements using sensors and transmits them to the server in real time.
[0907] Step 26:
[0908] The server issues instructions to make the character's avatar move naturally in response to the user's movement information.
[0909] Input: Real-time operation data from the device.
[0910] Output: Naruto's avatar moves within the virtual space.
[0911] Specific operation: Based on the received motion data, the server instructs the Naruto avatar to perform appropriate actions within VRChat, making it move naturally in real time in accordance with the user's movements.
[0912] Through these steps, users can register and log in to the system, interact with characters and idols, and enjoy real-time interactive experiences within the virtual space.
[0913] (Application Example 1)
[0914] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0915] Conventional voice interaction systems have limited functionality for users to interact with characters or idols, and providing a truly interactive experience in real time presented many challenges. Furthermore, character-like behavior and real-time reflection of actions in virtual space were difficult, preventing users from enjoying a consistent entertainment experience. Moreover, no system existed that integrated the entire process from speech recognition to generational AI and speech synthesis to achieve real-time responses within a virtual space.
[0916] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0917] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for displaying the voice response in a virtual space; means for controlling the character's actions in the virtual space; and means for playing the voice response back to the user. This enables the user to interact with characters and idols in real time in a virtual space and enjoy an interactive and consistent experience.
[0918] A "user" is a person who uses this system to interact with characters and idols in a virtual space.
[0919] "Character voice" refers to voice data corresponding to a specific character, and is used to reproduce the voice of the character selected by the user.
[0920] A "generative model" is a machine learning model designed to generate the voice of a character selected by the user.
[0921] "Speech recognition means" is a general term for hardware and software used to convert a user's voice into text data.
[0922] "Artificial intelligence means" refers to machine learning algorithms that generate appropriate responses from input text based on a generative model.
[0923] "Speech synthesis means" is a general term for hardware and software used to convert generated text responses into character voices.
[0924] A "virtual space" is a three-dimensional computer graphics space where users and characters engage in dialogue and interaction.
[0925] "Means for displaying voice responses" refers to the collective term for hardware and software used to visually represent voice responses generated within a virtual space.
[0926] "Means for controlling character-like behavior" refers to a general term for algorithms and systems used to control how characters in a virtual space behave naturally.
[0927] "Means for playing back voice responses" refers to an audio output device and associated software that allows the user to hear the character's voice responses.
[0928] One embodiment for carrying out this invention will be described. This system enables users to interact with characters and idols in a virtual space in real time and enjoy an interactive experience.
[0929] Hardware and software to be used
[0930] The main hardware used to carry out this invention includes:
[0931] Terminal: A device used by a user to access the system (smartphone, PC, head-mounted display)
[0932] Server: The central hub for speech recognition, generative AI, speech synthesis, and virtual space control.
[0933] Voice input device: A microphone that captures the user's voice.
[0934] Audio output device: A speaker that plays the character's voice responses.
[0935] The main software includes:
[0936] Speech recognition software: Google's speech recognition API
[0937] Generative AI Model: OpenAI GPT-3 Model
[0938] Speech synthesis software: pyttsx3 library
[0939] Virtual space management software: pyqtmetaverse library
[0940] Data processing and data calculation
[0941] 1. User voice input
[0942] The user speaks into the device. The voice input device captures the user's voice and sends it to the device.
[0943] 2. Speech recognition means
[0944] The device passes the captured audio data to speech recognition software, which then converts it into text data.
[0945] 3. Artificial intelligence tools
[0946] The server passes text data to the generative AI model, which then generates appropriate response text. This process creates prompt text for the generative AI.
[0947] Example prompt: "Generate a character's response to the question, 'What is the best product in this store?'"
[0948] 4. Speech synthesis means
[0949] The server passes the generated text response to speech synthesis software, which converts it into a character's voice.
[0950] 5. Display in virtual space
[0951] The virtual space management software displays the generated voice responses to the user in the virtual space both visually and audibly. It also controls the character's movements to allow the user to interact naturally.
[0952] 6. Playback of voice response
[0953] The device plays character voices to the user, providing an interactive dialogue experience.
[0954] Specific example
[0955] Suppose a user asks, "What is the best item in this store?" The server captures this audio and converts it into text, "What is the best item in this store?", through speech recognition software. Based on the prompt, "Generate a character response to 'What is the best item in this store?'", the generative AI model generates the appropriate response text, "Then it's the automatic ramen cooker! It's convenient and you can enjoy delicious ramen anytime." Speech synthesis software converts this response text into the character's voice. Finally, virtual space management software displays this voice within the virtual space, and the terminal plays the voice for the user.
[0956] This allows users to enjoy natural, real-time interactions with characters and idols in a virtual space.
[0957] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0958] Step 1:
[0959] The user speaks into a voice input device (microphone). The terminal captures this audio and prepares the data.
[0960] Input: User's voice
[0961] Output: Audio data
[0962] Specific operation: When a user says, "What is the best product you recommend in this store?", the microphone captures the audio.
[0963] Step 2:
[0964] The device passes the captured audio data to speech recognition software (Google's speech recognition API) and converts it into text data.
[0965] Input: Audio data
[0966] Output: Text data
[0967] Specific operation: The device sends the captured audio data to Google's speech recognition API, which converts it into the text, "What is the best product in this store?"
[0968] Step 3:
[0969] The server receives text data and passes it to a generative AI model (OpenAI GPT-3 model) to generate appropriate response text. During this process, it generates a prompt and sends it to the AI.
[0970] Input: Text data
[0971] Output: Response text
[0972] Specific operation: Based on the text "What is the best recommended item in this store?", the server generates a prompt message "Generate a character response to 'What is the best recommended item in this store?'". The generating AI model receives this and generates the response text "Then it's the automatic ramen cooker. It's convenient, and you can enjoy delicious ramen anytime."
[0973] Step 4:
[0974] The server passes the generated response text to speech synthesis software (pyttsx3) and converts it into a character's voice.
[0975] Input: Response text
[0976] Output: Audio data
[0977] Specific operation: The server passes the text "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime." to pyttsx3, which then converts it into character voice.
[0978] Step 5:
[0979] The server uses virtual space management software (pyqtmetaverse) to display the generated audio data within the virtual space and control the character's movements.
[0980] Input: Voice data, response text
[0981] Output: Character movements and sound display in the virtual space
[0982] Specific operation: The server passes the generated audio data and response text to the pyqtmetaverse, which then controls the movement of the character in the virtual space and displays it to the user visually and audibly.
[0983] Step 6:
[0984] The device plays the generated character's voice to the user through its speaker.
[0985] Input: Audio data
[0986] Output: Audio to be played
[0987] Specific operation: The device plays a voice message from a character through its speaker saying, "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime," to the user.
[0988] Through the above processing steps, users can enjoy real-time, interactive conversations with characters and idols in a virtual space.
[0989] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0990] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space. It also has a function to recognize the user's emotions and adjust its response accordingly.
[0991] User registration and login
[0992] The user first accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which then receives the data and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[0993] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[0994] Choosing your favorite character or idol
[0995] When the user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records this information in the database and loads the associated voice model and animation data.
[0996] Initiating conversations with AI and emotion recognition
[0997] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", that voice data is sent from the device to the server. The server receives the voice data and uses a speech recognition service to convert the voice to text. At the same time, an emotion engine analyzes the voice and the user's facial expressions to recognize the user's emotions (e.g., joy, sadness, surprise, etc.).
[0998] The server sends text data and recognized emotion information to the generating AI and requests it to generate an appropriate response. The generating AI generates an appropriate response text based on the input text and emotion information and returns it to the server. The server processes the response text through a speech synthesis engine and generates audio in the character's voice. This audio data is sent to the terminal, which then plays the generated audio for the user.
[0999] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generating AI will produce the response "Hi! You seem to be doing well!" which will then be played back to the user in Naruto's voice.
[1000] Providing a sense of realism in the metaverse
[1001] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. Conversation with the generating AI also takes place in parallel, and voice and corresponding animations are reproduced in the virtual space.
[1002] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space. This allows users to not only have a realistic interactive experience in the virtual space but also receive responses that are appropriate to their emotions.
[1003] This invention allows users to enjoy interacting with their favorite characters or idols while gaining a realistic experience even in a virtual space. Emotion recognition improves the quality of the interaction, enabling deeper engagement.
[1004] The following describes the processing flow.
[1005] Step 1:
[1006] Users access the system's registration page from their smartphones or PCs.
[1007] Step 2:
[1008] The device displays a registration form to the user.
[1009] Step 3:
[1010] The user enters their name, email address, and password and submits the form.
[1011] Step 4:
[1012] The terminal sends the input data to the server.
[1013] Step 5:
[1014] The server receives the transmitted data and checks for formatting and duplicates.
[1015] Step 6:
[1016] After the server verifies the data, it saves it to the database.
[1017] Step 7:
[1018] The server sends a registration completion message to the device.
[1019] Step 8:
[1020] The device displays a registration completion message to the user.
[1021] Step 9:
[1022] The user accesses the login page and enters their email address and password.
[1023] Step 10:
[1024] The terminal sends the input data to the server.
[1025] Step 11:
[1026] The server compares the information with the registered data in the database.
[1027] Step 12:
[1028] The server matches the user information and creates a login session.
[1029] Step 13:
[1030] The server sends the login session to the terminal.
[1031] Step 14:
[1032] The device displays the user's dashboard.
[1033] Step 15:
[1034] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[1035] Step 16:
[1036] The terminal sends a request to the server.
[1037] Step 17:
[1038] The server retrieves a list of available characters / idols from the database.
[1039] Step 18:
[1040] The server sends the list to the terminal.
[1041] Step 19:
[1042] The device displays the list to the user.
[1043] Step 20:
[1044] The user selects their desired character or idol.
[1045] Step 21:
[1046] The device sends the selection information to the server.
[1047] Step 22:
[1048] The server records the selection information in a database and loads the associated voice model and animation data.
[1049] Step 23:
[1050] The user clicks the "Start Conversation" button.
[1051] Step 24:
[1052] The device turns on the microphone and captures the user's voice and facial expressions.
[1053] Step 25:
[1054] The device sends the captured audio and facial expression data to the server.
[1055] Step 26:
[1056] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[1057] Step 27:
[1058] The server uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotions.
[1059] Step 28:
[1060] The server sends text data and recognized emotional information to the AI and requests it to generate a response.
[1061] Step 29:
[1062] The generation AI generates appropriate response text based on the input text and sentiment information, and sends it to the server.
[1063] Step 30:
[1064] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[1065] Step 31:
[1066] The server sends the generated audio file to the terminal.
[1067] Step 32:
[1068] The device plays generated audio and visual information for the user.
[1069] Step 33:
[1070] The user selects the "Metaverse Mode" option.
[1071] Step 34:
[1072] The terminal sends a request to the server.
[1073] Step 35:
[1074] The server connects to the metaverse platform and loads information from the virtual space.
[1075] Step 36:
[1076] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[1077] Step 37:
[1078] The user puts on a VR device and accesses a virtual space.
[1079] Step 38:
[1080] The device captures the user's movements and sends them to the server in real time.
[1081] Step 39:
[1082] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[1083] Step 40:
[1084] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space.
[1085] Step 41:
[1086] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[1087] Step 42:
[1088] The device provides an interactive experience between the user and their favorite character or idol.
[1089] (Example 2)
[1090] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1091] Conventional voice dialogue systems simply provide monotonous responses to user input, making it difficult to recognize emotions and engage in interactive conversations. Furthermore, even in virtual reality interactions, real-time control based on user actions is insufficient, failing to provide a realistic experience. This resulted in a decline in the quality of the user experience.
[1092] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1093] In this invention, the server includes means for providing a generative model for generating the voice of a virtual character selected by the user, speech recognition means for converting the user's voice into text, and artificial intelligence means for generating a response based on the text and emotional information. This makes it possible to recognize the user's emotions and adjust the response accordingly.
[1094] A "user" is an end-user who uses the system to interact with others and have experiences in a virtual space.
[1095] A "virtual character" is a virtual character or idol that users can choose to use in conversations and interactive experiences.
[1096] A "generative model" is an artificial intelligence model used to generate the voice and behavior of a virtual character selected by the user.
[1097] "Speech recognition means" refers to technologies and devices for converting a user's speech into text.
[1098] "Artificial intelligence means" refers to technologies and devices that generate appropriate responses based on text data and emotional information.
[1099] "Voice synthesis means" refers to technologies and devices for converting generated responses into the voice of a virtual character.
[1100] "Emotion recognition means" refers to technologies and devices that recognize emotions from a user's voice and facial expressions and adjust responses based on that information.
[1101] "Means for reproducing voice responses" refers to technologies or devices for reproducing generated voice responses to the user.
[1102] An "avatar" is a graphical representation displayed in a virtual space as a digital representation of a user or virtual character.
[1103] A "virtual space" is a computer-generated three-dimensional space that users can access through VR devices or similar means.
[1104] "Motion information" refers to data about the user's body movements and is used to control the avatar in the virtual space.
[1105] "Real-time" refers to a sense of time in which processing and responses occur almost instantly, meaning that delays to user operations and inputs are kept to a minimum.
[1106] Embodiments of this invention will now be described in detail. This system aims to allow users to interact with their favorite virtual characters or idols in real time and enjoy an interactive experience in a virtual space. It also has the function of recognizing the user's emotions and adjusting its response accordingly.
[1107] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who enters necessary information such as their name, email address, and password, and sends it from the device to the server. The server stores the received data in a database and notifies the user via the device that registration is complete. Next, the user accesses the login page and enters the registered email address and password. The device sends the entered data to the server, which verifies it against the information in the database. If authentication is successful, the server generates a login session and sends it to the device, displaying the user's dashboard.
[1108] When a user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of characters and idols from the database and sends it to the device. The device displays this list to the user, and when the user selects their desired character or idol, that information is sent to the server and recorded in the database.
[1109] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the audio data is sent from the device to the server. The server converts this audio data to text using the Google Cloud Speech-to-Text API and analyzes the text and emotion information with an emotion engine (e.g., Microsoft Azure Emotion API). If the analysis result is positive, the server sends the text data and emotion information to a generating AI model (e.g., OpenAI GPT-3) to generate a response such as "Hi! You seem to be doing well!". The generated response text is then synthesized into a character's voice using Amazon Polly and sent to the device for playback.
[1110] Furthermore, when the user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to a Unity-based metaverse platform and loads information for the virtual space. The server places the user's avatar and character / idol avatars in the virtual space, and the user accesses the virtual space by wearing a VR device (e.g., Oculus Rift). The device captures the user's movements in real time and sends them to the server. The server controls the movements of the virtual avatar through the Unity engine and simultaneously provides a more natural and realistic interactive experience through voice interaction via a generated AI model. In addition, an emotion engine recognizes the user's emotional changes in real time and reflects that information in the movements and facial expressions of the virtual avatar.
[1111] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generative AI model will generate the response "Hi! You seem to be doing well!". This response is synthesized into Naruto's voice and played back to the user through the speaker.
[1112] Examples of prompts for operating this system include:
[1113] Prompt to convert audio data to text:
[1114] Please convert the audio data 'Hello, Naruto!' to text.
[1115] Response generation prompt:
[1116] "Based on the text 'Hello, Naruto!' and the emotion 'Joyful,' please generate an appropriate response."
[1117] Text-to-speech prompt:
[1118] "Please convert the text 'Hey! You seem to be doing well!' into speech in Naruto's voice."
[1119] Through these processes, users can gain real-time interaction and an interactive experience in a virtual space.
[1120] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1121] Program processing flow
[1122] Step 1:
[1123] Users access the system's registration page from their smartphones or PCs.
[1124] Input: Name, email address, password
[1125] The terminal receives these inputs, displays them on the registration form, and waits for user input.
[1126] Output: Input data
[1127] The user enters their name, email address, and password, and this data is sent to the device.
[1128] Step 2:
[1129] The terminal sends the input data to the server.
[1130] Input: User registration information (name, email address, password)
[1131] The server receives the input data and saves it to the database.
[1132] Output: Registration complete message
[1133] The device displays "Registration complete!" to the user.
[1134] Step 3:
[1135] The user accesses the login page and enters their email address and password.
[1136] Enter: Email address, password
[1137] The terminal sends the input data to the server.
[1138] Output: Authentication result
[1139] The server compares the information with the registration details in the database, and if a match is found, it generates a login session.
[1140] A login session is sent to the device, and the device displays the user's dashboard.
[1141] Step 4:
[1142] The user clicks the "Select Favorite Character / Idol" button in the main menu.
[1143] Input: Click Event
[1144] The terminal sends a request to the server.
[1145] Output: List of Characters / Idols
[1146] The server retrieves a list of characters and idols from the database and sends it to the terminal.
[1147] The device displays the list to the user.
[1148] Step 5:
[1149] The user selects their desired character or idol.
[1150] Input: Selected character or idol
[1151] The device sends the selection information to the server.
[1152] Output: Database update results
[1153] The server records that information in the database.
[1154] Step 6:
[1155] The user clicks the "Start Conversation" button.
[1156] Input: Click Event
[1157] The device turns on the microphone and captures the user's voice.
[1158] The user says, "Hello, [Character Name]!"
[1159] Output: Audio data
[1160] The audio data is then sent from the terminal to the server.
[1161] Step 7:
[1162] The server sends the voice data to the speech recognition service.
[1163] Input: Audio data
[1164] The speech recognition service converts speech into text.
[1165] Output: Text data
[1166] The server receives this text data.
[1167] Step 8:
[1168] The server recognizes emotions from voice and user facial expression data through an emotion engine.
[1169] Input: Voice data, user facial expression data
[1170] Output: Emotional information
[1171] The server uses text data and recognized emotion information to send to a generative AI model.
[1172] Step 9:
[1173] The generative AI model generates appropriate responses based on text data and sentiment information.
[1174] Input: Text data, sentiment information
[1175] Output: Response text
[1176] The server receives the response text.
[1177] Step 10:
[1178] The server processes the response text into a speech synthesis engine.
[1179] Input: Response text
[1180] The speech synthesis engine generates speech using the voice of a virtual character.
[1181] Output: Generated audio data
[1182] The server sends the audio data to the terminal.
[1183] Step 11:
[1184] The device plays the generated audio to the user.
[1185] Input: Audio data
[1186] The device plays audio through its speaker, and the user listens to the response.
[1187] Output: Played audio
[1188] Step 12:
[1189] The user selects the "Metaverse Mode" option.
[1190] Input: Click Event
[1191] The terminal sends a request to the server.
[1192] The server connects to the metaverse platform and loads the virtual space.
[1193] Output: Display data for the virtual space
[1194] Step 13:
[1195] The server places user avatars and character / idol avatars in the virtual space.
[1196] Input: None (automatic processing)
[1197] Users wear VR devices to access a virtual space.
[1198] Output: Initial placement of user and character avatars
[1199] Step 14:
[1200] The device captures the user's movements and sends them to the server.
[1201] Input: User activity information
[1202] The server controls the avatar in real time via the Unity engine.
[1203] Output: Avatar movement data
[1204] This allows users to enjoy interactive experiences in a virtual space.
[1205] The above is the specific processing flow of this system.
[1206] (Application Example 2)
[1207] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1208] Conventional virtual space systems have limited interaction between users and virtual characters or idols, lacking interaction based on emotions and actions. Therefore, users find it difficult to have a realistic conversational experience, resulting in a less realistic virtual experience overall. Furthermore, the lack of interactive dialogue systems utilizing emotion recognition is another challenge.
[1209] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for providing a generation model for generating the voice of a virtual character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating an appropriate response based on the generation model; speech synthesis means for converting the generated response into the voice of a virtual character; emotion recognition means for analyzing the user's emotions; means for adjusting the response based on the emotion recognition results; and means for playing the generated voice response to the user. This makes it possible for the user to enjoy a realistic dialogue experience with their favorite virtual character or idol that responds to their emotions.
[1210] A "user" is someone who uses the system to enjoy interacting with virtual characters or idols.
[1211] A "virtual space" is a three-dimensional virtual environment created by a computer program.
[1212] A "virtual character" is a character or idol selected by a user within a virtual space, whose voice and actions are generated based on AI.
[1213] A "generative model" refers to the algorithms and datasets used to generate voices and actions for virtual characters.
[1214] "Speech recognition means" refers to a technical device or program for generating text from a user's speech.
[1215] An "artificial intelligence system" is a system that uses a pre-trained model to generate an appropriate response to user input.
[1216] "Speech synthesis means" refers to a device or program that converts text data into the voice of a virtual character and generates speech data.
[1217] "Emotion recognition means" refers to technologies and programs that analyze and recognize emotions from a user's voice, facial expressions, etc.
[1218] "Means for adjusting responses" refer to technologies or programs for appropriately modifying responses generated based on the user's emotion recognition results.
[1219] "Means for reproducing voice responses" refers to devices or programs that output generated voice data to the user.
[1220] The embodiments for carrying out this invention are described in detail below.
[1221] First, the system consists of the following main components:
[1222] 1. A generative model for generating voices for virtual characters based on user selections.
[1223] 2. Speech recognition means
[1224] 3. Artificial intelligence means for generating appropriate responses
[1225] 4. Speech synthesis means for converting the generated response into speech.
[1226] 5. Emotion recognition means for analyzing user emotions
[1227] 6. Means for adjusting responses based on emotion recognition results
[1228] 7. Means for playing voice responses to the user
[1229] Specifically, when a user accesses the system and enters their email address and password on the login page, the server compares this information with existing data in the database. If the login is successful, the user's dashboard is displayed. The user then selects their preferred virtual character from the dashboard, and the server loads the character's voice model and animation data.
[1230] Hardware and software configuration
[1231] The hardware and software used in this invention are as follows:
[1232] Hardware: Smartphone, tablet, or VR headset, microphone, speaker
[1233] software:
[1234] Google Cloud Speech-to-Text: Used as a speech recognition method.
[1235] Azure Text Analytics: Used as a means of emotion recognition
[1236] OpenAI GPT-3: Used as an artificial intelligence means to generate appropriate responses.
[1237] Amazon Polly: Used as a speech synthesis method.
[1238] Data adjustment and calculation flow
[1239] 1. Speech recognition:
[1240] The audio data spoken by the user through the microphone is sent to the server and converted into text using Google Cloud Speech-to-Text.
[1241] 2. Emotion recognition:
[1242] The data, converted to text, is analyzed by Azure Text Analytics to recognize the user's sentiment.
[1243] 3. Response generation:
[1244] The server inputs the recognized emotion and text data into OpenAI GPT-3 and generates an appropriate response text.
[1245] 4. Speech synthesis:
[1246] The generated response text is converted into a virtual character's voice by Amazon Polly.
[1247] 5. Playback of voice response:
[1248] Audio data generated from the server is sent to the user's device, and the device plays the audio.
[1249] Specific example
[1250] For example, a user says "Hello, Naruto!" into the microphone. This audio is sent to the server. Google Cloud Speech-to-Text converts the audio into text "Hello, Naruto!". Next, Azure Text Analytics recognizes the user's emotion from this text as "joy". The server sends this recognized emotion and text to OpenAI GPT-3 to generate an appropriate response, "Hi! You seem to be doing well!". This response is converted into speech by Amazon Polly and played back on the user's device in the voice of the virtual character Naruto.
[1251] Example of a prompt
[1252] Joyful user: Hello, Naruto! Virtual character:
[1253] In this way, the system of the present invention enables users to enjoy a realistic dialogue experience that responds to their emotions.
[1254] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1255] Step 1:
[1256] The server receives the email address and password entered by the user and compares them with existing information in the database. If the entered email address and password match, the server creates a login session and displays the user's dashboard.
[1257] Step 2:
[1258] The user selects a virtual character from the dashboard. The device sends the selection information to the server, which loads the corresponding voice model and animation data from the database and presents it to the user.
[1259] Step 3:
[1260] When the user presses the "Start Conversation" button, the device turns on the microphone and captures the user's voice data. This voice data is sent to the server. The server uses Google Cloud Speech-to-Text to convert this voice data into text.
[1261] Step 4:
[1262] The server sends the converted text data to Azure Text Analytics for user sentiment analysis. The sentiment analysis results are stored along with the text data.
[1263] Step 5:
[1264] The server sends text data and emotion recognition results to OpenAI GPT-3 to generate an appropriate response. Specifically, natural-sounding dialogue is generated based on the prompt text and the recognized emotion information.
[1265] Step 6:
[1266] The generated response text is converted into audio data by Amazon Polly. The server then sends the generated audio data to the device.
[1267] Step 7:
[1268] The user's device plays the received audio data. This provides the user with the experience of a virtual character responding to them verbally.
[1269] Step 8:
[1270] User movement information is also captured in real time and sent from the terminal to the server. The server controls the movement of the avatar in the virtual space and adjusts the avatar to respond to the user's movements.
[1271] This allows users to enjoy realistic, emotion-driven conversations and action-based interactions within the virtual space.
[1272] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1273] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1274] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1275] [Third Embodiment]
[1276] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1277] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1278] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1279] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1280] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1281] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1282] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1283] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1284] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1285] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1286] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1287] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1288] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space.
[1289] User registration and login
[1290] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which receives it and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[1291] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[1292] Choosing your favorite character or idol
[1293] When the user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records the selection information in the database and loads the associated voice model and animation data.
[1294] Starting a conversation with AI
[1295] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the voice data is sent from the device to the server. The server uses a speech recognition service to convert the voice data into text and sends the text to a generative AI. The generative AI generates an appropriate response text based on the input text, which is then returned to the server. The server then uses a speech synthesis engine to convert the response text into speech, generating the voice in the character's voice. This voice data is sent to the device, which then plays the voice for the user.
[1296] For example, if the user says "Hello, Naruto!", the AI will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[1297] Providing a sense of realism in the metaverse
[1298] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. This allows the user to enjoy a realistic interactive experience in the virtual space.
[1299] The following describes the processing flow.
[1300] Step 1:
[1301] Users access the system's registration page from their smartphones or PCs.
[1302] Step 2:
[1303] The device displays a registration form to the user.
[1304] Step 3:
[1305] The user enters their name, email address, and password and submits the form.
[1306] Step 4:
[1307] The terminal sends the input data to the server.
[1308] Step 5:
[1309] The server receives the transmitted data and checks for formatting and duplicates.
[1310] Step 6:
[1311] After the server verifies the data, it saves it to the database.
[1312] Step 7:
[1313] The server sends a registration completion message to the device.
[1314] Step 8:
[1315] The device displays a registration completion message to the user.
[1316] Step 9:
[1317] The user accesses the login page and enters their email address and password.
[1318] Step 10:
[1319] The terminal sends the input data to the server.
[1320] Step 11:
[1321] The server compares the information with the registered data in the database.
[1322] Step 12:
[1323] The server matches the user information and creates a login session.
[1324] Step 13:
[1325] The server sends the login session to the terminal.
[1326] Step 14:
[1327] The device displays the user's dashboard.
[1328] Step 15:
[1329] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[1330] Step 16:
[1331] The terminal sends a request to the server.
[1332] Step 17:
[1333] The server retrieves a list of available characters / idols from the database.
[1334] Step 18:
[1335] The server sends the list to the terminal.
[1336] Step 19:
[1337] The device displays the list to the user.
[1338] Step 20:
[1339] The user selects their desired character or idol.
[1340] Step 21:
[1341] The device sends the selection information to the server.
[1342] Step 22:
[1343] The server records the selection information in a database and loads the associated voice model and animation data.
[1344] Step 23:
[1345] The user clicks the "Start Conversation" button.
[1346] Step 24:
[1347] The device turns on the microphone and captures the user's voice.
[1348] Step 25:
[1349] The device sends the captured audio data to the server.
[1350] Step 26:
[1351] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[1352] Step 27:
[1353] The server sends text data to the AI generating the response and requests that it generate a response.
[1354] Step 28:
[1355] The AI generates appropriate response text and sends it to the server.
[1356] Step 29:
[1357] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[1358] Step 30:
[1359] The server sends the generated audio file to the terminal.
[1360] Step 31:
[1361] The device plays the generated audio for the user.
[1362] Step 32:
[1363] The user selects the "Metaverse Mode" option.
[1364] Step 33:
[1365] The terminal sends a request to the server.
[1366] Step 34:
[1367] The server connects to the metaverse platform and loads information from the virtual space.
[1368] Step 35:
[1369] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[1370] Step 36:
[1371] The user puts on a VR device and accesses a virtual space.
[1372] Step 37:
[1373] The device captures the user's movements and sends them to the server in real time.
[1374] Step 38:
[1375] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[1376] Step 39:
[1377] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[1378] Step 40:
[1379] The device provides an interactive experience between the user and their favorite character or idol.
[1380] (Example 1)
[1381] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1382] In modern virtual space systems and interactive character dialogue systems, it is difficult for users to engage in natural conversations and actions with characters within the virtual space. Furthermore, there is a lack of easy registration and login methods for users accessing the system for the first time, which leads to decreased user satisfaction. This invention aims to solve these problems and provide a system that allows users to interact with characters in real time and experience natural actions within the virtual space.
[1383] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1384] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for playing the voice response back to the user; registration means for the user to input registration information into the system; means for storing the registration information in a database; login means for the user to log in to the system; and means for matching the login information with information in the database. This allows the user to easily register and log in to the system, and subsequently engage in natural conversations with characters and interactive experiences in a virtual space in real time.
[1385] A "generative model" refers to a data model used to generate the voice of a character selected by the user.
[1386] "Voice recognition means" refers to technologies and devices for converting a user's voice into text data.
[1387] "Artificial intelligence means" refers to algorithms and programs for generating appropriate responses based on generative models.
[1388] "Speech synthesis means" refers to the technology that converts generated text data into the voice of a character.
[1389] "Means of playing voice responses to the user" refers to devices and technologies for playing the generated character's voice to the user.
[1390] "Registration method" refers to a series of processes and interfaces for a user to input their information into a system and for that information to be processed.
[1391] A "database" refers to a system used to store and manage user registration information and selection information.
[1392] "Login method" refers to the process or technology used to authenticate users when they access a system.
[1393] "Verification means" refers to technology used to compare login information entered by a user with registered information in a database and confirm that they match.
[1394] A "virtual space" refers to a digital environment where user and character avatars exist and which provides an interactive experience.
[1395] An "avatar" refers to a digital character that represents a user or character within a virtual space.
[1396] "Motion information" refers to information that detects the user's physical movements and captures them as digital data.
[1397] "Means of capturing" refers to devices and technologies for acquiring user behavior information in real time.
[1398] "Means of reflection" refers to technologies that reflect acquired user behavior information onto the movements of an avatar in a virtual space.
[1399] As a specific embodiment of this invention, a system is provided that allows users to interact in real time with characters or idols of their choice and enjoy an interactive experience in a virtual space.
[1400] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form for the user to enter their name, email address, and password. Once the user enters the required information and submits it, the device sends this data to the server, which receives the data and stores it in its database. Upon completion of registration, the device displays a registration completion message to the user.
[1401] Next, the user accesses the login page and enters their email address and password. The device sends this data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard. This process allows users to easily access the system.
[1402] When a user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays this list to the user, who then selects their desired character or idol. The device sends this selection information to the server, which records the selection information in the database and loads the associated voice model and animation data. This prepares the user to interact with their chosen character.
[1403] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello," the voice data is sent from the device to the server. The server uses a speech recognition service (e.g., Google Speech-to-Text) to convert the voice data into text. The server then sends the text to a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response text. The generated text is returned to the server, which uses a speech synthesis engine (e.g., Amazon Polly) to generate the character's voice. This generated voice is sent to the device and played back to the user. This process allows the user to experience a natural conversation with the character.
[1404] For example, if a user says "Hello, Naruto!", the generative AI model will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[1405] Furthermore, when a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on the VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character avatar to move naturally in response to the user's movements. Through this process, the user can enjoy a realistic interactive experience in the virtual space.
[1406] An example of a prompt message is, "I want to talk to Naruto for a bit. Turn on the microphone and speak to him." Based on this prompt, users can access the system from their smartphones or PCs and enjoy interacting with characters and experiencing virtual spaces.
[1407] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1408] Step 1:
[1409] Access the user's system registration page.
[1410] Input: The user enters the URL into their browser and accesses it.
[1411] Output: The registration page is displayed.
[1412] Specific operation: When a user enters the system's URL into their browser, the browser sends a request to the server, which generates a registration page and sends it back to the user's device.
[1413] Step 2:
[1414] The device displays a registration form to the user.
[1415] Input: Registration page data received from the server.
[1416] Output: The registration form is displayed on the screen.
[1417] Specific operation: The device uses the HTML and CSS included in the registration page to render the form on the user's screen.
[1418] Step 3:
[1419] The user enters the required information into the registration form and clicks the submit button.
[1420] Input: The action of a user entering their name, email address, and password and clicking the submit button.
[1421] Output: The input data is sent to the server.
[1422] Specific operation: The user enters, for example, "Taro Yamada", "example@mail.com", and "password123" into the form and clicks the "Submit" button. The input data is sent to the server as an HTTP POST request.
[1423] Step 4:
[1424] The server receives the input data and saves it to the database.
[1425] Input: User registration data received from the terminal.
[1426] Output: The registered data is saved in the database.
[1427] Specific operation: The server parses the received data and executes an INSERT query to save the data to the MySQL database.
[1428] Step 5:
[1429] The device displays a registration completion message to the user.
[1430] Input: Registration completion response from the server.
[1431] Output: A registration complete message is displayed on the screen.
[1432] Specific actions: The server notifies the terminal that saving to the database is complete and sends a response containing that information. The terminal displays the message "Registration complete!" on its screen.
[1433] Step 6:
[1434] The user accesses the login page and enters their email address and password.
[1435] Input: Access the login page URL, enter user information (email address, password).
[1436] Output: Your email address and password will be sent to the server.
[1437] Specific operation: The user enters and accesses the login page URL, enters "example@mail.com" and "password123" in the login form, and clicks the "Login" button. The data is sent via an HTTP POST request.
[1438] Step 7:
[1439] The server verifies the login information against the registered information in the database.
[1440] Input: Login data from the device.
[1441] Output: Login authentication result.
[1442] Specific operation: The server searches the "users" table in the database for a record with a matching email address and verifies the password. If a match is found, a login session ID is generated.
[1443] Step 8:
[1444] The server creates a login session and sends it to the terminal.
[1445] Input: Result indicating successful login.
[1446] Output: Session ID.
[1447] Specific operation: The server generates a session ID and sends it to the terminal.
[1448] Step 9:
[1449] The device displays the user's dashboard.
[1450] Input: Session ID response from the server.
[1451] Output: User's dashboard screen.
[1452] Specific action: The device receives the session ID, generates a dashboard page, and displays it on the screen. This page will include a message such as "Hello, Taro Yamada!".
[1453] Step 10:
[1454] The user clicks the "Select your favorite character / idol" option.
[1455] Input: User click operation.
[1456] Output: The request is sent to the server.
[1457] Specific operation: When the user clicks the "Select Favorite Character / Idol" button, the device sends an HTTP GET request to the server.
[1458] Step 11:
[1459] The server retrieves a list of character idols from the database and sends it to the terminal.
[1460] Input: Request from the terminal.
[1461] Output: Character / Idol List.
[1462] Specific operation: The server retrieves all records from the "characters" table and sends them to the terminal.
[1463] Step 12:
[1464] The device displays a list of characters / idols to the user.
[1465] Input: List data from the server.
[1466] Output: A list of character idols is displayed on the screen.
[1467] Specific operation: Based on the received list data, the terminal displays a list of characters / idols to the user.
[1468] Step 13:
[1469] The user selects their desired character or idol.
[1470] Input: User selection operation.
[1471] Output: The selected character / idol information is sent to the server.
[1472] Specific operation: The user clicks "Naruto" from the list, and the selection information is sent to the server via an HTTP POST request.
[1473] Step 14:
[1474] The server records the selection information in a database and loads the associated voice model and animation data.
[1475] Input: Character selection information from the device.
[1476] Output: Recording to a database, loading of voice models and animation data.
[1477] Specific operation: The server records the selection information in the "users" table and processes the loading of voice models and animation data related to "Naruto".
[1478] Step 15:
[1479] The user clicks the "Start Conversation" button.
[1480] Input: User click operation.
[1481] Output: The microphone turns on.
[1482] Specific operation: When the user clicks the "Start Conversation" button, the device turns on the microphone and prepares to capture audio.
[1483] Step 16:
[1484] The device captures the user's voice and sends it to the server.
[1485] Input: User's voice.
[1486] Output: Audio data is sent to the server.
[1487] Specific operation: The device records the user's voice saying "Hello, Naruto!" and sends that data to the server via an HTTP POST request.
[1488] Step 17:
[1489] The server uses a speech recognition service to convert the speech data into text.
[1490] Input: Audio data from the device.
[1491] Output: Text data.
[1492] Specific operation: The server uses the Google Speech-to-Text API to convert the audio data into the text "Hello, Naruto!".
[1493] Step 18:
[1494] The server sends text to the AI model and generates an appropriate response text.
[1495] Input: Text data converted by speech recognition.
[1496] Output: Response text data.
[1497] Specific operation: The server sends text data to OpenAI's GPT-3 and generates the response, "Hey! What's up?"
[1498] Step 19:
[1499] The server converts the response text into speech using a speech synthesis engine.
[1500] Input: Response text from a generating AI model.
[1501] Output: Audio data.
[1502] Specific operation: The server uses Amazon Polly to generate audio data that says, "Hey! What's up?"
[1503] Step 20:
[1504] The server sends the generated audio data to the terminal, and the terminal plays it for the user.
[1505] Input: Audio data from the server.
[1506] Output: The audio played to the user.
[1507] Specific action: The device plays the received Naruto voice, and the user hears the voice say, "Hey! What's up?"
[1508] Step 21:
[1509] The user selects the "Metaverse Mode" option.
[1510] Input: User click operation.
[1511] Output: The request is sent to the server.
[1512] Specific action: The user clicks the "Metaverse Mode" button on the dashboard, and the device sends an HTTP GET request to the server.
[1513] Step 22:
[1514] The server connects to the metaverse platform and loads information from the virtual space.
[1515] Input: Request from the terminal.
[1516] Output: Data in a virtual space.
[1517] Specific operation: The server uses the API of the metaverse platform (such as VRChat) to obtain the necessary virtual space data.
[1518] Step 23:
[1519] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[1520] Input: Data from the virtual space, user and character information.
[1521] Output: An avatar placed in a virtual space.
[1522] Specific operation: The server places the user's avatar and Naruto's avatar within VRChat.
[1523] Step 24:
[1524] The user puts on a VR device and accesses a virtual space.
[1525] Input: User action.
[1526] Output: Logged into the virtual space.
[1527] Specific actions: The user puts on an Oculus Quest, launches the VRChat app, and logs into the virtual space.
[1528] Step 25:
[1529] The device captures the user's movements and sends them to the server in real time.
[1530] Input: VR device operation information.
[1531] Output: Real-time operation data.
[1532] Specific operation: The terminal (VR device) captures the user's movements using sensors and transmits them to the server in real time.
[1533] Step 26:
[1534] The server issues instructions to make the character's avatar move naturally in response to the user's movement information.
[1535] Input: Real-time operation data from the device.
[1536] Output: Naruto's avatar moves within the virtual space.
[1537] Specific operation: Based on the received motion data, the server instructs the Naruto avatar to perform appropriate actions within VRChat, making it move naturally in real time in accordance with the user's movements.
[1538] Through these steps, users can register and log in to the system, interact with characters and idols, and enjoy real-time interactive experiences within the virtual space.
[1539] (Application Example 1)
[1540] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1541] Conventional voice interaction systems have limited functionality for users to interact with characters or idols, and providing a truly interactive experience in real time presented many challenges. Furthermore, character-like behavior and real-time reflection of actions in virtual space were difficult, preventing users from enjoying a consistent entertainment experience. Moreover, no system existed that integrated the entire process from speech recognition to generational AI and speech synthesis to achieve real-time responses within a virtual space.
[1542] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1543] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for displaying the voice response in a virtual space; means for controlling the character's actions in the virtual space; and means for playing the voice response back to the user. This enables the user to interact with characters and idols in real time in a virtual space and enjoy an interactive and consistent experience.
[1544] A "user" is a person who uses this system to interact with characters and idols in a virtual space.
[1545] "Character voice" refers to voice data corresponding to a specific character, and is used to reproduce the voice of the character selected by the user.
[1546] A "generative model" is a machine learning model designed to generate the voice of a character selected by the user.
[1547] "Speech recognition means" is a general term for hardware and software used to convert a user's voice into text data.
[1548] "Artificial intelligence means" refers to machine learning algorithms that generate appropriate responses from input text based on a generative model.
[1549] "Speech synthesis means" is a general term for hardware and software used to convert generated text responses into character voices.
[1550] A "virtual space" is a three-dimensional computer graphics space where users and characters engage in dialogue and interaction.
[1551] "Means for displaying voice responses" refers to the collective term for hardware and software used to visually represent voice responses generated within a virtual space.
[1552] "Means for controlling character-like behavior" refers to a general term for algorithms and systems used to control how characters in a virtual space behave naturally.
[1553] "Means for playing back voice responses" refers to an audio output device and associated software that allows the user to hear the character's voice responses.
[1554] One embodiment for carrying out this invention will be described. This system enables users to interact with characters and idols in a virtual space in real time and enjoy an interactive experience.
[1555] Hardware and software to be used
[1556] The main hardware used to carry out this invention includes:
[1557] Terminal: A device used by a user to access the system (smartphone, PC, head-mounted display)
[1558] Server: The central hub for speech recognition, generative AI, speech synthesis, and virtual space control.
[1559] Voice input device: A microphone that captures the user's voice.
[1560] Audio output device: A speaker that plays the character's voice responses.
[1561] The main software includes:
[1562] Speech recognition software: Google's speech recognition API
[1563] Generative AI Model: OpenAI GPT-3 Model
[1564] Speech synthesis software: pyttsx3 library
[1565] Virtual space management software: pyqtmetaverse library
[1566] Data processing and data calculation
[1567] 1. User voice input
[1568] The user speaks into the device. The voice input device captures the user's voice and sends it to the device.
[1569] 2. Speech recognition means
[1570] The device passes the captured audio data to speech recognition software, which then converts it into text data.
[1571] 3. Artificial intelligence tools
[1572] The server passes text data to the generative AI model, which then generates appropriate response text. This process creates prompt text for the generative AI.
[1573] Example prompt: "Generate a character's response to the question, 'What is the best product in this store?'"
[1574] 4. Speech synthesis means
[1575] The server passes the generated text response to speech synthesis software, which converts it into a character's voice.
[1576] 5. Display in virtual space
[1577] The virtual space management software displays the generated voice responses to the user in the virtual space both visually and audibly. It also controls the character's movements to allow the user to interact naturally.
[1578] 6. Playback of voice response
[1579] The device plays character voices to the user, providing an interactive dialogue experience.
[1580] Specific example
[1581] Suppose a user asks, "What is the best item in this store?" The server captures this audio and converts it into text, "What is the best item in this store?", through speech recognition software. Based on the prompt, "Generate a character response to 'What is the best item in this store?'", the generative AI model generates the appropriate response text, "Then it's the automatic ramen cooker! It's convenient and you can enjoy delicious ramen anytime." Speech synthesis software converts this response text into the character's voice. Finally, virtual space management software displays this voice within the virtual space, and the terminal plays the voice for the user.
[1582] This allows users to enjoy natural, real-time interactions with characters and idols in a virtual space.
[1583] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1584] Step 1:
[1585] The user speaks into a voice input device (microphone). The terminal captures this audio and prepares the data.
[1586] Input: User's voice
[1587] Output: Audio data
[1588] Specific operation: When a user says, "What is the best product you recommend in this store?", the microphone captures the audio.
[1589] Step 2:
[1590] The device passes the captured audio data to speech recognition software (Google's speech recognition API) and converts it into text data.
[1591] Input: Audio data
[1592] Output: Text data
[1593] Specific operation: The device sends the captured audio data to Google's speech recognition API, which converts it into the text, "What is the best product in this store?"
[1594] Step 3:
[1595] The server receives text data and passes it to a generative AI model (OpenAI GPT-3 model) to generate appropriate response text. During this process, it generates a prompt and sends it to the AI.
[1596] Input: Text data
[1597] Output: Response text
[1598] Specific operation: Based on the text "What is the best recommended item in this store?", the server generates a prompt message "Generate a character response to 'What is the best recommended item in this store?'". The generating AI model receives this and generates the response text "Then it's the automatic ramen cooker. It's convenient, and you can enjoy delicious ramen anytime."
[1599] Step 4:
[1600] The server passes the generated response text to speech synthesis software (pyttsx3) and converts it into a character's voice.
[1601] Input: Response text
[1602] Output: Audio data
[1603] Specific operation: The server passes the text "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime." to pyttsx3, which then converts it into character voice.
[1604] Step 5:
[1605] The server uses virtual space management software (pyqtmetaverse) to display the generated audio data within the virtual space and control the character's movements.
[1606] Input: Voice data, response text
[1607] Output: Character movements and sound display in the virtual space
[1608] Specific operation: The server passes the generated audio data and response text to the pyqtmetaverse, which then controls the movement of the character in the virtual space and displays it to the user visually and audibly.
[1609] Step 6:
[1610] The device plays the generated character's voice to the user through its speaker.
[1611] Input: Audio data
[1612] Output: Audio to be played
[1613] Specific operation: The device plays a voice message from a character through its speaker saying, "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime," to the user.
[1614] Through the above processing steps, users can enjoy real-time, interactive conversations with characters and idols in a virtual space.
[1615] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1616] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space. It also has a function to recognize the user's emotions and adjust its response accordingly.
[1617] User registration and login
[1618] The user first accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which then receives the data and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[1619] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[1620] Choosing your favorite character or idol
[1621] When the user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records this information in the database and loads the associated voice model and animation data.
[1622] Initiating conversations with AI and emotion recognition
[1623] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", that voice data is sent from the device to the server. The server receives the voice data and uses a speech recognition service to convert the voice to text. At the same time, an emotion engine analyzes the voice and the user's facial expressions to recognize the user's emotions (e.g., joy, sadness, surprise, etc.).
[1624] The server sends text data and recognized emotion information to the generating AI and requests it to generate an appropriate response. The generating AI generates an appropriate response text based on the input text and emotion information and returns it to the server. The server processes the response text through a speech synthesis engine and generates audio in the character's voice. This audio data is sent to the terminal, which then plays the generated audio for the user.
[1625] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generating AI will produce the response "Hi! You seem to be doing well!" which will then be played back to the user in Naruto's voice.
[1626] Providing a sense of realism in the metaverse
[1627] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. Conversation with the generating AI also takes place in parallel, and voice and corresponding animations are reproduced in the virtual space.
[1628] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space. This allows users to not only have a realistic interactive experience in the virtual space but also receive responses that are appropriate to their emotions.
[1629] This invention allows users to enjoy interacting with their favorite characters or idols while gaining a realistic experience even in a virtual space. Emotion recognition improves the quality of the interaction, enabling deeper engagement.
[1630] The following describes the processing flow.
[1631] Step 1:
[1632] Users access the system's registration page from their smartphones or PCs.
[1633] Step 2:
[1634] The device displays a registration form to the user.
[1635] Step 3:
[1636] The user enters their name, email address, and password and submits the form.
[1637] Step 4:
[1638] The terminal sends the input data to the server.
[1639] Step 5:
[1640] The server receives the transmitted data and checks for formatting and duplicates.
[1641] Step 6:
[1642] After the server verifies the data, it saves it to the database.
[1643] Step 7:
[1644] The server sends a registration completion message to the device.
[1645] Step 8:
[1646] The device displays a registration completion message to the user.
[1647] Step 9:
[1648] The user accesses the login page and enters their email address and password.
[1649] Step 10:
[1650] The terminal sends the input data to the server.
[1651] Step 11:
[1652] The server compares the information with the registered data in the database.
[1653] Step 12:
[1654] The server matches the user information and creates a login session.
[1655] Step 13:
[1656] The server sends the login session to the terminal.
[1657] Step 14:
[1658] The device displays the user's dashboard.
[1659] Step 15:
[1660] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[1661] Step 16:
[1662] The terminal sends a request to the server.
[1663] Step 17:
[1664] The server retrieves a list of available characters / idols from the database.
[1665] Step 18:
[1666] The server sends the list to the terminal.
[1667] Step 19:
[1668] The device displays the list to the user.
[1669] Step 20:
[1670] The user selects their desired character or idol.
[1671] Step 21:
[1672] The device sends the selection information to the server.
[1673] Step 22:
[1674] The server records the selection information in a database and loads the associated voice model and animation data.
[1675] Step 23:
[1676] The user clicks the "Start Conversation" button.
[1677] Step 24:
[1678] The device turns on the microphone and captures the user's voice and facial expressions.
[1679] Step 25:
[1680] The device sends the captured audio and facial expression data to the server.
[1681] Step 26:
[1682] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[1683] Step 27:
[1684] The server uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotions.
[1685] Step 28:
[1686] The server sends text data and recognized emotional information to the AI and requests it to generate a response.
[1687] Step 29:
[1688] The generation AI generates appropriate response text based on the input text and sentiment information, and sends it to the server.
[1689] Step 30:
[1690] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[1691] Step 31:
[1692] The server sends the generated audio file to the terminal.
[1693] Step 32:
[1694] The device plays generated audio and visual information for the user.
[1695] Step 33:
[1696] The user selects the "Metaverse Mode" option.
[1697] Step 34:
[1698] The terminal sends a request to the server.
[1699] Step 35:
[1700] The server connects to the metaverse platform and loads information from the virtual space.
[1701] Step 36:
[1702] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[1703] Step 37:
[1704] The user puts on a VR device and accesses a virtual space.
[1705] Step 38:
[1706] The device captures the user's movements and sends them to the server in real time.
[1707] Step 39:
[1708] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[1709] Step 40:
[1710] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space.
[1711] Step 41:
[1712] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[1713] Step 42:
[1714] The device provides an interactive experience between the user and their favorite character or idol.
[1715] (Example 2)
[1716] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1717] Conventional voice dialogue systems simply provide monotonous responses to user input, making it difficult to recognize emotions and engage in interactive conversations. Furthermore, even in virtual reality interactions, real-time control based on user actions is insufficient, failing to provide a realistic experience. This resulted in a decline in the quality of the user experience.
[1718] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1719] In this invention, the server includes means for providing a generative model for generating the voice of a virtual character selected by the user, speech recognition means for converting the user's voice into text, and artificial intelligence means for generating a response based on the text and emotional information. This makes it possible to recognize the user's emotions and adjust the response accordingly.
[1720] A "user" is an end-user who uses the system to interact with others and have experiences in a virtual space.
[1721] A "virtual character" is a virtual character or idol that users can choose to use in conversations and interactive experiences.
[1722] A "generative model" is an artificial intelligence model used to generate the voice and behavior of a virtual character selected by the user.
[1723] "Speech recognition means" refers to technologies and devices for converting a user's speech into text.
[1724] "Artificial intelligence means" refers to technologies and devices that generate appropriate responses based on text data and emotional information.
[1725] "Voice synthesis means" refers to technologies and devices for converting generated responses into the voice of a virtual character.
[1726] "Emotion recognition means" refers to technologies and devices that recognize emotions from a user's voice and facial expressions and adjust responses based on that information.
[1727] "Means for reproducing voice responses" refers to technologies or devices for reproducing generated voice responses to the user.
[1728] An "avatar" is a graphical representation displayed in a virtual space as a digital representation of a user or virtual character.
[1729] A "virtual space" is a computer-generated three-dimensional space that users can access through VR devices or similar means.
[1730] "Motion information" refers to data about the user's body movements and is used to control the avatar in the virtual space.
[1731] "Real-time" refers to a sense of time in which processing and responses occur almost instantly, meaning that delays to user operations and inputs are kept to a minimum.
[1732] Embodiments of this invention will now be described in detail. This system aims to allow users to interact with their favorite virtual characters or idols in real time and enjoy an interactive experience in a virtual space. It also has the function of recognizing the user's emotions and adjusting its response accordingly.
[1733] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who enters necessary information such as their name, email address, and password, and sends it from the device to the server. The server stores the received data in a database and notifies the user via the device that registration is complete. Next, the user accesses the login page and enters the registered email address and password. The device sends the entered data to the server, which verifies it against the information in the database. If authentication is successful, the server generates a login session and sends it to the device, displaying the user's dashboard.
[1734] When a user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of characters and idols from the database and sends it to the device. The device displays this list to the user, and when the user selects their desired character or idol, that information is sent to the server and recorded in the database.
[1735] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the audio data is sent from the device to the server. The server converts this audio data to text using the Google Cloud Speech-to-Text API and analyzes the text and emotion information with an emotion engine (e.g., Microsoft Azure Emotion API). If the analysis result is positive, the server sends the text data and emotion information to a generating AI model (e.g., OpenAI GPT-3) to generate a response such as "Hi! You seem to be doing well!". The generated response text is then synthesized into a character's voice using Amazon Polly and sent to the device for playback.
[1736] Furthermore, when the user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to a Unity-based metaverse platform and loads information for the virtual space. The server places the user's avatar and character / idol avatars in the virtual space, and the user accesses the virtual space by wearing a VR device (e.g., Oculus Rift). The device captures the user's movements in real time and sends them to the server. The server controls the movements of the virtual avatar through the Unity engine and simultaneously provides a more natural and realistic interactive experience through voice interaction via a generated AI model. In addition, an emotion engine recognizes the user's emotional changes in real time and reflects that information in the movements and facial expressions of the virtual avatar.
[1737] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generative AI model will generate the response "Hi! You seem to be doing well!". This response is synthesized into Naruto's voice and played back to the user through the speaker.
[1738] Examples of prompts for operating this system include:
[1739] Prompt to convert audio data to text:
[1740] Please convert the audio data 'Hello, Naruto!' to text.
[1741] Response generation prompt:
[1742] "Based on the text 'Hello, Naruto!' and the emotion 'Joyful,' please generate an appropriate response."
[1743] Text-to-speech prompt:
[1744] "Please convert the text 'Hey! You seem to be doing well!' into speech in Naruto's voice."
[1745] Through these processes, users can gain real-time interaction and an interactive experience in a virtual space.
[1746] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1747] Program processing flow
[1748] Step 1:
[1749] Users access the system's registration page from their smartphones or PCs.
[1750] Input: Name, email address, password
[1751] The terminal receives these inputs, displays them on the registration form, and waits for user input.
[1752] Output: Input data
[1753] The user enters their name, email address, and password, and this data is sent to the device.
[1754] Step 2:
[1755] The terminal sends the input data to the server.
[1756] Input: User registration information (name, email address, password)
[1757] The server receives the input data and saves it to the database.
[1758] Output: Registration complete message
[1759] The device displays "Registration complete!" to the user.
[1760] Step 3:
[1761] The user accesses the login page and enters their email address and password.
[1762] Enter: Email address, password
[1763] The terminal sends the input data to the server.
[1764] Output: Authentication result
[1765] The server compares the information with the registration details in the database, and if a match is found, it generates a login session.
[1766] A login session is sent to the device, and the device displays the user's dashboard.
[1767] Step 4:
[1768] The user clicks the "Select Favorite Character / Idol" button in the main menu.
[1769] Input: Click Event
[1770] The terminal sends a request to the server.
[1771] Output: List of Characters / Idols
[1772] The server retrieves a list of characters and idols from the database and sends it to the terminal.
[1773] The device displays the list to the user.
[1774] Step 5:
[1775] The user selects their desired character or idol.
[1776] Input: Selected character or idol
[1777] The device sends the selection information to the server.
[1778] Output: Database update results
[1779] The server records that information in the database.
[1780] Step 6:
[1781] The user clicks the "Start Conversation" button.
[1782] Input: Click Event
[1783] The device turns on the microphone and captures the user's voice.
[1784] The user says, "Hello, [Character Name]!"
[1785] Output: Audio data
[1786] The audio data is then sent from the terminal to the server.
[1787] Step 7:
[1788] The server sends the voice data to the speech recognition service.
[1789] Input: Audio data
[1790] The speech recognition service converts speech into text.
[1791] Output: Text data
[1792] The server receives this text data.
[1793] Step 8:
[1794] The server recognizes emotions from voice and user facial expression data through an emotion engine.
[1795] Input: Voice data, user facial expression data
[1796] Output: Emotional information
[1797] The server uses text data and recognized emotion information to send to a generative AI model.
[1798] Step 9:
[1799] The generative AI model generates appropriate responses based on text data and sentiment information.
[1800] Input: Text data, sentiment information
[1801] Output: Response text
[1802] The server receives the response text.
[1803] Step 10:
[1804] The server processes the response text into a speech synthesis engine.
[1805] Input: Response text
[1806] The speech synthesis engine generates speech using the voice of a virtual character.
[1807] Output: Generated audio data
[1808] The server sends the audio data to the terminal.
[1809] Step 11:
[1810] The device plays the generated audio to the user.
[1811] Input: Audio data
[1812] The device plays audio through its speaker, and the user listens to the response.
[1813] Output: Played audio
[1814] Step 12:
[1815] The user selects the "Metaverse Mode" option.
[1816] Input: Click Event
[1817] The terminal sends a request to the server.
[1818] The server connects to the metaverse platform and loads the virtual space.
[1819] Output: Display data for the virtual space
[1820] Step 13:
[1821] The server places user avatars and character / idol avatars in the virtual space.
[1822] Input: None (automatic processing)
[1823] Users wear VR devices to access a virtual space.
[1824] Output: Initial placement of user and character avatars
[1825] Step 14:
[1826] The device captures the user's movements and sends them to the server.
[1827] Input: User activity information
[1828] The server controls the avatar in real time via the Unity engine.
[1829] Output: Avatar movement data
[1830] This allows users to enjoy interactive experiences in a virtual space.
[1831] The above is the specific processing flow of this system.
[1832] (Application Example 2)
[1833] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1834] Conventional virtual space systems have limited interaction between users and virtual characters or idols, lacking interaction based on emotions and actions. Therefore, users find it difficult to have a realistic conversational experience, resulting in a less realistic virtual experience overall. Furthermore, the lack of interactive dialogue systems utilizing emotion recognition is another challenge.
[1835] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for providing a generation model for generating the voice of a virtual character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating an appropriate response based on the generation model; speech synthesis means for converting the generated response into the voice of a virtual character; emotion recognition means for analyzing the user's emotions; means for adjusting the response based on the emotion recognition results; and means for playing the generated voice response to the user. This makes it possible for the user to enjoy a realistic dialogue experience with their favorite virtual character or idol that responds to their emotions.
[1836] A "user" is someone who uses the system to enjoy interacting with virtual characters or idols.
[1837] A "virtual space" is a three-dimensional virtual environment created by a computer program.
[1838] A "virtual character" is a character or idol selected by a user within a virtual space, whose voice and actions are generated based on AI.
[1839] A "generative model" refers to the algorithms and datasets used to generate voices and actions for virtual characters.
[1840] "Speech recognition means" refers to a technical device or program for generating text from a user's speech.
[1841] An "artificial intelligence system" is a system that uses a pre-trained model to generate an appropriate response to user input.
[1842] "Speech synthesis means" refers to a device or program that converts text data into the voice of a virtual character and generates speech data.
[1843] "Emotion recognition means" refers to technologies and programs that analyze and recognize emotions from a user's voice, facial expressions, etc.
[1844] "Means for adjusting responses" refer to technologies or programs for appropriately modifying responses generated based on the user's emotion recognition results.
[1845] "Means for reproducing voice responses" refers to devices or programs that output generated voice data to the user.
[1846] The embodiments for carrying out this invention are described in detail below.
[1847] First, the system consists of the following main components:
[1848] 1. A generative model for generating voices for virtual characters based on user selections.
[1849] 2. Speech recognition means
[1850] 3. Artificial intelligence means for generating appropriate responses
[1851] 4. Speech synthesis means for converting the generated response into speech.
[1852] 5. Emotion recognition means for analyzing user emotions
[1853] 6. Means for adjusting responses based on emotion recognition results
[1854] 7. Means for playing voice responses to the user
[1855] Specifically, when a user accesses the system and enters their email address and password on the login page, the server compares this information with existing data in the database. If the login is successful, the user's dashboard is displayed. The user then selects their preferred virtual character from the dashboard, and the server loads the character's voice model and animation data.
[1856] Hardware and software configuration
[1857] The hardware and software used in this invention are as follows:
[1858] Hardware: Smartphone, tablet, or VR headset, microphone, speaker
[1859] software:
[1860] Google Cloud Speech-to-Text: Used as a speech recognition method.
[1861] Azure Text Analytics: Used as a means of emotion recognition
[1862] OpenAI GPT-3: Used as an artificial intelligence means to generate appropriate responses.
[1863] Amazon Polly: Used as a speech synthesis method.
[1864] Data adjustment and calculation flow
[1865] 1. Speech recognition:
[1866] The audio data spoken by the user through the microphone is sent to the server and converted into text using Google Cloud Speech-to-Text.
[1867] 2. Emotion recognition:
[1868] The data, converted to text, is analyzed by Azure Text Analytics to recognize the user's sentiment.
[1869] 3. Response generation:
[1870] The server inputs the recognized emotion and text data into OpenAI GPT-3 and generates an appropriate response text.
[1871] 4. Speech synthesis:
[1872] The generated response text is converted into a virtual character's voice by Amazon Polly.
[1873] 5. Playback of voice response:
[1874] Audio data generated from the server is sent to the user's device, and the device plays the audio.
[1875] Specific example
[1876] For example, a user says "Hello, Naruto!" into the microphone. This audio is sent to the server. Google Cloud Speech-to-Text converts the audio into text "Hello, Naruto!". Next, Azure Text Analytics recognizes the user's emotion from this text as "joy". The server sends this recognized emotion and text to OpenAI GPT-3 to generate an appropriate response, "Hi! You seem to be doing well!". This response is converted into speech by Amazon Polly and played back on the user's device in the voice of the virtual character Naruto.
[1877] Example of a prompt
[1878] Joyful user: Hello, Naruto! Virtual character:
[1879] In this way, the system of the present invention enables users to enjoy a realistic dialogue experience that responds to their emotions.
[1880] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1881] Step 1:
[1882] The server receives the email address and password entered by the user and compares them with existing information in the database. If the entered email address and password match, the server creates a login session and displays the user's dashboard.
[1883] Step 2:
[1884] The user selects a virtual character from the dashboard. The device sends the selection information to the server, which loads the corresponding voice model and animation data from the database and presents it to the user.
[1885] Step 3:
[1886] When the user presses the "Start Conversation" button, the device turns on the microphone and captures the user's voice data. This voice data is sent to the server. The server uses Google Cloud Speech-to-Text to convert this voice data into text.
[1887] Step 4:
[1888] The server sends the converted text data to Azure Text Analytics for user sentiment analysis. The sentiment analysis results are stored along with the text data.
[1889] Step 5:
[1890] The server sends text data and emotion recognition results to OpenAI GPT-3 to generate an appropriate response. Specifically, natural-sounding dialogue is generated based on the prompt text and the recognized emotion information.
[1891] Step 6:
[1892] The generated response text is converted into audio data by Amazon Polly. The server then sends the generated audio data to the device.
[1893] Step 7:
[1894] The user's device plays the received audio data. This provides the user with the experience of a virtual character responding to them verbally.
[1895] Step 8:
[1896] User movement information is also captured in real time and sent from the terminal to the server. The server controls the movement of the avatar in the virtual space and adjusts the avatar to respond to the user's movements.
[1897] This allows users to enjoy realistic, emotion-driven conversations and action-based interactions within the virtual space.
[1898] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1899] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1900] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1901] [Fourth Embodiment]
[1902] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1903] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1904] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1905] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1906] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1907] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1908] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1909] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1910] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1911] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1912] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1913] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1914] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1915] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space.
[1916] User registration and login
[1917] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which receives it and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[1918] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[1919] Choosing your favorite character or idol
[1920] When the user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records the selection information in the database and loads the associated voice model and animation data.
[1921] Starting a conversation with AI
[1922] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the voice data is sent from the device to the server. The server uses a speech recognition service to convert the voice data into text and sends the text to a generative AI. The generative AI generates an appropriate response text based on the input text, which is then returned to the server. The server then uses a speech synthesis engine to convert the response text into speech, generating the voice in the character's voice. This voice data is sent to the device, which then plays the voice for the user.
[1923] For example, if the user says "Hello, Naruto!", the AI will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[1924] Providing a sense of realism in the metaverse
[1925] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. This allows the user to enjoy a realistic interactive experience in the virtual space.
[1926] The following describes the processing flow.
[1927] Step 1:
[1928] Users access the system's registration page from their smartphones or PCs.
[1929] Step 2:
[1930] The device displays a registration form to the user.
[1931] Step 3:
[1932] The user enters their name, email address, and password and submits the form.
[1933] Step 4:
[1934] The terminal sends the input data to the server.
[1935] Step 5:
[1936] The server receives the transmitted data and checks for formatting and duplicates.
[1937] Step 6:
[1938] After the server verifies the data, it saves it to the database.
[1939] Step 7:
[1940] The server sends a registration completion message to the device.
[1941] Step 8:
[1942] The device displays a registration completion message to the user.
[1943] Step 9:
[1944] The user accesses the login page and enters their email address and password.
[1945] Step 10:
[1946] The terminal sends the input data to the server.
[1947] Step 11:
[1948] The server compares the information with the registered data in the database.
[1949] Step 12:
[1950] The server matches the user information and creates a login session.
[1951] Step 13:
[1952] The server sends the login session to the terminal.
[1953] Step 14:
[1954] The device displays the user's dashboard.
[1955] Step 15:
[1956] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[1957] Step 16:
[1958] The terminal sends a request to the server.
[1959] Step 17:
[1960] The server retrieves a list of available characters / idols from the database.
[1961] Step 18:
[1962] The server sends the list to the terminal.
[1963] Step 19:
[1964] The device displays the list to the user.
[1965] Step 20:
[1966] The user selects their desired character or idol.
[1967] Step 21:
[1968] The device sends the selection information to the server.
[1969] Step 22:
[1970] The server records the selection information in a database and loads the associated voice model and animation data.
[1971] Step 23:
[1972] The user clicks the "Start Conversation" button.
[1973] Step 24:
[1974] The device turns on the microphone and captures the user's voice.
[1975] Step 25:
[1976] The device sends the captured audio data to the server.
[1977] Step 26:
[1978] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[1979] Step 27:
[1980] The server sends text data to the AI generating the response and requests that it generate a response.
[1981] Step 28:
[1982] The AI generates appropriate response text and sends it to the server.
[1983] Step 29:
[1984] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[1985] Step 30:
[1986] The server sends the generated audio file to the terminal.
[1987] Step 31:
[1988] The device plays the generated audio for the user.
[1989] Step 32:
[1990] The user selects the "Metaverse Mode" option.
[1991] Step 33:
[1992] The terminal sends a request to the server.
[1993] Step 34:
[1994] The server connects to the metaverse platform and loads information from the virtual space.
[1995] Step 35:
[1996] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[1997] Step 36:
[1998] The user puts on a VR device and accesses a virtual space.
[1999] Step 37:
[2000] The device captures the user's movements and sends them to the server in real time.
[2001] Step 38:
[2002] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[2003] Step 39:
[2004] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[2005] Step 40:
[2006] The device provides an interactive experience between the user and their favorite character or idol.
[2007] (Example 1)
[2008] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2009] In modern virtual space systems and interactive character dialogue systems, it is difficult for users to engage in natural conversations and actions with characters within the virtual space. Furthermore, there is a lack of easy registration and login methods for users accessing the system for the first time, which leads to decreased user satisfaction. This invention aims to solve these problems and provide a system that allows users to interact with characters in real time and experience natural actions within the virtual space.
[2010] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[2011] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for playing the voice response back to the user; registration means for the user to input registration information into the system; means for storing the registration information in a database; login means for the user to log in to the system; and means for matching the login information with information in the database. This allows the user to easily register and log in to the system, and subsequently engage in natural conversations with characters and interactive experiences in a virtual space in real time.
[2012] A "generative model" refers to a data model used to generate the voice of a character selected by the user.
[2013] "Voice recognition means" refers to technologies and devices for converting a user's voice into text data.
[2014] "Artificial intelligence means" refers to algorithms and programs for generating appropriate responses based on generative models.
[2015] "Speech synthesis means" refers to the technology that converts generated text data into the voice of a character.
[2016] "Means of playing voice responses to the user" refers to devices and technologies for playing the generated character's voice to the user.
[2017] "Registration method" refers to a series of processes and interfaces for a user to input their information into a system and for that information to be processed.
[2018] A "database" refers to a system used to store and manage user registration information and selection information.
[2019] "Login method" refers to the process or technology used to authenticate users when they access a system.
[2020] "Verification means" refers to technology used to compare login information entered by a user with registered information in a database and confirm that they match.
[2021] A "virtual space" refers to a digital environment where user and character avatars exist and which provides an interactive experience.
[2022] An "avatar" refers to a digital character that represents a user or character within a virtual space.
[2023] "Motion information" refers to information that detects the user's physical movements and captures them as digital data.
[2024] "Means of capturing" refers to devices and technologies for acquiring user behavior information in real time.
[2025] "Means of reflection" refers to technologies that reflect acquired user behavior information onto the movements of an avatar in a virtual space.
[2026] As a specific embodiment of this invention, a system is provided that allows users to interact in real time with characters or idols of their choice and enjoy an interactive experience in a virtual space.
[2027] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form for the user to enter their name, email address, and password. Once the user enters the required information and submits it, the device sends this data to the server, which receives the data and stores it in its database. Upon completion of registration, the device displays a registration completion message to the user.
[2028] Next, the user accesses the login page and enters their email address and password. The device sends this data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard. This process allows users to easily access the system.
[2029] When a user clicks the "Select Favorite Character / Idol" option in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays this list to the user, who then selects their desired character or idol. The device sends this selection information to the server, which records the selection information in the database and loads the associated voice model and animation data. This prepares the user to interact with their chosen character.
[2030] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello," the voice data is sent from the device to the server. The server uses a speech recognition service (e.g., Google Speech-to-Text) to convert the voice data into text. The server then sends the text to a generative AI model (e.g., OpenAI's GPT-3) to generate an appropriate response text. The generated text is returned to the server, which uses a speech synthesis engine (e.g., Amazon Polly) to generate the character's voice. This generated voice is sent to the device and played back to the user. This process allows the user to experience a natural conversation with the character.
[2031] For example, if a user says "Hello, Naruto!", the generative AI model will generate the response "Hey! What's up?", which will then be played back to the user in Naruto's voice.
[2032] Furthermore, when a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on the VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character avatar to move naturally in response to the user's movements. Through this process, the user can enjoy a realistic interactive experience in the virtual space.
[2033] An example of a prompt message is, "I want to talk to Naruto for a bit. Turn on the microphone and speak to him." Based on this prompt, users can access the system from their smartphones or PCs and enjoy interacting with characters and experiencing virtual spaces.
[2034] The flow of the specific processing in Example 1 will be explained using Figure 11.
[2035] Step 1:
[2036] Access the user's system registration page.
[2037] Input: The user enters the URL into their browser and accesses it.
[2038] Output: The registration page is displayed.
[2039] Specific operation: When a user enters the system's URL into their browser, the browser sends a request to the server, which generates a registration page and sends it back to the user's device.
[2040] Step 2:
[2041] The device displays a registration form to the user.
[2042] Input: Registration page data received from the server.
[2043] Output: The registration form is displayed on the screen.
[2044] Specific operation: The device uses the HTML and CSS included in the registration page to render the form on the user's screen.
[2045] Step 3:
[2046] The user enters the required information into the registration form and clicks the submit button.
[2047] Input: The action of a user entering their name, email address, and password and clicking the submit button.
[2048] Output: The input data is sent to the server.
[2049] Specific operation: The user enters, for example, "Taro Yamada", "example@mail.com", and "password123" into the form and clicks the "Submit" button. The input data is sent to the server as an HTTP POST request.
[2050] Step 4:
[2051] The server receives the input data and saves it to the database.
[2052] Input: User registration data received from the terminal.
[2053] Output: The registered data is saved in the database.
[2054] Specific operation: The server parses the received data and executes an INSERT query to save the data to the MySQL database.
[2055] Step 5:
[2056] The device displays a registration completion message to the user.
[2057] Input: Registration completion response from the server.
[2058] Output: A registration complete message is displayed on the screen.
[2059] Specific actions: The server notifies the terminal that saving to the database is complete and sends a response containing that information. The terminal displays the message "Registration complete!" on its screen.
[2060] Step 6:
[2061] The user accesses the login page and enters their email address and password.
[2062] Input: Access the login page URL, enter user information (email address, password).
[2063] Output: Your email address and password will be sent to the server.
[2064] Specific operation: The user enters and accesses the login page URL, enters "example@mail.com" and "password123" in the login form, and clicks the "Login" button. The data is sent via an HTTP POST request.
[2065] Step 7:
[2066] The server verifies the login information against the registered information in the database.
[2067] Input: Login data from the device.
[2068] Output: Login authentication result.
[2069] Specific operation: The server searches the "users" table in the database for a record with a matching email address and verifies the password. If a match is found, a login session ID is generated.
[2070] Step 8:
[2071] The server creates a login session and sends it to the terminal.
[2072] Input: Result indicating successful login.
[2073] Output: Session ID.
[2074] Specific operation: The server generates a session ID and sends it to the terminal.
[2075] Step 9:
[2076] The device displays the user's dashboard.
[2077] Input: Session ID response from the server.
[2078] Output: User's dashboard screen.
[2079] Specific action: The device receives the session ID, generates a dashboard page, and displays it on the screen. This page will include a message such as "Hello, Taro Yamada!".
[2080] Step 10:
[2081] The user clicks the "Select your favorite character / idol" option.
[2082] Input: User click operation.
[2083] Output: The request is sent to the server.
[2084] Specific operation: When the user clicks the "Select Favorite Character / Idol" button, the device sends an HTTP GET request to the server.
[2085] Step 11:
[2086] The server retrieves a list of character idols from the database and sends it to the terminal.
[2087] Input: Request from the terminal.
[2088] Output: Character / Idol List.
[2089] Specific operation: The server retrieves all records from the "characters" table and sends them to the terminal.
[2090] Step 12:
[2091] The device displays a list of characters / idols to the user.
[2092] Input: List data from the server.
[2093] Output: A list of character idols is displayed on the screen.
[2094] Specific operation: Based on the received list data, the terminal displays a list of characters / idols to the user.
[2095] Step 13:
[2096] The user selects their desired character or idol.
[2097] Input: User selection operation.
[2098] Output: The selected character / idol information is sent to the server.
[2099] Specific operation: The user clicks "Naruto" from the list, and the selection information is sent to the server via an HTTP POST request.
[2100] Step 14:
[2101] The server records the selection information in a database and loads the associated voice model and animation data.
[2102] Input: Character selection information from the device.
[2103] Output: Recording to a database, loading of voice models and animation data.
[2104] Specific operation: The server records the selection information in the "users" table and processes the loading of voice models and animation data related to "Naruto".
[2105] Step 15:
[2106] The user clicks the "Start Conversation" button.
[2107] Input: User click operation.
[2108] Output: The microphone turns on.
[2109] Specific operation: When the user clicks the "Start Conversation" button, the device turns on the microphone and prepares to capture audio.
[2110] Step 16:
[2111] The device captures the user's voice and sends it to the server.
[2112] Input: User's voice.
[2113] Output: Audio data is sent to the server.
[2114] Specific operation: The device records the user's voice saying "Hello, Naruto!" and sends that data to the server via an HTTP POST request.
[2115] Step 17:
[2116] The server uses a speech recognition service to convert the speech data into text.
[2117] Input: Audio data from the device.
[2118] Output: Text data.
[2119] Specific operation: The server uses the Google Speech-to-Text API to convert the audio data into the text "Hello, Naruto!".
[2120] Step 18:
[2121] The server sends text to the AI model and generates an appropriate response text.
[2122] Input: Text data converted by speech recognition.
[2123] Output: Response text data.
[2124] Specific operation: The server sends text data to OpenAI's GPT-3 and generates the response, "Hey! What's up?"
[2125] Step 19:
[2126] The server converts the response text into speech using a speech synthesis engine.
[2127] Input: Response text from a generating AI model.
[2128] Output: Audio data.
[2129] Specific operation: The server uses Amazon Polly to generate audio data that says, "Hey! What's up?"
[2130] Step 20:
[2131] The server sends the generated audio data to the terminal, and the terminal plays it for the user.
[2132] Input: Audio data from the server.
[2133] Output: The audio played to the user.
[2134] Specific action: The device plays the received Naruto voice, and the user hears the voice say, "Hey! What's up?"
[2135] Step 21:
[2136] The user selects the "Metaverse Mode" option.
[2137] Input: User click operation.
[2138] Output: The request is sent to the server.
[2139] Specific action: The user clicks the "Metaverse Mode" button on the dashboard, and the device sends an HTTP GET request to the server.
[2140] Step 22:
[2141] The server connects to the metaverse platform and loads information from the virtual space.
[2142] Input: Request from the terminal.
[2143] Output: Data in a virtual space.
[2144] Specific operation: The server uses the API of the metaverse platform (such as VRChat) to obtain the necessary virtual space data.
[2145] Step 23:
[2146] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[2147] Input: Data from the virtual space, user and character information.
[2148] Output: An avatar placed in a virtual space.
[2149] Specific operation: The server places the user's avatar and Naruto's avatar within VRChat.
[2150] Step 24:
[2151] The user puts on a VR device and accesses a virtual space.
[2152] Input: User action.
[2153] Output: Logged into the virtual space.
[2154] Specific actions: The user puts on an Oculus Quest, launches the VRChat app, and logs into the virtual space.
[2155] Step 25:
[2156] The device captures the user's movements and sends them to the server in real time.
[2157] Input: VR device operation information.
[2158] Output: Real-time operation data.
[2159] Specific operation: The terminal (VR device) captures the user's movements using sensors and transmits them to the server in real time.
[2160] Step 26:
[2161] The server issues instructions to make the character's avatar move naturally in response to the user's movement information.
[2162] Input: Real-time operation data from the device.
[2163] Output: Naruto's avatar moves within the virtual space.
[2164] Specific operation: Based on the received motion data, the server instructs the Naruto avatar to perform appropriate actions within VRChat, making it move naturally in real time in accordance with the user's movements.
[2165] Through these steps, users can register and log in to the system, interact with characters and idols, and enjoy real-time interactive experiences within the virtual space.
[2166] (Application Example 1)
[2167] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2168] Conventional voice interaction systems have limited functionality for users to interact with characters or idols, and providing a truly interactive experience in real time presented many challenges. Furthermore, character-like behavior and real-time reflection of actions in virtual space were difficult, preventing users from enjoying a consistent entertainment experience. Moreover, no system existed that integrated the entire process from speech recognition to generational AI and speech synthesis to achieve real-time responses within a virtual space.
[2169] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[2170] In this invention, the server includes means for providing a generative model for generating the voice of a character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating a response based on the generative model; speech synthesis means for converting the generated response into the character's voice; means for displaying the voice response in a virtual space; means for controlling the character's actions in the virtual space; and means for playing the voice response back to the user. This enables the user to interact with characters and idols in real time in a virtual space and enjoy an interactive and consistent experience.
[2171] A "user" is a person who uses this system to interact with characters and idols in a virtual space.
[2172] "Character voice" refers to voice data corresponding to a specific character, and is used to reproduce the voice of the character selected by the user.
[2173] A "generative model" is a machine learning model designed to generate the voice of a character selected by the user.
[2174] "Speech recognition means" is a general term for hardware and software used to convert a user's voice into text data.
[2175] "Artificial intelligence means" refers to machine learning algorithms that generate appropriate responses from input text based on a generative model.
[2176] "Speech synthesis means" is a general term for hardware and software used to convert generated text responses into character voices.
[2177] A "virtual space" is a three-dimensional computer graphics space where users and characters engage in dialogue and interaction.
[2178] "Means for displaying voice responses" refers to the collective term for hardware and software used to visually represent voice responses generated within a virtual space.
[2179] "Means for controlling character-like behavior" refers to a general term for algorithms and systems used to control how characters in a virtual space behave naturally.
[2180] "Means for playing back voice responses" refers to an audio output device and associated software that allows the user to hear the character's voice responses.
[2181] One embodiment for carrying out this invention will be described. This system enables users to interact with characters and idols in a virtual space in real time and enjoy an interactive experience.
[2182] Hardware and software to be used
[2183] The main hardware used to carry out this invention includes:
[2184] Terminal: A device used by a user to access the system (smartphone, PC, head-mounted display)
[2185] Server: The central hub for speech recognition, generative AI, speech synthesis, and virtual space control.
[2186] Voice input device: A microphone that captures the user's voice.
[2187] Audio output device: A speaker that plays the character's voice responses.
[2188] The main software includes:
[2189] Speech recognition software: Google's speech recognition API
[2190] Generative AI Model: OpenAI GPT-3 Model
[2191] Speech synthesis software: pyttsx3 library
[2192] Virtual space management software: pyqtmetaverse library
[2193] Data processing and data calculation
[2194] 1. User voice input
[2195] The user speaks into the device. The voice input device captures the user's voice and sends it to the device.
[2196] 2. Speech recognition means
[2197] The device passes the captured audio data to speech recognition software, which then converts it into text data.
[2198] 3. Artificial intelligence tools
[2199] The server passes text data to the generative AI model, which then generates appropriate response text. This process creates prompt text for the generative AI.
[2200] Example prompt: "Generate a character's response to the question, 'What is the best product in this store?'"
[2201] 4. Speech synthesis means
[2202] The server passes the generated text response to speech synthesis software, which converts it into a character's voice.
[2203] 5. Display in virtual space
[2204] The virtual space management software displays the generated voice responses to the user in the virtual space both visually and audibly. It also controls the character's movements to allow the user to interact naturally.
[2205] 6. Playback of voice response
[2206] The device plays character voices to the user, providing an interactive dialogue experience.
[2207] Specific example
[2208] Suppose a user asks, "What is the best item in this store?" The server captures this audio and converts it into text, "What is the best item in this store?", through speech recognition software. Based on the prompt, "Generate a character response to 'What is the best item in this store?'", the generative AI model generates the appropriate response text, "Then it's the automatic ramen cooker! It's convenient and you can enjoy delicious ramen anytime." Speech synthesis software converts this response text into the character's voice. Finally, virtual space management software displays this voice within the virtual space, and the terminal plays the voice for the user.
[2209] This allows users to enjoy natural, real-time interactions with characters and idols in a virtual space.
[2210] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[2211] Step 1:
[2212] The user speaks into a voice input device (microphone). The terminal captures this audio and prepares the data.
[2213] Input: User's voice
[2214] Output: Audio data
[2215] Specific operation: When a user says, "What is the best product you recommend in this store?", the microphone captures the audio.
[2216] Step 2:
[2217] The device passes the captured audio data to speech recognition software (Google's speech recognition API) and converts it into text data.
[2218] Input: Audio data
[2219] Output: Text data
[2220] Specific operation: The device sends the captured audio data to Google's speech recognition API, which converts it into the text, "What is the best product in this store?"
[2221] Step 3:
[2222] The server receives text data and passes it to a generative AI model (OpenAI GPT-3 model) to generate appropriate response text. During this process, it generates a prompt and sends it to the AI.
[2223] Input: Text data
[2224] Output: Response text
[2225] Specific operation: Based on the text "What is the best recommended item in this store?", the server generates a prompt message "Generate a character response to 'What is the best recommended item in this store?'". The generating AI model receives this and generates the response text "Then it's the automatic ramen cooker. It's convenient, and you can enjoy delicious ramen anytime."
[2226] Step 4:
[2227] The server passes the generated response text to speech synthesis software (pyttsx3) and converts it into a character's voice.
[2228] Input: Response text
[2229] Output: Audio data
[2230] Specific operation: The server passes the text "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime." to pyttsx3, which then converts it into character voice.
[2231] Step 5:
[2232] The server uses virtual space management software (pyqtmetaverse) to display the generated audio data within the virtual space and control the character's movements.
[2233] Input: Voice data, response text
[2234] Output: Character movements and sound display in the virtual space
[2235] Specific operation: The server passes the generated audio data and response text to the pyqtmetaverse, which then controls the movement of the character in the virtual space and displays it to the user visually and audibly.
[2236] Step 6:
[2237] The device plays the generated character's voice to the user through its speaker.
[2238] Input: Audio data
[2239] Output: Audio to be played
[2240] Specific operation: The device plays a voice message from a character through its speaker saying, "Then it's an automatic ramen cooker. Convenient, and you can enjoy delicious ramen anytime," to the user.
[2241] Through the above processing steps, users can enjoy real-time, interactive conversations with characters and idols in a virtual space.
[2242] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[2243] A specific embodiment of this invention is shown below. This system aims to allow users to interact with their favorite characters or idols in real time and enjoy an interactive experience in a virtual space. It also has a function to recognize the user's emotions and adjust its response accordingly.
[2244] User registration and login
[2245] The user first accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who then enters the required information (name, email address, password) and submits it. The device sends the entered data to the server, which then receives the data and stores it in its database. Once registration is complete, the device displays a registration completion message to the user.
[2246] Next, the user accesses the login page and enters their email address and password. The device sends the entered data to the server, which verifies it against the registered information in the database. If they match, the server generates a login session and sends it to the device, displaying the user's dashboard.
[2247] Choosing your favorite character or idol
[2248] When the user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of available characters / idols from the database and sends it to the device. The device displays the list to the user, and the user selects their desired character or idol. The device sends the selection information to the server, which records this information in the database and loads the associated voice model and animation data.
[2249] Initiating conversations with AI and emotion recognition
[2250] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", that voice data is sent from the device to the server. The server receives the voice data and uses a speech recognition service to convert the voice to text. At the same time, an emotion engine analyzes the voice and the user's facial expressions to recognize the user's emotions (e.g., joy, sadness, surprise, etc.).
[2251] The server sends text data and recognized emotion information to the generating AI and requests it to generate an appropriate response. The generating AI generates an appropriate response text based on the input text and emotion information and returns it to the server. The server processes the response text through a speech synthesis engine and generates audio in the character's voice. This audio data is sent to the terminal, which then plays the generated audio for the user.
[2252] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generating AI will produce the response "Hi! You seem to be doing well!" which will then be played back to the user in Naruto's voice.
[2253] Providing a sense of realism in the metaverse
[2254] When a user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to the metaverse platform and loads information for the virtual space. The server places the user's avatar and the avatar of their favorite character or idol in the virtual space. When the user puts on a VR device and accesses the virtual space, the device captures the user's movements and sends them to the server in real time. The server instructs the character or idol avatar to move naturally in response to the user's actions. Conversation with the generating AI also takes place in parallel, and voice and corresponding animations are reproduced in the virtual space.
[2255] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space. This allows users to not only have a realistic interactive experience in the virtual space but also receive responses that are appropriate to their emotions.
[2256] This invention allows users to enjoy interacting with their favorite characters or idols while gaining a realistic experience even in a virtual space. Emotion recognition improves the quality of the interaction, enabling deeper engagement.
[2257] The following describes the processing flow.
[2258] Step 1:
[2259] Users access the system's registration page from their smartphones or PCs.
[2260] Step 2:
[2261] The device displays a registration form to the user.
[2262] Step 3:
[2263] The user enters their name, email address, and password and submits the form.
[2264] Step 4:
[2265] The terminal sends the input data to the server.
[2266] Step 5:
[2267] The server receives the transmitted data and checks for formatting and duplicates.
[2268] Step 6:
[2269] After the server verifies the data, it saves it to the database.
[2270] Step 7:
[2271] The server sends a registration completion message to the device.
[2272] Step 8:
[2273] The device displays a registration completion message to the user.
[2274] Step 9:
[2275] The user accesses the login page and enters their email address and password.
[2276] Step 10:
[2277] The terminal sends the input data to the server.
[2278] Step 11:
[2279] The server compares the information with the registered data in the database.
[2280] Step 12:
[2281] The server matches the user information and creates a login session.
[2282] Step 13:
[2283] The server sends the login session to the terminal.
[2284] Step 14:
[2285] The device displays the user's dashboard.
[2286] Step 15:
[2287] The user clicks the "Select Favorite Character / Idol" option in the main menu.
[2288] Step 16:
[2289] The terminal sends a request to the server.
[2290] Step 17:
[2291] The server retrieves a list of available characters / idols from the database.
[2292] Step 18:
[2293] The server sends the list to the terminal.
[2294] Step 19:
[2295] The device displays the list to the user.
[2296] Step 20:
[2297] The user selects their desired character or idol.
[2298] Step 21:
[2299] The device sends the selection information to the server.
[2300] Step 22:
[2301] The server records the selection information in a database and loads the associated voice model and animation data.
[2302] Step 23:
[2303] The user clicks the "Start Conversation" button.
[2304] Step 24:
[2305] The device turns on the microphone and captures the user's voice and facial expressions.
[2306] Step 25:
[2307] The device sends the captured audio and facial expression data to the server.
[2308] Step 26:
[2309] The server receives the audio data and uses a speech recognition service to convert the audio into text.
[2310] Step 27:
[2311] The server uses an emotion engine to analyze the user's voice and facial expressions to recognize their emotions.
[2312] Step 28:
[2313] The server sends text data and recognized emotional information to the AI and requests it to generate a response.
[2314] Step 29:
[2315] The generation AI generates appropriate response text based on the input text and sentiment information, and sends it to the server.
[2316] Step 30:
[2317] The server processes the response text into a speech synthesis engine and generates speech using the character's voice.
[2318] Step 31:
[2319] The server sends the generated audio file to the terminal.
[2320] Step 32:
[2321] The device plays generated audio and visual information for the user.
[2322] Step 33:
[2323] The user selects the "Metaverse Mode" option.
[2324] Step 34:
[2325] The terminal sends a request to the server.
[2326] Step 35:
[2327] The server connects to the metaverse platform and loads information from the virtual space.
[2328] Step 36:
[2329] The server places the user's avatar and the avatar of their favorite character or idol in the virtual space.
[2330] Step 37:
[2331] The user puts on a VR device and accesses a virtual space.
[2332] Step 38:
[2333] The device captures the user's movements and sends them to the server in real time.
[2334] Step 39:
[2335] The server issues instructions to the character / idol avatar to move naturally in response to the user's actions.
[2336] Step 40:
[2337] The emotion engine recognizes the user's emotional changes in real time and reflects that information in the avatar's movements and facial expressions within the virtual space.
[2338] Step 41:
[2339] The server also processes conversations with the generated AI in parallel, reproducing the audio and corresponding animations within the virtual space.
[2340] Step 42:
[2341] The device provides an interactive experience between the user and their favorite character or idol.
[2342] (Example 2)
[2343] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2344] Conventional voice dialogue systems simply provide monotonous responses to user input, making it difficult to recognize emotions and engage in interactive conversations. Furthermore, even in virtual reality interactions, real-time control based on user actions is insufficient, failing to provide a realistic experience. This resulted in a decline in the quality of the user experience.
[2345] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[2346] In this invention, the server includes means for providing a generative model for generating the voice of a virtual character selected by the user, speech recognition means for converting the user's voice into text, and artificial intelligence means for generating a response based on the text and emotional information. This makes it possible to recognize the user's emotions and adjust the response accordingly.
[2347] A "user" is an end-user who uses the system to interact with others and have experiences in a virtual space.
[2348] A "virtual character" is a virtual character or idol that users can choose to use in conversations and interactive experiences.
[2349] A "generative model" is an artificial intelligence model used to generate the voice and behavior of a virtual character selected by the user.
[2350] "Speech recognition means" refers to technologies and devices for converting a user's speech into text.
[2351] "Artificial intelligence means" refers to technologies and devices that generate appropriate responses based on text data and emotional information.
[2352] "Voice synthesis means" refers to technologies and devices for converting generated responses into the voice of a virtual character.
[2353] "Emotion recognition means" refers to technologies and devices that recognize emotions from a user's voice and facial expressions and adjust responses based on that information.
[2354] "Means for reproducing voice responses" refers to technologies or devices for reproducing generated voice responses to the user.
[2355] An "avatar" is a graphical representation displayed in a virtual space as a digital representation of a user or virtual character.
[2356] A "virtual space" is a computer-generated three-dimensional space that users can access through VR devices or similar means.
[2357] "Motion information" refers to data about the user's body movements and is used to control the avatar in the virtual space.
[2358] "Real-time" refers to a sense of time in which processing and responses occur almost instantly, meaning that delays to user operations and inputs are kept to a minimum.
[2359] Embodiments of this invention will now be described in detail. This system aims to allow users to interact with their favorite virtual characters or idols in real time and enjoy an interactive experience in a virtual space. It also has the function of recognizing the user's emotions and adjusting its response accordingly.
[2360] First, the user accesses the system's registration page from their smartphone or PC. The device displays a registration form to the user, who enters necessary information such as their name, email address, and password, and sends it from the device to the server. The server stores the received data in a database and notifies the user via the device that registration is complete. Next, the user accesses the login page and enters the registered email address and password. The device sends the entered data to the server, which verifies it against the information in the database. If authentication is successful, the server generates a login session and sends it to the device, displaying the user's dashboard.
[2361] When a user clicks the "Select Favorite Character / Idol" button in the main menu, the device sends a request to the server. The server retrieves a list of characters and idols from the database and sends it to the device. The device displays this list to the user, and when the user selects their desired character or idol, that information is sent to the server and recorded in the database.
[2362] When the user clicks the "Start Conversation" button, the device turns on the microphone and captures the user's voice. For example, if the user says "Hello, Naruto!", the audio data is sent from the device to the server. The server converts this audio data to text using the Google Cloud Speech-to-Text API and analyzes the text and emotion information with an emotion engine (e.g., Microsoft Azure Emotion API). If the analysis result is positive, the server sends the text data and emotion information to a generating AI model (e.g., OpenAI GPT-3) to generate a response such as "Hi! You seem to be doing well!". The generated response text is then synthesized into a character's voice using Amazon Polly and sent to the device for playback.
[2363] Furthermore, when the user selects the "Metaverse Mode" option, the device sends a request to the server. The server connects to a Unity-based metaverse platform and loads information for the virtual space. The server places the user's avatar and character / idol avatars in the virtual space, and the user accesses the virtual space by wearing a VR device (e.g., Oculus Rift). The device captures the user's movements in real time and sends them to the server. The server controls the movements of the virtual avatar through the Unity engine and simultaneously provides a more natural and realistic interactive experience through voice interaction via a generated AI model. In addition, an emotion engine recognizes the user's emotional changes in real time and reflects that information in the movements and facial expressions of the virtual avatar.
[2364] For example, if a user says "Hello, Naruto!" and the emotion engine recognizes this as joy, the generative AI model will generate the response "Hi! You seem to be doing well!". This response is synthesized into Naruto's voice and played back to the user through the speaker.
[2365] Examples of prompts for operating this system include:
[2366] Prompt to convert audio data to text:
[2367] Please convert the audio data 'Hello, Naruto!' to text.
[2368] Response generation prompt:
[2369] "Based on the text 'Hello, Naruto!' and the emotion 'Joyful,' please generate an appropriate response."
[2370] Text-to-speech prompt:
[2371] "Please convert the text 'Hey! You seem to be doing well!' into speech in Naruto's voice."
[2372] Through these processes, users can gain real-time interaction and an interactive experience in a virtual space.
[2373] The flow of the specific processing in Example 2 will be explained using Figure 13.
[2374] Program processing flow
[2375] Step 1:
[2376] Users access the system's registration page from their smartphones or PCs.
[2377] Input: Name, email address, password
[2378] The terminal receives these inputs, displays them on the registration form, and waits for user input.
[2379] Output: Input data
[2380] The user enters their name, email address, and password, and this data is sent to the device.
[2381] Step 2:
[2382] The terminal sends the input data to the server.
[2383] Input: User registration information (name, email address, password)
[2384] The server receives the input data and saves it to the database.
[2385] Output: Registration complete message
[2386] The device displays "Registration complete!" to the user.
[2387] Step 3:
[2388] The user accesses the login page and enters their email address and password.
[2389] Enter: Email address, password
[2390] The terminal sends the input data to the server.
[2391] Output: Authentication result
[2392] The server compares the information with the registration details in the database, and if a match is found, it generates a login session.
[2393] A login session is sent to the device, and the device displays the user's dashboard.
[2394] Step 4:
[2395] The user clicks the "Select Favorite Character / Idol" button in the main menu.
[2396] Input: Click Event
[2397] The terminal sends a request to the server.
[2398] Output: List of Characters / Idols
[2399] The server retrieves a list of characters and idols from the database and sends it to the terminal.
[2400] The device displays the list to the user.
[2401] Step 5:
[2402] The user selects their desired character or idol.
[2403] Input: Selected character or idol
[2404] The device sends the selection information to the server.
[2405] Output: Database update results
[2406] The server records that information in the database.
[2407] Step 6:
[2408] The user clicks the "Start Conversation" button.
[2409] Input: Click Event
[2410] The device turns on the microphone and captures the user's voice.
[2411] The user says, "Hello, [Character Name]!"
[2412] Output: Audio data
[2413] The audio data is then sent from the terminal to the server.
[2414] Step 7:
[2415] The server sends the voice data to the speech recognition service.
[2416] Input: Audio data
[2417] The speech recognition service converts speech into text.
[2418] Output: Text data
[2419] The server receives this text data.
[2420] Step 8:
[2421] The server recognizes emotions from voice and user facial expression data through an emotion engine.
[2422] Input: Voice data, user facial expression data
[2423] Output: Emotional information
[2424] The server uses text data and recognized emotion information to send to a generative AI model.
[2425] Step 9:
[2426] The generative AI model generates appropriate responses based on text data and sentiment information.
[2427] Input: Text data, sentiment information
[2428] Output: Response text
[2429] The server receives the response text.
[2430] Step 10:
[2431] The server processes the response text into a speech synthesis engine.
[2432] Input: Response text
[2433] The speech synthesis engine generates speech using the voice of a virtual character.
[2434] Output: Generated audio data
[2435] The server sends the audio data to the terminal.
[2436] Step 11:
[2437] The device plays the generated audio to the user.
[2438] Input: Audio data
[2439] The device plays audio through its speaker, and the user listens to the response.
[2440] Output: Played audio
[2441] Step 12:
[2442] The user selects the "Metaverse Mode" option.
[2443] Input: Click Event
[2444] The terminal sends a request to the server.
[2445] The server connects to the metaverse platform and loads the virtual space.
[2446] Output: Display data for the virtual space
[2447] Step 13:
[2448] The server places user avatars and character / idol avatars in the virtual space.
[2449] Input: None (automatic processing)
[2450] Users wear VR devices to access a virtual space.
[2451] Output: Initial placement of user and character avatars
[2452] Step 14:
[2453] The device captures the user's movements and sends them to the server.
[2454] Input: User activity information
[2455] The server controls the avatar in real time via the Unity engine.
[2456] Output: Avatar movement data
[2457] This allows users to enjoy interactive experiences in a virtual space.
[2458] The above is the specific processing flow of this system.
[2459] (Application Example 2)
[2460] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2461] Conventional virtual space systems have limited interaction between users and virtual characters or idols, lacking interaction based on emotions and actions. Therefore, users find it difficult to have a realistic conversational experience, resulting in a less realistic virtual experience overall. Furthermore, the lack of interactive dialogue systems utilizing emotion recognition is another challenge.
[2462] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for providing a generation model for generating the voice of a virtual character selected by the user; speech recognition means for converting the user's voice into text; artificial intelligence means for generating an appropriate response based on the generation model; speech synthesis means for converting the generated response into the voice of a virtual character; emotion recognition means for analyzing the user's emotions; means for adjusting the response based on the emotion recognition results; and means for playing the generated voice response to the user. This makes it possible for the user to enjoy a realistic dialogue experience with their favorite virtual character or idol that responds to their emotions.
[2463] A "user" is someone who uses the system to enjoy interacting with virtual characters or idols.
[2464] A "virtual space" is a three-dimensional virtual environment created by a computer program.
[2465] A "virtual character" is a character or idol selected by a user within a virtual space, whose voice and actions are generated based on AI.
[2466] A "generative model" refers to the algorithms and datasets used to generate voices and actions for virtual characters.
[2467] "Speech recognition means" refers to a technical device or program for generating text from a user's speech.
[2468] An "artificial intelligence system" is a system that uses a pre-trained model to generate an appropriate response to user input.
[2469] "Speech synthesis means" refers to a device or program that converts text data into the voice of a virtual character and generates speech data.
[2470] "Emotion recognition means" refers to technologies and programs that analyze and recognize emotions from a user's voice, facial expressions, etc.
[2471] "Means for adjusting responses" refer to technologies or programs for appropriately modifying responses generated based on the user's emotion recognition results.
[2472] "Means for reproducing voice responses" refers to devices or programs that output generated voice data to the user.
[2473] The embodiments for carrying out this invention are described in detail below.
[2474] First, the system consists of the following main components:
[2475] 1. A generative model for generating voices for virtual characters based on user selections.
[2476] 2. Speech recognition means
[2477] 3. Artificial intelligence means for generating appropriate responses
[2478] 4. Speech synthesis means for converting the generated response into speech.
[2479] 5. Emotion recognition means for analyzing user emotions
[2480] 6. Means for adjusting responses based on emotion recognition results
[2481] 7. Means for playing voice responses to the user
[2482] Specifically, when a user accesses the system and enters their email address and password on the login page, the server compares this information with existing data in the database. If the login is successful, the user's dashboard is displayed. The user then selects their preferred virtual character from the dashboard, and the server loads the character's voice model and animation data.
[2483] Hardware and software configuration
[2484] The hardware and software used in this invention are as follows:
[2485] Hardware: Smartphone, tablet, or VR headset, microphone, speaker
[2486] software:
[2487] Google Cloud Speech-to-Text: Used as a speech recognition method.
[2488] Azure Text Analytics: Used as a means of emotion recognition
[2489] OpenAI GPT-3: Used as an artificial intelligence means to generate appropriate responses.
[2490] Amazon Polly: Used as a speech synthesis method.
[2491] Data adjustment and calculation flow
[2492] 1. Speech recognition:
[2493] The audio data spoken by the user through the microphone is sent to the server and converted into text using Google Cloud Speech-to-Text.
[2494] 2. Emotion recognition:
[2495] The data, converted to text, is analyzed by Azure Text Analytics to recognize the user's sentiment.
[2496] 3. Response generation:
[2497] The server inputs the recognized emotion and text data into OpenAI GPT-3 and generates an appropriate response text.
[2498] 4. Speech synthesis:
[2499] The generated response text is converted into a virtual character's voice by Amazon Polly.
[2500] 5. Playback of voice response:
[2501] Audio data generated from the server is sent to the user's device, and the device plays the audio.
[2502] Specific example
[2503] For example, a user says "Hello, Naruto!" into the microphone. This audio is sent to the server. Google Cloud Speech-to-Text converts the audio into text "Hello, Naruto!". Next, Azure Text Analytics recognizes the user's emotion from this text as "joy". The server sends this recognized emotion and text to OpenAI GPT-3 to generate an appropriate response, "Hi! You seem to be doing well!". This response is converted into speech by Amazon Polly and played back on the user's device in the voice of the virtual character Naruto.
[2504] Example of a prompt
[2505] Joyful user: Hello, Naruto! Virtual character:
[2506] In this way, the system of the present invention enables users to enjoy a realistic dialogue experience that responds to their emotions.
[2507] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2508] Step 1:
[2509] The server receives the email address and password entered by the user and compares them with existing information in the database. If the entered email address and password match, the server creates a login session and displays the user's dashboard.
[2510] Step 2:
[2511] The user selects a virtual character from the dashboard. The device sends the selection information to the server, which loads the corresponding voice model and animation data from the database and presents it to the user.
[2512] Step 3:
[2513] When the user presses the "Start Conversation" button, the device turns on the microphone and captures the user's voice data. This voice data is sent to the server. The server uses Google Cloud Speech-to-Text to convert this voice data into text.
[2514] Step 4:
[2515] The server sends the converted text data to Azure Text Analytics for user sentiment analysis. The sentiment analysis results are stored along with the text data.
[2516] Step 5:
[2517] The server sends text data and emotion recognition results to OpenAI GPT-3 to generate an appropriate response. Specifically, natural-sounding dialogue is generated based on the prompt text and the recognized emotion information.
[2518] Step 6:
[2519] The generated response text is converted into audio data by Amazon Polly. The server then sends the generated audio data to the device.
[2520] Step 7:
[2521] The user's device plays the received audio data. This provides the user with the experience of a virtual character responding to them verbally.
[2522] Step 8:
[2523] User movement information is also captured in real time and sent from the terminal to the server. The server controls the movement of the avatar in the virtual space and adjusts the avatar to respond to the user's movements.
[2524] This allows users to enjoy realistic, emotion-driven conversations and action-based interactions within the virtual space.
[2525] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2526] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2527] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[2528] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[2529] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[2530] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[2531] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[2532] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[2533] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[2534] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[2535] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[2536] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[2537] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specifi...
Claims
1. A means for providing a generative model for generating the voice of a character selected by the user, A speech recognition means that converts user speech into text, An artificial intelligence means for generating responses based on a generative model, A speech synthesis means that converts the generated response into a character's voice, A means of playing a voice response to the user, A system that includes this.
2. The system according to claim 1, further comprising means for placing user and character avatars in a virtual space and controlling the avatars in response to user actions.
3. The system according to claim 1, further comprising means for capturing user action information in real time and reflecting it on an avatar in a virtual space.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A