System
The system simplifies the creation of personalized VTubers by using generative AI to generate and modify character images, addressing the complexity and cost barriers, allowing users to easily create and operate VTubers.
Patent Information
- Application Number
- JP2024118989
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Creating a virtual YouTuber (VTuber) requires specialized skills and high initial costs, making it difficult for ordinary users to easily start their VTuber careers, and the process of designing and setting up characters is complex without specialized knowledge.
A system that allows users to easily create personalized VTubers by receiving character settings, generating character images using generative AI, regenerating images based on user corrections, adding facial expressions and movements, and interacting through a chat-based interface, enabling users to download the final character.
Enables users to create high-quality VTuber characters without specialized knowledge, making it accessible and affordable for a wider audience.
Smart Images

Figure 2026017928000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Currently, creating a virtual YouTuber (VTuber) requires highly specialized skills and high initial costs, making it difficult for ordinary users to easily start working as a VTuber. Additionally, it is difficult to quickly and accurately design and set up characters without specialized knowledge. This has led to many potential VTubers giving up on their careers. [Means for solving the problem]
[0005] The present invention solves the above-mentioned problems by providing a system that allows users to easily create and activate personalized VTubers. The system of the present invention includes the following means:
[0006] 1. A means of receiving character settings from the user's device
[0007] 2. A method for generating character images using a generation AI based on the received character settings
[0008] 3. A method for sending the generated character image to the user's device
[0009] 4. A method to regenerate character images based on correction requests from users
[0010] 5. A way to save the final character image and allow users to download it
[0011] Furthermore, the system may include the following additional means:
[0012] A means of adding facial expressions and movements to the generated character images
[0013] A way to interact with users using a chat-based interface
[0014] This allows users to easily create high-quality VTuber characters and start their activities without the need for specialized knowledge.
[0015] "Generation processing device" refers to a system that receives input data from a user and performs processing based on that data.
[0016] "User's device" refers to electronic devices such as PCs and smartphones operated by users.
[0017] "Character settings" refers to information that indicates the appearance, personality, and other characteristics of the character that the user wants to create.
[0018] "Generative AI" refers to technology that uses artificial intelligence techniques to generate output based on specified input data.
[0019] "Character Image" refers to the visual representation of a virtual character generated by generative AI.
[0020] "Correction request" refers to a request input by a user to request corrections to a generated character image.
[0021] "Facial expression differences" refer to different facial expression patterns of a character.
[0022] "Movement" refers to the actions, such as animations and poses, that a character can take.
[0023] "Chat-based interface" refers to an interface that allows users and systems to communicate in the form of text chat.
[0024] Making the character image "downloadable" means that the user can save the generated character image to their own device via the Internet. [Brief explanation of the drawings]
[0025] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0026] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0027] First, the terms used in the following description will be explained.
[0028] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0029] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0030] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0031] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0032] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0033] [First embodiment]
[0034] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0035] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0036] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0037] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0038] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0039] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0040] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0041] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0042] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0043] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0044] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0045] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0046] This invention provides a system that utilizes generative AI to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, and a means of communication between the two.
[0047] 1. Enter character settings from the user's device
[0048] User:
[0049] Users access V-AI-Producers from a web browser on their device, and the system provides them with a chat-based interface.
[0050] For example, a user might enter text into a chat window saying, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input is sent from the user's device to the generation processing device as character settings.
[0051] 2. Character Generation
[0052] server:
[0053] The generation processing device receives the character settings sent by the user, and then generates a character image using a generation AI based on the settings.
[0054] For example, the generation processing device analyzes the settings such as "blue hair, cheerful personality, round eyes" and has the generation AI output a character image that matches those settings. This generated character image is temporarily saved and sent to the user's device.
[0055] 3. Character Modifications
[0056] user:
[0057] Users can check the preview image sent and enter correction requests if they are dissatisfied. For example, they can request that the eye color be changed to green.
[0058] Device:
[0059] The user's terminal transmits this modification request to the generation processing device.
[0060] server:
[0061] The generation processing unit receives the modification request and again instructs the generation AI to generate a character image with the eye color changed to green. The modified character image is again sent to the user for confirmation. This process is repeated until the user is satisfied.
[0062] 4. Adding facial expressions and actions
[0063] user:
[0064] Users can also add facial expressions and actions, for example, by requesting "add a smiling and surprised expression."
[0065] Device:
[0066] The user's terminal sends this request to the generation processing device.
[0067] server:
[0068] The generation processing device instructs the generation AI to generate facial expression differences and actions, and generates images or animations of new corresponding facial expressions and actions. These generated facial expression differences and actions are sent to the user for confirmation.
[0069] 5. Final check and download of character
[0070] user:
[0071] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[0072] Device:
[0073] The user's terminal transmits this final confirmation data to the generation processing device.
[0074] server:
[0075] The generation processor stores the final version of the character in a database and generates a download link for the user, which is sent to the user's device, allowing the user to save the generated character on their device.
[0076] This allows anyone, even those without specialized knowledge, to easily create a high-quality VTuber character and start activities with that character. This system is user-friendly and offered at an affordable price, so it is expected to be adopted by many users.
[0077] The processing flow will be explained below.
[0078] Step 1:
[0079] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[0080] Step 2:
[0081] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[0082] Step 3:
[0083] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[0084] Step 4:
[0085] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[0086] Step 5:
[0087] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[0088] Step 6:
[0089] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[0090] Step 7:
[0091] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[0092] Step 8:
[0093] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[0094] Step 9:
[0095] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[0096] Step 10:
[0097] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[0098] Step 11:
[0099] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[0100] Step 12:
[0101] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[0102] Step 13:
[0103] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[0104] Step 14:
[0105] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[0106] Example 1
[0107] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0108] Currently, there are several tools on the market for generating high-quality virtual characters (VTubers), but most of them require specialized knowledge and are difficult for average users to use. Furthermore, the character generation and modification process is often complex and time-consuming. Furthermore, the process of adding facial expressions and movements to the generated character image is also time-consuming, so a simple system that allows users to efficiently create characters that they are satisfied with is needed.
[0109] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0110] In this invention, the server includes means for receiving character settings from a user terminal, means for generating a character image using a generative AI model, means for sending the generated character image to the user terminal, means for generating a new character image based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for interacting with the user using a chat-based interface, means for generating and adding facial expressions and movements using the generative AI model, means for sending prompt text to the generative AI model based on a modification request from the user terminal, means for analyzing the prompt text and creating optimized input data for the generative AI model, means for temporarily saving the generated character image, and means for repeatedly generating and modifying the character image according to the user's satisfaction. This makes it possible to easily generate, modify, and add facial expressions and movements to personalized, high-quality character images without specialized knowledge.
[0111] "User terminal" means a device such as a computer, tablet, or smartphone that a user accesses and operates.
[0112] "Character settings" are information about the character's attributes and characteristics that the user specifies for the generated AI.
[0113] A "generative AI model" is an artificial intelligence algorithm or system that generates images or videos based on a given prompt.
[0114] A "prompt sentence" is text-based input data used to give instructions to a generative AI model.
[0115] A "generation processing device" is a device or server that receives character settings and generates images and videos using a generative AI model.
[0116] "Character Image" is an image of a virtual character created by a generative AI model based on user settings.
[0117] A "modification request" is an instruction sent by a user to request changes or modifications to a generated character image.
[0118] "Expressions and actions" refers to the changes in a character's emotions and actions that are added to the character image.
[0119] A "chat-based interface" is a means of communication that allows users and systems to interact through text messages.
[0120] "Downloadable" means that the final generated character images and data are provided in a format that can be saved on the user's device.
[0121] This invention relates to a system that allows users to easily generate and modify personalized virtual characters using generative AI models. The main components of this system are a generation processing device (server), a user terminal, and communication means between them.
[0122] First, the user accesses the system from a web browser on their own device. The user device is a computer device such as a PC, tablet, or smartphone. The system provides a chat-based interface through which the user can input text to configure their character. For example, the user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." The user's input data is sent to the server as the character configuration.
[0123] The server receives the character settings sent by the user and generates a character image using a generative AI model based on those settings. Examples of generative AI models used include DALL-E and GAN (generative adversarial network)-based AI. The generation processing device converts the received data into a format that the AI model can understand, and sends a generation prompt to the AI model to instruct it to generate an image. The generated character image is temporarily stored on the server and then sent to the user's device.
[0124] The user checks the preview image sent and inputs correction requests as needed. For example, they can input "Please change the eye color to green." The device then sends this correction request to the server. The server receives the correction request and again sends instructions to the generation AI model, regenerating the character image according to the corrections. The corrected character image is then saved again on the server and sent to the user's device.
[0125] Furthermore, users can add facial expressions and movements to their characters. For example, they can request the addition of a smiling and surprised expression. The device sends this request to the server. The server then instructs the AI model to add the new facial expressions and movements, generating images or animations of the new expressions and movements. The generated facial expressions and movements are then sent back to the user's device for confirmation.
[0126] In the final confirmation stage, the user checks the generated character, its expressions, and movements, and confirms by clicking "OK." The device then sends the final confirmation data to the server. The server then saves the final version of the character in a database and generates a link that the user can use to download it. This link is then sent to the user's device, allowing the user to save the generated character on their device.
[0127] This makes it possible to easily generate, modify, and add facial expressions and movements to high-quality virtual characters, even without specialized knowledge.
[0128] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0129] Step 1:
[0130] User inputs character settings
[0131] Users access the system through a web browser on their device, where a chat-based interface is displayed, allowing them to enter their character settings.
[0132] Input: Enter the following text: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[0133] Processing: The terminal receives the user's input, converts the data into JSON format, and sends it to the server. Specifically, the input data is converted into object format and sent over the network.
[0134] Output: Character setting data is sent to the server.
[0135] Step 2:
[0136] The server receives the character settings.
[0137] The server receives the character setting data sent from the user's terminal.
[0138] Input: Character configuration data sent from the user's device.
[0139] Processing: Analyzes the incoming data and converts it into a format that the generative AI model can understand. Specifically, it organizes the text data into prompt sentences and forms generative prompts.
[0140] Output: A generated prompt is sent to the generative AI model.
[0141] Step 3:
[0142] The server generates the character image
[0143] The server uses a generative AI model to generate a character image based on the received character settings.
[0144] Input: A generative prompt (e.g., "Her hair color is blue, her personality is cheerful, and her eyes are round").
[0145] Processing: The generated prompt is sent to the AI model, which then generates a corresponding character image. Specifically, the AI model analyzes the prompt and generates the corresponding image data. The generated image is temporarily stored in the server's storage.
[0146] Output: The generated character image is saved and sent to the user's device.
[0147] Step 4:
[0148] User checks the character image and enters correction request
[0149] The user checks the preview image sent and enters correction requests if necessary.
[0150] Input: Type your request: "Please change my eye color to green."
[0151] Processing: The device sends a correction request to the server. Specifically, the correction content is converted into JSON format and sent again to the server.
[0152] Output: The modified request data is sent to the server.
[0153] Step 5:
[0154] The server regenerates the modified request
[0155] The server receives the correction request and sends instructions to the generation AI model again, generating a character image with the eye color changed to green.
[0156] Input: Your correction request (e.g., "Please change my eye color to green").
[0157] Processing: A new prompt (e.g., "Change the eye color to green") is sent to the generative AI model to instruct it to regenerate. Specifically, the AI model regenerates the image data based on the modification request and stores it again in the server storage.
[0158] Output: The regenerated character image is sent to the user's device.
[0159] Step 6:
[0160] User requests for additional facial expressions and actions
[0161] Users also request the addition of different facial expressions and movements.
[0162] Input: "Add a smiling and surprised expression"
[0163] Processing: The device sends this request to the server. Specifically, it converts the request content into JSON format and sends it to the server.
[0164] Output: Requests for additional facial expressions and actions are sent to the server.
[0165] Step 7:
[0166] The server generates facial expression differences and actions
[0167] The server instructs the generative AI model to add facial expression differences and movements.
[0168] Input: Request for additional facial expressions or actions (e.g., "Please add a smiling and surprised expression").
[0169] Processing: New prompts are sent to the generative AI model to generate the corresponding facial expressions and actions. Specifically, the AI model generates new images and animations based on the additional prompts and saves them in server storage.
[0170] Output: The generated facial expression differences and movement data are sent to the user's device.
[0171] Step 8:
[0172] The user makes a final confirmation and downloads
[0173] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[0174] Enter: "OK" to confirm.
[0175] Processing: The device sends the final confirmation data to the server, which saves the final version of the character to the database and generates a download link.
[0176] Output: A downloadable link is generated and sent to the user's device.
[0177] Step 9:
[0178] User downloads character
[0179] Users save the generated characters on their devices.
[0180] Enter: Click on the download link.
[0181] Action: Download the linked file.
[0182] Output: The generated character is saved on the user's device.
[0183] (Application example 1)
[0184] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0185] In recent years, users have begun to demand online and virtual shopping experiences, and there is an increasing demand for personalized customer service and product explanations. However, traditional online shopping systems have difficulty providing the detailed product explanations and two-way dialogue that users desire. Furthermore, there is a lack of a way for users to create their own personalized characters and have them explain products, providing a more user-friendly shopping experience. This poses a challenge, making it difficult to improve user satisfaction and sales efficiency.
[0186] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0187] In this invention, the server includes a means for receiving character settings from a user's device, a means for generating a character image using a generation AI based on the character settings, and a means for transmitting the generated character image to the user's device. This allows users to easily create their own characters, receive product explanations through the characters, and receive answers to questions through dialogue. As a result, users can enjoy a more personalized shopping experience and overcome the drawbacks of traditional online shopping.
[0188] "Generation Processor" means a computer system for generating and modifying characters based on user input.
[0189] "User's terminal" refers to a device operated by a user, including a smartphone, tablet, PC, etc.
[0190] "Character settings" refers to information about the character's attributes and characteristics, such as hair color and personality, entered by the user.
[0191] "Generative AI" is an artificial intelligence technology that uses machine learning algorithms to generate character images.
[0192] "Character image" refers to image data of a visual character generated by a generation AI.
[0193] A "modification request" is an instruction sent by a user to modify a character's attributes or characteristics.
[0194] "Facial expression differences" are image data showing different facial expressions of a character.
[0195] "Action" is data that indicates the animation or movement that a character performs.
[0196] A "chat-based interface" is an interface that allows a user to interact with a system in a text-based manner.
[0197] "Product description" is information that the generated character uses to communicate the product's features and how to use it to the user.
[0198] "Answering questions through dialogue" is the process in which a character answers questions posed by a user through generated AI.
[0199] This invention provides a system that allows users to easily generate personalized characters and use them to improve shopping in virtual stores. The main components include a generation processing device, a user terminal, and means for communicating between them.
[0200] First, the user accesses the system using a web browser on their device (smartphone, tablet, PC, etc.). A chat-based interface is provided, through which the user can input character settings. For example, the user can input a prompt such as, "I want to create a character with blue hair, a cheerful personality, and round eyes." This information is sent to the generation processing device as the character settings.
[0201] The generation processing device analyzes the received character settings and generates a character image using a generation AI. The generation AI used is a high-performance model such as GPT-4 or Stable Diffusion. The generated character image is temporarily saved and sent to the user's device. The user can check this as a preview and, if necessary, send a request to make corrections. For example, they can request to change the eye color to green.
[0202] The modification request is sent to the generation processing device again, and the generation AI generates a new character image. This process is repeated until the user is satisfied. The final generated character image is saved in the database and available for download by the user.
[0203] Furthermore, users can use the generated character to check product descriptions and features. When the store clerk character explains a product, for example, they might say, "This smartwatch is equipped with the latest sensors and is useful for health management." The system also has a function where users can ask questions in a dialogue format, and the generated character will respond appropriately to the question. For example, if a user asks, "How do I use this product?", the character will respond, "This smartwatch is easy to use by operating the touchscreen."
[0204] The system integrates a server (including a database such as Firebase), user devices, and generative AI, and is provided as a cross-platform mobile application using React Native, enabling users to have an engaging and personalized virtual store experience.
[0205] Example prompt sentence:
[0206] Character generation: "Create a character with blonde hair, sporty personality, and round glasses"
[0207] Product description: "Describe the features of a newly released smartwatch in a friendly and detailed manner"
[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0209] Step 1:
[0210] Users access the system through a web browser on their own device. They use a chat-based interface to input their character settings, such as a specific prompt, such as "I want a character with blue hair, a cheerful personality, and round eyes." This input information is sent to the system as the character settings.
[0211] Input: Character settings entered by the user (text format)
[0212] Output: Character setting data sent to the server
[0213] Step 2:
[0214] The server receives the character setting data sent by the user and generates a character image using a generation AI (e.g., GPT-4 or Stable Diffusion).
[0215] Input: Character setting data (received from the user's device)
[0216] Output: Generated character image
[0217] Specific behavior:
[0218] Send a prompt to the generation AI: "Generate a character with blue hair, a cheerful personality, and round eyes."
[0219] The image is generated by the AI and temporarily saved.
[0220] Step 3:
[0221] The server sends the generated character image to the user's device, where the user can check the character image as a preview.
[0222] Input: Generated character image
[0223] Output: Preview of character image on user device
[0224] Step 4:
[0225] If the user is dissatisfied with the character image, they can input a request for correction, such as "I want the eye color to be changed to green," and send that request from their device to the server.
[0226] Input: Correction request (sent from the user's device)
[0227] Output: Modified request data sent to the server
[0228] Step 5:
[0229] The server receives the modification request and issues a prompt to the AI again. The AI then generates a new character image that reflects the specified modifications. This generation process is repeated until the user is satisfied.
[0230] Input: Correction request data
[0231] Output: Modified character image
[0232] Specific behavior:
[0233] Send the generator AI a new prompt: "Generate a character with blue hair, a cheerful personality, and green eyes."
[0234] Generate an image and save it again
[0235] Step 6:
[0236] The user finally confirms the character image they are satisfied with and requests a download. The server saves the final version of the character image in the database and generates a download link. The server sends this link to the user, who then downloads the character image.
[0237] Input: User confirmation and download request
[0238] Output: Generate and send a download link, final character image
[0239] Step 7:
[0240] The user checks the product description and features using the generated character, and the server displays the product description through the character and answers the user's questions in an interactive format.
[0241] Input: User question and selected product information
[0242] Output: Product description and answer by the character
[0243] Specific behavior:
[0244] Use generative AI to generate a product description by sending a prompt: "Please describe this product."
[0245] The character presents the generated explanation to the user.
[0246] Answering user questions
[0247] Through this entire process, users can create their own character and receive product information and purchasing assistance through that character.
[0248] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0249] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, an emotion engine, and communication means between these devices.
[0250] 1. Enter character settings from the user's device
[0251] User:
[0252] Users access "V-AI-Producers" through a web browser on their device, where they are presented with a chat-based interface through which they can input their character's characteristics and settings.
[0253] For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user terminal to the generation processing device as character settings.
[0254] 2. Character Generation
[0255] server:
[0256] The generation processing device analyzes the character settings received from the user and, based on the analysis results, sends a character generation request to the generation AI model.
[0257] For example, the AI analyzes the settings such as "blue hair, cheerful personality, round eyes" and generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[0258] 3. Emotion engine that recognizes user emotions
[0259] User:
[0260] Users can check the preview image and input correction requests if they are dissatisfied. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends that information to the server.
[0261] For example, if the user has a dissatisfied expression when entering a modification request such as "I want my eyes to be green," the emotion engine will send that emotional data to the server.
[0262] 4. Modifying characters based on emotional data
[0263] server:
[0264] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Based on the emotion data, the generation AI fine-tunes the character image.
[0265] For example, if the user has a dissatisfied expression, the generative AI will take the user's emotions into account and change their eye color to more closely match their desired look.
[0266] 5. Adding facial expressions and actions
[0267] User:
[0268] If the user wants to add facial expressions or actions, they can input a request. For example, they can request to add "smiling and surprised expressions."
[0269] server:
[0270] The server receives this request and issues instructions to the AI to generate new facial expressions and movements. The generated expressions and movements are stored on the server and sent to the user.
[0271] 6. Chat-based interface adjustments
[0272] server:
[0273] The server uses data from the emotion engine to tailor responses in the chat-based interface: if the user is stressed, for example, the system will take this into account and use kinder words.
[0274] 7. Final confirmation and download
[0275] user:
[0276] The user checks the final character, its facial expressions, and movements, and then gives a final confirmation by clicking "OK." The final confirmation data is then sent from the device to the server.
[0277] server:
[0278] The server saves the final character image and generates a downloadable link for the user, which is sent to the user's device and the user can click on the link to download the character image.
[0279] summary
[0280] This invention allows users to easily create high-quality VTuber characters without specialized knowledge, and utilizes an emotion engine to enjoy a more personalized character experience. The combination of the emotion engine increases user satisfaction and enables more natural and consistent character generation.
[0281] The processing flow will be explained below.
[0282] Step 1:
[0283] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[0284] Step 2:
[0285] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[0286] Step 3:
[0287] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[0288] Step 4:
[0289] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[0290] Step 5:
[0291] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[0292] Step 6:
[0293] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[0294] Step 7:
[0295] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[0296] Step 8:
[0297] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[0298] Step 9:
[0299] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[0300] Step 10:
[0301] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[0302] Step 11:
[0303] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[0304] Step 12:
[0305] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[0306] Step 13:
[0307] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[0308] Step 14:
[0309] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[0310] Step 15:
[0311] While the user is inputting their character settings, their emotions are recognized through the device's built-in camera and microphone. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server.
[0312] Step 16:
[0313] The server receives emotional data from the emotion engine and automatically adjusts the character design based on this data. For example, if the user has a dissatisfied expression, the server will instruct the generation AI to correct it and generate a character image that more closely matches the user's desire.
[0314] Step 17:
[0315] If the user wants to add facial expression differences or actions to the final character, the user inputs a request for this. For example, the user may request "I want a smiling and surprised expression added."
[0316] Step 18:
[0317] The server receives the user's request and has the AI generate new facial expressions and movements. The generated facial expressions and movements are stored on the server and sent to the user's device.
[0318] Step 19:
[0319] The server uses data from the emotion engine to tailor responses in the chat-based interface: if a user is feeling stressed, for example, the system will take this into account and use kinder words.
[0320] Example 2
[0321] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0322] While conventional VTuber character generation systems can create character settings based on user requests, they have problems in that they are unable to fully increase user satisfaction because they do not adequately accommodate requests for modifications or the reflection of emotions in the generated characters.Furthermore, there are few systems that can take user emotions into account when fine-tuning the generated characters or adding variations in facial expressions and movements, making it difficult to create characters that meet the user's intentions.
[0323] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from the user's information terminal, means for generating a character image using a generation artificial intelligence based on the character settings, means for transmitting the generated character image to the user's information terminal, means for regenerating the character image based on a modification request from the user, means for recognizing the user's emotions and fine-tuning the character image based on the recognition result, and means for saving the final character image and making it available for download by the user. This enables character generation and modification taking the user's emotions into consideration, thereby increasing user satisfaction. Furthermore, it is possible to add variations in facial expressions and movements to the generated character, providing a more personalized character experience.
[0324] A "generation processing device" is a device that receives character settings from a user's information terminal and generates, modifies, and saves a character image using generation artificial intelligence.
[0325] "User's information terminal" refers to the device through which the user accesses the system, inputs character settings, and checks and edits the generated character image. Specifically, this applies to a computer or smartphone.
[0326] "Character settings" refers to input information that specifically specifies the characteristics and personality of the virtual YouTuber created by the user, such as hair color, eye shape, and personality.
[0327] "Generative AI" is an AI technology for automatically generating character images based on character settings received from users.
[0328] "Character Image" refers to a visual image of a virtual YouTuber generated by generative artificial intelligence.
[0329] A "modification request" is an instruction from a user to change or modify a generated character image.
[0330] "Emotion recognition" is the process of detecting emotions from the user's facial expressions, voice, etc., and analyzing that information.
[0331] "Fine-tuning" refers to making small changes or improvements to existing character images based on user sentiment and correction requests.
[0332] "Variations in facial expressions and movements" enrich the character's expression by adding different facial expressions and movements to the generated character image.
[0333] "Downloadable Link" means the URL from which a user can download the final character image via the Internet.
[0334] This invention provides a system that utilizes generative artificial intelligence and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user's information terminal, an emotion engine, and communication means between these devices.
[0335] Users access this system from a web browser on their own information terminal. Once accessed, a chat-based interface is displayed, through which the user can input the characteristics and settings of their character. Specifically, the user inputs a prompt in text format, such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user's terminal to the generation processing device as the character settings.
[0336] The generation processing device receives and analyzes the character settings. Based on the analysis results, it sends a character generation request to the generative AI model. This generative AI model generates a character image using advanced generation techniques such as "GPT-3" or "DALL-E." The generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[0337] The user checks the preview image and, if dissatisfied, inputs a request for correction. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion data is analyzed by the emotion engine and sent to the server.
[0338] The server receives the user's modification request and emotion data from the emotion engine. Based on this, it issues instructions for regenerating the character, and the AI generator fine-tunes the character image. After the modification is complete, a preview image is sent to the user's device again. For example, if the user expresses dissatisfaction, the AI generator will take the user's emotion into account and change the eye color to match their desired color.
[0339] Users can also add different facial expressions and movements to the generated character. In this case, the user inputs a request such as "I want to add a smiling and surprised expression." Based on this request, the server issues instructions to the generation AI to generate new facial expressions and movements. The generated expressions and movements are saved on the server and made available to users.
[0340] The server also uses data from the emotion engine to tailor the chat interface's responses: if a user is stressed, for example, the system will adjust its responses to be more gentle.
[0341] The user checks the final character, its facial expressions, and movements, and gives a final confirmation by clicking "OK." This final confirmation data is sent from the user's device to the server. The server saves the character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[0342] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0343] Step 1:
[0344] User:
[0345] Users access the system through a web browser on an information terminal, where a chat-based interface is displayed and users input their character's characteristics and settings.
[0346] Input: Prompt: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[0347] Output: Character setting information
[0348] Specific operation: When the user inputs text into the interface and presses the send button, the character setting information is sent to the generation processing device.
[0349] Step 2:
[0350] server:
[0351] The server receives the character setting information sent from the user's terminal.
[0352] Input: Character setting information
[0353] Output: Input data to the generative AI model
[0354] Specific operation: The analysis module in the server analyzes the character setting information and converts it into a format that sends a character generation request to the generation AI model.
[0355] Step 3:
[0356] Generative AI models:
[0357] The generative AI model generates character images based on character setting information sent from the server.
[0358] Input: Character setting information
[0359] Output: Character image
[0360] Specific operation: A generative AI model (e.g., GPT-3 or DALL-E) runs an image generation algorithm based on the character settings to generate a character image.
[0361] Step 4:
[0362] server:
[0363] The server receives the generated character image, temporarily stores the image for preview, and sends it to the user's terminal.
[0364] Input: Character image
[0365] Output: Preview image
[0366] Specific operation: The generated character image is saved on the server, and its URL and image data are sent to the user's device.
[0367] Step 5:
[0368] User:
[0369] Users can view preview images and enter correction requests if necessary, and emotion data is collected using the device's camera and microphone.
[0370] Input: Preview image, correction request, emotion data
[0371] Output: Modification request and emotion data
[0372] How it works: When the user enters and submits the corrections in text, the data recorded by the camera and microphone is analyzed by the emotion engine.
[0373] Step 6:
[0374] server:
[0375] The server receives correction requests and emotion data from the user and issues correction instructions to the generative AI model.
[0376] Input: Modification request, emotion data
[0377] Output: Correction instructions
[0378] How it works: The server's algorithm analyzes the correction request and emotion data, converts it into the necessary correction instructions, and sends them to the generative AI model.
[0379] Step 7:
[0380] Generative AI models:
[0381] The generative AI model fine-tunes the character image based on correction instructions and generates a new image.
[0382] Input: Correction instructions
[0383] Output: Modified character image
[0384] Specific operation: The generative AI model runs the image generation algorithm again to generate a new character image that reflects the modifications.
[0385] Step 8:
[0386] server:
[0387] The server receives the modified character image and sends it back to the user's device.
[0388] Input: Modified character image
[0389] Output: Updated preview image
[0390] Specific operation: The modified character image is saved on the server and an updated preview image is sent to the user's device.
[0391] Step 9:
[0392] User:
[0393] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[0394] Input: Final confirmation data
[0395] Output: Sending final confirmation data
[0396] Specific operation: After the user confirms, he / she presses the OK button to send the final confirmation data.
[0397] Step 10:
[0398] server:
[0399] The server saves the final character image, generates a downloadable link for the user, and sends it to the user's device.
[0400] Input: Final confirmation data
[0401] Output: Download link
[0402] Specific operation: After final confirmation, the character image is saved on the server, and a download link is generated and sent to the user's device.
[0403] (Application example 2)
[0404] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0405] Previous virtual YouTuber (VTuber) generation systems were difficult to use if the user did not have specialized knowledge, and they lacked personalization based on the user's emotions. Furthermore, these systems have not yet been applied to products suggestions in virtual stores, making it impossible to provide an effective shopping experience for each individual user.
[0406] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from a user's terminal, means for generating a character image using a generation AI based on the character settings, means for sending the generated character image to the user's terminal, means for generating a character image again based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for analyzing the user's emotional data and modifying the character based on the emotion, and means for making product suggestions using the character. This allows users to easily generate personalized VTuber characters without specialized knowledge and enjoy emotion-based modifications and product suggestions.
[0407] A "generation processing device" is a device that generates a character image using a generation AI based on character settings, and generates the character image again in response to a modification request.
[0408] "User's terminal" refers to a device that a user uses to input character settings, check generated character images, and request modifications, and includes personal computers, smartphones, etc.
[0409] "Character settings" refers to information about the characteristics and attributes of a character entered by the user, and includes specific settings such as hair color, personality, and eye shape.
[0410] "Generation AI" is artificial intelligence that generates character images based on input setting information.
[0411] "Character Image" is a digital representation of a visual character created by generative AI.
[0412] "Emotional data" is information about emotions analyzed from the user's facial expressions, voice, etc.
[0413] "Product suggestion" refers to recommending suitable products to users within a virtual store based on their emotional data.
[0414] A "modification request" is a request for changes or improvements that a user makes to a generated character image.
[0415] A "chat-based interface" is an interface that allows a user and a system to interact through text or voice.
[0416] "Different facial expressions and movements" are variations of different facial expressions and movements that are added to the character image.
[0417] "Means for enabling users to download" refers to a method for allowing users to download the final generated character image to their terminal.
[0418] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber) or virtual store assistant. The main hardware required to implement this invention includes a user terminal, a generation processing device, an emotion engine, and communication means between these devices.
[0419] Enter character settings from the user's device
[0420] Users access the "V-Shop Assistant" from a web browser on their device. Once accessed, a chat-based interface is displayed. Through this interface, users can input their character's characteristics and settings. For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is sent from the user's device to the generation processing device as the character settings.
[0421] Examples:
[0422] - Prompt: "I want to create a character with blue hair, a bright personality, and round eyes."
[0423] Character Generation
[0424] The generation processing device analyzes the character settings received from the user. Based on the analysis results, it uses the generation AI model to send a request to generate a character image. For example, settings such as "blue hair color, cheerful personality, round eyes" are analyzed, and the generation AI generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[0425] Emotion engine that recognizes user emotions
[0426] The user checks the preview image and, if dissatisfied, inputs a request for correction. Furthermore, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice, and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion engine will send this emotional data to the server.
[0427] Examples:
[0428] - Prompt: "I want my eyes to be green."
[0429] Character modification based on emotional data
[0430] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Depending on the emotion data, the generation AI fine-tunes the character image. For example, if the user has a dissatisfied expression, the generation AI will take the user's emotion into account and change the eye color to more closely match the user's desired color.
[0431] Adding facial expressions and actions
[0432] If the user wants to add additional facial expressions or actions, they can input a request. For example, they can request to add a "smiling and surprised expression." The server receives this request and instructs the AI to generate new facial expressions and actions. The generated expressions and actions are stored on the server and sent to the user.
[0433] Final confirmation and download
[0434] The user checks the final character, its facial expressions, and movements, and then clicks "OK" to confirm. The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[0435] Product proposals in virtual stores
[0436] The generated personalized character will then make product suggestions based on the user's emotional data: for example, if the user is feeling stressed, the character will suggest relaxation-related products.
[0437] Through this series of processes, users can easily generate high-quality characters without any specialized knowledge, and utilize the emotion engine to enjoy a more personalized experience.
[0438] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0439] Step 1:
[0440] The user accesses the "V-Shop Assistant" using a terminal and inputs character settings. For example, they might input setting information such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is then sent from the user terminal to the generation processing device.
[0441] Input: Character settings (hair color, personality, eye shape)
[0442] Output: The configuration information is sent to the generation processing device.
[0443] Step 2:
[0444] The generation processing device analyzes the character settings received from the user and sends a character generation request to the generation AI. The generation AI model generates a character image based on the settings, and the character image is temporarily stored on the server.
[0445] Input: Character setting information
[0446] Output: The generated character image is saved on the server and a preview image is sent to the user's device.
[0447] Step 3:
[0448] The user checks the preview image and inputs a modification request. For example, they may request, "I want the eye color to be changed to green." The user's device then sends this modification request to the server. The device's camera and microphone are also used to collect the user's emotional data (facial expressions and voice).
[0449] Input: Correction request, emotional data (facial expression, voice)
[0450] Output: The modification request and emotion data are sent to the server.
[0451] Step 4:
[0452] The server analyzes the modification request and the emotion data, and instructs the generation AI model to regenerate the character. The generation AI modifies the character image taking into account the emotion data, saves the modified character image back to the server, and sends a preview image to the user.
[0453] Input: Correction request, emotion data
[0454] Output: The modified character image is saved to the server and a preview image is sent to the user.
[0455] Step 5:
[0456] The user makes a final check of the character image and confirms it by clicking "OK." The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download.
[0457] Input: Final confirmation data
[0458] Output: A downloadable link is generated and sent to the user
[0459] Step 6:
[0460] The user can then use the generated character to receive product suggestions in a virtual store. The character will then make personalized product suggestions based on emotional data. For example, if the user is feeling stressed, the character will suggest relaxation-related products.
[0461] Input: User emotion data
[0462] Output: Product proposals by characters
[0463] At each step, the server, device, and user interact with each other, using a generative AI model and emotion engine to generate personalized VTuber characters and recommend products. Based on input information, the generative AI model generates a character image, and the emotion engine analyzes the user's emotions, providing a more satisfying experience for the user.
[0464] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0465] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0466] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0467] [Second embodiment]
[0468] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0469] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0470] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0471] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0472] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0473] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0474] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0475] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0476] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0477] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0478] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0479] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0480] This invention provides a system that utilizes generative AI to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, and a means of communication between the two.
[0481] 1. Enter character settings from the user's device
[0482] User:
[0483] Users access V-AI-Producers from a web browser on their device, and the system provides them with a chat-based interface.
[0484] For example, a user might enter text into a chat window saying, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input is sent from the user's device to the generation processing device as character settings.
[0485] 2. Character Generation
[0486] server:
[0487] The generation processing device receives the character settings sent by the user, and then generates a character image using a generation AI based on the settings.
[0488] For example, the generation processing device analyzes the settings such as "blue hair, cheerful personality, round eyes" and has the generation AI output a character image that matches those settings. This generated character image is temporarily saved and sent to the user's device.
[0489] 3. Character Modifications
[0490] user:
[0491] Users can check the preview image sent and enter correction requests if they are dissatisfied. For example, they can request that the eye color be changed to green.
[0492] Device:
[0493] The user's terminal transmits this modification request to the generation processing device.
[0494] server:
[0495] The generation processing unit receives the modification request and again instructs the generation AI to generate a character image with the eye color changed to green. The modified character image is again sent to the user for confirmation. This process is repeated until the user is satisfied.
[0496] 4. Adding facial expressions and actions
[0497] user:
[0498] Users can also add facial expressions and actions, for example, by requesting "add a smiling and surprised expression."
[0499] Device:
[0500] The user's terminal sends this request to the generation processing device.
[0501] server:
[0502] The generation processing device instructs the generation AI to generate facial expression differences and actions, and generates images or animations of new corresponding facial expressions and actions. These generated facial expression differences and actions are sent to the user for confirmation.
[0503] 5. Final check and download of character
[0504] user:
[0505] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[0506] Device:
[0507] The user's terminal transmits this final confirmation data to the generation processing device.
[0508] server:
[0509] The generation processor stores the final version of the character in a database and generates a download link for the user, which is sent to the user's device, allowing the user to save the generated character on their device.
[0510] This allows anyone, even those without specialized knowledge, to easily create a high-quality VTuber character and start activities with that character. This system is user-friendly and offered at an affordable price, so it is expected to be adopted by many users.
[0511] The processing flow will be explained below.
[0512] Step 1:
[0513] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[0514] Step 2:
[0515] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[0516] Step 3:
[0517] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[0518] Step 4:
[0519] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[0520] Step 5:
[0521] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[0522] Step 6:
[0523] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[0524] Step 7:
[0525] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[0526] Step 8:
[0527] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[0528] Step 9:
[0529] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[0530] Step 10:
[0531] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[0532] Step 11:
[0533] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[0534] Step 12:
[0535] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[0536] Step 13:
[0537] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[0538] Step 14:
[0539] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[0540] Example 1
[0541] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0542] Currently, there are several tools on the market for generating high-quality virtual characters (VTubers), but most of them require specialized knowledge and are difficult for average users to use. Furthermore, the character generation and modification process is often complex and time-consuming. Furthermore, the process of adding facial expressions and movements to the generated character image is also time-consuming, so a simple system that allows users to efficiently create characters that they are satisfied with is needed.
[0543] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0544] In this invention, the server includes means for receiving character settings from a user terminal, means for generating a character image using a generative AI model, means for sending the generated character image to the user terminal, means for generating a new character image based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for interacting with the user using a chat-based interface, means for generating and adding facial expressions and movements using the generative AI model, means for sending prompt text to the generative AI model based on a modification request from the user terminal, means for analyzing the prompt text and creating optimized input data for the generative AI model, means for temporarily saving the generated character image, and means for repeatedly generating and modifying the character image according to the user's satisfaction. This makes it possible to easily generate, modify, and add facial expressions and movements to personalized, high-quality character images without specialized knowledge.
[0545] "User terminal" means a device such as a computer, tablet, or smartphone that a user accesses and operates.
[0546] "Character settings" are information about the character's attributes and characteristics that the user specifies for the generated AI.
[0547] A "generative AI model" is an artificial intelligence algorithm or system that generates images or videos based on a given prompt.
[0548] A "prompt sentence" is text-based input data used to give instructions to a generative AI model.
[0549] A "generation processing device" is a device or server that receives character settings and generates images and videos using a generative AI model.
[0550] "Character Image" is an image of a virtual character created by a generative AI model based on user settings.
[0551] A "modification request" is an instruction sent by a user to request changes or modifications to a generated character image.
[0552] "Expressions and actions" refers to the changes in a character's emotions and actions that are added to the character image.
[0553] A "chat-based interface" is a means of communication that allows users and systems to interact through text messages.
[0554] "Downloadable" means that the final generated character images and data are provided in a format that can be saved on the user's device.
[0555] This invention relates to a system that allows users to easily generate and modify personalized virtual characters using generative AI models. The main components of this system are a generation processing device (server), a user terminal, and communication means between them.
[0556] First, the user accesses the system from a web browser on their own device. The user device is a computer device such as a PC, tablet, or smartphone. The system provides a chat-based interface through which the user can input text to configure their character. For example, the user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." The user's input data is sent to the server as the character configuration.
[0557] The server receives the character settings sent by the user and generates a character image using a generative AI model based on those settings. Examples of generative AI models used include DALL-E and GAN (generative adversarial network)-based AI. The generation processing device converts the received data into a format that the AI model can understand, and sends a generation prompt to the AI model to instruct it to generate an image. The generated character image is temporarily stored on the server and then sent to the user's device.
[0558] The user checks the preview image sent and inputs correction requests as needed. For example, they can input "Please change the eye color to green." The device then sends this correction request to the server. The server receives the correction request and again sends instructions to the generation AI model, regenerating the character image according to the corrections. The corrected character image is then saved again on the server and sent to the user's device.
[0559] Furthermore, users can add facial expressions and movements to their characters. For example, they can request the addition of a smiling and surprised expression. The device sends this request to the server. The server then instructs the AI model to add the new facial expressions and movements, generating images or animations of the new expressions and movements. The generated facial expressions and movements are then sent back to the user's device for confirmation.
[0560] In the final confirmation stage, the user checks the generated character, its expressions, and movements, and confirms by clicking "OK." The device then sends the final confirmation data to the server. The server then saves the final version of the character in a database and generates a link that the user can use to download it. This link is then sent to the user's device, allowing the user to save the generated character on their device.
[0561] This makes it possible to easily generate, modify, and add facial expressions and movements to high-quality virtual characters, even without specialized knowledge.
[0562] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0563] Step 1:
[0564] User inputs character settings
[0565] Users access the system through a web browser on their device, where a chat-based interface is displayed, allowing them to enter their character settings.
[0566] Input: Enter the following text: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[0567] Processing: The terminal receives the user's input, converts the data into JSON format, and sends it to the server. Specifically, the input data is converted into object format and sent over the network.
[0568] Output: Character setting data is sent to the server.
[0569] Step 2:
[0570] The server receives the character settings.
[0571] The server receives the character setting data sent from the user's terminal.
[0572] Input: Character configuration data sent from the user's device.
[0573] Processing: Analyzes the incoming data and converts it into a format that the generative AI model can understand. Specifically, it organizes the text data into prompt sentences and forms generative prompts.
[0574] Output: A generated prompt is sent to the generative AI model.
[0575] Step 3:
[0576] The server generates the character image
[0577] The server uses a generative AI model to generate a character image based on the received character settings.
[0578] Input: A generative prompt (e.g., "Her hair color is blue, her personality is cheerful, and her eyes are round").
[0579] Processing: The generated prompt is sent to the AI model, which then generates a corresponding character image. Specifically, the AI model analyzes the prompt and generates the corresponding image data. The generated image is temporarily stored in the server's storage.
[0580] Output: The generated character image is saved and sent to the user's device.
[0581] Step 4:
[0582] User checks the character image and enters correction request
[0583] The user checks the preview image sent and enters correction requests if necessary.
[0584] Input: Type your request: "Please change my eye color to green."
[0585] Processing: The device sends a correction request to the server. Specifically, the correction content is converted into JSON format and sent again to the server.
[0586] Output: The modified request data is sent to the server.
[0587] Step 5:
[0588] The server regenerates the modified request
[0589] The server receives the correction request and sends instructions to the generation AI model again, generating a character image with the eye color changed to green.
[0590] Input: Your correction request (e.g., "Please change my eye color to green").
[0591] Processing: A new prompt (e.g., "Change the eye color to green") is sent to the generative AI model to instruct it to regenerate. Specifically, the AI model regenerates the image data based on the modification request and stores it again in the server storage.
[0592] Output: The regenerated character image is sent to the user's device.
[0593] Step 6:
[0594] User requests for additional facial expressions and actions
[0595] Users also request the addition of different facial expressions and movements.
[0596] Input: "Add a smiling and surprised expression"
[0597] Processing: The device sends this request to the server. Specifically, it converts the request content into JSON format and sends it to the server.
[0598] Output: Requests for additional facial expressions and actions are sent to the server.
[0599] Step 7:
[0600] The server generates facial expression differences and actions
[0601] The server instructs the generative AI model to add facial expression differences and movements.
[0602] Input: Request for additional facial expressions or actions (e.g., "Please add a smiling and surprised expression").
[0603] Processing: New prompts are sent to the generative AI model to generate the corresponding facial expressions and actions. Specifically, the AI model generates new images and animations based on the additional prompts and saves them in server storage.
[0604] Output: The generated facial expression differences and movement data are sent to the user's device.
[0605] Step 8:
[0606] The user makes a final confirmation and downloads
[0607] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[0608] Enter: "OK" to confirm.
[0609] Processing: The device sends the final confirmation data to the server, which saves the final version of the character to the database and generates a download link.
[0610] Output: A downloadable link is generated and sent to the user's device.
[0611] Step 9:
[0612] User downloads character
[0613] Users save the generated characters on their devices.
[0614] Enter: Click on the download link.
[0615] Action: Download the linked file.
[0616] Output: The generated character is saved on the user's device.
[0617] (Application example 1)
[0618] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0619] In recent years, users have begun to demand online and virtual shopping experiences, and there is an increasing demand for personalized customer service and product explanations. However, traditional online shopping systems have difficulty providing the detailed product explanations and two-way dialogue that users desire. Furthermore, there is a lack of a way for users to create their own personalized characters and have them explain products, providing a more user-friendly shopping experience. This poses a challenge, making it difficult to improve user satisfaction and sales efficiency.
[0620] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0621] In this invention, the server includes a means for receiving character settings from a user's device, a means for generating a character image using a generation AI based on the character settings, and a means for transmitting the generated character image to the user's device. This allows users to easily create their own characters, receive product explanations through the characters, and receive answers to questions through dialogue. As a result, users can enjoy a more personalized shopping experience and overcome the drawbacks of traditional online shopping.
[0622] "Generation Processor" means a computer system for generating and modifying characters based on user input.
[0623] "User's terminal" refers to a device operated by a user, including a smartphone, tablet, PC, etc.
[0624] "Character settings" refers to information about the character's attributes and characteristics, such as hair color and personality, entered by the user.
[0625] "Generative AI" is an artificial intelligence technology that uses machine learning algorithms to generate character images.
[0626] "Character image" refers to image data of a visual character generated by a generation AI.
[0627] A "modification request" is an instruction sent by a user to modify a character's attributes or characteristics.
[0628] "Facial expression differences" are image data showing different facial expressions of a character.
[0629] "Action" is data that indicates the animation or movement that a character performs.
[0630] A "chat-based interface" is an interface that allows a user to interact with a system in a text-based manner.
[0631] "Product description" is information that the generated character uses to communicate the product's features and how to use it to the user.
[0632] "Answering questions through dialogue" is the process in which a character answers questions posed by a user through generated AI.
[0633] This invention provides a system that allows users to easily generate personalized characters and use them to improve shopping in virtual stores. The main components include a generation processing device, a user terminal, and means for communicating between them.
[0634] First, the user accesses the system using a web browser on their device (smartphone, tablet, PC, etc.). A chat-based interface is provided, through which the user can input character settings. For example, the user can input a prompt such as, "I want to create a character with blue hair, a cheerful personality, and round eyes." This information is sent to the generation processing device as the character settings.
[0635] The generation processing device analyzes the received character settings and generates a character image using a generation AI. The generation AI used is a high-performance model such as GPT-4 or Stable Diffusion. The generated character image is temporarily saved and sent to the user's device. The user can check this as a preview and, if necessary, send a request to make corrections. For example, they can request to change the eye color to green.
[0636] The modification request is sent to the generation processing device again, and the generation AI generates a new character image. This process is repeated until the user is satisfied. The final generated character image is saved in the database and available for download by the user.
[0637] Furthermore, users can use the generated character to check product descriptions and features. When the store clerk character explains a product, for example, they might say, "This smartwatch is equipped with the latest sensors and is useful for health management." The system also has a function where users can ask questions in a dialogue format, and the generated character will respond appropriately to the question. For example, if a user asks, "How do I use this product?", the character will respond, "This smartwatch is easy to use by operating the touchscreen."
[0638] The system integrates a server (including a database such as Firebase), user devices, and generative AI, and is provided as a cross-platform mobile application using React Native, enabling users to have an engaging and personalized virtual store experience.
[0639] Example prompt sentence:
[0640] Character generation: "Create a character with blonde hair, sporty personality, and round glasses"
[0641] Product description: "Describe the features of a newly released smartwatch in a friendly and detailed manner"
[0642] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0643] Step 1:
[0644] Users access the system through a web browser on their own device. They use a chat-based interface to input their character settings, such as a specific prompt, such as "I want a character with blue hair, a cheerful personality, and round eyes." This input information is sent to the system as the character settings.
[0645] Input: Character settings entered by the user (text format)
[0646] Output: Character setting data sent to the server
[0647] Step 2:
[0648] The server receives the character setting data sent by the user and generates a character image using a generation AI (e.g., GPT-4 or Stable Diffusion).
[0649] Input: Character setting data (received from the user's device)
[0650] Output: Generated character image
[0651] Specific behavior:
[0652] Send a prompt to the generation AI: "Generate a character with blue hair, a cheerful personality, and round eyes."
[0653] The image is generated by the AI and temporarily saved.
[0654] Step 3:
[0655] The server sends the generated character image to the user's device, where the user can check the character image as a preview.
[0656] Input: Generated character image
[0657] Output: Preview of character image on user device
[0658] Step 4:
[0659] If the user is dissatisfied with the character image, they can input a request for correction, such as "I want the eye color to be changed to green," and send that request from their device to the server.
[0660] Input: Correction request (sent from the user's device)
[0661] Output: Modified request data sent to the server
[0662] Step 5:
[0663] The server receives the modification request and issues a prompt to the AI again. The AI then generates a new character image that reflects the specified modifications. This generation process is repeated until the user is satisfied.
[0664] Input: Correction request data
[0665] Output: Modified character image
[0666] Specific behavior:
[0667] Send the generator AI a new prompt: "Generate a character with blue hair, a cheerful personality, and green eyes."
[0668] Generate an image and save it again
[0669] Step 6:
[0670] The user finally confirms the character image they are satisfied with and requests a download. The server saves the final version of the character image in the database and generates a download link. The server sends this link to the user, who then downloads the character image.
[0671] Input: User confirmation and download request
[0672] Output: Generate and send a download link, final character image
[0673] Step 7:
[0674] The user checks the product description and features using the generated character, and the server displays the product description through the character and answers the user's questions in an interactive format.
[0675] Input: User question and selected product information
[0676] Output: Product description and answer by the character
[0677] Specific behavior:
[0678] Use generative AI to generate a product description by sending a prompt: "Please describe this product."
[0679] The character presents the generated explanation to the user.
[0680] Answering user questions
[0681] Through this entire process, users can create their own character and receive product information and purchasing assistance through that character.
[0682] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0683] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, an emotion engine, and communication means between these devices.
[0684] 1. Enter character settings from the user's device
[0685] User:
[0686] Users access "V-AI-Producers" through a web browser on their device, where they are presented with a chat-based interface through which they can input their character's characteristics and settings.
[0687] For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user terminal to the generation processing device as character settings.
[0688] 2. Character Generation
[0689] server:
[0690] The generation processing device analyzes the character settings received from the user and, based on the analysis results, sends a character generation request to the generation AI model.
[0691] For example, the AI analyzes the settings such as "blue hair, cheerful personality, round eyes" and generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[0692] 3. Emotion engine that recognizes user emotions
[0693] User:
[0694] Users can check the preview image and input correction requests if they are dissatisfied. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends that information to the server.
[0695] For example, if the user has a dissatisfied expression when entering a modification request such as "I want my eyes to be green," the emotion engine will send that emotional data to the server.
[0696] 4. Modifying characters based on emotional data
[0697] server:
[0698] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Based on the emotion data, the generation AI fine-tunes the character image.
[0699] For example, if the user has a dissatisfied expression, the generative AI will take the user's emotions into account and change their eye color to more closely match their desired look.
[0700] 5. Adding facial expressions and actions
[0701] User:
[0702] If the user wants to add facial expressions or actions, they can input a request. For example, they can request to add "smiling and surprised expressions."
[0703] server:
[0704] The server receives this request and issues instructions to the AI to generate new facial expressions and movements. The generated expressions and movements are stored on the server and sent to the user.
[0705] 6. Chat-based interface adjustments
[0706] server:
[0707] The server uses data from the emotion engine to tailor responses in the chat-based interface: if the user is stressed, for example, the system will take this into account and use kinder words.
[0708] 7. Final confirmation and download
[0709] user:
[0710] The user checks the final character, its facial expressions, and movements, and then gives a final confirmation by clicking "OK." The final confirmation data is then sent from the device to the server.
[0711] server:
[0712] The server saves the final character image and generates a downloadable link for the user, which is sent to the user's device and the user can click on the link to download the character image.
[0713] summary
[0714] This invention allows users to easily create high-quality VTuber characters without specialized knowledge, and utilizes an emotion engine to enjoy a more personalized character experience. The combination of the emotion engine increases user satisfaction and enables more natural and consistent character generation.
[0715] The processing flow will be explained below.
[0716] Step 1:
[0717] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[0718] Step 2:
[0719] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[0720] Step 3:
[0721] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[0722] Step 4:
[0723] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[0724] Step 5:
[0725] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[0726] Step 6:
[0727] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[0728] Step 7:
[0729] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[0730] Step 8:
[0731] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[0732] Step 9:
[0733] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[0734] Step 10:
[0735] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[0736] Step 11:
[0737] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[0738] Step 12:
[0739] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[0740] Step 13:
[0741] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[0742] Step 14:
[0743] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[0744] Step 15:
[0745] While the user is inputting their character settings, their emotions are recognized through the device's built-in camera and microphone. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server.
[0746] Step 16:
[0747] The server receives emotional data from the emotion engine and automatically adjusts the character design based on this data. For example, if the user has a dissatisfied expression, the server will instruct the generation AI to correct it and generate a character image that more closely matches the user's desire.
[0748] Step 17:
[0749] If the user wants to add facial expression differences or actions to the final character, the user inputs a request for this. For example, the user may request "I want a smiling and surprised expression added."
[0750] Step 18:
[0751] The server receives the user's request and has the AI generate new facial expressions and movements. The generated facial expressions and movements are stored on the server and sent to the user's device.
[0752] Step 19:
[0753] The server uses data from the emotion engine to tailor responses in the chat-based interface: if a user is feeling stressed, for example, the system will take this into account and use kinder words.
[0754] Example 2
[0755] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0756] While conventional VTuber character generation systems can create character settings based on user requests, they have problems in that they are unable to fully increase user satisfaction because they do not adequately accommodate requests for modifications or the reflection of emotions in the generated characters.Furthermore, there are few systems that can take user emotions into account when fine-tuning the generated characters or adding variations in facial expressions and movements, making it difficult to create characters that meet the user's intentions.
[0757] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from the user's information terminal, means for generating a character image using a generation artificial intelligence based on the character settings, means for transmitting the generated character image to the user's information terminal, means for regenerating the character image based on a modification request from the user, means for recognizing the user's emotions and fine-tuning the character image based on the recognition result, and means for saving the final character image and making it available for download by the user. This enables character generation and modification taking the user's emotions into consideration, thereby increasing user satisfaction. Furthermore, it is possible to add variations in facial expressions and movements to the generated character, providing a more personalized character experience.
[0758] A "generation processing device" is a device that receives character settings from a user's information terminal and generates, modifies, and saves a character image using generation artificial intelligence.
[0759] "User's information terminal" refers to the device through which the user accesses the system, inputs character settings, and checks and edits the generated character image. Specifically, this applies to a computer or smartphone.
[0760] "Character settings" refers to input information that specifically specifies the characteristics and personality of the virtual YouTuber created by the user, such as hair color, eye shape, and personality.
[0761] "Generative AI" is an AI technology for automatically generating character images based on character settings received from users.
[0762] "Character Image" refers to a visual image of a virtual YouTuber generated by generative artificial intelligence.
[0763] A "modification request" is an instruction from a user to change or modify a generated character image.
[0764] "Emotion recognition" is the process of detecting emotions from the user's facial expressions, voice, etc., and analyzing that information.
[0765] "Fine-tuning" refers to making small changes or improvements to existing character images based on user sentiment and correction requests.
[0766] "Variations in facial expressions and movements" enrich the character's expression by adding different facial expressions and movements to the generated character image.
[0767] "Downloadable Link" means the URL from which a user can download the final character image via the Internet.
[0768] This invention provides a system that utilizes generative artificial intelligence and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user's information terminal, an emotion engine, and communication means between these devices.
[0769] Users access this system from a web browser on their own information terminal. Once accessed, a chat-based interface is displayed, through which the user can input the characteristics and settings of their character. Specifically, the user inputs a prompt in text format, such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user's terminal to the generation processing device as the character settings.
[0770] The generation processing device receives and analyzes the character settings. Based on the analysis results, it sends a character generation request to the generative AI model. This generative AI model generates a character image using advanced generation techniques such as "GPT-3" or "DALL-E." The generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[0771] The user checks the preview image and, if dissatisfied, inputs a request for correction. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion data is analyzed by the emotion engine and sent to the server.
[0772] The server receives the user's modification request and emotion data from the emotion engine. Based on this, it issues instructions for regenerating the character, and the AI generator fine-tunes the character image. After the modification is complete, a preview image is sent to the user's device again. For example, if the user expresses dissatisfaction, the AI generator will take the user's emotion into account and change the eye color to match their desired color.
[0773] Users can also add different facial expressions and movements to the generated character. In this case, the user inputs a request such as "I want to add a smiling and surprised expression." Based on this request, the server issues instructions to the generation AI to generate new facial expressions and movements. The generated expressions and movements are saved on the server and made available to users.
[0774] The server also uses data from the emotion engine to tailor the chat interface's responses: if a user is stressed, for example, the system will adjust its responses to be more gentle.
[0775] The user checks the final character, its facial expressions, and movements, and gives a final confirmation by clicking "OK." This final confirmation data is sent from the user's device to the server. The server saves the character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[0776] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0777] Step 1:
[0778] User:
[0779] Users access the system through a web browser on an information terminal, where a chat-based interface is displayed and users input their character's characteristics and settings.
[0780] Input: Prompt: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[0781] Output: Character setting information
[0782] Specific operation: When the user inputs text into the interface and presses the send button, the character setting information is sent to the generation processing device.
[0783] Step 2:
[0784] server:
[0785] The server receives the character setting information sent from the user's terminal.
[0786] Input: Character setting information
[0787] Output: Input data to the generative AI model
[0788] Specific operation: The analysis module in the server analyzes the character setting information and converts it into a format that sends a character generation request to the generation AI model.
[0789] Step 3:
[0790] Generative AI models:
[0791] The generative AI model generates character images based on character setting information sent from the server.
[0792] Input: Character setting information
[0793] Output: Character image
[0794] Specific operation: A generative AI model (e.g., GPT-3 or DALL-E) runs an image generation algorithm based on the character settings to generate a character image.
[0795] Step 4:
[0796] server:
[0797] The server receives the generated character image, temporarily stores the image for preview, and sends it to the user's terminal.
[0798] Input: Character image
[0799] Output: Preview image
[0800] Specific operation: The generated character image is saved on the server, and its URL and image data are sent to the user's device.
[0801] Step 5:
[0802] User:
[0803] Users can view preview images and enter correction requests if necessary, and emotion data is collected using the device's camera and microphone.
[0804] Input: Preview image, correction request, emotion data
[0805] Output: Modification request and emotion data
[0806] How it works: When the user enters and submits the corrections in text, the data recorded by the camera and microphone is analyzed by the emotion engine.
[0807] Step 6:
[0808] server:
[0809] The server receives correction requests and emotion data from the user and issues correction instructions to the generative AI model.
[0810] Input: Modification request, emotion data
[0811] Output: Correction instructions
[0812] How it works: The server's algorithm analyzes the correction request and emotion data, converts it into the necessary correction instructions, and sends them to the generative AI model.
[0813] Step 7:
[0814] Generative AI models:
[0815] The generative AI model fine-tunes the character image based on correction instructions and generates a new image.
[0816] Input: Correction instructions
[0817] Output: Modified character image
[0818] Specific operation: The generative AI model runs the image generation algorithm again to generate a new character image that reflects the modifications.
[0819] Step 8:
[0820] server:
[0821] The server receives the modified character image and sends it back to the user's device.
[0822] Input: Modified character image
[0823] Output: Updated preview image
[0824] Specific operation: The modified character image is saved on the server and an updated preview image is sent to the user's device.
[0825] Step 9:
[0826] User:
[0827] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[0828] Input: Final confirmation data
[0829] Output: Sending final confirmation data
[0830] Specific operation: After the user confirms, he / she presses the OK button to send the final confirmation data.
[0831] Step 10:
[0832] server:
[0833] The server saves the final character image, generates a downloadable link for the user, and sends it to the user's device.
[0834] Input: Final confirmation data
[0835] Output: Download link
[0836] Specific operation: After final confirmation, the character image is saved on the server, and a download link is generated and sent to the user's device.
[0837] (Application example 2)
[0838] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0839] Previous virtual YouTuber (VTuber) generation systems were difficult to use if the user did not have specialized knowledge, and they lacked personalization based on the user's emotions. Furthermore, these systems have not yet been applied to products suggestions in virtual stores, making it impossible to provide an effective shopping experience for each individual user.
[0840] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from a user's terminal, means for generating a character image using a generation AI based on the character settings, means for sending the generated character image to the user's terminal, means for generating a character image again based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for analyzing the user's emotional data and modifying the character based on the emotion, and means for making product suggestions using the character. This allows users to easily generate personalized VTuber characters without specialized knowledge and enjoy emotion-based modifications and product suggestions.
[0841] A "generation processing device" is a device that generates a character image using a generation AI based on character settings, and generates the character image again in response to a modification request.
[0842] "User's terminal" refers to a device that a user uses to input character settings, check generated character images, and request modifications, and includes personal computers, smartphones, etc.
[0843] "Character settings" refers to information about the characteristics and attributes of a character entered by the user, and includes specific settings such as hair color, personality, and eye shape.
[0844] "Generation AI" is artificial intelligence that generates character images based on input setting information.
[0845] "Character Image" is a digital representation of a visual character created by generative AI.
[0846] "Emotional data" is information about emotions analyzed from the user's facial expressions, voice, etc.
[0847] "Product suggestion" refers to recommending suitable products to users within a virtual store based on their emotional data.
[0848] A "modification request" is a request for changes or improvements that a user makes to a generated character image.
[0849] A "chat-based interface" is an interface that allows a user and a system to interact through text or voice.
[0850] "Different facial expressions and movements" are variations of different facial expressions and movements that are added to the character image.
[0851] "Means for enabling users to download" refers to a method for allowing users to download the final generated character image to their terminal.
[0852] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber) or virtual store assistant. The main hardware required to implement this invention includes a user terminal, a generation processing device, an emotion engine, and communication means between these devices.
[0853] Enter character settings from the user's device
[0854] Users access the "V-Shop Assistant" from a web browser on their device. Once accessed, a chat-based interface is displayed. Through this interface, users can input their character's characteristics and settings. For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is sent from the user's device to the generation processing device as the character settings.
[0855] Examples:
[0856] - Prompt: "I want to create a character with blue hair, a bright personality, and round eyes."
[0857] Character Generation
[0858] The generation processing device analyzes the character settings received from the user. Based on the analysis results, it uses the generation AI model to send a request to generate a character image. For example, settings such as "blue hair color, cheerful personality, round eyes" are analyzed, and the generation AI generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[0859] Emotion engine that recognizes user emotions
[0860] The user checks the preview image and, if dissatisfied, inputs a request for correction. Furthermore, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice, and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion engine will send this emotional data to the server.
[0861] Examples:
[0862] - Prompt: "I want my eyes to be green."
[0863] Character modification based on emotional data
[0864] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Depending on the emotion data, the generation AI fine-tunes the character image. For example, if the user has a dissatisfied expression, the generation AI will take the user's emotion into account and change the eye color to more closely match the user's desired color.
[0865] Adding facial expressions and actions
[0866] If the user wants to add additional facial expressions or actions, they can input a request. For example, they can request to add a "smiling and surprised expression." The server receives this request and instructs the AI to generate new facial expressions and actions. The generated expressions and actions are stored on the server and sent to the user.
[0867] Final confirmation and download
[0868] The user checks the final character, its facial expressions, and movements, and then clicks "OK" to confirm. The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[0869] Product proposals in virtual stores
[0870] The generated personalized character will then make product suggestions based on the user's emotional data: for example, if the user is feeling stressed, the character will suggest relaxation-related products.
[0871] Through this series of processes, users can easily generate high-quality characters without any specialized knowledge, and utilize the emotion engine to enjoy a more personalized experience.
[0872] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0873] Step 1:
[0874] The user accesses the "V-Shop Assistant" using a terminal and inputs character settings. For example, they might input setting information such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is then sent from the user terminal to the generation processing device.
[0875] Input: Character settings (hair color, personality, eye shape)
[0876] Output: The configuration information is sent to the generation processing device.
[0877] Step 2:
[0878] The generation processing device analyzes the character settings received from the user and sends a character generation request to the generation AI. The generation AI model generates a character image based on the settings, and the character image is temporarily stored on the server.
[0879] Input: Character setting information
[0880] Output: The generated character image is saved on the server and a preview image is sent to the user's device.
[0881] Step 3:
[0882] The user checks the preview image and inputs a modification request. For example, they may request, "I want the eye color to be changed to green." The user's device then sends this modification request to the server. The device's camera and microphone are also used to collect the user's emotional data (facial expressions and voice).
[0883] Input: Correction request, emotional data (facial expression, voice)
[0884] Output: The modification request and emotion data are sent to the server.
[0885] Step 4:
[0886] The server analyzes the modification request and the emotion data, and instructs the generation AI model to regenerate the character. The generation AI modifies the character image taking into account the emotion data, saves the modified character image back to the server, and sends a preview image to the user.
[0887] Input: Correction request, emotion data
[0888] Output: The modified character image is saved to the server and a preview image is sent to the user.
[0889] Step 5:
[0890] The user makes a final check of the character image and confirms it by clicking "OK." The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download.
[0891] Input: Final confirmation data
[0892] Output: A downloadable link is generated and sent to the user
[0893] Step 6:
[0894] The user can then use the generated character to receive product suggestions in a virtual store. The character will then make personalized product suggestions based on emotional data. For example, if the user is feeling stressed, the character will suggest relaxation-related products.
[0895] Input: User emotion data
[0896] Output: Product proposals by characters
[0897] At each step, the server, device, and user interact with each other, using a generative AI model and emotion engine to generate personalized VTuber characters and recommend products. Based on input information, the generative AI model generates a character image, and the emotion engine analyzes the user's emotions, providing a more satisfying experience for the user.
[0898] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0899] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0900] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0901] [Third embodiment]
[0902] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0903] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0904] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0905] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0906] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0907] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0908] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0909] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0910] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0911] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0912] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0913] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0914] This invention provides a system that utilizes generative AI to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, and a means of communication between the two.
[0915] 1. Enter character settings from the user's device
[0916] User:
[0917] Users access V-AI-Producers from a web browser on their device, and the system provides them with a chat-based interface.
[0918] For example, a user might enter text into a chat window saying, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input is sent from the user's device to the generation processing device as character settings.
[0919] 2. Character Generation
[0920] server:
[0921] The generation processing device receives the character settings sent by the user, and then generates a character image using a generation AI based on the settings.
[0922] For example, the generation processing device analyzes the settings such as "blue hair, cheerful personality, round eyes" and has the generation AI output a character image that matches those settings. This generated character image is temporarily saved and sent to the user's device.
[0923] 3. Character Modifications
[0924] user:
[0925] Users can check the preview image sent and enter correction requests if they are dissatisfied. For example, they can request that the eye color be changed to green.
[0926] Device:
[0927] The user's terminal transmits this modification request to the generation processing device.
[0928] server:
[0929] The generation processing unit receives the modification request and again instructs the generation AI to generate a character image with the eye color changed to green. The modified character image is again sent to the user for confirmation. This process is repeated until the user is satisfied.
[0930] 4. Adding facial expressions and actions
[0931] user:
[0932] Users can also add facial expressions and actions, for example, by requesting "add a smiling and surprised expression."
[0933] Device:
[0934] The user's terminal sends this request to the generation processing device.
[0935] server:
[0936] The generation processing device instructs the generation AI to generate facial expression differences and actions, and generates images or animations of new corresponding facial expressions and actions. These generated facial expression differences and actions are sent to the user for confirmation.
[0937] 5. Final check and download of character
[0938] user:
[0939] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[0940] Device:
[0941] The user's terminal transmits this final confirmation data to the generation processing device.
[0942] server:
[0943] The generation processor stores the final version of the character in a database and generates a download link for the user, which is sent to the user's device, allowing the user to save the generated character on their device.
[0944] This allows anyone, even those without specialized knowledge, to easily create a high-quality VTuber character and start activities with that character. This system is user-friendly and offered at an affordable price, so it is expected to be adopted by many users.
[0945] The processing flow will be explained below.
[0946] Step 1:
[0947] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[0948] Step 2:
[0949] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[0950] Step 3:
[0951] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[0952] Step 4:
[0953] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[0954] Step 5:
[0955] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[0956] Step 6:
[0957] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[0958] Step 7:
[0959] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[0960] Step 8:
[0961] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[0962] Step 9:
[0963] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[0964] Step 10:
[0965] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[0966] Step 11:
[0967] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[0968] Step 12:
[0969] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[0970] Step 13:
[0971] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[0972] Step 14:
[0973] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[0974] Example 1
[0975] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0976] Currently, there are several tools on the market for generating high-quality virtual characters (VTubers), but most of them require specialized knowledge and are difficult for average users to use. Furthermore, the character generation and modification process is often complex and time-consuming. Furthermore, the process of adding facial expressions and movements to the generated character image is also time-consuming, so a simple system that allows users to efficiently create characters that they are satisfied with is needed.
[0977] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0978] In this invention, the server includes means for receiving character settings from a user terminal, means for generating a character image using a generative AI model, means for sending the generated character image to the user terminal, means for generating a new character image based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for interacting with the user using a chat-based interface, means for generating and adding facial expressions and movements using the generative AI model, means for sending prompt text to the generative AI model based on a modification request from the user terminal, means for analyzing the prompt text and creating optimized input data for the generative AI model, means for temporarily saving the generated character image, and means for repeatedly generating and modifying the character image according to the user's satisfaction. This makes it possible to easily generate, modify, and add facial expressions and movements to personalized, high-quality character images without specialized knowledge.
[0979] "User terminal" means a device such as a computer, tablet, or smartphone that a user accesses and operates.
[0980] "Character settings" are information about the character's attributes and characteristics that the user specifies for the generated AI.
[0981] A "generative AI model" is an artificial intelligence algorithm or system that generates images or videos based on a given prompt.
[0982] A "prompt sentence" is text-based input data used to give instructions to a generative AI model.
[0983] A "generation processing device" is a device or server that receives character settings and generates images and videos using a generative AI model.
[0984] "Character Image" is an image of a virtual character created by a generative AI model based on user settings.
[0985] A "modification request" is an instruction sent by a user to request changes or modifications to a generated character image.
[0986] "Expressions and actions" refers to the changes in a character's emotions and actions that are added to the character image.
[0987] A "chat-based interface" is a means of communication that allows users and systems to interact through text messages.
[0988] "Downloadable" means that the final generated character images and data are provided in a format that can be saved on the user's device.
[0989] This invention relates to a system that allows users to easily generate and modify personalized virtual characters using generative AI models. The main components of this system are a generation processing device (server), a user terminal, and communication means between them.
[0990] First, the user accesses the system from a web browser on their own device. The user device is a computer device such as a PC, tablet, or smartphone. The system provides a chat-based interface through which the user can input text to configure their character. For example, the user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." The user's input data is sent to the server as the character configuration.
[0991] The server receives the character settings sent by the user and generates a character image using a generative AI model based on those settings. Examples of generative AI models used include DALL-E and GAN (generative adversarial network)-based AI. The generation processing device converts the received data into a format that the AI model can understand, and sends a generation prompt to the AI model to instruct it to generate an image. The generated character image is temporarily stored on the server and then sent to the user's device.
[0992] The user checks the preview image sent and inputs correction requests as needed. For example, they can input "Please change the eye color to green." The device then sends this correction request to the server. The server receives the correction request and again sends instructions to the generation AI model, regenerating the character image according to the corrections. The corrected character image is then saved again on the server and sent to the user's device.
[0993] Furthermore, users can add facial expressions and movements to their characters. For example, they can request the addition of a smiling and surprised expression. The device sends this request to the server. The server then instructs the AI model to add the new facial expressions and movements, generating images or animations of the new expressions and movements. The generated facial expressions and movements are then sent back to the user's device for confirmation.
[0994] In the final confirmation stage, the user checks the generated character, its expressions, and movements, and confirms by clicking "OK." The device then sends the final confirmation data to the server. The server then saves the final version of the character in a database and generates a link that the user can use to download it. This link is then sent to the user's device, allowing the user to save the generated character on their device.
[0995] This makes it possible to easily generate, modify, and add facial expressions and movements to high-quality virtual characters, even without specialized knowledge.
[0996] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0997] Step 1:
[0998] User inputs character settings
[0999] Users access the system through a web browser on their device, where a chat-based interface is displayed, allowing them to enter their character settings.
[1000] Input: Enter the following text: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[1001] Processing: The terminal receives the user's input, converts the data into JSON format, and sends it to the server. Specifically, the input data is converted into object format and sent over the network.
[1002] Output: Character setting data is sent to the server.
[1003] Step 2:
[1004] The server receives the character settings.
[1005] The server receives the character setting data sent from the user's terminal.
[1006] Input: Character configuration data sent from the user's device.
[1007] Processing: Analyzes the incoming data and converts it into a format that the generative AI model can understand. Specifically, it organizes the text data into prompt sentences and forms generative prompts.
[1008] Output: A generated prompt is sent to the generative AI model.
[1009] Step 3:
[1010] The server generates the character image
[1011] The server uses a generative AI model to generate a character image based on the received character settings.
[1012] Input: A generative prompt (e.g., "Her hair color is blue, her personality is cheerful, and her eyes are round").
[1013] Processing: The generated prompt is sent to the AI model, which then generates a corresponding character image. Specifically, the AI model analyzes the prompt and generates the corresponding image data. The generated image is temporarily stored in the server's storage.
[1014] Output: The generated character image is saved and sent to the user's device.
[1015] Step 4:
[1016] User checks the character image and enters correction request
[1017] The user checks the preview image sent and enters correction requests if necessary.
[1018] Input: Type your request: "Please change my eye color to green."
[1019] Processing: The device sends a correction request to the server. Specifically, the correction content is converted into JSON format and sent again to the server.
[1020] Output: The modified request data is sent to the server.
[1021] Step 5:
[1022] The server regenerates the modified request
[1023] The server receives the correction request and sends instructions to the generation AI model again, generating a character image with the eye color changed to green.
[1024] Input: Your correction request (e.g., "Please change my eye color to green").
[1025] Processing: A new prompt (e.g., "Change the eye color to green") is sent to the generative AI model to instruct it to regenerate. Specifically, the AI model regenerates the image data based on the modification request and stores it again in the server storage.
[1026] Output: The regenerated character image is sent to the user's device.
[1027] Step 6:
[1028] User requests for additional facial expressions and actions
[1029] Users also request the addition of different facial expressions and movements.
[1030] Input: "Add a smiling and surprised expression"
[1031] Processing: The device sends this request to the server. Specifically, it converts the request content into JSON format and sends it to the server.
[1032] Output: Requests for additional facial expressions and actions are sent to the server.
[1033] Step 7:
[1034] The server generates facial expression differences and actions
[1035] The server instructs the generative AI model to add facial expression differences and movements.
[1036] Input: Request for additional facial expressions or actions (e.g., "Please add a smiling and surprised expression").
[1037] Processing: New prompts are sent to the generative AI model to generate the corresponding facial expressions and actions. Specifically, the AI model generates new images and animations based on the additional prompts and saves them in server storage.
[1038] Output: The generated facial expression differences and movement data are sent to the user's device.
[1039] Step 8:
[1040] The user makes a final confirmation and downloads
[1041] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[1042] Enter: "OK" to confirm.
[1043] Processing: The device sends the final confirmation data to the server, which saves the final version of the character to the database and generates a download link.
[1044] Output: A downloadable link is generated and sent to the user's device.
[1045] Step 9:
[1046] User downloads character
[1047] Users save the generated characters on their devices.
[1048] Enter: Click on the download link.
[1049] Action: Download the linked file.
[1050] Output: The generated character is saved on the user's device.
[1051] (Application example 1)
[1052] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1053] In recent years, users have begun to demand online and virtual shopping experiences, and there is an increasing demand for personalized customer service and product explanations. However, traditional online shopping systems have difficulty providing the detailed product explanations and two-way dialogue that users desire. Furthermore, there is a lack of a way for users to create their own personalized characters and have them explain products, providing a more user-friendly shopping experience. This poses a challenge, making it difficult to improve user satisfaction and sales efficiency.
[1054] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1055] In this invention, the server includes a means for receiving character settings from a user's device, a means for generating a character image using a generation AI based on the character settings, and a means for transmitting the generated character image to the user's device. This allows users to easily create their own characters, receive product explanations through the characters, and receive answers to questions through dialogue. As a result, users can enjoy a more personalized shopping experience and overcome the drawbacks of traditional online shopping.
[1056] "Generation Processor" means a computer system for generating and modifying characters based on user input.
[1057] "User's terminal" refers to a device operated by a user, including a smartphone, tablet, PC, etc.
[1058] "Character settings" refers to information about the character's attributes and characteristics, such as hair color and personality, entered by the user.
[1059] "Generative AI" is an artificial intelligence technology that uses machine learning algorithms to generate character images.
[1060] "Character image" refers to image data of a visual character generated by a generation AI.
[1061] A "modification request" is an instruction sent by a user to modify a character's attributes or characteristics.
[1062] "Facial expression differences" are image data showing different facial expressions of a character.
[1063] "Action" is data that indicates the animation or movement that a character performs.
[1064] A "chat-based interface" is an interface that allows a user to interact with a system in a text-based manner.
[1065] "Product description" is information that the generated character uses to communicate the product's features and how to use it to the user.
[1066] "Answering questions through dialogue" is the process in which a character answers questions posed by a user through generated AI.
[1067] This invention provides a system that allows users to easily generate personalized characters and use them to improve shopping in virtual stores. The main components include a generation processing device, a user terminal, and means for communicating between them.
[1068] First, the user accesses the system using a web browser on their device (smartphone, tablet, PC, etc.). A chat-based interface is provided, through which the user can input character settings. For example, the user can input a prompt such as, "I want to create a character with blue hair, a cheerful personality, and round eyes." This information is sent to the generation processing device as the character settings.
[1069] The generation processing device analyzes the received character settings and generates a character image using a generation AI. The generation AI used is a high-performance model such as GPT-4 or Stable Diffusion. The generated character image is temporarily saved and sent to the user's device. The user can check this as a preview and, if necessary, send a request to make corrections. For example, they can request to change the eye color to green.
[1070] The modification request is sent to the generation processing device again, and the generation AI generates a new character image. This process is repeated until the user is satisfied. The final generated character image is saved in the database and available for download by the user.
[1071] Furthermore, users can use the generated character to check product descriptions and features. When the store clerk character explains a product, for example, they might say, "This smartwatch is equipped with the latest sensors and is useful for health management." The system also has a function where users can ask questions in a dialogue format, and the generated character will respond appropriately to the question. For example, if a user asks, "How do I use this product?", the character will respond, "This smartwatch is easy to use by operating the touchscreen."
[1072] The system integrates a server (including a database such as Firebase), user devices, and generative AI, and is provided as a cross-platform mobile application using React Native, enabling users to have an engaging and personalized virtual store experience.
[1073] Example prompt sentence:
[1074] Character generation: "Create a character with blonde hair, sporty personality, and round glasses"
[1075] Product description: "Describe the features of a newly released smartwatch in a friendly and detailed manner"
[1076] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1077] Step 1:
[1078] Users access the system through a web browser on their own device. They use a chat-based interface to input their character settings, such as a specific prompt, such as "I want a character with blue hair, a cheerful personality, and round eyes." This input information is sent to the system as the character settings.
[1079] Input: Character settings entered by the user (text format)
[1080] Output: Character setting data sent to the server
[1081] Step 2:
[1082] The server receives the character setting data sent by the user and generates a character image using a generation AI (e.g., GPT-4 or Stable Diffusion).
[1083] Input: Character setting data (received from the user's device)
[1084] Output: Generated character image
[1085] Specific behavior:
[1086] Send a prompt to the generation AI: "Generate a character with blue hair, a cheerful personality, and round eyes."
[1087] The image is generated by the AI and temporarily saved.
[1088] Step 3:
[1089] The server sends the generated character image to the user's device, where the user can check the character image as a preview.
[1090] Input: Generated character image
[1091] Output: Preview of character image on user device
[1092] Step 4:
[1093] If the user is dissatisfied with the character image, they can input a request for correction, such as "I want the eye color to be changed to green," and send that request from their device to the server.
[1094] Input: Correction request (sent from the user's device)
[1095] Output: Modified request data sent to the server
[1096] Step 5:
[1097] The server receives the modification request and issues a prompt to the AI again. The AI then generates a new character image that reflects the specified modifications. This generation process is repeated until the user is satisfied.
[1098] Input: Correction request data
[1099] Output: Modified character image
[1100] Specific behavior:
[1101] Send the generator AI a new prompt: "Generate a character with blue hair, a cheerful personality, and green eyes."
[1102] Generate an image and save it again
[1103] Step 6:
[1104] The user finally confirms the character image they are satisfied with and requests a download. The server saves the final version of the character image in the database and generates a download link. The server sends this link to the user, who then downloads the character image.
[1105] Input: User confirmation and download request
[1106] Output: Generate and send a download link, final character image
[1107] Step 7:
[1108] The user checks the product description and features using the generated character, and the server displays the product description through the character and answers the user's questions in an interactive format.
[1109] Input: User question and selected product information
[1110] Output: Product description and answer by the character
[1111] Specific behavior:
[1112] Use generative AI to generate a product description by sending a prompt: "Please describe this product."
[1113] The character presents the generated explanation to the user.
[1114] Answering user questions
[1115] Through this entire process, users can create their own character and receive product information and purchasing assistance through that character.
[1116] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1117] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, an emotion engine, and communication means between these devices.
[1118] 1. Enter character settings from the user's device
[1119] User:
[1120] Users access "V-AI-Producers" through a web browser on their device, where they are presented with a chat-based interface through which they can input their character's characteristics and settings.
[1121] For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user terminal to the generation processing device as character settings.
[1122] 2. Character Generation
[1123] server:
[1124] The generation processing device analyzes the character settings received from the user and, based on the analysis results, sends a character generation request to the generation AI model.
[1125] For example, the AI analyzes the settings such as "blue hair, cheerful personality, round eyes" and generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[1126] 3. Emotion engine that recognizes user emotions
[1127] User:
[1128] Users can check the preview image and input correction requests if they are dissatisfied. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends that information to the server.
[1129] For example, if the user has a dissatisfied expression when entering a modification request such as "I want my eyes to be green," the emotion engine will send that emotional data to the server.
[1130] 4. Modifying characters based on emotional data
[1131] server:
[1132] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Based on the emotion data, the generation AI fine-tunes the character image.
[1133] For example, if the user has a dissatisfied expression, the generative AI will take the user's emotions into account and change their eye color to more closely match their desired look.
[1134] 5. Adding facial expressions and actions
[1135] User:
[1136] If the user wants to add facial expressions or actions, they can input a request. For example, they can request to add "smiling and surprised expressions."
[1137] server:
[1138] The server receives this request and issues instructions to the AI to generate new facial expressions and movements. The generated expressions and movements are stored on the server and sent to the user.
[1139] 6. Chat-based interface adjustments
[1140] server:
[1141] The server uses data from the emotion engine to tailor responses in the chat-based interface: if the user is stressed, for example, the system will take this into account and use kinder words.
[1142] 7. Final confirmation and download
[1143] user:
[1144] The user checks the final character, its facial expressions, and movements, and then gives a final confirmation by clicking "OK." The final confirmation data is then sent from the device to the server.
[1145] server:
[1146] The server saves the final character image and generates a downloadable link for the user, which is sent to the user's device and the user can click on the link to download the character image.
[1147] summary
[1148] This invention allows users to easily create high-quality VTuber characters without specialized knowledge, and utilizes an emotion engine to enjoy a more personalized character experience. The combination of the emotion engine increases user satisfaction and enables more natural and consistent character generation.
[1149] The processing flow will be explained below.
[1150] Step 1:
[1151] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[1152] Step 2:
[1153] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[1154] Step 3:
[1155] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[1156] Step 4:
[1157] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[1158] Step 5:
[1159] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[1160] Step 6:
[1161] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[1162] Step 7:
[1163] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[1164] Step 8:
[1165] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[1166] Step 9:
[1167] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[1168] Step 10:
[1169] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[1170] Step 11:
[1171] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[1172] Step 12:
[1173] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[1174] Step 13:
[1175] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[1176] Step 14:
[1177] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[1178] Step 15:
[1179] While the user is inputting their character settings, their emotions are recognized through the device's built-in camera and microphone. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server.
[1180] Step 16:
[1181] The server receives emotional data from the emotion engine and automatically adjusts the character design based on this data. For example, if the user has a dissatisfied expression, the server will instruct the generation AI to correct it and generate a character image that more closely matches the user's desire.
[1182] Step 17:
[1183] If the user wants to add facial expression differences or actions to the final character, the user inputs a request for this. For example, the user may request "I want a smiling and surprised expression added."
[1184] Step 18:
[1185] The server receives the user's request and has the AI generate new facial expressions and movements. The generated facial expressions and movements are stored on the server and sent to the user's device.
[1186] Step 19:
[1187] The server uses data from the emotion engine to tailor responses in the chat-based interface: if a user is feeling stressed, for example, the system will take this into account and use kinder words.
[1188] Example 2
[1189] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1190] While conventional VTuber character generation systems can create character settings based on user requests, they have problems in that they are unable to fully increase user satisfaction because they do not adequately accommodate requests for modifications or the reflection of emotions in the generated characters.Furthermore, there are few systems that can take user emotions into account when fine-tuning the generated characters or adding variations in facial expressions and movements, making it difficult to create characters that meet the user's intentions.
[1191] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from the user's information terminal, means for generating a character image using a generation artificial intelligence based on the character settings, means for transmitting the generated character image to the user's information terminal, means for regenerating the character image based on a modification request from the user, means for recognizing the user's emotions and fine-tuning the character image based on the recognition result, and means for saving the final character image and making it available for download by the user. This enables character generation and modification taking the user's emotions into consideration, thereby increasing user satisfaction. Furthermore, it is possible to add variations in facial expressions and movements to the generated character, providing a more personalized character experience.
[1192] A "generation processing device" is a device that receives character settings from a user's information terminal and generates, modifies, and saves a character image using generation artificial intelligence.
[1193] "User's information terminal" refers to the device through which the user accesses the system, inputs character settings, and checks and edits the generated character image. Specifically, this applies to a computer or smartphone.
[1194] "Character settings" refers to input information that specifically specifies the characteristics and personality of the virtual YouTuber created by the user, such as hair color, eye shape, and personality.
[1195] "Generative AI" is an AI technology for automatically generating character images based on character settings received from users.
[1196] "Character Image" refers to a visual image of a virtual YouTuber generated by generative artificial intelligence.
[1197] A "modification request" is an instruction from a user to change or modify a generated character image.
[1198] "Emotion recognition" is the process of detecting emotions from the user's facial expressions, voice, etc., and analyzing that information.
[1199] "Fine-tuning" refers to making small changes or improvements to existing character images based on user sentiment and correction requests.
[1200] "Variations in facial expressions and movements" enrich the character's expression by adding different facial expressions and movements to the generated character image.
[1201] "Downloadable Link" means the URL from which a user can download the final character image via the Internet.
[1202] This invention provides a system that utilizes generative artificial intelligence and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user's information terminal, an emotion engine, and communication means between these devices.
[1203] Users access this system from a web browser on their own information terminal. Once accessed, a chat-based interface is displayed, through which the user can input the characteristics and settings of their character. Specifically, the user inputs a prompt in text format, such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user's terminal to the generation processing device as the character settings.
[1204] The generation processing device receives and analyzes the character settings. Based on the analysis results, it sends a character generation request to the generative AI model. This generative AI model generates a character image using advanced generation techniques such as "GPT-3" or "DALL-E." The generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[1205] The user checks the preview image and, if dissatisfied, inputs a request for correction. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion data is analyzed by the emotion engine and sent to the server.
[1206] The server receives the user's modification request and emotion data from the emotion engine. Based on this, it issues instructions for regenerating the character, and the AI generator fine-tunes the character image. After the modification is complete, a preview image is sent to the user's device again. For example, if the user expresses dissatisfaction, the AI generator will take the user's emotion into account and change the eye color to match their desired color.
[1207] Users can also add different facial expressions and movements to the generated character. In this case, the user inputs a request such as "I want to add a smiling and surprised expression." Based on this request, the server issues instructions to the generation AI to generate new facial expressions and movements. The generated expressions and movements are saved on the server and made available to users.
[1208] The server also uses data from the emotion engine to tailor the chat interface's responses: if a user is stressed, for example, the system will adjust its responses to be more gentle.
[1209] The user checks the final character, its facial expressions, and movements, and gives a final confirmation by clicking "OK." This final confirmation data is sent from the user's device to the server. The server saves the character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[1210] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1211] Step 1:
[1212] User:
[1213] Users access the system through a web browser on an information terminal, where a chat-based interface is displayed and users input their character's characteristics and settings.
[1214] Input: Prompt: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[1215] Output: Character setting information
[1216] Specific operation: When the user inputs text into the interface and presses the send button, the character setting information is sent to the generation processing device.
[1217] Step 2:
[1218] server:
[1219] The server receives the character setting information sent from the user's terminal.
[1220] Input: Character setting information
[1221] Output: Input data to the generative AI model
[1222] Specific operation: The analysis module in the server analyzes the character setting information and converts it into a format that sends a character generation request to the generation AI model.
[1223] Step 3:
[1224] Generative AI models:
[1225] The generative AI model generates character images based on character setting information sent from the server.
[1226] Input: Character setting information
[1227] Output: Character image
[1228] Specific operation: A generative AI model (e.g., GPT-3 or DALL-E) runs an image generation algorithm based on the character settings to generate a character image.
[1229] Step 4:
[1230] server:
[1231] The server receives the generated character image, temporarily stores the image for preview, and sends it to the user's terminal.
[1232] Input: Character image
[1233] Output: Preview image
[1234] Specific operation: The generated character image is saved on the server, and its URL and image data are sent to the user's device.
[1235] Step 5:
[1236] User:
[1237] Users can view preview images and enter correction requests if necessary, and emotion data is collected using the device's camera and microphone.
[1238] Input: Preview image, correction request, emotion data
[1239] Output: Modification request and emotion data
[1240] How it works: When the user enters and submits the corrections in text, the data recorded by the camera and microphone is analyzed by the emotion engine.
[1241] Step 6:
[1242] server:
[1243] The server receives correction requests and emotion data from the user and issues correction instructions to the generative AI model.
[1244] Input: Modification request, emotion data
[1245] Output: Correction instructions
[1246] How it works: The server's algorithm analyzes the correction request and emotion data, converts it into the necessary correction instructions, and sends them to the generative AI model.
[1247] Step 7:
[1248] Generative AI models:
[1249] The generative AI model fine-tunes the character image based on correction instructions and generates a new image.
[1250] Input: Correction instructions
[1251] Output: Modified character image
[1252] Specific operation: The generative AI model runs the image generation algorithm again to generate a new character image that reflects the modifications.
[1253] Step 8:
[1254] server:
[1255] The server receives the modified character image and sends it back to the user's device.
[1256] Input: Modified character image
[1257] Output: Updated preview image
[1258] Specific operation: The modified character image is saved on the server and an updated preview image is sent to the user's device.
[1259] Step 9:
[1260] User:
[1261] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[1262] Input: Final confirmation data
[1263] Output: Sending final confirmation data
[1264] Specific operation: After the user confirms, he / she presses the OK button to send the final confirmation data.
[1265] Step 10:
[1266] server:
[1267] The server saves the final character image, generates a downloadable link for the user, and sends it to the user's device.
[1268] Input: Final confirmation data
[1269] Output: Download link
[1270] Specific operation: After final confirmation, the character image is saved on the server, and a download link is generated and sent to the user's device.
[1271] (Application example 2)
[1272] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1273] Previous virtual YouTuber (VTuber) generation systems were difficult to use if the user did not have specialized knowledge, and they lacked personalization based on the user's emotions. Furthermore, these systems have not yet been applied to products suggestions in virtual stores, making it impossible to provide an effective shopping experience for each individual user.
[1274] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from a user's terminal, means for generating a character image using a generation AI based on the character settings, means for sending the generated character image to the user's terminal, means for generating a character image again based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for analyzing the user's emotional data and modifying the character based on the emotion, and means for making product suggestions using the character. This allows users to easily generate personalized VTuber characters without specialized knowledge and enjoy emotion-based modifications and product suggestions.
[1275] A "generation processing device" is a device that generates a character image using a generation AI based on character settings, and generates the character image again in response to a modification request.
[1276] "User's terminal" refers to a device that a user uses to input character settings, check generated character images, and request modifications, and includes personal computers, smartphones, etc.
[1277] "Character settings" refers to information about the characteristics and attributes of a character entered by the user, and includes specific settings such as hair color, personality, and eye shape.
[1278] "Generation AI" is artificial intelligence that generates character images based on input setting information.
[1279] "Character Image" is a digital representation of a visual character created by generative AI.
[1280] "Emotional data" is information about emotions analyzed from the user's facial expressions, voice, etc.
[1281] "Product suggestion" refers to recommending suitable products to users within a virtual store based on their emotional data.
[1282] A "modification request" is a request for changes or improvements that a user makes to a generated character image.
[1283] A "chat-based interface" is an interface that allows a user and a system to interact through text or voice.
[1284] "Different facial expressions and movements" are variations of different facial expressions and movements that are added to the character image.
[1285] "Means for enabling users to download" refers to a method for allowing users to download the final generated character image to their terminal.
[1286] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber) or virtual store assistant. The main hardware required to implement this invention includes a user terminal, a generation processing device, an emotion engine, and communication means between these devices.
[1287] Enter character settings from the user's device
[1288] Users access the "V-Shop Assistant" from a web browser on their device. Once accessed, a chat-based interface is displayed. Through this interface, users can input their character's characteristics and settings. For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is sent from the user's device to the generation processing device as the character settings.
[1289] Examples:
[1290] - Prompt: "I want to create a character with blue hair, a bright personality, and round eyes."
[1291] Character Generation
[1292] The generation processing device analyzes the character settings received from the user. Based on the analysis results, it uses the generation AI model to send a request to generate a character image. For example, settings such as "blue hair color, cheerful personality, round eyes" are analyzed, and the generation AI generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[1293] Emotion engine that recognizes user emotions
[1294] The user checks the preview image and, if dissatisfied, inputs a request for correction. Furthermore, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice, and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion engine will send this emotional data to the server.
[1295] Examples:
[1296] - Prompt: "I want my eyes to be green."
[1297] Character modification based on emotional data
[1298] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Depending on the emotion data, the generation AI fine-tunes the character image. For example, if the user has a dissatisfied expression, the generation AI will take the user's emotion into account and change the eye color to more closely match the user's desired color.
[1299] Adding facial expressions and actions
[1300] If the user wants to add additional facial expressions or actions, they can input a request. For example, they can request to add a "smiling and surprised expression." The server receives this request and instructs the AI to generate new facial expressions and actions. The generated expressions and actions are stored on the server and sent to the user.
[1301] Final confirmation and download
[1302] The user checks the final character, its facial expressions, and movements, and then clicks "OK" to confirm. The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[1303] Product proposals in virtual stores
[1304] The generated personalized character will then make product suggestions based on the user's emotional data: for example, if the user is feeling stressed, the character will suggest relaxation-related products.
[1305] Through this series of processes, users can easily generate high-quality characters without any specialized knowledge, and utilize the emotion engine to enjoy a more personalized experience.
[1306] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1307] Step 1:
[1308] The user accesses the "V-Shop Assistant" using a terminal and inputs character settings. For example, they might input setting information such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is then sent from the user terminal to the generation processing device.
[1309] Input: Character settings (hair color, personality, eye shape)
[1310] Output: The configuration information is sent to the generation processing device.
[1311] Step 2:
[1312] The generation processing device analyzes the character settings received from the user and sends a character generation request to the generation AI. The generation AI model generates a character image based on the settings, and the character image is temporarily stored on the server.
[1313] Input: Character setting information
[1314] Output: The generated character image is saved on the server and a preview image is sent to the user's device.
[1315] Step 3:
[1316] The user checks the preview image and inputs a modification request. For example, they may request, "I want the eye color to be changed to green." The user's device then sends this modification request to the server. The device's camera and microphone are also used to collect the user's emotional data (facial expressions and voice).
[1317] Input: Correction request, emotional data (facial expression, voice)
[1318] Output: The modification request and emotion data are sent to the server.
[1319] Step 4:
[1320] The server analyzes the modification request and the emotion data, and instructs the generation AI model to regenerate the character. The generation AI modifies the character image taking into account the emotion data, saves the modified character image back to the server, and sends a preview image to the user.
[1321] Input: Correction request, emotion data
[1322] Output: The modified character image is saved to the server and a preview image is sent to the user.
[1323] Step 5:
[1324] The user makes a final check of the character image and confirms it by clicking "OK." The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download.
[1325] Input: Final confirmation data
[1326] Output: A downloadable link is generated and sent to the user
[1327] Step 6:
[1328] The user can then use the generated character to receive product suggestions in a virtual store. The character will then make personalized product suggestions based on emotional data. For example, if the user is feeling stressed, the character will suggest relaxation-related products.
[1329] Input: User emotion data
[1330] Output: Product proposals by characters
[1331] At each step, the server, device, and user interact with each other, using a generative AI model and emotion engine to generate personalized VTuber characters and recommend products. Based on input information, the generative AI model generates a character image, and the emotion engine analyzes the user's emotions, providing a more satisfying experience for the user.
[1332] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1333] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1334] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1335] [Fourth embodiment]
[1336] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1337] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1338] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1339] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1340] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1341] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1342] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1343] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1344] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1345] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1346] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1347] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1348] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1349] This invention provides a system that utilizes generative AI to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, and a means of communication between the two.
[1350] 1. Enter character settings from the user's device
[1351] User:
[1352] Users access V-AI-Producers from a web browser on their device, and the system provides them with a chat-based interface.
[1353] For example, a user might enter text into a chat window saying, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input is sent from the user's device to the generation processing device as character settings.
[1354] 2. Character Generation
[1355] server:
[1356] The generation processing device receives the character settings sent by the user, and then generates a character image using a generation AI based on the settings.
[1357] For example, the generation processing device analyzes the settings such as "blue hair, cheerful personality, round eyes" and has the generation AI output a character image that matches those settings. This generated character image is temporarily saved and sent to the user's device.
[1358] 3. Character Modifications
[1359] user:
[1360] Users can check the preview image sent and enter correction requests if they are dissatisfied. For example, they can request that the eye color be changed to green.
[1361] Device:
[1362] The user's terminal transmits this modification request to the generation processing device.
[1363] server:
[1364] The generation processing unit receives the modification request and again instructs the generation AI to generate a character image with the eye color changed to green. The modified character image is again sent to the user for confirmation. This process is repeated until the user is satisfied.
[1365] 4. Adding facial expressions and actions
[1366] user:
[1367] Users can also add facial expressions and actions, for example, by requesting "add a smiling and surprised expression."
[1368] Device:
[1369] The user's terminal sends this request to the generation processing device.
[1370] server:
[1371] The generation processing device instructs the generation AI to generate facial expression differences and actions, and generates images or animations of new corresponding facial expressions and actions. These generated facial expression differences and actions are sent to the user for confirmation.
[1372] 5. Final check and download of character
[1373] user:
[1374] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[1375] Device:
[1376] The user's terminal transmits this final confirmation data to the generation processing device.
[1377] server:
[1378] The generation processor stores the final version of the character in a database and generates a download link for the user, which is sent to the user's device, allowing the user to save the generated character on their device.
[1379] This allows anyone, even those without specialized knowledge, to easily create a high-quality VTuber character and start activities with that character. This system is user-friendly and offered at an affordable price, so it is expected to be adopted by many users.
[1380] The processing flow will be explained below.
[1381] Step 1:
[1382] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[1383] Step 2:
[1384] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[1385] Step 3:
[1386] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[1387] Step 4:
[1388] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[1389] Step 5:
[1390] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[1391] Step 6:
[1392] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[1393] Step 7:
[1394] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[1395] Step 8:
[1396] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[1397] Step 9:
[1398] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[1399] Step 10:
[1400] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[1401] Step 11:
[1402] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[1403] Step 12:
[1404] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[1405] Step 13:
[1406] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[1407] Step 14:
[1408] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[1409] Example 1
[1410] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1411] Currently, there are several tools on the market for generating high-quality virtual characters (VTubers), but most of them require specialized knowledge and are difficult for average users to use. Furthermore, the character generation and modification process is often complex and time-consuming. Furthermore, the process of adding facial expressions and movements to the generated character image is also time-consuming, so a simple system that allows users to efficiently create characters that they are satisfied with is needed.
[1412] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1413] In this invention, the server includes means for receiving character settings from a user terminal, means for generating a character image using a generative AI model, means for sending the generated character image to the user terminal, means for generating a new character image based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for interacting with the user using a chat-based interface, means for generating and adding facial expressions and movements using the generative AI model, means for sending prompt text to the generative AI model based on a modification request from the user terminal, means for analyzing the prompt text and creating optimized input data for the generative AI model, means for temporarily saving the generated character image, and means for repeatedly generating and modifying the character image according to the user's satisfaction. This makes it possible to easily generate, modify, and add facial expressions and movements to personalized, high-quality character images without specialized knowledge.
[1414] "User terminal" means a device such as a computer, tablet, or smartphone that a user accesses and operates.
[1415] "Character settings" are information about the character's attributes and characteristics that the user specifies for the generated AI.
[1416] A "generative AI model" is an artificial intelligence algorithm or system that generates images or videos based on a given prompt.
[1417] A "prompt sentence" is text-based input data used to give instructions to a generative AI model.
[1418] A "generation processing device" is a device or server that receives character settings and generates images and videos using a generative AI model.
[1419] "Character Image" is an image of a virtual character created by a generative AI model based on user settings.
[1420] A "modification request" is an instruction sent by a user to request changes or modifications to a generated character image.
[1421] "Expressions and actions" refers to the changes in a character's emotions and actions that are added to the character image.
[1422] A "chat-based interface" is a means of communication that allows users and systems to interact through text messages.
[1423] "Downloadable" means that the final generated character images and data are provided in a format that can be saved on the user's device.
[1424] This invention relates to a system that allows users to easily generate and modify personalized virtual characters using generative AI models. The main components of this system are a generation processing device (server), a user terminal, and communication means between them.
[1425] First, the user accesses the system from a web browser on their own device. The user device is a computer device such as a PC, tablet, or smartphone. The system provides a chat-based interface through which the user can input text to configure their character. For example, the user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." The user's input data is sent to the server as the character configuration.
[1426] The server receives the character settings sent by the user and generates a character image using a generative AI model based on those settings. Examples of generative AI models used include DALL-E and GAN (generative adversarial network)-based AI. The generation processing device converts the received data into a format that the AI model can understand, and sends a generation prompt to the AI model to instruct it to generate an image. The generated character image is temporarily stored on the server and then sent to the user's device.
[1427] The user checks the preview image sent and inputs correction requests as needed. For example, they can input "Please change the eye color to green." The device then sends this correction request to the server. The server receives the correction request and again sends instructions to the generation AI model, regenerating the character image according to the corrections. The corrected character image is then saved again on the server and sent to the user's device.
[1428] Furthermore, users can add facial expressions and movements to their characters. For example, they can request the addition of a smiling and surprised expression. The device sends this request to the server. The server then instructs the AI model to add the new facial expressions and movements, generating images or animations of the new expressions and movements. The generated facial expressions and movements are then sent back to the user's device for confirmation.
[1429] In the final confirmation stage, the user checks the generated character, its expressions, and movements, and confirms by clicking "OK." The device then sends the final confirmation data to the server. The server then saves the final version of the character in a database and generates a link that the user can use to download it. This link is then sent to the user's device, allowing the user to save the generated character on their device.
[1430] This makes it possible to easily generate, modify, and add facial expressions and movements to high-quality virtual characters, even without specialized knowledge.
[1431] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1432] Step 1:
[1433] User inputs character settings
[1434] Users access the system through a web browser on their device, where a chat-based interface is displayed, allowing them to enter their character settings.
[1435] Input: Enter the following text: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[1436] Processing: The terminal receives the user's input, converts the data into JSON format, and sends it to the server. Specifically, the input data is converted into object format and sent over the network.
[1437] Output: Character setting data is sent to the server.
[1438] Step 2:
[1439] The server receives the character settings.
[1440] The server receives the character setting data sent from the user's terminal.
[1441] Input: Character configuration data sent from the user's device.
[1442] Processing: Analyzes the incoming data and converts it into a format that the generative AI model can understand. Specifically, it organizes the text data into prompt sentences and forms generative prompts.
[1443] Output: A generated prompt is sent to the generative AI model.
[1444] Step 3:
[1445] The server generates the character image
[1446] The server uses a generative AI model to generate a character image based on the received character settings.
[1447] Input: A generative prompt (e.g., "Her hair color is blue, her personality is cheerful, and her eyes are round").
[1448] Processing: The generated prompt is sent to the AI model, which then generates a corresponding character image. Specifically, the AI model analyzes the prompt and generates the corresponding image data. The generated image is temporarily stored in the server's storage.
[1449] Output: The generated character image is saved and sent to the user's device.
[1450] Step 4:
[1451] User checks the character image and enters correction request
[1452] The user checks the preview image sent and enters correction requests if necessary.
[1453] Input: Type your request: "Please change my eye color to green."
[1454] Processing: The device sends a correction request to the server. Specifically, the correction content is converted into JSON format and sent again to the server.
[1455] Output: The modified request data is sent to the server.
[1456] Step 5:
[1457] The server regenerates the modified request
[1458] The server receives the correction request and sends instructions to the generation AI model again, generating a character image with the eye color changed to green.
[1459] Input: Your correction request (e.g., "Please change my eye color to green").
[1460] Processing: A new prompt (e.g., "Change the eye color to green") is sent to the generative AI model to instruct it to regenerate. Specifically, the AI model regenerates the image data based on the modification request and stores it again in the server storage.
[1461] Output: The regenerated character image is sent to the user's device.
[1462] Step 6:
[1463] User requests for additional facial expressions and actions
[1464] Users also request the addition of different facial expressions and movements.
[1465] Input: "Add a smiling and surprised expression"
[1466] Processing: The device sends this request to the server. Specifically, it converts the request content into JSON format and sends it to the server.
[1467] Output: Requests for additional facial expressions and actions are sent to the server.
[1468] Step 7:
[1469] The server generates facial expression differences and actions
[1470] The server instructs the generative AI model to add facial expression differences and movements.
[1471] Input: Request for additional facial expressions or actions (e.g., "Please add a smiling and surprised expression").
[1472] Processing: New prompts are sent to the generative AI model to generate the corresponding facial expressions and actions. Specifically, the AI model generates new images and animations based on the additional prompts and saves them in server storage.
[1473] Output: The generated facial expression differences and movement data are sent to the user's device.
[1474] Step 8:
[1475] The user makes a final confirmation and downloads
[1476] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[1477] Enter: "OK" to confirm.
[1478] Processing: The device sends the final confirmation data to the server, which saves the final version of the character to the database and generates a download link.
[1479] Output: A downloadable link is generated and sent to the user's device.
[1480] Step 9:
[1481] User downloads character
[1482] Users save the generated characters on their devices.
[1483] Enter: Click on the download link.
[1484] Action: Download the linked file.
[1485] Output: The generated character is saved on the user's device.
[1486] (Application example 1)
[1487] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1488] In recent years, users have begun to demand online and virtual shopping experiences, and there is an increasing demand for personalized customer service and product explanations. However, traditional online shopping systems have difficulty providing the detailed product explanations and two-way dialogue that users desire. Furthermore, there is a lack of a way for users to create their own personalized characters and have them explain products, providing a more user-friendly shopping experience. This poses a challenge, making it difficult to improve user satisfaction and sales efficiency.
[1489] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1490] In this invention, the server includes a means for receiving character settings from a user's device, a means for generating a character image using a generation AI based on the character settings, and a means for transmitting the generated character image to the user's device. This allows users to easily create their own characters, receive product explanations through the characters, and receive answers to questions through dialogue. As a result, users can enjoy a more personalized shopping experience and overcome the drawbacks of traditional online shopping.
[1491] "Generation Processor" means a computer system for generating and modifying characters based on user input.
[1492] "User's terminal" refers to a device operated by a user, including a smartphone, tablet, PC, etc.
[1493] "Character settings" refers to information about the character's attributes and characteristics, such as hair color and personality, entered by the user.
[1494] "Generative AI" is an artificial intelligence technology that uses machine learning algorithms to generate character images.
[1495] "Character image" refers to image data of a visual character generated by a generation AI.
[1496] A "modification request" is an instruction sent by a user to modify a character's attributes or characteristics.
[1497] "Facial expression differences" are image data showing different facial expressions of a character.
[1498] "Action" is data that indicates the animation or movement that a character performs.
[1499] A "chat-based interface" is an interface that allows a user to interact with a system in a text-based manner.
[1500] "Product description" is information that the generated character uses to communicate the product's features and how to use it to the user.
[1501] "Answering questions through dialogue" is the process in which a character answers questions posed by a user through generated AI.
[1502] This invention provides a system that allows users to easily generate personalized characters and use them to improve shopping in virtual stores. The main components include a generation processing device, a user terminal, and means for communicating between them.
[1503] First, the user accesses the system using a web browser on their device (smartphone, tablet, PC, etc.). A chat-based interface is provided, through which the user can input character settings. For example, the user can input a prompt such as, "I want to create a character with blue hair, a cheerful personality, and round eyes." This information is sent to the generation processing device as the character settings.
[1504] The generation processing device analyzes the received character settings and generates a character image using a generation AI. The generation AI used is a high-performance model such as GPT-4 or Stable Diffusion. The generated character image is temporarily saved and sent to the user's device. The user can check this as a preview and, if necessary, send a request to make corrections. For example, they can request to change the eye color to green.
[1505] The modification request is sent to the generation processing device again, and the generation AI generates a new character image. This process is repeated until the user is satisfied. The final generated character image is saved in the database and available for download by the user.
[1506] Furthermore, users can use the generated character to check product descriptions and features. When the store clerk character explains a product, for example, they might say, "This smartwatch is equipped with the latest sensors and is useful for health management." The system also has a function where users can ask questions in a dialogue format, and the generated character will respond appropriately to the question. For example, if a user asks, "How do I use this product?", the character will respond, "This smartwatch is easy to use by operating the touchscreen."
[1507] The system integrates a server (including a database such as Firebase), user devices, and generative AI, and is provided as a cross-platform mobile application using React Native, enabling users to have an engaging and personalized virtual store experience.
[1508] Example prompt sentence:
[1509] Character generation: "Create a character with blonde hair, sporty personality, and round glasses"
[1510] Product description: "Describe the features of a newly released smartwatch in a friendly and detailed manner"
[1511] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1512] Step 1:
[1513] Users access the system through a web browser on their own device. They use a chat-based interface to input their character settings, such as a specific prompt, such as "I want a character with blue hair, a cheerful personality, and round eyes." This input information is sent to the system as the character settings.
[1514] Input: Character settings entered by the user (text format)
[1515] Output: Character setting data sent to the server
[1516] Step 2:
[1517] The server receives the character setting data sent by the user and generates a character image using a generation AI (e.g., GPT-4 or Stable Diffusion).
[1518] Input: Character setting data (received from the user's device)
[1519] Output: Generated character image
[1520] Specific behavior:
[1521] Send a prompt to the generation AI: "Generate a character with blue hair, a cheerful personality, and round eyes."
[1522] The image is generated by the AI and temporarily saved.
[1523] Step 3:
[1524] The server sends the generated character image to the user's device, where the user can check the character image as a preview.
[1525] Input: Generated character image
[1526] Output: Preview of character image on user device
[1527] Step 4:
[1528] If the user is dissatisfied with the character image, they can input a request for correction, such as "I want the eye color to be changed to green," and send that request from their device to the server.
[1529] Input: Correction request (sent from the user's device)
[1530] Output: Modified request data sent to the server
[1531] Step 5:
[1532] The server receives the modification request and issues a prompt to the AI again. The AI then generates a new character image that reflects the specified modifications. This generation process is repeated until the user is satisfied.
[1533] Input: Correction request data
[1534] Output: Modified character image
[1535] Specific behavior:
[1536] Send the generator AI a new prompt: "Generate a character with blue hair, a cheerful personality, and green eyes."
[1537] Generate an image and save it again
[1538] Step 6:
[1539] The user finally confirms the character image they are satisfied with and requests a download. The server saves the final version of the character image in the database and generates a download link. The server sends this link to the user, who then downloads the character image.
[1540] Input: User confirmation and download request
[1541] Output: Generate and send a download link, final character image
[1542] Step 7:
[1543] The user checks the product description and features using the generated character, and the server displays the product description through the character and answers the user's questions in an interactive format.
[1544] Input: User question and selected product information
[1545] Output: Product description and answer by the character
[1546] Specific behavior:
[1547] Use generative AI to generate a product description by sending a prompt: "Please describe this product."
[1548] The character presents the generated explanation to the user.
[1549] Answering user questions
[1550] Through this entire process, users can create their own character and receive product information and purchasing assistance through that character.
[1551] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1552] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user terminal, an emotion engine, and communication means between these devices.
[1553] 1. Enter character settings from the user's device
[1554] User:
[1555] Users access "V-AI-Producers" through a web browser on their device, where they are presented with a chat-based interface through which they can input their character's characteristics and settings.
[1556] For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user terminal to the generation processing device as character settings.
[1557] 2. Character Generation
[1558] server:
[1559] The generation processing device analyzes the character settings received from the user and, based on the analysis results, sends a character generation request to the generation AI model.
[1560] For example, the AI analyzes the settings such as "blue hair, cheerful personality, round eyes" and generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[1561] 3. Emotion engine that recognizes user emotions
[1562] User:
[1563] Users can check the preview image and input correction requests if they are dissatisfied. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends that information to the server.
[1564] For example, if the user has a dissatisfied expression when entering a modification request such as "I want my eyes to be green," the emotion engine will send that emotional data to the server.
[1565] 4. Modifying characters based on emotional data
[1566] server:
[1567] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Based on the emotion data, the generation AI fine-tunes the character image.
[1568] For example, if the user has a dissatisfied expression, the generative AI will take the user's emotions into account and change their eye color to more closely match their desired look.
[1569] 5. Adding facial expressions and actions
[1570] User:
[1571] If the user wants to add facial expressions or actions, they can input a request. For example, they can request to add "smiling and surprised expressions."
[1572] server:
[1573] The server receives this request and issues instructions to the AI to generate new facial expressions and movements. The generated expressions and movements are stored on the server and sent to the user.
[1574] 6. Chat-based interface adjustments
[1575] server:
[1576] The server uses data from the emotion engine to tailor responses in the chat-based interface: if the user is stressed, for example, the system will take this into account and use kinder words.
[1577] 7. Final confirmation and download
[1578] user:
[1579] The user checks the final character, its facial expressions, and movements, and then gives a final confirmation by clicking "OK." The final confirmation data is then sent from the device to the server.
[1580] server:
[1581] The server saves the final character image and generates a downloadable link for the user, which is sent to the user's device and the user can click on the link to download the character image.
[1582] summary
[1583] This invention allows users to easily create high-quality VTuber characters without specialized knowledge, and utilizes an emotion engine to enjoy a more personalized character experience. The combination of the emotion engine increases user satisfaction and enables more natural and consistent character generation.
[1584] The processing flow will be explained below.
[1585] Step 1:
[1586] The user opens their device (PC or smartphone) and accesses "V-AI-Producers" from a web browser. The device sends this request to the server, requesting that the interface be displayed.
[1587] Step 2:
[1588] The server receives an access request from a user, generates a chat-based interface, and sends it to the device.
[1589] Step 3:
[1590] The device then displays a chat-based interface to the user, who then enters their character's characteristics and preferences into the chat window.
[1591] Step 4:
[1592] The user inputs in text format, for example, "I want to create a character with blue hair, a cheerful personality, and round eyes." The device then sends this setting data to the server.
[1593] Step 5:
[1594] The server analyzes the character settings received from the user and sends a character generation request to the generative AI model based on the analysis results.
[1595] Step 6:
[1596] The generation AI generates a character image based on the specified character settings and returns the generation results to the server.
[1597] Step 7:
[1598] The server temporarily stores the character image received from the generation AI and sends a preview image of it to the user's device.
[1599] Step 8:
[1600] The device displays the received preview image to the user and asks for confirmation, after which the user can decide whether or not the preview image is satisfactory.
[1601] Step 9:
[1602] If the user is dissatisfied with the preview image, the user inputs a correction request, for example, "I want the eye color to be changed to green," and the terminal transmits the request to the server.
[1603] Step 10:
[1604] The server receives the correction request and again issues instructions to the generation AI to generate the corrected character image. The corrected character image is then returned to the server.
[1605] Step 11:
[1606] The server then sends the revised character image to the user's device and asks for confirmation again, and this process is repeated until the user is satisfied.
[1607] Step 12:
[1608] When the user is satisfied with the final character image displayed, he or she confirms by clicking "OK." This final confirmation data is sent from the terminal to the server.
[1609] Step 13:
[1610] The server stores the final character image in a database and generates a downloadable link for the user, which is then sent to the device.
[1611] Step 14:
[1612] The device will display the received download link to the user, who can click the link to save the generated character image to their device.
[1613] Step 15:
[1614] While the user is inputting their character settings, their emotions are recognized through the device's built-in camera and microphone. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server.
[1615] Step 16:
[1616] The server receives emotional data from the emotion engine and automatically adjusts the character design based on this data. For example, if the user has a dissatisfied expression, the server will instruct the generation AI to correct it and generate a character image that more closely matches the user's desire.
[1617] Step 17:
[1618] If the user wants to add facial expression differences or actions to the final character, the user inputs a request for this. For example, the user may request "I want a smiling and surprised expression added."
[1619] Step 18:
[1620] The server receives the user's request and has the AI generate new facial expressions and movements. The generated facial expressions and movements are stored on the server and sent to the user's device.
[1621] Step 19:
[1622] The server uses data from the emotion engine to tailor responses in the chat-based interface: if a user is feeling stressed, for example, the system will take this into account and use kinder words.
[1623] Example 2
[1624] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1625] While conventional VTuber character generation systems can create character settings based on user requests, they have problems in that they are unable to fully increase user satisfaction because they do not adequately accommodate requests for modifications or the reflection of emotions in the generated characters.Furthermore, there are few systems that can take user emotions into account when fine-tuning the generated characters or adding variations in facial expressions and movements, making it difficult to create characters that meet the user's intentions.
[1626] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from the user's information terminal, means for generating a character image using a generation artificial intelligence based on the character settings, means for transmitting the generated character image to the user's information terminal, means for regenerating the character image based on a modification request from the user, means for recognizing the user's emotions and fine-tuning the character image based on the recognition result, and means for saving the final character image and making it available for download by the user. This enables character generation and modification taking the user's emotions into consideration, thereby increasing user satisfaction. Furthermore, it is possible to add variations in facial expressions and movements to the generated character, providing a more personalized character experience.
[1627] A "generation processing device" is a device that receives character settings from a user's information terminal and generates, modifies, and saves a character image using generation artificial intelligence.
[1628] "User's information terminal" refers to the device through which the user accesses the system, inputs character settings, and checks and edits the generated character image. Specifically, this applies to a computer or smartphone.
[1629] "Character settings" refers to input information that specifically specifies the characteristics and personality of the virtual YouTuber created by the user, such as hair color, eye shape, and personality.
[1630] "Generative AI" is an AI technology for automatically generating character images based on character settings received from users.
[1631] "Character Image" refers to a visual image of a virtual YouTuber generated by generative artificial intelligence.
[1632] A "modification request" is an instruction from a user to change or modify a generated character image.
[1633] "Emotion recognition" is the process of detecting emotions from the user's facial expressions, voice, etc., and analyzing that information.
[1634] "Fine-tuning" refers to making small changes or improvements to existing character images based on user sentiment and correction requests.
[1635] "Variations in facial expressions and movements" enrich the character's expression by adding different facial expressions and movements to the generated character image.
[1636] "Downloadable Link" means the URL from which a user can download the final character image via the Internet.
[1637] This invention provides a system that utilizes generative artificial intelligence and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber). The system mainly consists of a generation processing device, a user's information terminal, an emotion engine, and communication means between these devices.
[1638] Users access this system from a web browser on their own information terminal. Once accessed, a chat-based interface is displayed, through which the user can input the characteristics and settings of their character. Specifically, the user inputs a prompt in text format, such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This input information is sent from the user's terminal to the generation processing device as the character settings.
[1639] The generation processing device receives and analyzes the character settings. Based on the analysis results, it sends a character generation request to the generative AI model. This generative AI model generates a character image using advanced generation techniques such as "GPT-3" or "DALL-E." The generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[1640] The user checks the preview image and, if dissatisfied, inputs a request for correction. In addition, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion data is analyzed by the emotion engine and sent to the server.
[1641] The server receives the user's modification request and emotion data from the emotion engine. Based on this, it issues instructions for regenerating the character, and the AI generator fine-tunes the character image. After the modification is complete, a preview image is sent to the user's device again. For example, if the user expresses dissatisfaction, the AI generator will take the user's emotion into account and change the eye color to match their desired color.
[1642] Users can also add different facial expressions and movements to the generated character. In this case, the user inputs a request such as "I want to add a smiling and surprised expression." Based on this request, the server issues instructions to the generation AI to generate new facial expressions and movements. The generated expressions and movements are saved on the server and made available to users.
[1643] The server also uses data from the emotion engine to tailor the chat interface's responses: if a user is stressed, for example, the system will adjust its responses to be more gentle.
[1644] The user checks the final character, its facial expressions, and movements, and gives a final confirmation by clicking "OK." This final confirmation data is sent from the user's device to the server. The server saves the character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[1645] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1646] Step 1:
[1647] User:
[1648] Users access the system through a web browser on an information terminal, where a chat-based interface is displayed and users input their character's characteristics and settings.
[1649] Input: Prompt: "I want to create a character with blue hair, a cheerful personality, and round eyes."
[1650] Output: Character setting information
[1651] Specific operation: When the user inputs text into the interface and presses the send button, the character setting information is sent to the generation processing device.
[1652] Step 2:
[1653] server:
[1654] The server receives the character setting information sent from the user's terminal.
[1655] Input: Character setting information
[1656] Output: Input data to the generative AI model
[1657] Specific operation: The analysis module in the server analyzes the character setting information and converts it into a format that sends a character generation request to the generation AI model.
[1658] Step 3:
[1659] Generative AI models:
[1660] The generative AI model generates character images based on character setting information sent from the server.
[1661] Input: Character setting information
[1662] Output: Character image
[1663] Specific operation: A generative AI model (e.g., GPT-3 or DALL-E) runs an image generation algorithm based on the character settings to generate a character image.
[1664] Step 4:
[1665] server:
[1666] The server receives the generated character image, temporarily stores the image for preview, and sends it to the user's terminal.
[1667] Input: Character image
[1668] Output: Preview image
[1669] Specific operation: The generated character image is saved on the server, and its URL and image data are sent to the user's device.
[1670] Step 5:
[1671] User:
[1672] Users can view preview images and enter correction requests if necessary, and emotion data is collected using the device's camera and microphone.
[1673] Input: Preview image, correction request, emotion data
[1674] Output: Modification request and emotion data
[1675] How it works: When the user enters and submits the corrections in text, the data recorded by the camera and microphone is analyzed by the emotion engine.
[1676] Step 6:
[1677] server:
[1678] The server receives correction requests and emotion data from the user and issues correction instructions to the generative AI model.
[1679] Input: Modification request, emotion data
[1680] Output: Correction instructions
[1681] How it works: The server's algorithm analyzes the correction request and emotion data, converts it into the necessary correction instructions, and sends them to the generative AI model.
[1682] Step 7:
[1683] Generative AI models:
[1684] The generative AI model fine-tunes the character image based on correction instructions and generates a new image.
[1685] Input: Correction instructions
[1686] Output: Modified character image
[1687] Specific operation: The generative AI model runs the image generation algorithm again to generate a new character image that reflects the modifications.
[1688] Step 8:
[1689] server:
[1690] The server receives the modified character image and sends it back to the user's device.
[1691] Input: Modified character image
[1692] Output: Updated preview image
[1693] Specific operation: The modified character image is saved on the server and an updated preview image is sent to the user's device.
[1694] Step 9:
[1695] User:
[1696] The user checks the final character, its expressions and movements, and gives a final confirmation by clicking "OK."
[1697] Input: Final confirmation data
[1698] Output: Sending final confirmation data
[1699] Specific operation: After the user confirms, he / she presses the OK button to send the final confirmation data.
[1700] Step 10:
[1701] server:
[1702] The server saves the final character image, generates a downloadable link for the user, and sends it to the user's device.
[1703] Input: Final confirmation data
[1704] Output: Download link
[1705] Specific operation: After final confirmation, the character image is saved on the server, and a download link is generated and sent to the user's device.
[1706] (Application example 2)
[1707] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1708] Previous virtual YouTuber (VTuber) generation systems were difficult to use if the user did not have specialized knowledge, and they lacked personalization based on the user's emotions. Furthermore, these systems have not yet been applied to products suggestions in virtual stores, making it impossible to provide an effective shopping experience for each individual user.
[1709] The specification processing by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for receiving character settings from a user's terminal, means for generating a character image using a generation AI based on the character settings, means for sending the generated character image to the user's terminal, means for generating a character image again based on a modification request from the user, means for saving the final character image and making it available for download by the user, means for analyzing the user's emotional data and modifying the character based on the emotion, and means for making product suggestions using the character. This allows users to easily generate personalized VTuber characters without specialized knowledge and enjoy emotion-based modifications and product suggestions.
[1710] A "generation processing device" is a device that generates a character image using a generation AI based on character settings, and generates the character image again in response to a modification request.
[1711] "User's terminal" refers to a device that a user uses to input character settings, check generated character images, and request modifications, and includes personal computers, smartphones, etc.
[1712] "Character settings" refers to information about the characteristics and attributes of a character entered by the user, and includes specific settings such as hair color, personality, and eye shape.
[1713] "Generation AI" is artificial intelligence that generates character images based on input setting information.
[1714] "Character Image" is a digital representation of a visual character created by generative AI.
[1715] "Emotional data" is information about emotions analyzed from the user's facial expressions, voice, etc.
[1716] "Product suggestion" refers to recommending suitable products to users within a virtual store based on their emotional data.
[1717] A "modification request" is a request for changes or improvements that a user makes to a generated character image.
[1718] A "chat-based interface" is an interface that allows a user and a system to interact through text or voice.
[1719] "Different facial expressions and movements" are variations of different facial expressions and movements that are added to the character image.
[1720] "Means for enabling users to download" refers to a method for allowing users to download the final generated character image to their terminal.
[1721] This invention provides a system that utilizes generative AI and an emotion engine to allow anyone to easily generate and operate a personalized virtual YouTuber (VTuber) or virtual store assistant. The main hardware required to implement this invention includes a user terminal, a generation processing device, an emotion engine, and communication means between these devices.
[1722] Enter character settings from the user's device
[1723] Users access the "V-Shop Assistant" from a web browser on their device. Once accessed, a chat-based interface is displayed. Through this interface, users can input their character's characteristics and settings. For example, a user might input, "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is sent from the user's device to the generation processing device as the character settings.
[1724] Examples:
[1725] - Prompt: "I want to create a character with blue hair, a bright personality, and round eyes."
[1726] Character Generation
[1727] The generation processing device analyzes the character settings received from the user. Based on the analysis results, it uses the generation AI model to send a request to generate a character image. For example, settings such as "blue hair color, cheerful personality, round eyes" are analyzed, and the generation AI generates a character image based on those settings. This generated character image is temporarily stored on the server and sent to the user's device as a preview image.
[1728] Emotion engine that recognizes user emotions
[1729] The user checks the preview image and, if dissatisfied, inputs a request for correction. Furthermore, the user's device is equipped with a camera and microphone, which are used to recognize the user's emotions. The emotion engine analyzes the user's emotions from their facial expressions and voice, and sends this information to the server. For example, if the user has a dissatisfied expression when inputting a correction request such as "I want my eyes to be green," the emotion engine will send this emotional data to the server.
[1730] Examples:
[1731] - Prompt: "I want my eyes to be green."
[1732] Character modification based on emotional data
[1733] The server receives correction requests from the user and emotion data from the emotion engine. Based on the received data, it instructs the generation AI to regenerate the character. Depending on the emotion data, the generation AI fine-tunes the character image. For example, if the user has a dissatisfied expression, the generation AI will take the user's emotion into account and change the eye color to more closely match the user's desired color.
[1734] Adding facial expressions and actions
[1735] If the user wants to add additional facial expressions or actions, they can input a request. For example, they can request to add a "smiling and surprised expression." The server receives this request and instructs the AI to generate new facial expressions and actions. The generated expressions and actions are stored on the server and sent to the user.
[1736] Final confirmation and download
[1737] The user checks the final character, its facial expressions, and movements, and then clicks "OK" to confirm. The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download. This link is sent to the user's device, and the user can click the link to download the character image.
[1738] Product proposals in virtual stores
[1739] The generated personalized character will then make product suggestions based on the user's emotional data: for example, if the user is feeling stressed, the character will suggest relaxation-related products.
[1740] Through this series of processes, users can easily generate high-quality characters without any specialized knowledge, and utilize the emotion engine to enjoy a more personalized experience.
[1741] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1742] Step 1:
[1743] The user accesses the "V-Shop Assistant" using a terminal and inputs character settings. For example, they might input setting information such as "I want to create a character with blue hair, a cheerful personality, and round eyes." This setting information is then sent from the user terminal to the generation processing device.
[1744] Input: Character settings (hair color, personality, eye shape)
[1745] Output: The configuration information is sent to the generation processing device.
[1746] Step 2:
[1747] The generation processing device analyzes the character settings received from the user and sends a character generation request to the generation AI. The generation AI model generates a character image based on the settings, and the character image is temporarily stored on the server.
[1748] Input: Character setting information
[1749] Output: The generated character image is saved on the server and a preview image is sent to the user's device.
[1750] Step 3:
[1751] The user checks the preview image and inputs a modification request. For example, they may request, "I want the eye color to be changed to green." The user's device then sends this modification request to the server. The device's camera and microphone are also used to collect the user's emotional data (facial expressions and voice).
[1752] Input: Correction request, emotional data (facial expression, voice)
[1753] Output: The modification request and emotion data are sent to the server.
[1754] Step 4:
[1755] The server analyzes the modification request and the emotion data, and instructs the generation AI model to regenerate the character. The generation AI modifies the character image taking into account the emotion data, saves the modified character image back to the server, and sends a preview image to the user.
[1756] Input: Correction request, emotion data
[1757] Output: The modified character image is saved to the server and a preview image is sent to the user.
[1758] Step 5:
[1759] The user makes a final check of the character image and confirms it by clicking "OK." The final confirmation data is sent from the device to the server. The server saves the final character image and generates a link that the user can download.
[1760] Input: Final confirmation data
[1761] Output: A downloadable link is generated and sent to the user
[1762] Step 6:
[1763] The user can then use the generated character to receive product suggestions in a virtual store. The character will then make personalized product suggestions based on emotional data. For example, if the user is feeling stressed, the character will suggest relaxation-related products.
[1764] Input: User emotion data
[1765] Output: Product proposals by characters
[1766] At each step, the server, device, and user interact with each other, using a generative AI model and emotion engine to generate personalized VTuber characters and recommend products. Based on input information, the generative AI model generates a character image, and the emotion engine analyzes the user's emotions, providing a more satisfying experience for the user.
[1767] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1768] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1769] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1770] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1771] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1772] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1773] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1774] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1775] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1776] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1777] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1778] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1779] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1780] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1781] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1782] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1783] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1784] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1785] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1786] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1787] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1788] The following is further disclosed regarding the above embodiment.
[1789] (Claim 1)
[1790] The production processing device includes:
[1791] a means for receiving character settings from a user's device;
[1792] A means for generating a character image using a generation AI based on the character settings;
[1793] means for transmitting the generated character image to a user's terminal;
[1794] A method for regenerating character images based on correction requests from users;
[1795] The system includes a means to save the final character image and make it available for user download.
[1796] (Claim 2)
[1797] The production processing device includes:
[1798] The system according to claim 1, further comprising means for adding facial expression differences and actions to the generated character image.
[1799] (Claim 3)
[1800] The production processing device includes:
[1801] 10. The system of claim 1, further comprising means for interacting with the user using a chat-based interface.
[1802] "Example 1"
[1803] (Claim 1)
[1804] a means for receiving character settings from a user device;
[1805] A means for generating a character image using a generative AI model based on the character settings;
[1806] means for transmitting the generated character image to a user terminal;
[1807] A method for regenerating character images based on correction requests from users;
[1808] A means to save the final character image and make it available for download by users;
[1809] a means for interacting with a user using a chat-based interface;
[1810] A means for generating and adding facial expressions and movements using a generative AI model;
[1811] A means for sending a prompt sentence to the generative AI model based on a correction request from the user device;
[1812] A means for analyzing prompt sentences and creating optimized input data for the generative AI model;
[1813] A means for temporarily saving the generated character image;
[1814] A means to iteratively generate and modify according to user satisfaction
[1815] A system including:
[1816] (Claim 2)
[1817] The system according to claim 1, further comprising means for adding facial expression differences and actions to the generated character image.
[1818] (Claim 3)
[1819] 10. The system of claim 1, further comprising means for interacting with the user using a chat-based interface.
[1820] "Application Example 1"
[1821] (Claim 1)
[1822] The production processing device includes:
[1823] a means for receiving character settings from a user's device;
[1824] A means for generating a character image using a generation AI based on the character settings;
[1825] means for transmitting the generated character image to a user's terminal;
[1826] A method for regenerating character images based on correction requests from users;
[1827] In addition to saving the final character image and allowing users to download it,
[1828] A means for displaying product descriptions and features using the generated character;
[1829] A system that includes a means for users to ask product questions through dialogue and provide answers to those questions.
[1830] (Claim 2)
[1831] A means for adding facial expression differences and actions to the generated character image; and
[1832] Depending on the product you select, the character will explain the product.
[1833] 10. The system of claim 1, further comprising means for providing responses to the user through dialogue.
[1834] (Claim 3)
[1835] means for interacting with the user using a chat-based interface; and
[1836] 10. The system of claim 1, further comprising means for regenerating the character image in response to a user's character modification request.
[1837] "Example 2: Combining Emotion Engines"
[1838] (Claim 1)
[1839] The production processing device includes:
[1840] A means for receiving character settings from a user's information terminal;
[1841] A means for generating a character image using a generation artificial intelligence based on the character setting;
[1842] means for transmitting the generated character image to a user's information terminal;
[1843] A method for regenerating character images based on correction requests from users;
[1844] means for recognizing the user's emotion and fine-tuning the character image based on the recognition result;
[1845] The system includes a means to save the final character image and make it available for user download.
[1846] (Claim 2)
[1847] 2. The system according to claim 1, further comprising means for adding variations in facial expressions and movements to the generated character image.
[1848] (Claim 3)
[1849] 10. The system of claim 1, further comprising means for interacting with the user using a chat-based interface.
[1850] "Application example 2 when combining emotion engines"
[1851] (Claim 1)
[1852] The production processing device includes:
[1853] a means for receiving character settings from a user's device;
[1854] A means for generating a character image using a generation AI based on the character settings;
[1855] means for transmitting the generated character image to a user's terminal;
[1856] A method for regenerating character images based on correction requests from users;
[1857] A means to save the final character image and make it available for download by users;
[1858] A means for analyzing the user's emotional data and modifying the character based on the emotion;
[1859] A means of using characters to propose products;
[1860] A system including:
[1861] (Claim 2)
[1862] The system according to claim 1, further comprising means for adding facial expression differences and actions to the generated character image.
[1863] (Claim 3)
[1864] 10. The system of claim 1, further comprising means for interacting with the user using a chat-based interface. [Explanation of symbols]
[1865] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. The production processing device includes: a means for receiving character settings from a user's device; A means for generating a character image using a generation AI based on the character settings; means for transmitting the generated character image to a user's terminal; A method for regenerating character images based on correction requests from users; The system includes a means to save the final character image and make it available for user download.
2. The production processing device includes: The system according to claim 1, further comprising means for adding facial expression differences and actions to the generated character image.
3. The production processing device includes:
10. The system of claim 1, further comprising means for interacting with the user using a chat-based interface.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A