System
The AI-powered platform simplifies virtual character creation by generating avatars, synthesizing speech, and integrating subtitles, allowing users to create high-quality content without technical knowledge.
Patent Information
- Application Number
- JP2024123980
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Existing AI-based virtual character creation systems require advanced programming skills, making them inaccessible to beginners and non-technical users, and involve complex processes like avatar customization, speech content generation, and subtitle integration, which are time-consuming and tedious.
A platform that utilizes AI to generate 3D or Live2D avatars, synthesize speech, synchronize motion, and automatically create subtitles, providing intuitive customization options and management tools for users to easily create high-quality virtual characters without specialized knowledge.
Enables users to quickly and efficiently generate high-quality virtual characters and related content, simplifying the process and eliminating the need for programming expertise.
Smart Images

Figure 2026022463000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The "Problem that the invention aims to solve" and "Means for solving the problem" of the document are shown below.
[0005] Currently, AI-based virtual character creation is widely accepted in the market, but it requires advanced programming skills, creating barriers to entry that prevent many users with ideas from participating. This makes it difficult for beginners and non-technical people to create content using virtual characters. Furthermore, existing systems require complex processes such as avatar customization, speech content, voice synthesis, and subtitle generation, which consumes a great deal of time and effort from users. There is a need to solve these issues and provide an intuitive and easy-to-use system that can be used by a wide range of users. [Means for solving the problem]
[0006] This invention provides a platform that utilizes AI to enable users to easily create virtual characters. This system integrates a means for generating a 3D or Live2D avatar based on information input by the user, a means for generating speech content based on an input theme, a means for synthesizing the generated speech content with speech, a means for synchronizing the audio file with the avatar's movements, and a means for automatically generating subtitles based on the speech content and integrating them into the video. Furthermore, the system provides multiple customization options that users can select when generating the avatar, and includes a management tool that allows users to request corrections to the video they have generated, creating an environment that users can operate intuitively. This allows even beginners and non-technical users to easily create content using high-quality virtual characters.
[0007] "User" refers to the person or organization that operates the system to create and manage virtual characters.
[0008] "Platform" refers to the software environment or system used by users to create virtual characters.
[0009] A "3D avatar" refers to a virtual character represented in three-dimensional space.
[0010] A "Live2D avatar" is a virtual character whose movements and expressions are animated on a flat surface.
[0011] "Input information" refers to data and configuration parameters provided to the system by a user.
[0012] "Image generation AI module" refers to an artificial intelligence program for generating images of virtual characters based on input information.
[0013] "LLM (Large-scale Language Model)" refers to an artificial intelligence model that learns from massive amounts of data and performs natural language processing.
[0014] "Speech content" refers to text information that the user wants the virtual character to express.
[0015] "Speech synthesis" refers to the process of generating corresponding speech based on text information of the speech content.
[0016] "Audio file" refers to digital audio data generated by voice synthesis.
[0017] "Lip sync" refers to the technique of synchronizing a virtual character's mouth movements with generated audio.
[0018] "Body language" refers to the physical movements and gestures of a virtual character.
[0019] "Motion video" refers to video in which a virtual character moves based on generated voices and movements.
[0020] "Subtitles" refers to character information for displaying text corresponding to what is being said within a video.
[0021] "Administration tools" refers to the interfaces and functions within the system that allow users to perform creation processes and request modifications. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0024] First, the terms used in the following description will be explained.
[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0030] [First embodiment]
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0043] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech with motion, and generates and integrates subtitles into the video.
[0044] System configuration
[0045] 1. 3D / Live2D avatar generation
[0046] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[0047] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[0048] 2. Generating speech content
[0049] The server passes the topic information entered by the user to the LLM, generates a message, and sends the generated message to the terminal for the user to confirm.
[0050] The user confirms and modifies the generated message and sends the finalized message to the server.
[0051] 3. Speech synthesis
[0052] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[0053] The user checks the audio file and requests corrections if necessary.
[0054] 4. Audio and Motion Mapping
[0055] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[0056] The user checks the motion image and requests corrections if necessary.
[0057] 5. Creating subtitles
[0058] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file with the video, and sends the completed video to the device.
[0059] The user checks the video and subtitles and confirms them after a final check.
[0060] Specific examples
[0061] Example 1: When a user creates an introductory video for an AITuber
[0062] 1. Create an avatar
[0063] The user clicks the "Create Avatar" button on the device's UI and inputs the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[0064] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[0065] 2. Generating speech content
[0066] The user enters the topic "Introducing AITubers" into a form on the device. The server uses LLM to generate comments based on the topic and sends them to the device.
[0067] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[0068] 3. Speech Synthesis
[0069] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[0070] The user checks the audio file and confirms it if there are no particular problems.
[0071] 4. Audio and Motion Mapping
[0072] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[0073] The user checks the motion image and confirms it if there are no particular problems.
[0074] 5. Creating subtitles
[0075] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[0076] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[0077] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system enables anyone to easily generate content that utilizes high-quality virtual characters, without requiring programming knowledge.
[0078] The processing flow will be explained below.
[0079] 1. Creating a 3D / Live2D avatar
[0080] Step 1:
[0081] The user clicks the "Create Avatar" button on the device.
[0082] Step 2:
[0083] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[0084] Step 3:
[0085] The user enters the required basic status information and clicks the "Generate" button.
[0086] Step 4:
[0087] The server receives the information entered by the user.
[0088] Step 5:
[0089] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[0090] Step 6:
[0091] The server sends a preview image of the generated avatar to the terminal.
[0092] Step 7:
[0093] The user checks the preview image displayed on the terminal.
[0094] Step 8:
[0095] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button.
[0096] Step 9:
[0097] The terminal transmits the final avatar information to the server.
[0098] 2. Generating speech content
[0099] Step 1:
[0100] The user clicks the "Generate Message" button on the terminal.
[0101] Step 2:
[0102] The terminal displays a form for the user to input the topic or theme of the comment.
[0103] Step 3:
[0104] The user enters a topic or theme and clicks the "Generate" button.
[0105] Step 4:
[0106] The server receives topic information input by the user.
[0107] Step 5:
[0108] The server launches the LLM and generates comments based on the input topic.
[0109] Step 6:
[0110] The server transmits the generated speech content to the terminal.
[0111] Step 7:
[0112] The user checks the message displayed on the terminal.
[0113] Step 8:
[0114] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[0115] Step 9:
[0116] The terminal transmits the confirmed statement content to the server.
[0117] 3. Speech synthesis
[0118] Step 1:
[0119] The user sends a request for speech synthesis based on the confirmed utterance content.
[0120] Step 2:
[0121] The server receives the confirmed statement content.
[0122] Step 3:
[0123] The server performs speech synthesis through a natural language processing engine.
[0124] Step 4:
[0125] The server transmits the generated audio file to the terminal.
[0126] Step 5:
[0127] The user plays the audio file on the device and checks the content.
[0128] Step 6:
[0129] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[0130] 4. Audio and Motion Mapping
[0131] Step 1:
[0132] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[0133] Step 2:
[0134] The server receives the audio file and avatar data.
[0135] Step 3:
[0136] The server generates motion sequences to achieve natural lip sync and body language.
[0137] Step 4:
[0138] The server transmits the completed motion image to the terminal.
[0139] Step 5:
[0140] The user plays and checks the generated motion video on the terminal.
[0141] Step 6:
[0142] The user requests corrections to movements and facial expressions as needed.
[0143] Step 7:
[0144] The user saves the finalized motion image.
[0145] 5. Creating subtitles
[0146] Step 1:
[0147] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[0148] Step 2:
[0149] The server receives the text of the confirmed utterance.
[0150] Step 3:
[0151] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[0152] Step 4:
[0153] The server overlays and integrates the subtitle file onto the motion video.
[0154] Step 5:
[0155] The server sends the completed video to the terminal.
[0156] Step 6:
[0157] The user checks the video and subtitles on the device.
[0158] Step 7:
[0159] The user checks whether there are any problems during playback of the video and subtitles and requests corrections if necessary.
[0160] Step 8:
[0161] The user saves the finalized video and publishes or shares it.
[0162] The above is a specific program process divided into each processing step.
[0163] Example 1
[0164] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0165] Conventional virtual character creation systems require advanced programming knowledge and specialized skills, making them difficult for average users to use. Furthermore, multiple processes, such as generating speech content, synthesizing voices, and synchronizing motions, are required, making the process time-consuming and tedious. There is a need for a system that can solve these problems and enable anyone to easily create high-quality virtual characters.
[0166] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0167] In this invention, the server includes: means for generating a 3D or Live2D avatar based on information input by a user; means for the user to perform detailed customization of the avatar; means for using a generative AI model to generate speech content based on an input theme; means for providing a user interface for confirming and correcting the generated speech content; means for voice-synthesizing the generated speech content; means for synchronizing a voice file with lip sync and body language based on the confirmed speech content; means for generating a motion sequence corresponding to the voice file; and means for automatically generating subtitles based on the speech content and integrating them into video. This enables users to quickly and efficiently create high-quality virtual characters without requiring special technical knowledge.
[0168] "User" refers to any individual or organization that uses the Platform to create and customize virtual characters.
[0169] "Server" refers to a computer system that performs processes such as generating virtual characters, generating speech content, synthesizing voice, and integrating video and subtitles.
[0170] "Avatar" refers to a 3D or Live2D virtual character generated based on user input.
[0171] "3D avatar" refers to a virtual character that is visually represented in three-dimensional space.
[0172] A "Live2D avatar" is a virtual character that dynamically expresses two-dimensional image data and appears three-dimensional.
[0173] A "generative AI model" refers to an artificial intelligence model that performs natural language processing based on topic information and text information from users.
[0174] "Comment content" refers to the text data generated by the generative AI model based on the topic specified by the user.
[0175] "Speech synthesis" refers to the technology of generating a voice file based on the text data of a speech.
[0176] "Lip sync" refers to the technology of synchronizing the mouth movements of a virtual character with an audio file.
[0177] "Body language" refers to the technology that generates the physical movements and gestures of a virtual character in accordance with the voice and content of what is being said.
[0178] "Motion sequence" refers to a series of movement data including lip sync and body language.
[0179] "Subtitles" refers to text data generated based on what is being said and displayed to viewers within the video.
[0180] "Customization options" refers to settings that a user can select to adjust the appearance and behavior of their avatar.
[0181] "Administration tool" refers to software that provides an interface for users to request modifications to the videos they generate.
[0182] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. This system enables anyone to generate high-quality virtual characters and create related content without requiring specialized skills or programming knowledge.
[0183] System configuration
[0184] The system includes the following main components:
[0185] 1. Avatar generation
[0186] The server receives information input by the user (e.g., gender, age, style, etc.) and activates an image generation AI module (e.g., Stable Diffusion, DALL-E 2) to generate a 3D or Live2D avatar. The generated avatar is then sent to the device.
[0187] On the terminal, the user customizes the details of the avatar (e.g., hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[0188] 2. Generating speech content
[0189] Users enter topic information into a form on their device (e.g., "Introducing AITubers"). The server uses a generative AI model (e.g., GPT-3, ChatGPT) to generate a post based on the topic. The generated post is sent to the device for the user to review.
[0190] The user confirms and modifies the generated message and sends the finalized message to the server.
[0191] 3. Speech synthesis
[0192] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly), and sends the generated audio file to the device.
[0193] The user checks the audio file and requests corrections if necessary.
[0194] 4. Audio and Motion Mapping
[0195] The server receives the audio file and avatar data, creates a motion sequence to generate lip sync and body language (e.g., Live2D Cubism, Adobe Character Animator), and sends the generated motion image to the device.
[0196] The user checks the motion image and requests corrections if necessary.
[0197] 5. Creating subtitles
[0198] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech. The subtitle file is integrated into the video, and the completed video is sent to the terminal.
[0199] The user checks the video and subtitles and confirms them after a final check.
[0200] Specific examples
[0201] Example 1: When a user creates an introductory video for an AITuber
[0202] 1. Create an avatar
[0203] The user clicks the "Create Avatar" button on the device's UI and enters the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[0204] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[0205] 2. Generating speech content
[0206] The user enters the topic "Introducing AITubers" into a form on their device. The server uses a generative AI model to generate comments based on the topic and sends them to the device.
[0207] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[0208] 3. Speech Synthesis
[0209] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[0210] The user checks the audio file and confirms it if there are no particular problems.
[0211] 4. Audio and Motion Mapping
[0212] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[0213] The user checks the motion image and confirms it if there are no particular problems.
[0214] 5. Creating subtitles
[0215] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[0216] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[0217] Prompt Sentence Examples
[0218] Generate a 3D character based on the following information: Gender is female, age is in her 20s, and she wears casual clothing. For background information, the character is an AITuber on YouTube. The statement should read, "Hello, today I'd like to introduce my channel."
[0219] The above system allows users to quickly and easily generate virtual characters and create compelling content without any special technical knowledge.
[0220] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0221] Step 1:
[0222] The user clicks the "Create Avatar" button on the device's UI and enters basic status information (e.g., gender, age, style, etc.).
[0223] Input: The user fills in a form with basic information such as gender, age, and style.
[0224] Action: Sends the entered information to the server.
[0225] Output: Basic status information is entered into the server.
[0226] Step 2:
[0227] The server launches an image generation AI module (e.g., Stable Diffusion, DALL-E 2) based on the received basic status information and generates a 3D or Live2D avatar.
[0228] Input: Basic status information.
[0229] How it works: Passes information to an image generation AI module, which applies an algorithm to generate an avatar.
[0230] Output: An initial avatar image is generated and sent to the device as a preview.
[0231] Step 3:
[0232] On the device, the user receives a preview of the generated avatar and can then perform detailed customization (e.g., hairstyle, clothing, etc.).
[0233] Input: Previewed avatar image and customization options.
[0234] How it works: Use the customization UI to adjust the details using sliders and dropdown menus.
[0235] Output: Generates the finalized avatar configuration.
[0236] Step 4:
[0237] When the user confirms the avatar, the data of the confirmed avatar is transmitted to the server.
[0238] Input: Avatar settings adjusted by the user.
[0239] Action: Press the Confirm button to send the avatar configuration data to the server.
[0240] Output: The confirmed avatar configuration data is sent to the server.
[0241] Step 5:
[0242] Users enter topic information into a form on their device (e.g., "Introducing AITubers").
[0243] Input: Topic information.
[0244] Action: Enter a topic in the text area and press the send button.
[0245] Output: The topic information is sent to the server.
[0246] Step 6:
[0247] The server passes the received topic information to a generative AI model (e.g., GPT-3, ChatGPT) to generate the content of the comments.
[0248] Input: Topic information.
[0249] How it works: Provide topic information as prompts to a generative AI model and receive the generated utterances.
[0250] Output: The generated speech is sent to the terminal.
[0251] Step 7:
[0252] The user checks the generated comment and corrects it if necessary.
[0253] Input: The generated utterance.
[0254] Action: Edit the comment in the editor and press the send button after editing.
[0255] Output: The corrected statement is sent to the server.
[0256] Step 8:
[0257] The user confirms the content of the statement and sends it to the server.
[0258] Input: Confirmed statement.
[0259] Action: Press the Confirm button to send the final revised statement to the server.
[0260] Output: The confirmed statement is sent to the server.
[0261] Step 9:
[0262] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly).
[0263] Input: Confirmed statement.
[0264] What it does: Passes what is said to a speech synthesis engine, generating an audio file.
[0265] Output: An audio file is generated and sent to the device.
[0266] Step 10:
[0267] The user checks the generated audio file and requests corrections if necessary.
[0268] Input: An audio file.
[0269] Action: Plays an audio file and sends feedback for corrections if needed.
[0270] Output: Finalized audio file.
[0271] Step 11:
[0272] The server receives the audio file and avatar data and generates lip sync and body language.
[0273] Input: Audio file and avatar data.
[0274] Movement: Generate avatar mouth and body movements in sync with the audio to create a motion sequence.
[0275] Output: A motion image is generated and sent to the device.
[0276] Step 12:
[0277] The user checks the generated motion image and requests corrections if necessary.
[0278] Input: Motion footage.
[0279] How it works: Plays motion footage and provides feedback for corrections if there are any issues.
[0280] Output: Finalized motion footage.
[0281] Step 13:
[0282] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech.
[0283] Input: The text of what was said.
[0284] Operation: Converts text data into subtitle file format according to the time axis.
[0285] Output: Subtitle files are generated and integrated into the video.
[0286] Step 14:
[0287] The user can check the automatically generated subtitled video and request corrections if necessary.
[0288] Input: Video with subtitles.
[0289] Action: Play the video to check the timing and content of the subtitles, and send correction requests if necessary.
[0290] Output: Finalized video with subtitles.
[0291] The above are the specific steps of the system's program processing, which allows users to efficiently create virtual characters and generate related content.
[0292] (Application example 1)
[0293] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0294] Previous virtual character generation platforms required a lot of effort for users to create characters, making them inconvenient, especially when using smartphones. Furthermore, they lacked assistant functions for real-time user interaction and product introductions, making them difficult to apply to virtual stores. Furthermore, they lacked management tools that could flexibly accommodate user requests for corrections to the videos they generated.
[0295] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0296] In this invention, the server includes means for generating a 3D or Live2D avatar based on information input by a user, means for generating speech content based on the input theme, means for synthesizing the generated speech content with speech, means for synchronizing the speech file with the avatar's movements, means for automatically generating subtitles based on the speech content and integrating them into video, means for using the avatar as a virtual assistant character on a smartphone, and means for the virtual assistant character to introduce products and respond to user inquiries in real time. This allows users to easily create a virtual assistant character on their smartphone and realize real-time conversations, product introductions, and customer support.
[0297] "User-inputted information" refers to basic status information and thematic input data that a user provides to the system.
[0298] A "3D or Live2D avatar" is a three-dimensional or two-dimensional character generated based on user input.
[0299] "Comment content" is text information generated based on the theme entered by the user.
[0300] "Speech synthesis" is the process of creating an audio file based on generated utterances.
[0301] "Synchronizing the audio file with the avatar's movements" is the process of moving the character's mouth and body in sync with the generated audio.
[0302] "Automatic subtitling" is the process of generating text based on what is being said and integrating it with the video.
[0303] "Use on a smartphone" refers to using a virtual assistant character on a smartphone application.
[0304] A "virtual assistant character" is a digital character created by a user that introduces products and responds to user inquiries in real time.
[0305] The "product introduction function" is a function in which a virtual assistant character explains specific product information to the user.
[0306] "Real-time response" refers to the ability to respond immediately to user questions and requests.
[0307] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Furthermore, this system can be used as a virtual assistant character on smartphones, and has the ability to introduce products and respond to users in real time.
[0308] System configuration
[0309] 1. 3D / Live2D avatar generation
[0310] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[0311] On the device (smartphone), the user customizes the avatar in detail (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[0312] 2. Generating speech content
[0313] The server passes the topic information entered by the user to the generative AI model to generate the utterance content, which is then sent to the device for the user to confirm.
[0314] The user confirms and modifies the generated message and sends the finalized message to the server.
[0315] 3. Speech synthesis
[0316] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[0317] The user checks the audio file and requests corrections if necessary.
[0318] 4. Audio and Motion Mapping
[0319] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[0320] The user checks the motion image and requests corrections if necessary.
[0321] 5. Creating subtitles
[0322] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file into the video, and sends the completed video to the device.
[0323] The user checks the video and subtitles and confirms them after a final check.
[0324] 6. Use as a virtual assistant
[0325] The virtual assistant character generated on the device (smartphone) introduces products and responds to user inquiries in real time, realizing an interactive dialogue with the user.
[0326] Users obtain specific product information or inquiries through the virtual assistant.
[0327] Hardware and software used
[0328] Hardware: Smartphones, servers
[0329] Software: OpenAI API (generative AI model), pyttsx3 (speech synthesis), OpenCV (video and motion processing)
[0330] Specific examples
[0331] Example 1: A user creates a virtual assistant character to introduce a new product, and the character explains the company's product through a smartphone app.
[0332] Example prompt sentence:
[0333] "Please explain the features of your new product, including its target market and its benefits."
[0334] Sample output: "This new product is a casual-style smartwatch designed for young people. Key features include lightweight design, long battery life, and the latest fitness tracking capabilities. It's especially suitable for consumers who lead active lifestyles."
[0335] This system allows users to easily create and use virtual assistant characters on their smartphones, allowing them to introduce products and respond in real time.
[0336] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0337] Step 1:
[0338] Obtaining input information
[0339] Users use a smartphone app to input basic status information (gender, age, style, etc.) to create an avatar.
[0340] The server receives information entered by the user and launches the image generation AI module.
[0341] Output: Basic status information for the user.
[0342] Step 2:
[0343] Avatar generation
[0344] The server generates a 3D or Live2D avatar based on the user's basic status information received in step 1.
[0345] The server sends a preview of the generated avatar to the terminal.
[0346] Output: The generated avatar and its preview image.
[0347] Step 3:
[0348] Avatar Customization
[0349] Users customize their avatar's details (hairstyle, clothing, etc.) using a smartphone app.
[0350] The terminal sends the customized results to the server.
[0351] Output: Customized avatar information.
[0352] Step 4:
[0353] Generate speech content
[0354] The user uses the app to input the topic of the comment (e.g., new product introduction).
[0355] The server generates speech content based on the input topic using a generative AI model (OpenAI API).
[0356] The server sends the generated message to the terminal, where the user can confirm and edit it.
[0357] Output: The generated utterance.
[0358] Step 5:
[0359] Confirmation and confirmation of statements
[0360] The user checks the generated comment and corrects it if necessary.
[0361] The user sends the finalized content of the message to the server.
[0362] Output: Confirmed statement.
[0363] Step 6:
[0364] Speech synthesis
[0365] The server passes the confirmed utterance content to a natural language processing engine (pyttsx3) and performs speech synthesis.
[0366] The server transmits the generated audio file to the terminal.
[0367] Output: The generated audio file.
[0368] Step 7:
[0369] Audio and Motion Mapping
[0370] The server receives the audio files and avatar data and creates motion sequences to generate lip sync and body language.
[0371] The server transmits the generated motion image to the terminal.
[0372] Output: The generated motion footage.
[0373] Step 8:
[0374] Subtitle creation and integration
[0375] The server automatically generates a subtitle file based on the text of the confirmed speech content.
[0376] The server integrates the subtitle file into the video and sends the final video to the terminal.
[0377] Output: A merged video file.
[0378] Step 9:
[0379] Use as a virtual assistant
[0380] The device (smartphone) uses the generated virtual assistant character to introduce products and respond to user inquiries in real time.
[0381] Users can use the virtual assistant to obtain product information and make inquiries.
[0382] Output: Product information and support provided to the user.
[0383] The above is a detailed flow of the processing steps of this system. By performing these operations, users can easily create a virtual assistant character and introduce products and respond in real time on their smartphones.
[0384] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0385] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. By combining it with an emotion engine, it also has the ability to dynamically adjust the avatar's expressions based on the user's emotions. This system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Additionally, by incorporating an emotion engine that recognizes the user's emotions, it dynamically changes facial expressions and movements, and adjusts the tone and pitch of the voice, based on the user's emotions.
[0386] System configuration
[0387] 1. 3D / Live2D avatar generation
[0388] The server receives basic status information (gender, age, style, etc.) entered by the user and activates an image generation AI module to generate a 3D or Live2D avatar.
[0389] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[0390] 2. Generating speech content
[0391] The server passes the topic information entered by the user to the LLM to generate the content of the comment, which is then sent to the terminal.
[0392] The user confirms and modifies the generated message and sends the finalized message to the server.
[0393] 3. Speech synthesis
[0394] The server synthesizes the confirmed speech content through a natural language processing engine and transmits the generated audio file to the terminal.
[0395] The user checks the audio file and requests corrections if necessary.
[0396] 4. Audio and Motion Mapping
[0397] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and then transmits the generated motion image to the device.
[0398] The user checks the motion image and requests corrections if necessary.
[0399] 5. Creating subtitles
[0400] The server automatically generates a subtitle file based on the text of the confirmed remarks, integrates it with the video, and sends it to the terminal.
[0401] The user checks the video and subtitles and confirms them after a final check.
[0402] 6. Leveraging Emotional Engines
[0403] The device uses an emotion engine to analyze the user's facial expressions, tone of voice, and gestures in real time to recognize the user's emotions.
[0404] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[0405] The server adjusts the tone and pitch of the generated speech synthesis according to the user's emotions.
[0406] Specific examples
[0407] Example 1: A user creates a video to express their emotions
[0408] 1. Create an avatar
[0409] The user clicks the "Create Avatar" button and enters the avatar's basic status (e.g., male, young man, casual style, etc.). The server generates an avatar based on the entered information and sends a preview to the device.
[0410] The user can view the preview, customize the details and finally confirm the avatar.
[0411] 2. Generating speech content
[0412] The user inputs the topic "Emotional Expression" as the theme. The server uses LLM to generate utterances based on the theme and sends them to the device.
[0413] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[0414] 3. Speech Synthesis
[0415] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[0416] The user checks the audio and confirms it if there is no particular problem.
[0417] 4. Audio and Motion Mapping
[0418] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[0419] The user checks the motion image and confirms it if there are no particular problems.
[0420] 5. Creating subtitles
[0421] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[0422] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[0423] 6. Leveraging Emotional Engines
[0424] The device senses the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine.
[0425] The server dynamically adjusts lip sync, facial expressions, and body language based on the obtained emotional data, and appropriately changes the tone and pitch of the voice.
[0426] The user checks the video of the generated emotional expression, makes final adjustments and confirms it.
[0427] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system allows users to easily generate high-quality virtual character content that reflects their own emotions, without requiring programming knowledge.
[0428] The processing flow will be explained below.
[0429] 1. Creating a 3D / Live2D avatar
[0430] Step 1:
[0431] The user clicks the "Create Avatar" button on the device.
[0432] Step 2:
[0433] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[0434] Step 3:
[0435] The user enters the required basic status information and clicks the "Generate" button.
[0436] Step 4:
[0437] The server receives the information entered by the user.
[0438] Step 5:
[0439] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[0440] Step 6:
[0441] The server sends a preview image of the generated avatar to the terminal.
[0442] Step 7:
[0443] The user checks the preview image displayed on the terminal.
[0444] Step 8:
[0445] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button again.
[0446] Step 9:
[0447] The terminal transmits the final avatar information to the server.
[0448] 2. Generating speech content
[0449] Step 1:
[0450] The user clicks the "Generate Message" button on the terminal.
[0451] Step 2:
[0452] The terminal displays a form for the user to input the topic or theme of the comment.
[0453] Step 3:
[0454] The user enters a topic or theme and clicks the "Generate" button.
[0455] Step 4:
[0456] The server receives topic information input by the user.
[0457] Step 5:
[0458] The server launches the LLM and generates comments based on the input topic.
[0459] Step 6:
[0460] The server transmits the generated speech content to the terminal.
[0461] Step 7:
[0462] The user checks the message displayed on the terminal.
[0463] Step 8:
[0464] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[0465] Step 9:
[0466] The terminal transmits the confirmed statement content to the server.
[0467] 3. Speech synthesis
[0468] Step 1:
[0469] The user sends a request for speech synthesis based on the confirmed utterance content.
[0470] Step 2:
[0471] The server receives the confirmed statement content.
[0472] Step 3:
[0473] The server performs speech synthesis through a natural language processing engine.
[0474] Step 4:
[0475] The server transmits the generated audio file to the terminal.
[0476] Step 5:
[0477] The user plays the audio file on the device and checks the content.
[0478] Step 6:
[0479] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[0480] 4. Audio and Motion Mapping
[0481] Step 1:
[0482] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[0483] Step 2:
[0484] The server receives the audio file and avatar data.
[0485] Step 3:
[0486] The server generates motion sequences to achieve natural lip sync and body language.
[0487] Step 4:
[0488] The server transmits the completed motion image to the terminal.
[0489] Step 5:
[0490] The user plays and checks the generated motion video on the terminal.
[0491] Step 6:
[0492] The user requests corrections to movements and facial expressions as needed.
[0493] Step 7:
[0494] The user saves the finalized motion image.
[0495] 5. Creating subtitles
[0496] Step 1:
[0497] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[0498] Step 2:
[0499] The server receives the text of the confirmed utterance.
[0500] Step 3:
[0501] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[0502] Step 4:
[0503] The server overlays and integrates the subtitle file onto the motion video.
[0504] Step 5:
[0505] The server sends the completed video to the terminal.
[0506] Step 6:
[0507] The user checks the video and subtitles on the device.
[0508] Step 7:
[0509] The user saves the finalized video and publishes or shares it.
[0510] 6. Leveraging Emotional Engines
[0511] Step 1:
[0512] When a user plays back the video they created, the device's camera and microphone are used to detect emotional data such as facial expressions and tone of voice in real time.
[0513] Step 2:
[0514] The terminal transmits the emotion data acquired in real time to the emotion engine, and transmits the analysis results to the server.
[0515] Step 3:
[0516] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[0517] Step 4:
[0518] When the server synthesizes voice from the generated speech, it adjusts the tone and pitch of the voice according to the user's emotions.
[0519] Step 5:
[0520] The server generates the final video based on the emotion data and sends it to the terminal.
[0521] Step 6:
[0522] The user checks the video in which the emotional expression is reflected and requests corrections as necessary.
[0523] Step 7:
[0524] The user saves the finalized video and publishes or shares it.
[0525] The above are the specific processing steps in the program processing of a system that utilizes an emotion engine.
[0526] Example 2
[0527] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0528] Conventional virtual character creation platforms make it difficult for users to express emotions or easily generate speech content. Synchronizing video and audio and generating subtitles is also time-consuming, making the overall process complicated and making it difficult to easily create high-quality content.
[0529] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a 3D or 2D avatar based on information input by a user, a means for generating speech content based on an input theme, a means for voice-synthesizing the generated speech content, a means for synchronizing the speech file with the avatar's movement, a means for automatically generating subtitles based on the speech content and integrating them into video, and a means for collecting user emotion data using an emotion analysis engine and dynamically adjusting the avatar's expression. This allows a user to easily create high-quality virtual character content and generate video including real-time expressions that reflect the user's emotions.
[0530] "User" refers to the entity that uses this system to create a virtual character, customize it, generate comments, etc.
[0531] "Server" refers to a central processing unit that receives input information from a user and performs various processes.
[0532] "3D or 2D Avatar" means a virtual character created by a user and represented in three-dimensional or two-dimensional graphics.
[0533] "Input information" refers to data that a user provides to the system, such as basic status information such as gender, age, and style.
[0534] A "theme" refers to a topic that a user provides to the system for generating comment content.
[0535] "Speech content" refers to the lines and phrases of the character generated based on the input theme.
[0536] "Speech synthesis" is a technology that converts text data into voice data, and refers to a means of outputting the generated utterances as voice.
[0537] "Synchronizing" refers to the process of matching the audio and character movements.
[0538] "Subtitles" are texts that are automatically generated based on what is being said and are integrated into video.
[0539] "Integrating into video" refers to the process of incorporating the generated subtitles into a single video file in conjunction with the audio and avatar movements.
[0540] An "emotion analysis engine" refers to software or a system that analyzes a user's facial expressions, tone of voice, and gestures to recognize emotions.
[0541] "Emotion data" refers to data that indicates the user's emotional state as recognized by the emotion analysis engine.
[0542] "Dynamic adjustment" refers to the act of changing an avatar's facial expressions, movements, voice tone, pitch, etc. in real time.
[0543] "Customization options" refers to settings that allow users to change details about their avatar.
[0544] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters, and by combining it with an emotion analysis engine, it has the ability to dynamically adjust the avatar's expression according to the user's emotions.
[0545] System configuration
[0546] The system consists of the following main components:
[0547] 1. Server: A central processing unit that receives information from users and performs various processes.
[0548] 2. Terminal: A device operated by a user that provides an interface for user input, display, customization, etc.
[0549] 3. Sentiment analysis engine: Software for analyzing user emotions and generating data.
[0550] Technology used
[0551] 1. Image generation AI module (e.g. DeepArt): Generates a 3D or 2D avatar based on basic status information (gender, age, style, etc.) entered by the user.
[0552] 2. Large-scale language models (LLMs) (e.g., GPT-4): Generate utterances based on themes entered by the user.
[0553] 3. Natural language processing engine (e.g., Google Text-to-Speech): Converts the generated speech into audio.
[0554] 4. Emotion analysis engine (e.g., Affectiva SDK): Analyzes the user's facial expressions, tone of voice, and gestures to collect emotional data.
[0555] System Operation
[0556] The system operates in the following steps:
[0557] 1. Create your avatar:
[0558] The user clicks the "Create Avatar" button on the device and enters basic status information (gender, age, style, etc.).
[0559] The terminal transmits this information to the server.
[0560] The server uses an image generation AI module to generate a 3D or 2D avatar and sends a preview to the device.
[0561] The user checks the preview image, performs detailed customization (hairstyle, clothing, etc.), and confirms it.
[0562] The terminal transmits the determined avatar information to the server.
[0563] 2. Speech generation:
[0564] The user inputs topic information and sends it to the server.
[0565] The server uses LLM to generate speech content and sends it to the terminal.
[0566] The user confirms and corrects the content of the comment and sends the finalized content to the server.
[0567] 3. Speech synthesis:
[0568] The server synthesizes the confirmed speech content through a natural language processing engine.
[0569] The server transmits the generated audio file to the terminal.
[0570] The user checks the audio file and requests corrections if necessary.
[0571] 4. Audio and Motion Mapping:
[0572] The server passes the audio file and avatar data to a lip sync engine to generate lip sync and body language.
[0573] The server transmits the generated motion sequence to the terminal.
[0574] The user checks the motion image and requests corrections if necessary.
[0575] 5. Creating subtitles:
[0576] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[0577] The user checks the video and subtitles and confirms them after a final check.
[0578] 6. Utilizing sentiment analysis engines:
[0579] The device detects the user's facial expressions and tone of voice and passes them to an emotion analysis engine in real time.
[0580] The terminal transmits the emotion analysis results obtained to the server.
[0581] The server dynamically adjusts the avatar's facial expressions and movements based on the emotional data, and also appropriately changes the tone and pitch of the voice.
[0582] The user checks the final image, makes adjustments and confirms them.
[0583] Specific examples
[0584] Example 1: A user creates a video to express their emotions
[0585] 1. Create your avatar:
[0586] The user clicks the "Create Avatar" button and inputs the avatar's basic status (e.g., male, young man, casual style, etc.).
[0587] The terminal displays an input form and accepts input from the user.
[0588] When the user inputs information and presses the send button, the terminal sends the information to the server.
[0589] The server receives the information, activates the image generation AI module, and generates the avatar.
[0590] The server sends the generated avatar to the device as a preview.
[0591] The device displays a preview, and the user can customize the details (hairstyle, clothing, etc.) and finally confirm the avatar.
[0592] The terminal transmits the determined avatar information to the server.
[0593] 2. Speech generation:
[0594] The user opens the "Comment Content" input form and inputs the topic "Emotional Expressions."
[0595] The terminal transmits the topic information to the server.
[0596] The server generates the message using LLM and sends it to the terminal.
[0597] The user confirms and modifies the generated message and sends the finalized message to the server.
[0598] 3. Speech synthesis:
[0599] The server uses a natural language processing engine to synthesize speech based on the confirmed content of the speech.
[0600] The server transmits the generated audio file to the terminal.
[0601] The user checks the audio file and confirms it if there are no problems.
[0602] 4. Audio and Motion Mapping:
[0603] The server generates lip sync and body language based on the audio file and the determined avatar.
[0604] The server transmits the motion video to the terminal and provides it to the user.
[0605] The user checks the motion image and requests corrections if necessary.
[0606] 5. Creating subtitles:
[0607] The server automatically generates subtitles based on the confirmed remarks, integrates them into the video, and sends them to the terminal.
[0608] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[0609] 6. Utilizing sentiment analysis engines:
[0610] The device senses the user's facial expressions and tone of voice in real time and sends the analysis results to the emotion engine.
[0611] The terminal transmits the emotion analysis result to the server.
[0612] The server dynamically adjusts lip sync, facial expressions, and movements based on emotional data, and also appropriately changes the tone and pitch of the voice.
[0613] The user checks the final image, makes adjustments and confirms them.
[0614] By following the above steps, users can easily create high-quality virtual character content. This system also enables real-time expression according to emotions, making it possible to generate a wide variety of content.
[0615] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0616] Step 1: User begins creating an avatar
[0617] The user clicks the "Create Avatar" button on the device.
[0618] Input: User input (basic status information such as gender, age, style, etc.)
[0619] The terminal displays an input form and accepts input from the user.
[0620] How it works: The user enters the required information and presses the submit button.
[0621] Output: The information entered
[0622] Step 2: The server generates the avatar
[0623] The server launches an image generation AI module based on the user's basic status information received from the device.
[0624] Input: Basic status information entered
[0625] How it works: The server uses an image generation AI module to generate a 3D or 2D avatar.
[0626] Output: A preview image of the generated avatar
[0627] Step 3: User customizes avatar
[0628] The terminal displays a preview of the generated avatar to the user.
[0629] Input: A preview image of the generated avatar
[0630] How it works: The user customizes the avatar's details (hairstyle, clothing, etc.) and finally finalizes the avatar.
[0631] Output: A customized confirmed avatar
[0632] Step 4: Enter and submit topic information
[0633] The user opens the "Comment Generation" input form on the terminal and inputs a topic (e.g., "About emotional expression").
[0634] Input: Topic information
[0635] How it works: The user enters a topic and presses the send button.
[0636] Output: Input topic information
[0637] Step 5: Generate speech
[0638] The server receives topic information and generates utterances using a large-scale language model (LLM).
[0639] Input: Topic information
[0640] How it works: The server passes the prompt to the generative AI model to generate the speech.
[0641] Output: Generated speech
[0642] Step 6: Revise and confirm your statement
[0643] The terminal displays the generated comment content to the user.
[0644] Input: Generated speech
[0645] Action: The user confirms and corrects what has been said, and finally confirms it.
[0646] Output: Confirmed statement
[0647] Step 7: Speech synthesis of what is being said
[0648] The server passes the confirmed utterance content to a natural language processing engine (e.g., Google Text-to-Speech) for speech synthesis.
[0649] Input: Confirmed statement
[0650] How it works: The server uses a speech synthesis engine to generate an audio file.
[0651] Output: Generated audio file
[0652] Step 8: Check and correct the audio file
[0653] The terminal provides the generated audio file to the user.
[0654] Input: Generated audio file
[0655] Action: The user reviews the audio file and requests corrections if necessary.
[0656] Output: Finalized audio file
[0657] Step 9: Mapping Audio and Motion
[0658] The server passes the audio file and the determined avatar to a lip sync engine to generate lip sync and body language.
[0659] Input: Audio file, confirmed avatar
[0660] How it works: The server generates lip sync and motion data.
[0661] Output: Generated motion sequence
[0662] Step 10: Check and correct motion footage
[0663] The terminal displays the generated motion image to the user.
[0664] Input: Generated motion sequence
[0665] Action: The user reviews the motion footage and requests corrections if necessary.
[0666] Output: Confirmed motion footage
[0667] Step 11: Creating and merging subtitles
[0668] The server automatically generates subtitles based on the confirmed remarks and integrates them into the video.
[0669] Input: Confirmed speech content, confirmed motion image
[0670] How it works: The server generates a subtitle file and integrates it into the video.
[0671] Output: Integrated video (audio, motion, subtitles)
[0672] Step 12: Reflecting the results of sentiment analysis
[0673] The device passes the user's facial expressions and tone of voice to an emotion analysis engine.
[0674] Input: Real-time facial expressions, tone of voice
[0675] Operation: The device obtains the emotion analysis results and sends them to the server.
[0676] Output: Emotion data
[0677] Step 13: Adjust based on sentiment data
[0678] The server dynamically adjusts the avatar's facial expressions and movements based on the emotional data, and changes the tone and pitch of the voice.
[0679] Input: Emotion data, confirmed motion video, audio file
[0680] Movement: The server adjusts the avatar's facial expressions, movements, and voice tone and pitch.
[0681] Output: Adjusted video and audio
[0682] Step 14: Final review and confirmation
[0683] The terminal provides the final video to the user for confirmation.
[0684] Input: Adjusted video and audio
[0685] Action: The user performs a final review, makes any necessary adjustments, and confirms.
[0686] Output: Final video
[0687] (Application example 2)
[0688] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0689] Currently, it is difficult for users to easily create interactive, emotionally expressive virtual characters. It is also difficult to dynamically adjust the virtual character's facial expressions and voice tone in response to the user's real-time emotions. This can result in a lack of a natural and engaging experience for the user, leading to lower user satisfaction. The present invention aims to solve these problems and provide a system that enables a virtual character to express natural emotions in response to the user's emotions.
[0690] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating a 3D or Live2D avatar based on information input by the user, means for generating speech content based on the input theme, means for voice-synthesizing the generated speech content, means for synchronizing the audio file with the avatar's movements, means for automatically generating subtitles based on the speech content and integrating them into video, means for sensing the user's voice and facial expressions in real time and analyzing emotions using an emotion engine, and means for dynamically adjusting the avatar's facial expressions and voice tone based on the emotion data. This allows users to easily create virtual characters and enjoy natural expressions that are dynamically adjusted according to emotions.
[0691] "3D or Live2D Avatar" means a virtual character that is represented in three or two dimensions based on a user-selected appearance and style.
[0692] "Means for generating utterances" refers to a system or algorithm that generates appropriate responses or utterances based on themes or questions entered by users.
[0693] "Means for speech synthesis" refers to the technology for generating speech from text, and the function for converting the generated speech into natural-sounding speech.
[0694] "Means for synchronizing audio files with avatar movements" refers to a technology that synchronizes the mouth movements and facial expressions of an avatar with the generated audio.
[0695] "Means for automatically generating subtitles and integrating them into video" refers to technology that automatically generates subtitles based on what is being said and embeds them into the video.
[0696] The "emotion engine" is a system that analyzes the user's tone of voice and facial expressions to recognize the user's emotional state in real time.
[0697] "Means for dynamically adjusting the avatar's facial expression and voice tone based on emotional data" is a function that appropriately changes the avatar's facial expression and voice tone based on the results of analysis by the emotion engine.
[0698] This invention provides a platform that allows users to easily create 3D or Live2D avatars, which then dynamically change their expressions based on the user's emotions. Based on user input, this system generates and adjusts real-time avatar movements that integrate speech, audio, subtitles, and emotional expressions.
[0699] System configuration
[0700] 1. Avatar generation
[0701] The server receives basic information entered by the user (such as gender, age, and style) and generates a 3D or Live2D avatar using an image generation AI module.
[0702] The terminal provides an interface that allows the user to customize the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[0703] 2. Generating speech content
[0704] The server uses a generative AI model (e.g., GPT-4) to generate a message based on the themes entered by the user. This message is then sent to the device, where it is reviewed and revised by the user before being resent.
[0705] 3. Speech Synthesis
[0706] The server generates audio based on the confirmed utterance using a natural language processing engine (e.g., Google Cloud Text-to-Speech) and sends the generated audio file to the device.
[0707] 4. Audio and Motion Synchronization
[0708] The server creates a motion sequence to generate lip sync and body language based on the audio file and avatar data, and sends it to the device.
[0709] 5. Creating subtitles
[0710] The server automatically generates a subtitle file based on the text of the confirmed speech and integrates it into the video, which is then sent to the device.
[0711] 6. Leveraging Emotional Engines
[0712] The device detects the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine (e.g., Microsoft Azure Emotion API).
[0713] The server dynamically adjusts the avatar's facial expressions, movements, and voice tone based on the emotional data obtained from the emotion engine.
[0714] Specific examples
[0715] Example 1: Virtual Shopping Assistant
[0716] The user puts on the smart glasses and says, "Start shopping assistant." The device sends the command to the server.
[0717] The server generates a 3D avatar based on basic information provided by the user and sends it to the device. The user then customizes the avatar in detail (e.g., glasses, hairstyle) and finalizes the avatar.
[0718] The user asks, "What material is this shirt made of?" The server passes the question to a generative AI model (e.g., GPT-4), which generates an appropriate response and performs speech synthesis.
[0719] The device detects the user's facial expressions and tone of voice, analyzes them using an emotion engine, and displays and plays responses that dynamically adjust the avatar's facial expressions and tone of voice according to the emotional data.
[0720] Prompt Sentence Examples
[0721] User: What material is this shirt made of?
[0722] Virtual Character: Okay. This shirt is made of 100% cotton. It feels amazing!
[0723] This system allows users to easily generate high-quality virtual character content that reflects their own emotions, without requiring any programming knowledge.
[0724] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0725] Step 1:
[0726] Entering and creating basic avatar information
[0727] The user inputs basic information (gender, age, style, etc.) from the terminal, which is then sent to the server.
[0728] The server generates a 3D or Live2D avatar using an image generation AI module based on the received basic information, and a preview image of the generated avatar is sent to the device.
[0729] Output: A preview image of the avatar generated based on the basic information.
[0730] Step 2:
[0731] Detailed avatar customization
[0732] The user customizes the avatar in detail (hairstyle, clothing, etc.) from the terminal and finally confirms the avatar. The detailed customization data is sent to the server.
[0733] The server stores the determined avatar data, generates the final avatar image, and transmits it to the terminal.
[0734] Output: Final user-customized avatar image and data.
[0735] Step 3:
[0736] Generate speech content
[0737] The user inputs a specific topic or question (everyday conversation, product description, etc.) into the terminal and sends it to the server.
[0738] The server uses a generative AI model (e.g., GPT-4) to generate utterances based on the themes entered by the user. The generated utterances are then sent to the device.
[0739] Output: The speech text generated by the generative AI model.
[0740] Step 4:
[0741] Speech synthesis
[0742] The user checks and corrects the generated text and sends the finalized text to the server.
[0743] The server uses a natural language processing engine (e.g., Google Cloud Text-to-Speech) to convert what is said into audio, and the resulting audio file is sent to the device.
[0744] Output: Synthesized audio file.
[0745] Step 5:
[0746] Audio and motion synchronization
[0747] The server creates a motion sequence to generate lip sync and body language based on the audio file and avatar data, and sends it to the device.
[0748] The device will synchronize the motion sequence with the audio file and display a preview.
[0749] Output: Animated sequence with synchronized lip sync and body language.
[0750] Step 6:
[0751] Creating subtitles
[0752] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates it into the video, and sends the video file to the device.
[0753] The device will display a preview of the video and subtitles for the user to review.
[0754] Output: Video file with integrated subtitles.
[0755] Step 7:
[0756] Emotion analysis
[0757] The device senses the user's facial expressions and tone of voice in real time and transmits emotional data to the server.
[0758] The server analyzes the emotion data using an emotion engine (e.g., Microsoft Azure Emotion API).
[0759] Output: Parsed emotion data.
[0760] Step 8:
[0761] Dynamic adjustment based on emotions
[0762] The server generates sequences to dynamically adjust lip sync, facial expressions, and body language based on the emotional data and sends them to the device.
[0763] The device displays the adjusted animation and the user gives a final confirmation.
[0764] Output: Dynamically adjusted animation depending on the emotion.
[0765] Examples and prompts
[0766] Prompt Sentence Examples
[0767] User: What material is this shirt made of?
[0768] Virtual Character: Okay. This shirt is made of 100% cotton. It feels amazing!
[0769] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0770] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0771] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0772] [Second embodiment]
[0773] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0774] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0775] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0776] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0777] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0778] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0779] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0780] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0781] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0782] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0783] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0784] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0785] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech with motion, and generates and integrates subtitles into the video.
[0786] System configuration
[0787] 1. 3D / Live2D avatar generation
[0788] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[0789] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[0790] 2. Generating speech content
[0791] The server passes the topic information entered by the user to the LLM, generates a message, and sends the generated message to the terminal for the user to confirm.
[0792] The user confirms and modifies the generated message and sends the finalized message to the server.
[0793] 3. Speech synthesis
[0794] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[0795] The user checks the audio file and requests corrections if necessary.
[0796] 4. Audio and Motion Mapping
[0797] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[0798] The user checks the motion image and requests corrections if necessary.
[0799] 5. Creating subtitles
[0800] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file with the video, and sends the completed video to the device.
[0801] The user checks the video and subtitles and confirms them after a final check.
[0802] Specific examples
[0803] Example 1: When a user creates an introductory video for an AITuber
[0804] 1. Create an avatar
[0805] The user clicks the "Create Avatar" button on the device's UI and inputs the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[0806] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[0807] 2. Generating speech content
[0808] The user enters the topic "Introducing AITubers" into a form on the device. The server uses LLM to generate comments based on the topic and sends them to the device.
[0809] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[0810] 3. Speech Synthesis
[0811] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[0812] The user checks the audio file and confirms it if there are no particular problems.
[0813] 4. Audio and Motion Mapping
[0814] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[0815] The user checks the motion image and confirms it if there are no particular problems.
[0816] 5. Creating subtitles
[0817] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[0818] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[0819] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system enables anyone to easily generate content that utilizes high-quality virtual characters, without requiring programming knowledge.
[0820] The processing flow will be explained below.
[0821] 1. Creating a 3D / Live2D avatar
[0822] Step 1:
[0823] The user clicks the "Create Avatar" button on the device.
[0824] Step 2:
[0825] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[0826] Step 3:
[0827] The user enters the required basic status information and clicks the "Generate" button.
[0828] Step 4:
[0829] The server receives the information entered by the user.
[0830] Step 5:
[0831] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[0832] Step 6:
[0833] The server sends a preview image of the generated avatar to the terminal.
[0834] Step 7:
[0835] The user checks the preview image displayed on the terminal.
[0836] Step 8:
[0837] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button.
[0838] Step 9:
[0839] The terminal transmits the final avatar information to the server.
[0840] 2. Generating speech content
[0841] Step 1:
[0842] The user clicks the "Generate Message" button on the terminal.
[0843] Step 2:
[0844] The terminal displays a form for the user to input the topic or theme of the comment.
[0845] Step 3:
[0846] The user enters a topic or theme and clicks the "Generate" button.
[0847] Step 4:
[0848] The server receives topic information input by the user.
[0849] Step 5:
[0850] The server launches the LLM and generates comments based on the input topic.
[0851] Step 6:
[0852] The server transmits the generated speech content to the terminal.
[0853] Step 7:
[0854] The user checks the message displayed on the terminal.
[0855] Step 8:
[0856] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[0857] Step 9:
[0858] The terminal transmits the confirmed statement content to the server.
[0859] 3. Speech synthesis
[0860] Step 1:
[0861] The user sends a request for speech synthesis based on the confirmed utterance content.
[0862] Step 2:
[0863] The server receives the confirmed statement content.
[0864] Step 3:
[0865] The server performs speech synthesis through a natural language processing engine.
[0866] Step 4:
[0867] The server transmits the generated audio file to the terminal.
[0868] Step 5:
[0869] The user plays the audio file on the device and checks the content.
[0870] Step 6:
[0871] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[0872] 4. Audio and Motion Mapping
[0873] Step 1:
[0874] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[0875] Step 2:
[0876] The server receives the audio file and avatar data.
[0877] Step 3:
[0878] The server generates motion sequences to achieve natural lip sync and body language.
[0879] Step 4:
[0880] The server transmits the completed motion image to the terminal.
[0881] Step 5:
[0882] The user plays and checks the generated motion video on the terminal.
[0883] Step 6:
[0884] The user requests corrections to movements and facial expressions as needed.
[0885] Step 7:
[0886] The user saves the finalized motion image.
[0887] 5. Creating subtitles
[0888] Step 1:
[0889] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[0890] Step 2:
[0891] The server receives the text of the confirmed utterance.
[0892] Step 3:
[0893] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[0894] Step 4:
[0895] The server overlays and integrates the subtitle file onto the motion video.
[0896] Step 5:
[0897] The server sends the completed video to the terminal.
[0898] Step 6:
[0899] The user checks the video and subtitles on the device.
[0900] Step 7:
[0901] The user checks whether there are any problems during playback of the video and subtitles and requests corrections if necessary.
[0902] Step 8:
[0903] The user saves the finalized video and publishes or shares it.
[0904] The above is a specific program process divided into each processing step.
[0905] Example 1
[0906] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0907] Conventional virtual character creation systems require advanced programming knowledge and specialized skills, making them difficult for average users to use. Furthermore, multiple processes, such as generating speech content, synthesizing voices, and synchronizing motions, are required, making the process time-consuming and tedious. There is a need for a system that can solve these problems and enable anyone to easily create high-quality virtual characters.
[0908] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0909] In this invention, the server includes: means for generating a 3D or Live2D avatar based on information input by a user; means for the user to perform detailed customization of the avatar; means for using a generative AI model to generate speech content based on an input theme; means for providing a user interface for confirming and correcting the generated speech content; means for voice-synthesizing the generated speech content; means for synchronizing a voice file with lip sync and body language based on the confirmed speech content; means for generating a motion sequence corresponding to the voice file; and means for automatically generating subtitles based on the speech content and integrating them into video. This enables users to quickly and efficiently create high-quality virtual characters without requiring special technical knowledge.
[0910] "User" refers to any individual or organization that uses the Platform to create and customize virtual characters.
[0911] "Server" refers to a computer system that performs processes such as generating virtual characters, generating speech content, synthesizing voice, and integrating video and subtitles.
[0912] "Avatar" refers to a 3D or Live2D virtual character generated based on user input.
[0913] "3D avatar" refers to a virtual character that is visually represented in three-dimensional space.
[0914] A "Live2D avatar" is a virtual character that dynamically expresses two-dimensional image data and appears three-dimensional.
[0915] A "generative AI model" refers to an artificial intelligence model that performs natural language processing based on topic information and text information from users.
[0916] "Comment content" refers to the text data generated by the generative AI model based on the topic specified by the user.
[0917] "Speech synthesis" refers to the technology of generating a voice file based on the text data of a speech.
[0918] "Lip sync" refers to the technology of synchronizing the mouth movements of a virtual character with an audio file.
[0919] "Body language" refers to the technology that generates the physical movements and gestures of a virtual character in accordance with the voice and content of what is being said.
[0920] "Motion sequence" refers to a series of movement data including lip sync and body language.
[0921] "Subtitles" refers to text data generated based on what is being said and displayed to viewers within the video.
[0922] "Customization options" refers to settings that a user can select to adjust the appearance and behavior of their avatar.
[0923] "Administration tool" refers to software that provides an interface for users to request modifications to the videos they generate.
[0924] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. This system enables anyone to generate high-quality virtual characters and create related content without requiring specialized skills or programming knowledge.
[0925] System configuration
[0926] The system includes the following main components:
[0927] 1. Avatar generation
[0928] The server receives information input by the user (e.g., gender, age, style, etc.) and activates an image generation AI module (e.g., Stable Diffusion, DALL-E 2) to generate a 3D or Live2D avatar. The generated avatar is then sent to the device.
[0929] On the terminal, the user customizes the details of the avatar (e.g., hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[0930] 2. Generating speech content
[0931] Users enter topic information into a form on their device (e.g., "Introducing AITubers"). The server uses a generative AI model (e.g., GPT-3, ChatGPT) to generate a post based on the topic. The generated post is sent to the device for the user to review.
[0932] The user confirms and modifies the generated message and sends the finalized message to the server.
[0933] 3. Speech synthesis
[0934] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly), and sends the generated audio file to the device.
[0935] The user checks the audio file and requests corrections if necessary.
[0936] 4. Audio and Motion Mapping
[0937] The server receives the audio file and avatar data, creates a motion sequence to generate lip sync and body language (e.g., Live2D Cubism, Adobe Character Animator), and sends the generated motion image to the device.
[0938] The user checks the motion image and requests corrections if necessary.
[0939] 5. Creating subtitles
[0940] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech. The subtitle file is integrated into the video, and the completed video is sent to the terminal.
[0941] The user checks the video and subtitles and confirms them after a final check.
[0942] Specific examples
[0943] Example 1: When a user creates an introductory video for an AITuber
[0944] 1. Create an avatar
[0945] The user clicks the "Create Avatar" button on the device's UI and enters the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[0946] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[0947] 2. Generating speech content
[0948] The user enters the topic "Introducing AITubers" into a form on their device. The server uses a generative AI model to generate comments based on the topic and sends them to the device.
[0949] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[0950] 3. Speech Synthesis
[0951] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[0952] The user checks the audio file and confirms it if there are no particular problems.
[0953] 4. Audio and Motion Mapping
[0954] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[0955] The user checks the motion image and confirms it if there are no particular problems.
[0956] 5. Creating subtitles
[0957] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[0958] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[0959] Prompt Sentence Examples
[0960] Generate a 3D character based on the following information: Gender is female, age is in her 20s, and she wears casual clothing. For background information, the character is an AITuber on YouTube. The statement should read, "Hello, today I'd like to introduce my channel."
[0961] The above system allows users to quickly and easily generate virtual characters and create compelling content without any special technical knowledge.
[0962] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0963] Step 1:
[0964] The user clicks the "Create Avatar" button on the device's UI and enters basic status information (e.g., gender, age, style, etc.).
[0965] Input: The user fills in a form with basic information such as gender, age, and style.
[0966] Action: Sends the entered information to the server.
[0967] Output: Basic status information is entered into the server.
[0968] Step 2:
[0969] The server launches an image generation AI module (e.g., Stable Diffusion, DALL-E 2) based on the received basic status information and generates a 3D or Live2D avatar.
[0970] Input: Basic status information.
[0971] How it works: Passes information to an image generation AI module, which applies an algorithm to generate an avatar.
[0972] Output: An initial avatar image is generated and sent to the device as a preview.
[0973] Step 3:
[0974] On the device, the user receives a preview of the generated avatar and can then perform detailed customization (e.g., hairstyle, clothing, etc.).
[0975] Input: Previewed avatar image and customization options.
[0976] How it works: Use the customization UI to adjust the details using sliders and dropdown menus.
[0977] Output: Generates the finalized avatar configuration.
[0978] Step 4:
[0979] When the user confirms the avatar, the data of the confirmed avatar is transmitted to the server.
[0980] Input: Avatar settings adjusted by the user.
[0981] Action: Press the Confirm button to send the avatar configuration data to the server.
[0982] Output: The confirmed avatar configuration data is sent to the server.
[0983] Step 5:
[0984] Users enter topic information into a form on their device (e.g., "Introducing AITubers").
[0985] Input: Topic information.
[0986] Action: Enter a topic in the text area and press the send button.
[0987] Output: The topic information is sent to the server.
[0988] Step 6:
[0989] The server passes the received topic information to a generative AI model (e.g., GPT-3, ChatGPT) to generate the content of the comments.
[0990] Input: Topic information.
[0991] How it works: Provide topic information as prompts to a generative AI model and receive the generated utterances.
[0992] Output: The generated speech is sent to the terminal.
[0993] Step 7:
[0994] The user checks the generated comment and corrects it if necessary.
[0995] Input: The generated utterance.
[0996] Action: Edit the comment in the editor and press the send button after editing.
[0997] Output: The corrected statement is sent to the server.
[0998] Step 8:
[0999] The user confirms the content of the statement and sends it to the server.
[1000] Input: Confirmed statement.
[1001] Action: Press the Confirm button to send the final revised statement to the server.
[1002] Output: The confirmed statement is sent to the server.
[1003] Step 9:
[1004] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly).
[1005] Input: Confirmed statement.
[1006] What it does: Passes what is said to a speech synthesis engine, generating an audio file.
[1007] Output: An audio file is generated and sent to the device.
[1008] Step 10:
[1009] The user checks the generated audio file and requests corrections if necessary.
[1010] Input: An audio file.
[1011] Action: Plays an audio file and sends feedback for corrections if needed.
[1012] Output: Finalized audio file.
[1013] Step 11:
[1014] The server receives the audio file and avatar data and generates lip sync and body language.
[1015] Input: Audio file and avatar data.
[1016] Movement: Generate avatar mouth and body movements in sync with the audio to create a motion sequence.
[1017] Output: A motion image is generated and sent to the device.
[1018] Step 12:
[1019] The user checks the generated motion image and requests corrections if necessary.
[1020] Input: Motion footage.
[1021] How it works: Plays motion footage and provides feedback for corrections if there are any issues.
[1022] Output: Finalized motion footage.
[1023] Step 13:
[1024] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech.
[1025] Input: The text of what was said.
[1026] Operation: Converts text data into subtitle file format according to the time axis.
[1027] Output: Subtitle files are generated and integrated into the video.
[1028] Step 14:
[1029] The user can check the automatically generated subtitled video and request corrections if necessary.
[1030] Input: Video with subtitles.
[1031] Action: Play the video to check the timing and content of the subtitles, and send correction requests if necessary.
[1032] Output: Finalized video with subtitles.
[1033] The above are the specific steps of the system's program processing, which allows users to efficiently create virtual characters and generate related content.
[1034] (Application example 1)
[1035] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1036] Previous virtual character generation platforms required a lot of effort for users to create characters, making them inconvenient, especially when using smartphones. Furthermore, they lacked assistant functions for real-time user interaction and product introductions, making them difficult to apply to virtual stores. Furthermore, they lacked management tools that could flexibly accommodate user requests for corrections to the videos they generated.
[1037] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1038] In this invention, the server includes means for generating a 3D or Live2D avatar based on information input by a user, means for generating speech content based on the input theme, means for synthesizing the generated speech content with speech, means for synchronizing the speech file with the avatar's movements, means for automatically generating subtitles based on the speech content and integrating them into video, means for using the avatar as a virtual assistant character on a smartphone, and means for the virtual assistant character to introduce products and respond to user inquiries in real time. This allows users to easily create a virtual assistant character on their smartphone and realize real-time conversations, product introductions, and customer support.
[1039] "User-inputted information" refers to basic status information and thematic input data that a user provides to the system.
[1040] A "3D or Live2D avatar" is a three-dimensional or two-dimensional character generated based on user input.
[1041] "Comment content" is text information generated based on the theme entered by the user.
[1042] "Speech synthesis" is the process of creating an audio file based on generated utterances.
[1043] "Synchronizing the audio file with the avatar's movements" is the process of moving the character's mouth and body in sync with the generated audio.
[1044] "Automatic subtitling" is the process of generating text based on what is being said and integrating it with the video.
[1045] "Use on a smartphone" refers to using a virtual assistant character on a smartphone application.
[1046] A "virtual assistant character" is a digital character created by a user that introduces products and responds to user inquiries in real time.
[1047] The "product introduction function" is a function in which a virtual assistant character explains specific product information to the user.
[1048] "Real-time response" refers to the ability to respond immediately to user questions and requests.
[1049] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Furthermore, this system can be used as a virtual assistant character on smartphones, and has the ability to introduce products and respond to users in real time.
[1050] System configuration
[1051] 1. 3D / Live2D avatar generation
[1052] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[1053] On the device (smartphone), the user customizes the avatar in detail (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[1054] 2. Generating speech content
[1055] The server passes the topic information entered by the user to the generative AI model to generate the utterance content, which is then sent to the device for the user to confirm.
[1056] The user confirms and modifies the generated message and sends the finalized message to the server.
[1057] 3. Speech synthesis
[1058] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[1059] The user checks the audio file and requests corrections if necessary.
[1060] 4. Audio and Motion Mapping
[1061] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[1062] The user checks the motion image and requests corrections if necessary.
[1063] 5. Creating subtitles
[1064] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file into the video, and sends the completed video to the device.
[1065] The user checks the video and subtitles and confirms them after a final check.
[1066] 6. Use as a virtual assistant
[1067] The virtual assistant character generated on the device (smartphone) introduces products and responds to user inquiries in real time, realizing an interactive dialogue with the user.
[1068] Users obtain specific product information or inquiries through the virtual assistant.
[1069] Hardware and software used
[1070] Hardware: Smartphones, servers
[1071] Software: OpenAI API (generative AI model), pyttsx3 (speech synthesis), OpenCV (video and motion processing)
[1072] Specific examples
[1073] Example 1: A user creates a virtual assistant character to introduce a new product, and the character explains the company's product through a smartphone app.
[1074] Example prompt sentence:
[1075] "Please explain the features of your new product, including its target market and its benefits."
[1076] Sample output: "This new product is a casual-style smartwatch designed for young people. Key features include lightweight design, long battery life, and the latest fitness tracking capabilities. It's especially suitable for consumers who lead active lifestyles."
[1077] This system allows users to easily create and use virtual assistant characters on their smartphones, allowing them to introduce products and respond in real time.
[1078] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1079] Step 1:
[1080] Obtaining input information
[1081] Users use a smartphone app to input basic status information (gender, age, style, etc.) to create an avatar.
[1082] The server receives information entered by the user and launches the image generation AI module.
[1083] Output: Basic status information for the user.
[1084] Step 2:
[1085] Avatar generation
[1086] The server generates a 3D or Live2D avatar based on the user's basic status information received in step 1.
[1087] The server sends a preview of the generated avatar to the terminal.
[1088] Output: The generated avatar and its preview image.
[1089] Step 3:
[1090] Avatar Customization
[1091] Users customize their avatar's details (hairstyle, clothing, etc.) using a smartphone app.
[1092] The terminal sends the customized results to the server.
[1093] Output: Customized avatar information.
[1094] Step 4:
[1095] Generate speech content
[1096] The user uses the app to input the topic of the comment (e.g., new product introduction).
[1097] The server generates speech content based on the input topic using a generative AI model (OpenAI API).
[1098] The server sends the generated message to the terminal, where the user can confirm and edit it.
[1099] Output: The generated utterance.
[1100] Step 5:
[1101] Confirmation and confirmation of statements
[1102] The user checks the generated comment and corrects it if necessary.
[1103] The user sends the finalized content of the message to the server.
[1104] Output: Confirmed statement.
[1105] Step 6:
[1106] Speech synthesis
[1107] The server passes the confirmed utterance content to a natural language processing engine (pyttsx3) and performs speech synthesis.
[1108] The server transmits the generated audio file to the terminal.
[1109] Output: The generated audio file.
[1110] Step 7:
[1111] Audio and Motion Mapping
[1112] The server receives the audio files and avatar data and creates motion sequences to generate lip sync and body language.
[1113] The server transmits the generated motion image to the terminal.
[1114] Output: The generated motion footage.
[1115] Step 8:
[1116] Subtitle creation and integration
[1117] The server automatically generates a subtitle file based on the text of the confirmed speech content.
[1118] The server integrates the subtitle file into the video and sends the final video to the terminal.
[1119] Output: A merged video file.
[1120] Step 9:
[1121] Use as a virtual assistant
[1122] The device (smartphone) uses the generated virtual assistant character to introduce products and respond to user inquiries in real time.
[1123] Users can use the virtual assistant to obtain product information and make inquiries.
[1124] Output: Product information and support provided to the user.
[1125] The above is a detailed flow of the processing steps of this system. By performing these operations, users can easily create a virtual assistant character and introduce products and respond in real time on their smartphones.
[1126] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1127] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. By combining it with an emotion engine, it also has the ability to dynamically adjust the avatar's expressions based on the user's emotions. This system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Additionally, by incorporating an emotion engine that recognizes the user's emotions, it dynamically changes facial expressions and movements, and adjusts the tone and pitch of the voice, based on the user's emotions.
[1128] System configuration
[1129] 1. 3D / Live2D avatar generation
[1130] The server receives basic status information (gender, age, style, etc.) entered by the user and activates an image generation AI module to generate a 3D or Live2D avatar.
[1131] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[1132] 2. Generating speech content
[1133] The server passes the topic information entered by the user to the LLM to generate the content of the comment, which is then sent to the terminal.
[1134] The user confirms and modifies the generated message and sends the finalized message to the server.
[1135] 3. Speech synthesis
[1136] The server synthesizes the confirmed speech content through a natural language processing engine and transmits the generated audio file to the terminal.
[1137] The user checks the audio file and requests corrections if necessary.
[1138] 4. Audio and Motion Mapping
[1139] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and then transmits the generated motion image to the device.
[1140] The user checks the motion image and requests corrections if necessary.
[1141] 5. Creating subtitles
[1142] The server automatically generates a subtitle file based on the text of the confirmed remarks, integrates it with the video, and sends it to the terminal.
[1143] The user checks the video and subtitles and confirms them after a final check.
[1144] 6. Leveraging Emotional Engines
[1145] The device uses an emotion engine to analyze the user's facial expressions, tone of voice, and gestures in real time to recognize the user's emotions.
[1146] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[1147] The server adjusts the tone and pitch of the generated speech synthesis according to the user's emotions.
[1148] Specific examples
[1149] Example 1: A user creates a video to express their emotions
[1150] 1. Create an avatar
[1151] The user clicks the "Create Avatar" button and enters the avatar's basic status (e.g., male, young man, casual style, etc.). The server generates an avatar based on the entered information and sends a preview to the device.
[1152] The user can view the preview, customize the details and finally confirm the avatar.
[1153] 2. Generating speech content
[1154] The user inputs the topic "Emotional Expression" as the theme. The server uses LLM to generate utterances based on the theme and sends them to the device.
[1155] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[1156] 3. Speech Synthesis
[1157] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[1158] The user checks the audio and confirms it if there is no particular problem.
[1159] 4. Audio and Motion Mapping
[1160] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[1161] The user checks the motion image and confirms it if there are no particular problems.
[1162] 5. Creating subtitles
[1163] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[1164] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[1165] 6. Leveraging Emotional Engines
[1166] The device senses the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine.
[1167] The server dynamically adjusts lip sync, facial expressions, and body language based on the obtained emotional data, and appropriately changes the tone and pitch of the voice.
[1168] The user checks the video of the generated emotional expression, makes final adjustments and confirms it.
[1169] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system allows users to easily generate high-quality virtual character content that reflects their own emotions, without requiring programming knowledge.
[1170] The processing flow will be explained below.
[1171] 1. Creating a 3D / Live2D avatar
[1172] Step 1:
[1173] The user clicks the "Create Avatar" button on the device.
[1174] Step 2:
[1175] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[1176] Step 3:
[1177] The user enters the required basic status information and clicks the "Generate" button.
[1178] Step 4:
[1179] The server receives the information entered by the user.
[1180] Step 5:
[1181] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[1182] Step 6:
[1183] The server sends a preview image of the generated avatar to the terminal.
[1184] Step 7:
[1185] The user checks the preview image displayed on the terminal.
[1186] Step 8:
[1187] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button again.
[1188] Step 9:
[1189] The terminal transmits the final avatar information to the server.
[1190] 2. Generating speech content
[1191] Step 1:
[1192] The user clicks the "Generate Message" button on the terminal.
[1193] Step 2:
[1194] The terminal displays a form for the user to input the topic or theme of the comment.
[1195] Step 3:
[1196] The user enters a topic or theme and clicks the "Generate" button.
[1197] Step 4:
[1198] The server receives topic information input by the user.
[1199] Step 5:
[1200] The server launches the LLM and generates comments based on the input topic.
[1201] Step 6:
[1202] The server transmits the generated speech content to the terminal.
[1203] Step 7:
[1204] The user checks the message displayed on the terminal.
[1205] Step 8:
[1206] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[1207] Step 9:
[1208] The terminal transmits the confirmed statement content to the server.
[1209] 3. Speech synthesis
[1210] Step 1:
[1211] The user sends a request for speech synthesis based on the confirmed utterance content.
[1212] Step 2:
[1213] The server receives the confirmed statement content.
[1214] Step 3:
[1215] The server performs speech synthesis through a natural language processing engine.
[1216] Step 4:
[1217] The server transmits the generated audio file to the terminal.
[1218] Step 5:
[1219] The user plays the audio file on the device and checks the content.
[1220] Step 6:
[1221] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[1222] 4. Audio and Motion Mapping
[1223] Step 1:
[1224] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[1225] Step 2:
[1226] The server receives the audio file and avatar data.
[1227] Step 3:
[1228] The server generates motion sequences to achieve natural lip sync and body language.
[1229] Step 4:
[1230] The server transmits the completed motion image to the terminal.
[1231] Step 5:
[1232] The user plays and checks the generated motion video on the terminal.
[1233] Step 6:
[1234] The user requests corrections to movements and facial expressions as needed.
[1235] Step 7:
[1236] The user saves the finalized motion image.
[1237] 5. Creating subtitles
[1238] Step 1:
[1239] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[1240] Step 2:
[1241] The server receives the text of the confirmed utterance.
[1242] Step 3:
[1243] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[1244] Step 4:
[1245] The server overlays and integrates the subtitle file onto the motion video.
[1246] Step 5:
[1247] The server sends the completed video to the terminal.
[1248] Step 6:
[1249] The user checks the video and subtitles on the device.
[1250] Step 7:
[1251] The user saves the finalized video and publishes or shares it.
[1252] 6. Leveraging Emotional Engines
[1253] Step 1:
[1254] When a user plays back the video they created, the device's camera and microphone are used to detect emotional data such as facial expressions and tone of voice in real time.
[1255] Step 2:
[1256] The terminal transmits the emotion data acquired in real time to the emotion engine, and transmits the analysis results to the server.
[1257] Step 3:
[1258] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[1259] Step 4:
[1260] When the server synthesizes voice from the generated speech, it adjusts the tone and pitch of the voice according to the user's emotions.
[1261] Step 5:
[1262] The server generates the final video based on the emotion data and sends it to the terminal.
[1263] Step 6:
[1264] The user checks the video in which the emotional expression is reflected and requests corrections as necessary.
[1265] Step 7:
[1266] The user saves the finalized video and publishes or shares it.
[1267] The above are the specific processing steps in the program processing of a system that utilizes an emotion engine.
[1268] Example 2
[1269] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1270] Conventional virtual character creation platforms make it difficult for users to express emotions or easily generate speech content. Synchronizing video and audio and generating subtitles is also time-consuming, making the overall process complicated and making it difficult to easily create high-quality content.
[1271] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a 3D or 2D avatar based on information input by a user, a means for generating speech content based on an input theme, a means for voice-synthesizing the generated speech content, a means for synchronizing the speech file with the avatar's movement, a means for automatically generating subtitles based on the speech content and integrating them into video, and a means for collecting user emotion data using an emotion analysis engine and dynamically adjusting the avatar's expression. This allows a user to easily create high-quality virtual character content and generate video including real-time expressions that reflect the user's emotions.
[1272] "User" refers to the entity that uses this system to create a virtual character, customize it, generate comments, etc.
[1273] "Server" refers to a central processing unit that receives input information from a user and performs various processes.
[1274] "3D or 2D Avatar" means a virtual character created by a user and represented in three-dimensional or two-dimensional graphics.
[1275] "Input information" refers to data that a user provides to the system, such as basic status information such as gender, age, and style.
[1276] A "theme" refers to a topic that a user provides to the system for generating comment content.
[1277] "Speech content" refers to the lines and phrases of the character generated based on the input theme.
[1278] "Speech synthesis" is a technology that converts text data into voice data, and refers to a means of outputting the generated utterances as voice.
[1279] "Synchronizing" refers to the process of matching the audio and character movements.
[1280] "Subtitles" are texts that are automatically generated based on what is being said and are integrated into video.
[1281] "Integrating into video" refers to the process of incorporating the generated subtitles into a single video file in conjunction with the audio and avatar movements.
[1282] An "emotion analysis engine" refers to software or a system that analyzes a user's facial expressions, tone of voice, and gestures to recognize emotions.
[1283] "Emotion data" refers to data that indicates the user's emotional state as recognized by the emotion analysis engine.
[1284] "Dynamic adjustment" refers to the act of changing an avatar's facial expressions, movements, voice tone, pitch, etc. in real time.
[1285] "Customization options" refers to settings that allow users to change details about their avatar.
[1286] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters, and by combining it with an emotion analysis engine, it has the ability to dynamically adjust the avatar's expression according to the user's emotions.
[1287] System configuration
[1288] The system consists of the following main components:
[1289] 1. Server: A central processing unit that receives information from users and performs various processes.
[1290] 2. Terminal: A device operated by a user that provides an interface for user input, display, customization, etc.
[1291] 3. Sentiment analysis engine: Software for analyzing user emotions and generating data.
[1292] Technology used
[1293] 1. Image generation AI module (e.g. DeepArt): Generates a 3D or 2D avatar based on basic status information (gender, age, style, etc.) entered by the user.
[1294] 2. Large-scale language models (LLMs) (e.g., GPT-4): Generate utterances based on themes entered by the user.
[1295] 3. Natural language processing engine (e.g., Google Text-to-Speech): Converts the generated speech into audio.
[1296] 4. Emotion analysis engine (e.g., Affectiva SDK): Analyzes the user's facial expressions, tone of voice, and gestures to collect emotional data.
[1297] System Operation
[1298] The system operates in the following steps:
[1299] 1. Create your avatar:
[1300] The user clicks the "Create Avatar" button on the device and enters basic status information (gender, age, style, etc.).
[1301] The terminal transmits this information to the server.
[1302] The server uses an image generation AI module to generate a 3D or 2D avatar and sends a preview to the device.
[1303] The user checks the preview image, performs detailed customization (hairstyle, clothing, etc.), and confirms it.
[1304] The terminal transmits the determined avatar information to the server.
[1305] 2. Speech generation:
[1306] The user inputs topic information and sends it to the server.
[1307] The server uses LLM to generate speech content and sends it to the terminal.
[1308] The user confirms and corrects the content of the comment and sends the finalized content to the server.
[1309] 3. Speech synthesis:
[1310] The server synthesizes the confirmed speech content through a natural language processing engine.
[1311] The server transmits the generated audio file to the terminal.
[1312] The user checks the audio file and requests corrections if necessary.
[1313] 4. Audio and Motion Mapping:
[1314] The server passes the audio file and avatar data to a lip sync engine to generate lip sync and body language.
[1315] The server transmits the generated motion sequence to the terminal.
[1316] The user checks the motion image and requests corrections if necessary.
[1317] 5. Creating subtitles:
[1318] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[1319] The user checks the video and subtitles and confirms them after a final check.
[1320] 6. Utilizing sentiment analysis engines:
[1321] The device detects the user's facial expressions and tone of voice and passes them to an emotion analysis engine in real time.
[1322] The terminal transmits the emotion analysis results obtained to the server.
[1323] The server dynamically adjusts the avatar's facial expressions and movements based on the emotional data, and also appropriately changes the tone and pitch of the voice.
[1324] The user checks the final image, makes adjustments and confirms them.
[1325] Specific examples
[1326] Example 1: A user creates a video to express their emotions
[1327] 1. Create your avatar:
[1328] The user clicks the "Create Avatar" button and inputs the avatar's basic status (e.g., male, young man, casual style, etc.).
[1329] The terminal displays an input form and accepts input from the user.
[1330] When the user inputs information and presses the send button, the terminal sends the information to the server.
[1331] The server receives the information, activates the image generation AI module, and generates the avatar.
[1332] The server sends the generated avatar to the device as a preview.
[1333] The device displays a preview, and the user can customize the details (hairstyle, clothing, etc.) and finally confirm the avatar.
[1334] The terminal transmits the determined avatar information to the server.
[1335] 2. Speech generation:
[1336] The user opens the "Comment Content" input form and inputs the topic "Emotional Expressions."
[1337] The terminal transmits the topic information to the server.
[1338] The server generates the message using LLM and sends it to the terminal.
[1339] The user confirms and modifies the generated message and sends the finalized message to the server.
[1340] 3. Speech synthesis:
[1341] The server uses a natural language processing engine to synthesize speech based on the confirmed content of the speech.
[1342] The server transmits the generated audio file to the terminal.
[1343] The user checks the audio file and confirms it if there are no problems.
[1344] 4. Audio and Motion Mapping:
[1345] The server generates lip sync and body language based on the audio file and the determined avatar.
[1346] The server transmits the motion video to the terminal and provides it to the user.
[1347] The user checks the motion image and requests corrections if necessary.
[1348] 5. Creating subtitles:
[1349] The server automatically generates subtitles based on the confirmed remarks, integrates them into the video, and sends them to the terminal.
[1350] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[1351] 6. Utilizing sentiment analysis engines:
[1352] The device senses the user's facial expressions and tone of voice in real time and sends the analysis results to the emotion engine.
[1353] The terminal transmits the emotion analysis result to the server.
[1354] The server dynamically adjusts lip sync, facial expressions, and movements based on emotional data, and also appropriately changes the tone and pitch of the voice.
[1355] The user checks the final image, makes adjustments and confirms them.
[1356] By following the above steps, users can easily create high-quality virtual character content. This system also enables real-time expression according to emotions, making it possible to generate a wide variety of content.
[1357] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1358] Step 1: User begins creating an avatar
[1359] The user clicks the "Create Avatar" button on the device.
[1360] Input: User input (basic status information such as gender, age, style, etc.)
[1361] The terminal displays an input form and accepts input from the user.
[1362] How it works: The user enters the required information and presses the submit button.
[1363] Output: The information entered
[1364] Step 2: The server generates the avatar
[1365] The server launches an image generation AI module based on the user's basic status information received from the device.
[1366] Input: Basic status information entered
[1367] How it works: The server uses an image generation AI module to generate a 3D or 2D avatar.
[1368] Output: A preview image of the generated avatar
[1369] Step 3: User customizes avatar
[1370] The terminal displays a preview of the generated avatar to the user.
[1371] Input: A preview image of the generated avatar
[1372] How it works: The user customizes the avatar's details (hairstyle, clothing, etc.) and finally finalizes the avatar.
[1373] Output: A customized confirmed avatar
[1374] Step 4: Enter and submit topic information
[1375] The user opens the "Comment Generation" input form on the terminal and inputs a topic (e.g., "About emotional expression").
[1376] Input: Topic information
[1377] How it works: The user enters a topic and presses the send button.
[1378] Output: Input topic information
[1379] Step 5: Generate speech
[1380] The server receives topic information and generates utterances using a large-scale language model (LLM).
[1381] Input: Topic information
[1382] How it works: The server passes the prompt to the generative AI model to generate the speech.
[1383] Output: Generated speech
[1384] Step 6: Revise and confirm your statement
[1385] The terminal displays the generated comment content to the user.
[1386] Input: Generated speech
[1387] Action: The user confirms and corrects what has been said, and finally confirms it.
[1388] Output: Confirmed statement
[1389] Step 7: Speech synthesis of what is being said
[1390] The server passes the confirmed utterance content to a natural language processing engine (e.g., Google Text-to-Speech) for speech synthesis.
[1391] Input: Confirmed statement
[1392] How it works: The server uses a speech synthesis engine to generate an audio file.
[1393] Output: Generated audio file
[1394] Step 8: Check and correct the audio file
[1395] The terminal provides the generated audio file to the user.
[1396] Input: Generated audio file
[1397] Action: The user reviews the audio file and requests corrections if necessary.
[1398] Output: Finalized audio file
[1399] Step 9: Mapping Audio and Motion
[1400] The server passes the audio file and the determined avatar to a lip sync engine to generate lip sync and body language.
[1401] Input: Audio file, confirmed avatar
[1402] How it works: The server generates lip sync and motion data.
[1403] Output: Generated motion sequence
[1404] Step 10: Check and correct motion footage
[1405] The terminal displays the generated motion image to the user.
[1406] Input: Generated motion sequence
[1407] Action: The user reviews the motion footage and requests corrections if necessary.
[1408] Output: Confirmed motion footage
[1409] Step 11: Creating and merging subtitles
[1410] The server automatically generates subtitles based on the confirmed remarks and integrates them into the video.
[1411] Input: Confirmed speech content, confirmed motion image
[1412] How it works: The server generates a subtitle file and integrates it into the video.
[1413] Output: Integrated video (audio, motion, subtitles)
[1414] Step 12: Reflecting the results of sentiment analysis
[1415] The device passes the user's facial expressions and tone of voice to an emotion analysis engine.
[1416] Input: Real-time facial expressions, tone of voice
[1417] Operation: The device obtains the emotion analysis results and sends them to the server.
[1418] Output: Emotion data
[1419] Step 13: Adjust based on sentiment data
[1420] The server dynamically adjusts the avatar's facial expressions and movements based on the emotional data, and changes the tone and pitch of the voice.
[1421] Input: Emotion data, confirmed motion video, audio file
[1422] Movement: The server adjusts the avatar's facial expressions, movements, and voice tone and pitch.
[1423] Output: Adjusted video and audio
[1424] Step 14: Final review and confirmation
[1425] The terminal provides the final video to the user for confirmation.
[1426] Input: Adjusted video and audio
[1427] Action: The user performs a final review, makes any necessary adjustments, and confirms.
[1428] Output: Final video
[1429] (Application example 2)
[1430] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1431] Currently, it is difficult for users to easily create interactive, emotionally expressive virtual characters. It is also difficult to dynamically adjust the virtual character's facial expressions and voice tone in response to the user's real-time emotions. This can result in a lack of a natural and engaging experience for the user, leading to lower user satisfaction. The present invention aims to solve these problems and provide a system that enables a virtual character to express natural emotions in response to the user's emotions.
[1432] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating a 3D or Live2D avatar based on information input by the user, means for generating speech content based on the input theme, means for voice-synthesizing the generated speech content, means for synchronizing the audio file with the avatar's movements, means for automatically generating subtitles based on the speech content and integrating them into video, means for sensing the user's voice and facial expressions in real time and analyzing emotions using an emotion engine, and means for dynamically adjusting the avatar's facial expressions and voice tone based on the emotion data. This allows users to easily create virtual characters and enjoy natural expressions that are dynamically adjusted according to emotions.
[1433] "3D or Live2D Avatar" means a virtual character that is represented in three or two dimensions based on a user-selected appearance and style.
[1434] "Means for generating utterances" refers to a system or algorithm that generates appropriate responses or utterances based on themes or questions entered by users.
[1435] "Means for speech synthesis" refers to the technology for generating speech from text, and the function for converting the generated speech into natural-sounding speech.
[1436] "Means for synchronizing audio files with avatar movements" refers to a technology that synchronizes the mouth movements and facial expressions of an avatar with the generated audio.
[1437] "Means for automatically generating subtitles and integrating them into video" refers to technology that automatically generates subtitles based on what is being said and embeds them into the video.
[1438] The "emotion engine" is a system that analyzes the user's tone of voice and facial expressions to recognize the user's emotional state in real time.
[1439] "Means for dynamically adjusting the avatar's facial expression and voice tone based on emotional data" is a function that appropriately changes the avatar's facial expression and voice tone based on the results of analysis by the emotion engine.
[1440] This invention provides a platform that allows users to easily create 3D or Live2D avatars, which then dynamically change their expressions based on the user's emotions. Based on user input, this system generates and adjusts real-time avatar movements that integrate speech, audio, subtitles, and emotional expressions.
[1441] System configuration
[1442] 1. Avatar generation
[1443] The server receives basic information entered by the user (such as gender, age, and style) and generates a 3D or Live2D avatar using an image generation AI module.
[1444] The terminal provides an interface that allows the user to customize the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[1445] 2. Generating speech content
[1446] The server uses a generative AI model (e.g., GPT-4) to generate a message based on the themes entered by the user. This message is then sent to the device, where it is reviewed and revised by the user before being resent.
[1447] 3. Speech Synthesis
[1448] The server generates audio based on the confirmed utterance using a natural language processing engine (e.g., Google Cloud Text-to-Speech) and sends the generated audio file to the device.
[1449] 4. Audio and Motion Synchronization
[1450] The server creates a motion sequence to generate lip sync and body language based on the audio file and avatar data, and sends it to the device.
[1451] 5. Creating subtitles
[1452] The server automatically generates a subtitle file based on the text of the confirmed speech and integrates it into the video, which is then sent to the device.
[1453] 6. Leveraging Emotional Engines
[1454] The device detects the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine (e.g., Microsoft Azure Emotion API).
[1455] The server dynamically adjusts the avatar's facial expressions, movements, and voice tone based on the emotional data obtained from the emotion engine.
[1456] Specific examples
[1457] Example 1: Virtual Shopping Assistant
[1458] The user puts on the smart glasses and says, "Start shopping assistant." The device sends the command to the server.
[1459] The server generates a 3D avatar based on basic information provided by the user and sends it to the device. The user then customizes the avatar in detail (e.g., glasses, hairstyle) and finalizes the avatar.
[1460] The user asks, "What material is this shirt made of?" The server passes the question to a generative AI model (e.g., GPT-4), which generates an appropriate response and performs speech synthesis.
[1461] The device detects the user's facial expressions and tone of voice, analyzes them using an emotion engine, and displays and plays responses that dynamically adjust the avatar's facial expressions and tone of voice according to the emotional data.
[1462] Prompt Sentence Examples
[1463] User: What material is this shirt made of?
[1464] Virtual Character: Okay. This shirt is made of 100% cotton. It feels amazing!
[1465] This system allows users to easily generate high-quality virtual character content that reflects their own emotions, without requiring any programming knowledge.
[1466] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1467] Step 1:
[1468] Entering and creating basic avatar information
[1469] The user inputs basic information (gender, age, style, etc.) from the terminal, which is then sent to the server.
[1470] The server generates a 3D or Live2D avatar using an image generation AI module based on the received basic information, and a preview image of the generated avatar is sent to the device.
[1471] Output: A preview image of the avatar generated based on the basic information.
[1472] Step 2:
[1473] Detailed avatar customization
[1474] The user customizes the avatar in detail (hairstyle, clothing, etc.) from the terminal and finally confirms the avatar. The detailed customization data is sent to the server.
[1475] The server stores the determined avatar data, generates the final avatar image, and transmits it to the terminal.
[1476] Output: Final user-customized avatar image and data.
[1477] Step 3:
[1478] Generate speech content
[1479] The user inputs a specific topic or question (everyday conversation, product description, etc.) into the terminal and sends it to the server.
[1480] The server uses a generative AI model (e.g., GPT-4) to generate utterances based on the themes entered by the user. The generated utterances are then sent to the device.
[1481] Output: The speech text generated by the generative AI model.
[1482] Step 4:
[1483] Speech synthesis
[1484] The user checks and corrects the generated text and sends the finalized text to the server.
[1485] The server uses a natural language processing engine (e.g., Google Cloud Text-to-Speech) to convert what is said into audio, and the resulting audio file is sent to the device.
[1486] Output: Synthesized audio file.
[1487] Step 5:
[1488] Audio and motion synchronization
[1489] The server creates a motion sequence to generate lip sync and body language based on the audio file and avatar data, and sends it to the device.
[1490] The device will synchronize the motion sequence with the audio file and display a preview.
[1491] Output: Animated sequence with synchronized lip sync and body language.
[1492] Step 6:
[1493] Creating subtitles
[1494] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates it into the video, and sends the video file to the device.
[1495] The device will display a preview of the video and subtitles for the user to review.
[1496] Output: Video file with integrated subtitles.
[1497] Step 7:
[1498] Emotion analysis
[1499] The device senses the user's facial expressions and tone of voice in real time and transmits emotional data to the server.
[1500] The server analyzes the emotion data using an emotion engine (e.g., Microsoft Azure Emotion API).
[1501] Output: Parsed emotion data.
[1502] Step 8:
[1503] Dynamic adjustment based on emotions
[1504] The server generates sequences to dynamically adjust lip sync, facial expressions, and body language based on the emotional data and sends them to the device.
[1505] The device displays the adjusted animation and the user gives a final confirmation.
[1506] Output: Dynamically adjusted animation depending on the emotion.
[1507] Examples and prompts
[1508] Prompt Sentence Examples
[1509] User: What material is this shirt made of?
[1510] Virtual Character: Okay. This shirt is made of 100% cotton. It feels amazing!
[1511] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1512] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1513] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1514] [Third embodiment]
[1515] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1516] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1517] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1518] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1519] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1520] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1521] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1522] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1523] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1524] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1525] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1526] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1527] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech with motion, and generates and integrates subtitles into the video.
[1528] System configuration
[1529] 1. 3D / Live2D avatar generation
[1530] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[1531] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[1532] 2. Generating speech content
[1533] The server passes the topic information entered by the user to the LLM, generates a message, and sends the generated message to the terminal for the user to confirm.
[1534] The user confirms and modifies the generated message and sends the finalized message to the server.
[1535] 3. Speech synthesis
[1536] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[1537] The user checks the audio file and requests corrections if necessary.
[1538] 4. Audio and Motion Mapping
[1539] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[1540] The user checks the motion image and requests corrections if necessary.
[1541] 5. Creating subtitles
[1542] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file with the video, and sends the completed video to the device.
[1543] The user checks the video and subtitles and confirms them after a final check.
[1544] Specific examples
[1545] Example 1: When a user creates an introductory video for an AITuber
[1546] 1. Create an avatar
[1547] The user clicks the "Create Avatar" button on the device's UI and inputs the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[1548] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[1549] 2. Generating speech content
[1550] The user enters the topic "Introducing AITubers" into a form on the device. The server uses LLM to generate comments based on the topic and sends them to the device.
[1551] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[1552] 3. Speech Synthesis
[1553] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[1554] The user checks the audio file and confirms it if there are no particular problems.
[1555] 4. Audio and Motion Mapping
[1556] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[1557] The user checks the motion image and confirms it if there are no particular problems.
[1558] 5. Creating subtitles
[1559] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[1560] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[1561] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system enables anyone to easily generate content that utilizes high-quality virtual characters, without requiring programming knowledge.
[1562] The processing flow will be explained below.
[1563] 1. Creating a 3D / Live2D avatar
[1564] Step 1:
[1565] The user clicks the "Create Avatar" button on the device.
[1566] Step 2:
[1567] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[1568] Step 3:
[1569] The user enters the required basic status information and clicks the "Generate" button.
[1570] Step 4:
[1571] The server receives the information entered by the user.
[1572] Step 5:
[1573] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[1574] Step 6:
[1575] The server sends a preview image of the generated avatar to the terminal.
[1576] Step 7:
[1577] The user checks the preview image displayed on the terminal.
[1578] Step 8:
[1579] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button.
[1580] Step 9:
[1581] The terminal transmits the final avatar information to the server.
[1582] 2. Generating speech content
[1583] Step 1:
[1584] The user clicks the "Generate Message" button on the terminal.
[1585] Step 2:
[1586] The terminal displays a form for the user to input the topic or theme of the comment.
[1587] Step 3:
[1588] The user enters a topic or theme and clicks the "Generate" button.
[1589] Step 4:
[1590] The server receives topic information input by the user.
[1591] Step 5:
[1592] The server launches the LLM and generates comments based on the input topic.
[1593] Step 6:
[1594] The server transmits the generated speech content to the terminal.
[1595] Step 7:
[1596] The user checks the message displayed on the terminal.
[1597] Step 8:
[1598] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[1599] Step 9:
[1600] The terminal transmits the confirmed statement content to the server.
[1601] 3. Speech synthesis
[1602] Step 1:
[1603] The user sends a request for speech synthesis based on the confirmed utterance content.
[1604] Step 2:
[1605] The server receives the confirmed statement content.
[1606] Step 3:
[1607] The server performs speech synthesis through a natural language processing engine.
[1608] Step 4:
[1609] The server transmits the generated audio file to the terminal.
[1610] Step 5:
[1611] The user plays the audio file on the device and checks the content.
[1612] Step 6:
[1613] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[1614] 4. Audio and Motion Mapping
[1615] Step 1:
[1616] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[1617] Step 2:
[1618] The server receives the audio file and avatar data.
[1619] Step 3:
[1620] The server generates motion sequences to achieve natural lip sync and body language.
[1621] Step 4:
[1622] The server transmits the completed motion image to the terminal.
[1623] Step 5:
[1624] The user plays and checks the generated motion video on the terminal.
[1625] Step 6:
[1626] The user requests corrections to movements and facial expressions as needed.
[1627] Step 7:
[1628] The user saves the finalized motion image.
[1629] 5. Creating subtitles
[1630] Step 1:
[1631] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[1632] Step 2:
[1633] The server receives the text of the confirmed utterance.
[1634] Step 3:
[1635] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[1636] Step 4:
[1637] The server overlays and integrates the subtitle file onto the motion video.
[1638] Step 5:
[1639] The server sends the completed video to the terminal.
[1640] Step 6:
[1641] The user checks the video and subtitles on the device.
[1642] Step 7:
[1643] The user checks whether there are any problems during playback of the video and subtitles and requests corrections if necessary.
[1644] Step 8:
[1645] The user saves the finalized video and publishes or shares it.
[1646] The above is a specific program process divided into each processing step.
[1647] Example 1
[1648] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1649] Conventional virtual character creation systems require advanced programming knowledge and specialized skills, making them difficult for average users to use. Furthermore, multiple processes, such as generating speech content, synthesizing voices, and synchronizing motions, are required, making the process time-consuming and tedious. There is a need for a system that can solve these problems and enable anyone to easily create high-quality virtual characters.
[1650] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1651] In this invention, the server includes: means for generating a 3D or Live2D avatar based on information input by a user; means for the user to perform detailed customization of the avatar; means for using a generative AI model to generate speech content based on an input theme; means for providing a user interface for confirming and correcting the generated speech content; means for voice-synthesizing the generated speech content; means for synchronizing a voice file with lip sync and body language based on the confirmed speech content; means for generating a motion sequence corresponding to the voice file; and means for automatically generating subtitles based on the speech content and integrating them into video. This enables users to quickly and efficiently create high-quality virtual characters without requiring special technical knowledge.
[1652] "User" refers to any individual or organization that uses the Platform to create and customize virtual characters.
[1653] "Server" refers to a computer system that performs processes such as generating virtual characters, generating speech content, synthesizing voice, and integrating video and subtitles.
[1654] "Avatar" refers to a 3D or Live2D virtual character generated based on user input.
[1655] "3D avatar" refers to a virtual character that is visually represented in three-dimensional space.
[1656] A "Live2D avatar" is a virtual character that dynamically expresses two-dimensional image data and appears three-dimensional.
[1657] A "generative AI model" refers to an artificial intelligence model that performs natural language processing based on topic information and text information from users.
[1658] "Comment content" refers to the text data generated by the generative AI model based on the topic specified by the user.
[1659] "Speech synthesis" refers to the technology of generating a voice file based on the text data of a speech.
[1660] "Lip sync" refers to the technology of synchronizing the mouth movements of a virtual character with an audio file.
[1661] "Body language" refers to the technology that generates the physical movements and gestures of a virtual character in accordance with the voice and content of what is being said.
[1662] "Motion sequence" refers to a series of movement data including lip sync and body language.
[1663] "Subtitles" refers to text data generated based on what is being said and displayed to viewers within the video.
[1664] "Customization options" refers to settings that a user can select to adjust the appearance and behavior of their avatar.
[1665] "Administration tool" refers to software that provides an interface for users to request modifications to the videos they generate.
[1666] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. This system enables anyone to generate high-quality virtual characters and create related content without requiring specialized skills or programming knowledge.
[1667] System configuration
[1668] The system includes the following main components:
[1669] 1. Avatar generation
[1670] The server receives information input by the user (e.g., gender, age, style, etc.) and activates an image generation AI module (e.g., Stable Diffusion, DALL-E 2) to generate a 3D or Live2D avatar. The generated avatar is then sent to the device.
[1671] On the terminal, the user customizes the details of the avatar (e.g., hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[1672] 2. Generating speech content
[1673] Users enter topic information into a form on their device (e.g., "Introducing AITubers"). The server uses a generative AI model (e.g., GPT-3, ChatGPT) to generate a post based on the topic. The generated post is sent to the device for the user to review.
[1674] The user confirms and modifies the generated message and sends the finalized message to the server.
[1675] 3. Speech synthesis
[1676] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly), and sends the generated audio file to the device.
[1677] The user checks the audio file and requests corrections if necessary.
[1678] 4. Audio and Motion Mapping
[1679] The server receives the audio file and avatar data, creates a motion sequence to generate lip sync and body language (e.g., Live2D Cubism, Adobe Character Animator), and sends the generated motion image to the device.
[1680] The user checks the motion image and requests corrections if necessary.
[1681] 5. Creating subtitles
[1682] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech. The subtitle file is integrated into the video, and the completed video is sent to the terminal.
[1683] The user checks the video and subtitles and confirms them after a final check.
[1684] Specific examples
[1685] Example 1: When a user creates an introductory video for an AITuber
[1686] 1. Create an avatar
[1687] The user clicks the "Create Avatar" button on the device's UI and enters the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[1688] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[1689] 2. Generating speech content
[1690] The user enters the topic "Introducing AITubers" into a form on their device. The server uses a generative AI model to generate comments based on the topic and sends them to the device.
[1691] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[1692] 3. Speech Synthesis
[1693] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[1694] The user checks the audio file and confirms it if there are no particular problems.
[1695] 4. Audio and Motion Mapping
[1696] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[1697] The user checks the motion image and confirms it if there are no particular problems.
[1698] 5. Creating subtitles
[1699] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[1700] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[1701] Prompt Sentence Examples
[1702] Generate a 3D character based on the following information: Gender is female, age is in her 20s, and she wears casual clothing. For background information, the character is an AITuber on YouTube. The statement should read, "Hello, today I'd like to introduce my channel."
[1703] The above system allows users to quickly and easily generate virtual characters and create compelling content without any special technical knowledge.
[1704] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1705] Step 1:
[1706] The user clicks the "Create Avatar" button on the device's UI and enters basic status information (e.g., gender, age, style, etc.).
[1707] Input: The user fills in a form with basic information such as gender, age, and style.
[1708] Action: Sends the entered information to the server.
[1709] Output: Basic status information is entered into the server.
[1710] Step 2:
[1711] The server launches an image generation AI module (e.g., Stable Diffusion, DALL-E 2) based on the received basic status information and generates a 3D or Live2D avatar.
[1712] Input: Basic status information.
[1713] How it works: Passes information to an image generation AI module, which applies an algorithm to generate an avatar.
[1714] Output: An initial avatar image is generated and sent to the device as a preview.
[1715] Step 3:
[1716] On the device, the user receives a preview of the generated avatar and can then perform detailed customization (e.g., hairstyle, clothing, etc.).
[1717] Input: Previewed avatar image and customization options.
[1718] How it works: Use the customization UI to adjust the details using sliders and dropdown menus.
[1719] Output: Generates the finalized avatar configuration.
[1720] Step 4:
[1721] When the user confirms the avatar, the data of the confirmed avatar is transmitted to the server.
[1722] Input: Avatar settings adjusted by the user.
[1723] Action: Press the Confirm button to send the avatar configuration data to the server.
[1724] Output: The confirmed avatar configuration data is sent to the server.
[1725] Step 5:
[1726] Users enter topic information into a form on their device (e.g., "Introducing AITubers").
[1727] Input: Topic information.
[1728] Action: Enter a topic in the text area and press the send button.
[1729] Output: The topic information is sent to the server.
[1730] Step 6:
[1731] The server passes the received topic information to a generative AI model (e.g., GPT-3, ChatGPT) to generate the content of the comments.
[1732] Input: Topic information.
[1733] How it works: Provide topic information as prompts to a generative AI model and receive the generated utterances.
[1734] Output: The generated speech is sent to the terminal.
[1735] Step 7:
[1736] The user checks the generated comment and corrects it if necessary.
[1737] Input: The generated utterance.
[1738] Action: Edit the comment in the editor and press the send button after editing.
[1739] Output: The corrected statement is sent to the server.
[1740] Step 8:
[1741] The user confirms the content of the statement and sends it to the server.
[1742] Input: Confirmed statement.
[1743] Action: Press the Confirm button to send the final revised statement to the server.
[1744] Output: The confirmed statement is sent to the server.
[1745] Step 9:
[1746] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly).
[1747] Input: Confirmed statement.
[1748] What it does: Passes what is said to a speech synthesis engine, generating an audio file.
[1749] Output: An audio file is generated and sent to the device.
[1750] Step 10:
[1751] The user checks the generated audio file and requests corrections if necessary.
[1752] Input: An audio file.
[1753] Action: Plays an audio file and sends feedback for corrections if needed.
[1754] Output: Finalized audio file.
[1755] Step 11:
[1756] The server receives the audio file and avatar data and generates lip sync and body language.
[1757] Input: Audio file and avatar data.
[1758] Movement: Generate avatar mouth and body movements in sync with the audio to create a motion sequence.
[1759] Output: A motion image is generated and sent to the device.
[1760] Step 12:
[1761] The user checks the generated motion image and requests corrections if necessary.
[1762] Input: Motion footage.
[1763] How it works: Plays motion footage and provides feedback for corrections if there are any issues.
[1764] Output: Finalized motion footage.
[1765] Step 13:
[1766] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech.
[1767] Input: The text of what was said.
[1768] Operation: Converts text data into subtitle file format according to the time axis.
[1769] Output: Subtitle files are generated and integrated into the video.
[1770] Step 14:
[1771] The user can check the automatically generated subtitled video and request corrections if necessary.
[1772] Input: Video with subtitles.
[1773] Action: Play the video to check the timing and content of the subtitles, and send correction requests if necessary.
[1774] Output: Finalized video with subtitles.
[1775] The above are the specific steps of the system's program processing, which allows users to efficiently create virtual characters and generate related content.
[1776] (Application example 1)
[1777] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1778] Previous virtual character generation platforms required a lot of effort for users to create characters, making them inconvenient, especially when using smartphones. Furthermore, they lacked assistant functions for real-time user interaction and product introductions, making them difficult to apply to virtual stores. Furthermore, they lacked management tools that could flexibly accommodate user requests for corrections to the videos they generated.
[1779] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1780] In this invention, the server includes means for generating a 3D or Live2D avatar based on information input by a user, means for generating speech content based on the input theme, means for synthesizing the generated speech content with speech, means for synchronizing the speech file with the avatar's movements, means for automatically generating subtitles based on the speech content and integrating them into video, means for using the avatar as a virtual assistant character on a smartphone, and means for the virtual assistant character to introduce products and respond to user inquiries in real time. This allows users to easily create a virtual assistant character on their smartphone and realize real-time conversations, product introductions, and customer support.
[1781] "User-inputted information" refers to basic status information and thematic input data that a user provides to the system.
[1782] A "3D or Live2D avatar" is a three-dimensional or two-dimensional character generated based on user input.
[1783] "Comment content" is text information generated based on the theme entered by the user.
[1784] "Speech synthesis" is the process of creating an audio file based on generated utterances.
[1785] "Synchronizing the audio file with the avatar's movements" is the process of moving the character's mouth and body in sync with the generated audio.
[1786] "Automatic subtitling" is the process of generating text based on what is being said and integrating it with the video.
[1787] "Use on a smartphone" refers to using a virtual assistant character on a smartphone application.
[1788] A "virtual assistant character" is a digital character created by a user that introduces products and responds to user inquiries in real time.
[1789] The "product introduction function" is a function in which a virtual assistant character explains specific product information to the user.
[1790] "Real-time response" refers to the ability to respond immediately to user questions and requests.
[1791] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Furthermore, this system can be used as a virtual assistant character on smartphones, and has the ability to introduce products and respond to users in real time.
[1792] System configuration
[1793] 1. 3D / Live2D avatar generation
[1794] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[1795] On the device (smartphone), the user customizes the avatar in detail (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[1796] 2. Generating speech content
[1797] The server passes the topic information entered by the user to the generative AI model to generate the utterance content, which is then sent to the device for the user to confirm.
[1798] The user confirms and modifies the generated message and sends the finalized message to the server.
[1799] 3. Speech synthesis
[1800] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[1801] The user checks the audio file and requests corrections if necessary.
[1802] 4. Audio and Motion Mapping
[1803] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[1804] The user checks the motion image and requests corrections if necessary.
[1805] 5. Creating subtitles
[1806] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file into the video, and sends the completed video to the device.
[1807] The user checks the video and subtitles and confirms them after a final check.
[1808] 6. Use as a virtual assistant
[1809] The virtual assistant character generated on the device (smartphone) introduces products and responds to user inquiries in real time, realizing an interactive dialogue with the user.
[1810] Users obtain specific product information or inquiries through the virtual assistant.
[1811] Hardware and software used
[1812] Hardware: Smartphones, servers
[1813] Software: OpenAI API (generative AI model), pyttsx3 (speech synthesis), OpenCV (video and motion processing)
[1814] Specific examples
[1815] Example 1: A user creates a virtual assistant character to introduce a new product, and the character explains the company's product through a smartphone app.
[1816] Example prompt sentence:
[1817] "Please explain the features of your new product, including its target market and its benefits."
[1818] Sample output: "This new product is a casual-style smartwatch designed for young people. Key features include lightweight design, long battery life, and the latest fitness tracking capabilities. It's especially suitable for consumers who lead active lifestyles."
[1819] This system allows users to easily create and use virtual assistant characters on their smartphones, allowing them to introduce products and respond in real time.
[1820] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1821] Step 1:
[1822] Obtaining input information
[1823] Users use a smartphone app to input basic status information (gender, age, style, etc.) to create an avatar.
[1824] The server receives information entered by the user and launches the image generation AI module.
[1825] Output: Basic status information for the user.
[1826] Step 2:
[1827] Avatar generation
[1828] The server generates a 3D or Live2D avatar based on the user's basic status information received in step 1.
[1829] The server sends a preview of the generated avatar to the terminal.
[1830] Output: The generated avatar and its preview image.
[1831] Step 3:
[1832] Avatar Customization
[1833] Users customize their avatar's details (hairstyle, clothing, etc.) using a smartphone app.
[1834] The terminal sends the customized results to the server.
[1835] Output: Customized avatar information.
[1836] Step 4:
[1837] Generate speech content
[1838] The user uses the app to input the topic of the comment (e.g., new product introduction).
[1839] The server generates speech content based on the input topic using a generative AI model (OpenAI API).
[1840] The server sends the generated message to the terminal, where the user can confirm and edit it.
[1841] Output: The generated utterance.
[1842] Step 5:
[1843] Confirmation and confirmation of statements
[1844] The user checks the generated comment and corrects it if necessary.
[1845] The user sends the finalized content of the message to the server.
[1846] Output: Confirmed statement.
[1847] Step 6:
[1848] Speech synthesis
[1849] The server passes the confirmed utterance content to a natural language processing engine (pyttsx3) and performs speech synthesis.
[1850] The server transmits the generated audio file to the terminal.
[1851] Output: The generated audio file.
[1852] Step 7:
[1853] Audio and Motion Mapping
[1854] The server receives the audio files and avatar data and creates motion sequences to generate lip sync and body language.
[1855] The server transmits the generated motion image to the terminal.
[1856] Output: The generated motion footage.
[1857] Step 8:
[1858] Subtitle creation and integration
[1859] The server automatically generates a subtitle file based on the text of the confirmed speech content.
[1860] The server integrates the subtitle file into the video and sends the final video to the terminal.
[1861] Output: A merged video file.
[1862] Step 9:
[1863] Use as a virtual assistant
[1864] The device (smartphone) uses the generated virtual assistant character to introduce products and respond to user inquiries in real time.
[1865] Users can use the virtual assistant to obtain product information and make inquiries.
[1866] Output: Product information and support provided to the user.
[1867] The above is a detailed flow of the processing steps of this system. By performing these operations, users can easily create a virtual assistant character and introduce products and respond in real time on their smartphones.
[1868] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1869] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. By combining it with an emotion engine, it also has the ability to dynamically adjust the avatar's expressions based on the user's emotions. This system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Additionally, by incorporating an emotion engine that recognizes the user's emotions, it dynamically changes facial expressions and movements, and adjusts the tone and pitch of the voice, based on the user's emotions.
[1870] System configuration
[1871] 1. 3D / Live2D avatar generation
[1872] The server receives basic status information (gender, age, style, etc.) entered by the user and activates an image generation AI module to generate a 3D or Live2D avatar.
[1873] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[1874] 2. Generating speech content
[1875] The server passes the topic information entered by the user to the LLM to generate the content of the comment, which is then sent to the terminal.
[1876] The user confirms and modifies the generated message and sends the finalized message to the server.
[1877] 3. Speech synthesis
[1878] The server synthesizes the confirmed speech content through a natural language processing engine and transmits the generated audio file to the terminal.
[1879] The user checks the audio file and requests corrections if necessary.
[1880] 4. Audio and Motion Mapping
[1881] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and then transmits the generated motion image to the device.
[1882] The user checks the motion image and requests corrections if necessary.
[1883] 5. Creating subtitles
[1884] The server automatically generates a subtitle file based on the text of the confirmed remarks, integrates it with the video, and sends it to the terminal.
[1885] The user checks the video and subtitles and confirms them after a final check.
[1886] 6. Leveraging Emotional Engines
[1887] The device uses an emotion engine to analyze the user's facial expressions, tone of voice, and gestures in real time to recognize the user's emotions.
[1888] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[1889] The server adjusts the tone and pitch of the generated speech synthesis according to the user's emotions.
[1890] Specific examples
[1891] Example 1: A user creates a video to express their emotions
[1892] 1. Create an avatar
[1893] The user clicks the "Create Avatar" button and enters the avatar's basic status (e.g., male, young man, casual style, etc.). The server generates an avatar based on the entered information and sends a preview to the device.
[1894] The user can view the preview, customize the details and finally confirm the avatar.
[1895] 2. Generating speech content
[1896] The user inputs the topic "Emotional Expression" as the theme. The server uses LLM to generate utterances based on the theme and sends them to the device.
[1897] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[1898] 3. Speech Synthesis
[1899] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[1900] The user checks the audio and confirms it if there is no particular problem.
[1901] 4. Audio and Motion Mapping
[1902] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[1903] The user checks the motion image and confirms it if there are no particular problems.
[1904] 5. Creating subtitles
[1905] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[1906] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[1907] 6. Leveraging Emotional Engines
[1908] The device senses the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine.
[1909] The server dynamically adjusts lip sync, facial expressions, and body language based on the obtained emotional data, and appropriately changes the tone and pitch of the voice.
[1910] The user checks the video of the generated emotional expression, makes final adjustments and confirms it.
[1911] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system allows users to easily generate high-quality virtual character content that reflects their own emotions, without requiring programming knowledge.
[1912] The processing flow will be explained below.
[1913] 1. Creating a 3D / Live2D avatar
[1914] Step 1:
[1915] The user clicks the "Create Avatar" button on the device.
[1916] Step 2:
[1917] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[1918] Step 3:
[1919] The user enters the required basic status information and clicks the "Generate" button.
[1920] Step 4:
[1921] The server receives the information entered by the user.
[1922] Step 5:
[1923] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[1924] Step 6:
[1925] The server sends a preview image of the generated avatar to the terminal.
[1926] Step 7:
[1927] The user checks the preview image displayed on the terminal.
[1928] Step 8:
[1929] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button again.
[1930] Step 9:
[1931] The terminal transmits the final avatar information to the server.
[1932] 2. Generating speech content
[1933] Step 1:
[1934] The user clicks the "Generate Message" button on the terminal.
[1935] Step 2:
[1936] The terminal displays a form for the user to input the topic or theme of the comment.
[1937] Step 3:
[1938] The user enters a topic or theme and clicks the "Generate" button.
[1939] Step 4:
[1940] The server receives topic information input by the user.
[1941] Step 5:
[1942] The server launches the LLM and generates comments based on the input topic.
[1943] Step 6:
[1944] The server transmits the generated speech content to the terminal.
[1945] Step 7:
[1946] The user checks the message displayed on the terminal.
[1947] Step 8:
[1948] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[1949] Step 9:
[1950] The terminal transmits the confirmed statement content to the server.
[1951] 3. Speech synthesis
[1952] Step 1:
[1953] The user sends a request for speech synthesis based on the confirmed utterance content.
[1954] Step 2:
[1955] The server receives the confirmed statement content.
[1956] Step 3:
[1957] The server performs speech synthesis through a natural language processing engine.
[1958] Step 4:
[1959] The server transmits the generated audio file to the terminal.
[1960] Step 5:
[1961] The user plays the audio file on the device and checks the content.
[1962] Step 6:
[1963] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[1964] 4. Audio and Motion Mapping
[1965] Step 1:
[1966] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[1967] Step 2:
[1968] The server receives the audio file and avatar data.
[1969] Step 3:
[1970] The server generates motion sequences to achieve natural lip sync and body language.
[1971] Step 4:
[1972] The server transmits the completed motion image to the terminal.
[1973] Step 5:
[1974] The user plays and checks the generated motion video on the terminal.
[1975] Step 6:
[1976] The user requests corrections to movements and facial expressions as needed.
[1977] Step 7:
[1978] The user saves the finalized motion image.
[1979] 5. Creating subtitles
[1980] Step 1:
[1981] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[1982] Step 2:
[1983] The server receives the text of the confirmed utterance.
[1984] Step 3:
[1985] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[1986] Step 4:
[1987] The server overlays and integrates the subtitle file onto the motion video.
[1988] Step 5:
[1989] The server sends the completed video to the terminal.
[1990] Step 6:
[1991] The user checks the video and subtitles on the device.
[1992] Step 7:
[1993] The user saves the finalized video and publishes or shares it.
[1994] 6. Leveraging Emotional Engines
[1995] Step 1:
[1996] When a user plays back the video they created, the device's camera and microphone are used to detect emotional data such as facial expressions and tone of voice in real time.
[1997] Step 2:
[1998] The terminal transmits the emotion data acquired in real time to the emotion engine, and transmits the analysis results to the server.
[1999] Step 3:
[2000] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[2001] Step 4:
[2002] When the server synthesizes voice from the generated speech, it adjusts the tone and pitch of the voice according to the user's emotions.
[2003] Step 5:
[2004] The server generates the final video based on the emotion data and sends it to the terminal.
[2005] Step 6:
[2006] The user checks the video in which the emotional expression is reflected and requests corrections as necessary.
[2007] Step 7:
[2008] The user saves the finalized video and publishes or shares it.
[2009] The above are the specific processing steps in the program processing of a system that utilizes an emotion engine.
[2010] Example 2
[2011] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2012] Conventional virtual character creation platforms make it difficult for users to express emotions or easily generate speech content. Synchronizing video and audio and generating subtitles is also time-consuming, making the overall process complicated and making it difficult to easily create high-quality content.
[2013] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a 3D or 2D avatar based on information input by a user, a means for generating speech content based on an input theme, a means for voice-synthesizing the generated speech content, a means for synchronizing the speech file with the avatar's movement, a means for automatically generating subtitles based on the speech content and integrating them into video, and a means for collecting user emotion data using an emotion analysis engine and dynamically adjusting the avatar's expression. This allows a user to easily create high-quality virtual character content and generate video including real-time expressions that reflect the user's emotions.
[2014] "User" refers to the entity that uses this system to create a virtual character, customize it, generate comments, etc.
[2015] "Server" refers to a central processing unit that receives input information from a user and performs various processes.
[2016] "3D or 2D Avatar" means a virtual character created by a user and represented in three-dimensional or two-dimensional graphics.
[2017] "Input information" refers to data that a user provides to the system, such as basic status information such as gender, age, and style.
[2018] A "theme" refers to a topic that a user provides to the system for generating comment content.
[2019] "Speech content" refers to the lines and phrases of the character generated based on the input theme.
[2020] "Speech synthesis" is a technology that converts text data into voice data, and refers to a means of outputting the generated utterances as voice.
[2021] "Synchronizing" refers to the process of matching the audio and character movements.
[2022] "Subtitles" are texts that are automatically generated based on what is being said and are integrated into video.
[2023] "Integrating into video" refers to the process of incorporating the generated subtitles into a single video file in conjunction with the audio and avatar movements.
[2024] An "emotion analysis engine" refers to software or a system that analyzes a user's facial expressions, tone of voice, and gestures to recognize emotions.
[2025] "Emotion data" refers to data that indicates the user's emotional state as recognized by the emotion analysis engine.
[2026] "Dynamic adjustment" refers to the act of changing an avatar's facial expressions, movements, voice tone, pitch, etc. in real time.
[2027] "Customization options" refers to settings that allow users to change details about their avatar.
[2028] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters, and by combining it with an emotion analysis engine, it has the ability to dynamically adjust the avatar's expression according to the user's emotions.
[2029] System configuration
[2030] The system consists of the following main components:
[2031] 1. Server: A central processing unit that receives information from users and performs various processes.
[2032] 2. Terminal: A device operated by a user that provides an interface for user input, display, customization, etc.
[2033] 3. Sentiment analysis engine: Software for analyzing user emotions and generating data.
[2034] Technology used
[2035] 1. Image generation AI module (e.g. DeepArt): Generates a 3D or 2D avatar based on basic status information (gender, age, style, etc.) entered by the user.
[2036] 2. Large-scale language models (LLMs) (e.g., GPT-4): Generate utterances based on themes entered by the user.
[2037] 3. Natural language processing engine (e.g., Google Text-to-Speech): Converts the generated speech into audio.
[2038] 4. Emotion analysis engine (e.g., Affectiva SDK): Analyzes the user's facial expressions, tone of voice, and gestures to collect emotional data.
[2039] System Operation
[2040] The system operates in the following steps:
[2041] 1. Create your avatar:
[2042] The user clicks the "Create Avatar" button on the device and enters basic status information (gender, age, style, etc.).
[2043] The terminal transmits this information to the server.
[2044] The server uses an image generation AI module to generate a 3D or 2D avatar and sends a preview to the device.
[2045] The user checks the preview image, performs detailed customization (hairstyle, clothing, etc.), and confirms it.
[2046] The terminal transmits the determined avatar information to the server.
[2047] 2. Speech generation:
[2048] The user inputs topic information and sends it to the server.
[2049] The server uses LLM to generate speech content and sends it to the terminal.
[2050] The user confirms and corrects the content of the comment and sends the finalized content to the server.
[2051] 3. Speech synthesis:
[2052] The server synthesizes the confirmed speech content through a natural language processing engine.
[2053] The server transmits the generated audio file to the terminal.
[2054] The user checks the audio file and requests corrections if necessary.
[2055] 4. Audio and Motion Mapping:
[2056] The server passes the audio file and avatar data to a lip sync engine to generate lip sync and body language.
[2057] The server transmits the generated motion sequence to the terminal.
[2058] The user checks the motion image and requests corrections if necessary.
[2059] 5. Creating subtitles:
[2060] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[2061] The user checks the video and subtitles and confirms them after a final check.
[2062] 6. Utilizing sentiment analysis engines:
[2063] The device detects the user's facial expressions and tone of voice and passes them to an emotion analysis engine in real time.
[2064] The terminal transmits the emotion analysis results obtained to the server.
[2065] The server dynamically adjusts the avatar's facial expressions and movements based on the emotional data, and also appropriately changes the tone and pitch of the voice.
[2066] The user checks the final image, makes adjustments and confirms them.
[2067] Specific examples
[2068] Example 1: A user creates a video to express their emotions
[2069] 1. Create your avatar:
[2070] The user clicks the "Create Avatar" button and inputs the avatar's basic status (e.g., male, young man, casual style, etc.).
[2071] The terminal displays an input form and accepts input from the user.
[2072] When the user inputs information and presses the send button, the terminal sends the information to the server.
[2073] The server receives the information, activates the image generation AI module, and generates the avatar.
[2074] The server sends the generated avatar to the device as a preview.
[2075] The device displays a preview, and the user can customize the details (hairstyle, clothing, etc.) and finally confirm the avatar.
[2076] The terminal transmits the determined avatar information to the server.
[2077] 2. Speech generation:
[2078] The user opens the "Comment Content" input form and inputs the topic "Emotional Expressions."
[2079] The terminal transmits the topic information to the server.
[2080] The server generates the message using LLM and sends it to the terminal.
[2081] The user confirms and modifies the generated message and sends the finalized message to the server.
[2082] 3. Speech synthesis:
[2083] The server uses a natural language processing engine to synthesize speech based on the confirmed content of the speech.
[2084] The server transmits the generated audio file to the terminal.
[2085] The user checks the audio file and confirms it if there are no problems.
[2086] 4. Audio and Motion Mapping:
[2087] The server generates lip sync and body language based on the audio file and the determined avatar.
[2088] The server transmits the motion video to the terminal and provides it to the user.
[2089] The user checks the motion image and requests corrections if necessary.
[2090] 5. Creating subtitles:
[2091] The server automatically generates subtitles based on the confirmed remarks, integrates them into the video, and sends them to the terminal.
[2092] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[2093] 6. Utilizing sentiment analysis engines:
[2094] The device senses the user's facial expressions and tone of voice in real time and sends the analysis results to the emotion engine.
[2095] The terminal transmits the emotion analysis result to the server.
[2096] The server dynamically adjusts lip sync, facial expressions, and movements based on emotional data, and also appropriately changes the tone and pitch of the voice.
[2097] The user checks the final image, makes adjustments and confirms them.
[2098] By following the above steps, users can easily create high-quality virtual character content. This system also enables real-time expression according to emotions, making it possible to generate a wide variety of content.
[2099] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2100] Step 1: User begins creating an avatar
[2101] The user clicks the "Create Avatar" button on the device.
[2102] Input: User input (basic status information such as gender, age, style, etc.)
[2103] The terminal displays an input form and accepts input from the user.
[2104] How it works: The user enters the required information and presses the submit button.
[2105] Output: The information entered
[2106] Step 2: The server generates the avatar
[2107] The server launches an image generation AI module based on the user's basic status information received from the device.
[2108] Input: Basic status information entered
[2109] How it works: The server uses an image generation AI module to generate a 3D or 2D avatar.
[2110] Output: A preview image of the generated avatar
[2111] Step 3: User customizes avatar
[2112] The terminal displays a preview of the generated avatar to the user.
[2113] Input: A preview image of the generated avatar
[2114] How it works: The user customizes the avatar's details (hairstyle, clothing, etc.) and finally finalizes the avatar.
[2115] Output: A customized confirmed avatar
[2116] Step 4: Enter and submit topic information
[2117] The user opens the "Comment Generation" input form on the terminal and inputs a topic (e.g., "About emotional expression").
[2118] Input: Topic information
[2119] How it works: The user enters a topic and presses the send button.
[2120] Output: Input topic information
[2121] Step 5: Generate speech
[2122] The server receives topic information and generates utterances using a large-scale language model (LLM).
[2123] Input: Topic information
[2124] How it works: The server passes the prompt to the generative AI model to generate the speech.
[2125] Output: Generated speech
[2126] Step 6: Revise and confirm your statement
[2127] The terminal displays the generated comment content to the user.
[2128] Input: Generated speech
[2129] Action: The user confirms and corrects what has been said, and finally confirms it.
[2130] Output: Confirmed statement
[2131] Step 7: Speech synthesis of what is being said
[2132] The server passes the confirmed utterance content to a natural language processing engine (e.g., Google Text-to-Speech) for speech synthesis.
[2133] Input: Confirmed statement
[2134] How it works: The server uses a speech synthesis engine to generate an audio file.
[2135] Output: Generated audio file
[2136] Step 8: Check and correct the audio file
[2137] The terminal provides the generated audio file to the user.
[2138] Input: Generated audio file
[2139] Action: The user reviews the audio file and requests corrections if necessary.
[2140] Output: Finalized audio file
[2141] Step 9: Mapping Audio and Motion
[2142] The server passes the audio file and the determined avatar to a lip sync engine to generate lip sync and body language.
[2143] Input: Audio file, confirmed avatar
[2144] How it works: The server generates lip sync and motion data.
[2145] Output: Generated motion sequence
[2146] Step 10: Check and correct motion footage
[2147] The terminal displays the generated motion image to the user.
[2148] Input: Generated motion sequence
[2149] Action: The user reviews the motion footage and requests corrections if necessary.
[2150] Output: Confirmed motion footage
[2151] Step 11: Creating and merging subtitles
[2152] The server automatically generates subtitles based on the confirmed remarks and integrates them into the video.
[2153] Input: Confirmed speech content, confirmed motion image
[2154] How it works: The server generates a subtitle file and integrates it into the video.
[2155] Output: Integrated video (audio, motion, subtitles)
[2156] Step 12: Reflecting the results of sentiment analysis
[2157] The device passes the user's facial expressions and tone of voice to an emotion analysis engine.
[2158] Input: Real-time facial expressions, tone of voice
[2159] Operation: The device obtains the emotion analysis results and sends them to the server.
[2160] Output: Emotion data
[2161] Step 13: Adjust based on sentiment data
[2162] The server dynamically adjusts the avatar's facial expressions and movements based on the emotional data, and changes the tone and pitch of the voice.
[2163] Input: Emotion data, confirmed motion video, audio file
[2164] Movement: The server adjusts the avatar's facial expressions, movements, and voice tone and pitch.
[2165] Output: Adjusted video and audio
[2166] Step 14: Final review and confirmation
[2167] The terminal provides the final video to the user for confirmation.
[2168] Input: Adjusted video and audio
[2169] Action: The user performs a final review, makes any necessary adjustments, and confirms.
[2170] Output: Final video
[2171] (Application example 2)
[2172] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2173] Currently, it is difficult for users to easily create interactive, emotionally expressive virtual characters. It is also difficult to dynamically adjust the virtual character's facial expressions and voice tone in response to the user's real-time emotions. This can result in a lack of a natural and engaging experience for the user, leading to lower user satisfaction. The present invention aims to solve these problems and provide a system that enables a virtual character to express natural emotions in response to the user's emotions.
[2174] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for generating a 3D or Live2D avatar based on information input by the user, means for generating speech content based on the input theme, means for voice-synthesizing the generated speech content, means for synchronizing the audio file with the avatar's movements, means for automatically generating subtitles based on the speech content and integrating them into video, means for sensing the user's voice and facial expressions in real time and analyzing emotions using an emotion engine, and means for dynamically adjusting the avatar's facial expressions and voice tone based on the emotion data. This allows users to easily create virtual characters and enjoy natural expressions that are dynamically adjusted according to emotions.
[2175] "3D or Live2D Avatar" means a virtual character that is represented in three or two dimensions based on a user-selected appearance and style.
[2176] "Means for generating utterances" refers to a system or algorithm that generates appropriate responses or utterances based on themes or questions entered by users.
[2177] "Means for speech synthesis" refers to the technology for generating speech from text, and the function for converting the generated speech into natural-sounding speech.
[2178] "Means for synchronizing audio files with avatar movements" refers to a technology that synchronizes the mouth movements and facial expressions of an avatar with the generated audio.
[2179] "Means for automatically generating subtitles and integrating them into video" refers to technology that automatically generates subtitles based on what is being said and embeds them into the video.
[2180] The "emotion engine" is a system that analyzes the user's tone of voice and facial expressions to recognize the user's emotional state in real time.
[2181] "Means for dynamically adjusting the avatar's facial expression and voice tone based on emotional data" is a function that appropriately changes the avatar's facial expression and voice tone based on the results of analysis by the emotion engine.
[2182] This invention provides a platform that allows users to easily create 3D or Live2D avatars, which then dynamically change their expressions based on the user's emotions. Based on user input, this system generates and adjusts real-time avatar movements that integrate speech, audio, subtitles, and emotional expressions.
[2183] System configuration
[2184] 1. Avatar generation
[2185] The server receives basic information entered by the user (such as gender, age, and style) and generates a 3D or Live2D avatar using an image generation AI module.
[2186] The terminal provides an interface that allows the user to customize the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[2187] 2. Generating speech content
[2188] The server uses a generative AI model (e.g., GPT-4) to generate a message based on the themes entered by the user. This message is then sent to the device, where it is reviewed and revised by the user before being resent.
[2189] 3. Speech Synthesis
[2190] The server generates audio based on the confirmed utterance using a natural language processing engine (e.g., Google Cloud Text-to-Speech) and sends the generated audio file to the device.
[2191] 4. Audio and Motion Synchronization
[2192] The server creates a motion sequence to generate lip sync and body language based on the audio file and avatar data, and sends it to the device.
[2193] 5. Creating subtitles
[2194] The server automatically generates a subtitle file based on the text of the confirmed speech and integrates it into the video, which is then sent to the device.
[2195] 6. Leveraging Emotional Engines
[2196] The device detects the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine (e.g., Microsoft Azure Emotion API).
[2197] The server dynamically adjusts the avatar's facial expressions, movements, and voice tone based on the emotional data obtained from the emotion engine.
[2198] Specific examples
[2199] Example 1: Virtual Shopping Assistant
[2200] The user puts on the smart glasses and says, "Start shopping assistant." The device sends the command to the server.
[2201] The server generates a 3D avatar based on basic information provided by the user and sends it to the device. The user then customizes the avatar in detail (e.g., glasses, hairstyle) and finalizes the avatar.
[2202] The user asks, "What material is this shirt made of?" The server passes the question to a generative AI model (e.g., GPT-4), which generates an appropriate response and performs speech synthesis.
[2203] The device detects the user's facial expressions and tone of voice, analyzes them using an emotion engine, and displays and plays responses that dynamically adjust the avatar's facial expressions and tone of voice according to the emotional data.
[2204] Prompt Sentence Examples
[2205] User: What material is this shirt made of?
[2206] Virtual Character: Okay. This shirt is made of 100% cotton. It feels amazing!
[2207] This system allows users to easily generate high-quality virtual character content that reflects their own emotions, without requiring any programming knowledge.
[2208] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2209] Step 1:
[2210] Entering and creating basic avatar information
[2211] The user inputs basic information (gender, age, style, etc.) from the terminal, which is then sent to the server.
[2212] The server generates a 3D or Live2D avatar using an image generation AI module based on the received basic information, and a preview image of the generated avatar is sent to the device.
[2213] Output: A preview image of the avatar generated based on the basic information.
[2214] Step 2:
[2215] Detailed avatar customization
[2216] The user customizes the avatar in detail (hairstyle, clothing, etc.) from the terminal and finally confirms the avatar. The detailed customization data is sent to the server.
[2217] The server stores the determined avatar data, generates the final avatar image, and transmits it to the terminal.
[2218] Output: Final user-customized avatar image and data.
[2219] Step 3:
[2220] Generate speech content
[2221] The user inputs a specific topic or question (everyday conversation, product description, etc.) into the terminal and sends it to the server.
[2222] The server uses a generative AI model (e.g., GPT-4) to generate utterances based on the themes entered by the user. The generated utterances are then sent to the device.
[2223] Output: The speech text generated by the generative AI model.
[2224] Step 4:
[2225] Speech synthesis
[2226] The user checks and corrects the generated text and sends the finalized text to the server.
[2227] The server uses a natural language processing engine (e.g., Google Cloud Text-to-Speech) to convert what is said into audio, and the resulting audio file is sent to the device.
[2228] Output: Synthesized audio file.
[2229] Step 5:
[2230] Audio and motion synchronization
[2231] The server creates a motion sequence to generate lip sync and body language based on the audio file and avatar data, and sends it to the device.
[2232] The device will synchronize the motion sequence with the audio file and display a preview.
[2233] Output: Animated sequence with synchronized lip sync and body language.
[2234] Step 6:
[2235] Creating subtitles
[2236] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates it into the video, and sends the video file to the device.
[2237] The device will display a preview of the video and subtitles for the user to review.
[2238] Output: Video file with integrated subtitles.
[2239] Step 7:
[2240] Emotion analysis
[2241] The device senses the user's facial expressions and tone of voice in real time and transmits emotional data to the server.
[2242] The server analyzes the emotion data using an emotion engine (e.g., Microsoft Azure Emotion API).
[2243] Output: Parsed emotion data.
[2244] Step 8:
[2245] Dynamic adjustment based on emotions
[2246] The server generates sequences to dynamically adjust lip sync, facial expressions, and body language based on the emotional data and sends them to the device.
[2247] The device displays the adjusted animation and the user gives a final confirmation.
[2248] Output: Dynamically adjusted animation depending on the emotion.
[2249] Examples and prompts
[2250] Prompt Sentence Examples
[2251] User: What material is this shirt made of?
[2252] Virtual Character: Okay. This shirt is made of 100% cotton. It feels amazing!
[2253] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2254] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2255] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2256] [Fourth embodiment]
[2257] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2258] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2259] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2260] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2261] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2262] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2263] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2264] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2265] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2266] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2267] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2268] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2269] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2270] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech with motion, and generates and integrates subtitles into the video.
[2271] System configuration
[2272] 1. 3D / Live2D avatar generation
[2273] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[2274] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[2275] 2. Generating speech content
[2276] The server passes the topic information entered by the user to the LLM, generates a message, and sends the generated message to the terminal for the user to confirm.
[2277] The user confirms and modifies the generated message and sends the finalized message to the server.
[2278] 3. Speech synthesis
[2279] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[2280] The user checks the audio file and requests corrections if necessary.
[2281] 4. Audio and Motion Mapping
[2282] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[2283] The user checks the motion image and requests corrections if necessary.
[2284] 5. Creating subtitles
[2285] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file with the video, and sends the completed video to the device.
[2286] The user checks the video and subtitles and confirms them after a final check.
[2287] Specific examples
[2288] Example 1: When a user creates an introductory video for an AITuber
[2289] 1. Create an avatar
[2290] The user clicks the "Create Avatar" button on the device's UI and inputs the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[2291] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[2292] 2. Generating speech content
[2293] The user enters the topic "Introducing AITubers" into a form on the device. The server uses LLM to generate comments based on the topic and sends them to the device.
[2294] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[2295] 3. Speech Synthesis
[2296] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[2297] The user checks the audio file and confirms it if there are no particular problems.
[2298] 4. Audio and Motion Mapping
[2299] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[2300] The user checks the motion image and confirms it if there are no particular problems.
[2301] 5. Creating subtitles
[2302] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[2303] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[2304] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system enables anyone to easily generate content that utilizes high-quality virtual characters, without requiring programming knowledge.
[2305] The processing flow will be explained below.
[2306] 1. Creating a 3D / Live2D avatar
[2307] Step 1:
[2308] The user clicks the "Create Avatar" button on the device.
[2309] Step 2:
[2310] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[2311] Step 3:
[2312] The user enters the required basic status information and clicks the "Generate" button.
[2313] Step 4:
[2314] The server receives the information entered by the user.
[2315] Step 5:
[2316] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[2317] Step 6:
[2318] The server sends a preview image of the generated avatar to the terminal.
[2319] Step 7:
[2320] The user checks the preview image displayed on the terminal.
[2321] Step 8:
[2322] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button.
[2323] Step 9:
[2324] The terminal transmits the final avatar information to the server.
[2325] 2. Generating speech content
[2326] Step 1:
[2327] The user clicks the "Generate Message" button on the terminal.
[2328] Step 2:
[2329] The terminal displays a form for the user to input the topic or theme of the comment.
[2330] Step 3:
[2331] The user enters a topic or theme and clicks the "Generate" button.
[2332] Step 4:
[2333] The server receives topic information input by the user.
[2334] Step 5:
[2335] The server launches the LLM and generates comments based on the input topic.
[2336] Step 6:
[2337] The server transmits the generated speech content to the terminal.
[2338] Step 7:
[2339] The user checks the message displayed on the terminal.
[2340] Step 8:
[2341] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[2342] Step 9:
[2343] The terminal transmits the confirmed statement content to the server.
[2344] 3. Speech synthesis
[2345] Step 1:
[2346] The user sends a request for speech synthesis based on the confirmed utterance content.
[2347] Step 2:
[2348] The server receives the confirmed statement content.
[2349] Step 3:
[2350] The server performs speech synthesis through a natural language processing engine.
[2351] Step 4:
[2352] The server transmits the generated audio file to the terminal.
[2353] Step 5:
[2354] The user plays the audio file on the device and checks the content.
[2355] Step 6:
[2356] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[2357] 4. Audio and Motion Mapping
[2358] Step 1:
[2359] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[2360] Step 2:
[2361] The server receives the audio file and avatar data.
[2362] Step 3:
[2363] The server generates motion sequences to achieve natural lip sync and body language.
[2364] Step 4:
[2365] The server transmits the completed motion image to the terminal.
[2366] Step 5:
[2367] The user plays and checks the generated motion video on the terminal.
[2368] Step 6:
[2369] The user requests corrections to movements and facial expressions as needed.
[2370] Step 7:
[2371] The user saves the finalized motion image.
[2372] 5. Creating subtitles
[2373] Step 1:
[2374] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[2375] Step 2:
[2376] The server receives the text of the confirmed utterance.
[2377] Step 3:
[2378] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[2379] Step 4:
[2380] The server overlays and integrates the subtitle file onto the motion video.
[2381] Step 5:
[2382] The server sends the completed video to the terminal.
[2383] Step 6:
[2384] The user checks the video and subtitles on the device.
[2385] Step 7:
[2386] The user checks whether there are any problems during playback of the video and subtitles and requests corrections if necessary.
[2387] Step 8:
[2388] The user saves the finalized video and publishes or shares it.
[2389] The above is a specific program process divided into each processing step.
[2390] Example 1
[2391] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2392] Conventional virtual character creation systems require advanced programming knowledge and specialized skills, making them difficult for average users to use. Furthermore, multiple processes, such as generating speech content, synthesizing voices, and synchronizing motions, are required, making the process time-consuming and tedious. There is a need for a system that can solve these problems and enable anyone to easily create high-quality virtual characters.
[2393] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2394] In this invention, the server includes: means for generating a 3D or Live2D avatar based on information input by a user; means for the user to perform detailed customization of the avatar; means for using a generative AI model to generate speech content based on an input theme; means for providing a user interface for confirming and correcting the generated speech content; means for voice-synthesizing the generated speech content; means for synchronizing a voice file with lip sync and body language based on the confirmed speech content; means for generating a motion sequence corresponding to the voice file; and means for automatically generating subtitles based on the speech content and integrating them into video. This enables users to quickly and efficiently create high-quality virtual characters without requiring special technical knowledge.
[2395] "User" refers to any individual or organization that uses the Platform to create and customize virtual characters.
[2396] "Server" refers to a computer system that performs processes such as generating virtual characters, generating speech content, synthesizing voice, and integrating video and subtitles.
[2397] "Avatar" refers to a 3D or Live2D virtual character generated based on user input.
[2398] "3D avatar" refers to a virtual character that is visually represented in three-dimensional space.
[2399] A "Live2D avatar" is a virtual character that dynamically expresses two-dimensional image data and appears three-dimensional.
[2400] A "generative AI model" refers to an artificial intelligence model that performs natural language processing based on topic information and text information from users.
[2401] "Comment content" refers to the text data generated by the generative AI model based on the topic specified by the user.
[2402] "Speech synthesis" refers to the technology of generating a voice file based on the text data of a speech.
[2403] "Lip sync" refers to the technology of synchronizing the mouth movements of a virtual character with an audio file.
[2404] "Body language" refers to the technology that generates the physical movements and gestures of a virtual character in accordance with the voice and content of what is being said.
[2405] "Motion sequence" refers to a series of movement data including lip sync and body language.
[2406] "Subtitles" refers to text data generated based on what is being said and displayed to viewers within the video.
[2407] "Customization options" refers to settings that a user can select to adjust the appearance and behavior of their avatar.
[2408] "Administration tool" refers to software that provides an interface for users to request modifications to the videos they generate.
[2409] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. This system enables anyone to generate high-quality virtual characters and create related content without requiring specialized skills or programming knowledge.
[2410] System configuration
[2411] The system includes the following main components:
[2412] 1. Avatar generation
[2413] The server receives information input by the user (e.g., gender, age, style, etc.) and activates an image generation AI module (e.g., Stable Diffusion, DALL-E 2) to generate a 3D or Live2D avatar. The generated avatar is then sent to the device.
[2414] On the terminal, the user customizes the details of the avatar (e.g., hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[2415] 2. Generating speech content
[2416] Users enter topic information into a form on their device (e.g., "Introducing AITubers"). The server uses a generative AI model (e.g., GPT-3, ChatGPT) to generate a post based on the topic. The generated post is sent to the device for the user to review.
[2417] The user confirms and modifies the generated message and sends the finalized message to the server.
[2418] 3. Speech synthesis
[2419] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly), and sends the generated audio file to the device.
[2420] The user checks the audio file and requests corrections if necessary.
[2421] 4. Audio and Motion Mapping
[2422] The server receives the audio file and avatar data, creates a motion sequence to generate lip sync and body language (e.g., Live2D Cubism, Adobe Character Animator), and sends the generated motion image to the device.
[2423] The user checks the motion image and requests corrections if necessary.
[2424] 5. Creating subtitles
[2425] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech. The subtitle file is integrated into the video, and the completed video is sent to the terminal.
[2426] The user checks the video and subtitles and confirms them after a final check.
[2427] Specific examples
[2428] Example 1: When a user creates an introductory video for an AITuber
[2429] 1. Create an avatar
[2430] The user clicks the "Create Avatar" button on the device's UI and enters the avatar's basic status (e.g., female, young, casual style, etc.). The server generates an avatar based on this and sends a preview to the device.
[2431] The user can customize the avatar's hairstyle and clothing in detail while viewing the preview, and finally finalize the avatar.
[2432] 2. Generating speech content
[2433] The user enters the topic "Introducing AITubers" into a form on their device. The server uses a generative AI model to generate comments based on the topic and sends them to the device.
[2434] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[2435] 3. Speech Synthesis
[2436] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[2437] The user checks the audio file and confirms it if there are no particular problems.
[2438] 4. Audio and Motion Mapping
[2439] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[2440] The user checks the motion image and confirms it if there are no particular problems.
[2441] 5. Creating subtitles
[2442] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[2443] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[2444] Prompt Sentence Examples
[2445] Generate a 3D character based on the following information: Gender is female, age is in her 20s, and she wears casual clothing. For background information, the character is an AITuber on YouTube. The statement should read, "Hello, today I'd like to introduce my channel."
[2446] The above system allows users to quickly and easily generate virtual characters and create compelling content without any special technical knowledge.
[2447] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2448] Step 1:
[2449] The user clicks the "Create Avatar" button on the device's UI and enters basic status information (e.g., gender, age, style, etc.).
[2450] Input: The user fills in a form with basic information such as gender, age, and style.
[2451] Action: Sends the entered information to the server.
[2452] Output: Basic status information is entered into the server.
[2453] Step 2:
[2454] The server launches an image generation AI module (e.g., Stable Diffusion, DALL-E 2) based on the received basic status information and generates a 3D or Live2D avatar.
[2455] Input: Basic status information.
[2456] How it works: Passes information to an image generation AI module, which applies an algorithm to generate an avatar.
[2457] Output: An initial avatar image is generated and sent to the device as a preview.
[2458] Step 3:
[2459] On the device, the user receives a preview of the generated avatar and can then perform detailed customization (e.g., hairstyle, clothing, etc.).
[2460] Input: Previewed avatar image and customization options.
[2461] How it works: Use the customization UI to adjust the details using sliders and dropdown menus.
[2462] Output: Generates the finalized avatar configuration.
[2463] Step 4:
[2464] When the user confirms the avatar, the data of the confirmed avatar is transmitted to the server.
[2465] Input: Avatar settings adjusted by the user.
[2466] Action: Press the Confirm button to send the avatar configuration data to the server.
[2467] Output: The confirmed avatar configuration data is sent to the server.
[2468] Step 5:
[2469] Users enter topic information into a form on their device (e.g., "Introducing AITubers").
[2470] Input: Topic information.
[2471] Action: Enter a topic in the text area and press the send button.
[2472] Output: The topic information is sent to the server.
[2473] Step 6:
[2474] The server passes the received topic information to a generative AI model (e.g., GPT-3, ChatGPT) to generate the content of the comments.
[2475] Input: Topic information.
[2476] How it works: Provide topic information as prompts to a generative AI model and receive the generated utterances.
[2477] Output: The generated speech is sent to the terminal.
[2478] Step 7:
[2479] The user checks the generated comment and corrects it if necessary.
[2480] Input: The generated utterance.
[2481] Action: Edit the comment in the editor and press the send button after editing.
[2482] Output: The corrected statement is sent to the server.
[2483] Step 8:
[2484] The user confirms the content of the statement and sends it to the server.
[2485] Input: Confirmed statement.
[2486] Action: Press the Confirm button to send the final revised statement to the server.
[2487] Output: The confirmed statement is sent to the server.
[2488] Step 9:
[2489] The server synthesizes the confirmed speech using a natural language processing engine (e.g., Google Text-to-Speech, Amazon Polly).
[2490] Input: Confirmed statement.
[2491] What it does: Passes what is said to a speech synthesis engine, generating an audio file.
[2492] Output: An audio file is generated and sent to the device.
[2493] Step 10:
[2494] The user checks the generated audio file and requests corrections if necessary.
[2495] Input: An audio file.
[2496] Action: Plays an audio file and sends feedback for corrections if needed.
[2497] Output: Finalized audio file.
[2498] Step 11:
[2499] The server receives the audio file and avatar data and generates lip sync and body language.
[2500] Input: Audio file and avatar data.
[2501] Movement: Generate avatar mouth and body movements in sync with the audio to create a motion sequence.
[2502] Output: A motion image is generated and sent to the device.
[2503] Step 12:
[2504] The user checks the generated motion image and requests corrections if necessary.
[2505] Input: Motion footage.
[2506] How it works: Plays motion footage and provides feedback for corrections if there are any issues.
[2507] Output: Finalized motion footage.
[2508] Step 13:
[2509] The server automatically generates a subtitle file (e.g., an SRT file) based on the text of the confirmed speech.
[2510] Input: The text of what was said.
[2511] Operation: Converts text data into subtitle file format according to the time axis.
[2512] Output: Subtitle files are generated and integrated into the video.
[2513] Step 14:
[2514] The user can check the automatically generated subtitled video and request corrections if necessary.
[2515] Input: Video with subtitles.
[2516] Action: Play the video to check the timing and content of the subtitles, and send correction requests if necessary.
[2517] Output: Finalized video with subtitles.
[2518] The above are the specific steps of the system's program processing, which allows users to efficiently create virtual characters and generate related content.
[2519] (Application example 1)
[2520] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2521] Previous virtual character generation platforms required a lot of effort for users to create characters, making them inconvenient, especially when using smartphones. Furthermore, they lacked assistant functions for real-time user interaction and product introductions, making them difficult to apply to virtual stores. Furthermore, they lacked management tools that could flexibly accommodate user requests for corrections to the videos they generated.
[2522] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2523] In this invention, the server includes means for generating a 3D or Live2D avatar based on information input by a user, means for generating speech content based on the input theme, means for synthesizing the generated speech content with speech, means for synchronizing the speech file with the avatar's movements, means for automatically generating subtitles based on the speech content and integrating them into video, means for using the avatar as a virtual assistant character on a smartphone, and means for the virtual assistant character to introduce products and respond to user inquiries in real time. This allows users to easily create a virtual assistant character on their smartphone and realize real-time conversations, product introductions, and customer support.
[2524] "User-inputted information" refers to basic status information and thematic input data that a user provides to the system.
[2525] A "3D or Live2D avatar" is a three-dimensional or two-dimensional character generated based on user input.
[2526] "Comment content" is text information generated based on the theme entered by the user.
[2527] "Speech synthesis" is the process of creating an audio file based on generated utterances.
[2528] "Synchronizing the audio file with the avatar's movements" is the process of moving the character's mouth and body in sync with the generated audio.
[2529] "Automatic subtitling" is the process of generating text based on what is being said and integrating it with the video.
[2530] "Use on a smartphone" refers to using a virtual assistant character on a smartphone application.
[2531] A "virtual assistant character" is a digital character created by a user that introduces products and responds to user inquiries in real time.
[2532] The "product introduction function" is a function in which a virtual assistant character explains specific product information to the user.
[2533] "Real-time response" refers to the ability to respond immediately to user questions and requests.
[2534] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters. The system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Furthermore, this system can be used as a virtual assistant character on smartphones, and has the ability to introduce products and respond to users in real time.
[2535] System configuration
[2536] 1. 3D / Live2D avatar generation
[2537] The server receives basic status information (gender, age, style, etc.) input by the user and activates the image generation AI module, which generates a 3D or Live2D avatar based on the input information.
[2538] On the device (smartphone), the user customizes the avatar in detail (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[2539] 2. Generating speech content
[2540] The server passes the topic information entered by the user to the generative AI model to generate the utterance content, which is then sent to the device for the user to confirm.
[2541] The user confirms and modifies the generated message and sends the finalized message to the server.
[2542] 3. Speech synthesis
[2543] The server synthesizes the confirmed speech using a natural language processing engine, and the generated voice file is sent to the device.
[2544] The user checks the audio file and requests corrections if necessary.
[2545] 4. Audio and Motion Mapping
[2546] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and transmits the generated motion image to the device.
[2547] The user checks the motion image and requests corrections if necessary.
[2548] 5. Creating subtitles
[2549] The server automatically generates a subtitle file based on the text of the confirmed speech, integrates the subtitle file into the video, and sends the completed video to the device.
[2550] The user checks the video and subtitles and confirms them after a final check.
[2551] 6. Use as a virtual assistant
[2552] The virtual assistant character generated on the device (smartphone) introduces products and responds to user inquiries in real time, realizing an interactive dialogue with the user.
[2553] Users obtain specific product information or inquiries through the virtual assistant.
[2554] Hardware and software used
[2555] Hardware: Smartphones, servers
[2556] Software: OpenAI API (generative AI model), pyttsx3 (speech synthesis), OpenCV (video and motion processing)
[2557] Specific examples
[2558] Example 1: A user creates a virtual assistant character to introduce a new product, and the character explains the company's product through a smartphone app.
[2559] Example prompt sentence:
[2560] "Please explain the features of your new product, including its target market and its benefits."
[2561] Sample output: "This new product is a casual-style smartwatch designed for young people. Key features include lightweight design, long battery life, and the latest fitness tracking capabilities. It's especially suitable for consumers who lead active lifestyles."
[2562] This system allows users to easily create and use virtual assistant characters on their smartphones, allowing them to introduce products and respond in real time.
[2563] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2564] Step 1:
[2565] Obtaining input information
[2566] Users use a smartphone app to input basic status information (gender, age, style, etc.) to create an avatar.
[2567] The server receives information entered by the user and launches the image generation AI module.
[2568] Output: Basic status information for the user.
[2569] Step 2:
[2570] Avatar generation
[2571] The server generates a 3D or Live2D avatar based on the user's basic status information received in step 1.
[2572] The server sends a preview of the generated avatar to the terminal.
[2573] Output: The generated avatar and its preview image.
[2574] Step 3:
[2575] Avatar Customization
[2576] Users customize their avatar's details (hairstyle, clothing, etc.) using a smartphone app.
[2577] The terminal sends the customized results to the server.
[2578] Output: Customized avatar information.
[2579] Step 4:
[2580] Generate speech content
[2581] The user uses the app to input the topic of the comment (e.g., new product introduction).
[2582] The server generates speech content based on the input topic using a generative AI model (OpenAI API).
[2583] The server sends the generated message to the terminal, where the user can confirm and edit it.
[2584] Output: The generated utterance.
[2585] Step 5:
[2586] Confirmation and confirmation of statements
[2587] The user checks the generated comment and corrects it if necessary.
[2588] The user sends the finalized content of the message to the server.
[2589] Output: Confirmed statement.
[2590] Step 6:
[2591] Speech synthesis
[2592] The server passes the confirmed utterance content to a natural language processing engine (pyttsx3) and performs speech synthesis.
[2593] The server transmits the generated audio file to the terminal.
[2594] Output: The generated audio file.
[2595] Step 7:
[2596] Audio and Motion Mapping
[2597] The server receives the audio files and avatar data and creates motion sequences to generate lip sync and body language.
[2598] The server transmits the generated motion image to the terminal.
[2599] Output: The generated motion footage.
[2600] Step 8:
[2601] Subtitle creation and integration
[2602] The server automatically generates a subtitle file based on the text of the confirmed speech content.
[2603] The server integrates the subtitle file into the video and sends the final video to the terminal.
[2604] Output: A merged video file.
[2605] Step 9:
[2606] Use as a virtual assistant
[2607] The device (smartphone) uses the generated virtual assistant character to introduce products and respond to user inquiries in real time.
[2608] Users can use the virtual assistant to obtain product information and make inquiries.
[2609] Output: Product information and support provided to the user.
[2610] The above is a detailed flow of the processing steps of this system. By performing these operations, users can easily create a virtual assistant character and introduce products and respond in real time on their smartphones.
[2611] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2612] This invention provides a platform that utilizes AI to allow users to easily create virtual characters. By combining it with an emotion engine, it also has the ability to dynamically adjust the avatar's expressions based on the user's emotions. This system generates a 3D or Live2D avatar based on user input, synthesizes speech for the avatar, synchronizes speech and motion, and generates and integrates subtitles into the video. Additionally, by incorporating an emotion engine that recognizes the user's emotions, it dynamically changes facial expressions and movements, and adjusts the tone and pitch of the voice, based on the user's emotions.
[2613] System configuration
[2614] 1. 3D / Live2D avatar generation
[2615] The server receives basic status information (gender, age, style, etc.) entered by the user and activates an image generation AI module to generate a 3D or Live2D avatar.
[2616] On the terminal, the user customizes the details of the avatar (hairstyle, clothing, etc.), and finally finalizes the avatar and sends it to the server.
[2617] 2. Generating speech content
[2618] The server passes the topic information entered by the user to the LLM to generate the content of the comment, which is then sent to the terminal.
[2619] The user confirms and modifies the generated message and sends the finalized message to the server.
[2620] 3. Speech synthesis
[2621] The server synthesizes the confirmed speech content through a natural language processing engine and transmits the generated audio file to the terminal.
[2622] The user checks the audio file and requests corrections if necessary.
[2623] 4. Audio and Motion Mapping
[2624] The server receives the audio file and avatar data, creates a motion sequence for lip sync and body language generation, and then transmits the generated motion image to the device.
[2625] The user checks the motion image and requests corrections if necessary.
[2626] 5. Creating subtitles
[2627] The server automatically generates a subtitle file based on the text of the confirmed remarks, integrates it with the video, and sends it to the terminal.
[2628] The user checks the video and subtitles and confirms them after a final check.
[2629] 6. Leveraging Emotional Engines
[2630] The device uses an emotion engine to analyze the user's facial expressions, tone of voice, and gestures in real time to recognize the user's emotions.
[2631] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[2632] The server adjusts the tone and pitch of the generated speech synthesis according to the user's emotions.
[2633] Specific examples
[2634] Example 1: A user creates a video to express their emotions
[2635] 1. Create an avatar
[2636] The user clicks the "Create Avatar" button and enters the avatar's basic status (e.g., male, young man, casual style, etc.). The server generates an avatar based on the entered information and sends a preview to the device.
[2637] The user can view the preview, customize the details and finally confirm the avatar.
[2638] 2. Generating speech content
[2639] The user inputs the topic "Emotional Expression" as the theme. The server uses LLM to generate utterances based on the theme and sends them to the device.
[2640] The user checks and modifies the generated comment, finalizes it, and sends it to the server.
[2641] 3. Speech Synthesis
[2642] The server synthesizes voice based on the confirmed speech content and transmits the generated voice file to the terminal.
[2643] The user checks the audio and confirms it if there is no particular problem.
[2644] 4. Audio and Motion Mapping
[2645] The server generates lip sync and body language based on the audio file and the confirmed avatar, creates motion video, and sends it to the terminal.
[2646] The user checks the motion image and confirms it if there are no particular problems.
[2647] 5. Creating subtitles
[2648] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[2649] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[2650] 6. Leveraging Emotional Engines
[2651] The device senses the user's facial expressions and tone of voice in real time and analyzes them using an emotion engine.
[2652] The server dynamically adjusts lip sync, facial expressions, and body language based on the obtained emotional data, and appropriately changes the tone and pitch of the voice.
[2653] The user checks the video of the generated emotional expression, makes final adjustments and confirms it.
[2654] This concludes the detailed description of the "Mode for Carrying Out the Invention." This system allows users to easily generate high-quality virtual character content that reflects their own emotions, without requiring programming knowledge.
[2655] The processing flow will be explained below.
[2656] 1. Creating a 3D / Live2D avatar
[2657] Step 1:
[2658] The user clicks the "Create Avatar" button on the device.
[2659] Step 2:
[2660] The terminal displays a form for the user to input the basic status of the avatar (gender, age, style, etc.).
[2661] Step 3:
[2662] The user enters the required basic status information and clicks the "Generate" button.
[2663] Step 4:
[2664] The server receives the information entered by the user.
[2665] Step 5:
[2666] The server launches an image generation AI module and generates a 3D or Live2D avatar based on the input parameters.
[2667] Step 6:
[2668] The server sends a preview image of the generated avatar to the terminal.
[2669] Step 7:
[2670] The user checks the preview image displayed on the terminal.
[2671] Step 8:
[2672] The user makes detailed customizations (hairstyle, clothing, etc.) and clicks the "Confirm" button again.
[2673] Step 9:
[2674] The terminal transmits the final avatar information to the server.
[2675] 2. Generating speech content
[2676] Step 1:
[2677] The user clicks the "Generate Message" button on the terminal.
[2678] Step 2:
[2679] The terminal displays a form for the user to input the topic or theme of the comment.
[2680] Step 3:
[2681] The user enters a topic or theme and clicks the "Generate" button.
[2682] Step 4:
[2683] The server receives topic information input by the user.
[2684] Step 5:
[2685] The server launches the LLM and generates comments based on the input topic.
[2686] Step 6:
[2687] The server transmits the generated speech content to the terminal.
[2688] Step 7:
[2689] The user checks the message displayed on the terminal.
[2690] Step 8:
[2691] The user corrects the content of the statement as necessary and finally confirms the content of the statement.
[2692] Step 9:
[2693] The terminal transmits the confirmed statement content to the server.
[2694] 3. Speech synthesis
[2695] Step 1:
[2696] The user sends a request for speech synthesis based on the confirmed utterance content.
[2697] Step 2:
[2698] The server receives the confirmed statement content.
[2699] Step 3:
[2700] The server performs speech synthesis through a natural language processing engine.
[2701] Step 4:
[2702] The server transmits the generated audio file to the terminal.
[2703] Step 5:
[2704] The user plays the audio file on the device and checks the content.
[2705] Step 6:
[2706] If the user finds no problem with the audio, he / she confirms it and requests corrections if necessary.
[2707] 4. Audio and Motion Mapping
[2708] Step 1:
[2709] The user sends a request for audio and motion synchronization based on the confirmed audio file.
[2710] Step 2:
[2711] The server receives the audio file and avatar data.
[2712] Step 3:
[2713] The server generates motion sequences to achieve natural lip sync and body language.
[2714] Step 4:
[2715] The server transmits the completed motion image to the terminal.
[2716] Step 5:
[2717] The user plays and checks the generated motion video on the terminal.
[2718] Step 6:
[2719] The user requests corrections to movements and facial expressions as needed.
[2720] Step 7:
[2721] The user saves the finalized motion image.
[2722] 5. Creating subtitles
[2723] Step 1:
[2724] The user sends a request for subtitle generation based on the confirmed speech content and motion video.
[2725] Step 2:
[2726] The server receives the text of the confirmed utterance.
[2727] Step 3:
[2728] The server uses voice recognition technology to automatically generate subtitle files that correspond to what is being said.
[2729] Step 4:
[2730] The server overlays and integrates the subtitle file onto the motion video.
[2731] Step 5:
[2732] The server sends the completed video to the terminal.
[2733] Step 6:
[2734] The user checks the video and subtitles on the device.
[2735] Step 7:
[2736] The user saves the finalized video and publishes or shares it.
[2737] 6. Leveraging Emotional Engines
[2738] Step 1:
[2739] When a user plays back the video they created, the device's camera and microphone are used to detect emotional data such as facial expressions and tone of voice in real time.
[2740] Step 2:
[2741] The terminal transmits the emotion data acquired in real time to the emotion engine, and transmits the analysis results to the server.
[2742] Step 3:
[2743] The server dynamically adjusts the avatar's facial expressions and movements based on the emotion data obtained from the emotion engine.
[2744] Step 4:
[2745] When the server synthesizes voice from the generated speech, it adjusts the tone and pitch of the voice according to the user's emotions.
[2746] Step 5:
[2747] The server generates the final video based on the emotion data and sends it to the terminal.
[2748] Step 6:
[2749] The user checks the video in which the emotional expression is reflected and requests corrections as necessary.
[2750] Step 7:
[2751] The user saves the finalized video and publishes or shares it.
[2752] The above are the specific processing steps in the program processing of a system that utilizes an emotion engine.
[2753] Example 2
[2754] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2755] Conventional virtual character creation platforms make it difficult for users to express emotions or easily generate speech content. Synchronizing video and audio and generating subtitles is also time-consuming, making the overall process complicated and making it difficult to easily create high-quality content.
[2756] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for generating a 3D or 2D avatar based on information input by a user, a means for generating speech content based on an input theme, a means for voice-synthesizing the generated speech content, a means for synchronizing the speech file with the avatar's movement, a means for automatically generating subtitles based on the speech content and integrating them into video, and a means for collecting user emotion data using an emotion analysis engine and dynamically adjusting the avatar's expression. This allows a user to easily create high-quality virtual character content and generate video including real-time expressions that reflect the user's emotions.
[2757] "User" refers to the entity that uses this system to create a virtual character, customize it, generate comments, etc.
[2758] "Server" refers to a central processing unit that receives input information from a user and performs various processes.
[2759] "3D or 2D Avatar" means a virtual character created by a user and represented in three-dimensional or two-dimensional graphics.
[2760] "Input information" refers to data that a user provides to the system, such as basic status information such as gender, age, and style.
[2761] A "theme" refers to a topic that a user provides to the system for generating comment content.
[2762] "Speech content" refers to the lines and phrases of the character generated based on the input theme.
[2763] "Speech synthesis" is a technology that converts text data into voice data, and refers to a means of outputting the generated utterances as voice.
[2764] "Synchronizing" refers to the process of matching the audio and character movements.
[2765] "Subtitles" are texts that are automatically generated based on what is being said and are integrated into video.
[2766] "Integrating into video" refers to the process of incorporating the generated subtitles into a single video file in conjunction with the audio and avatar movements.
[2767] An "emotion analysis engine" refers to software or a system that analyzes a user's facial expressions, tone of voice, and gestures to recognize emotions.
[2768] "Emotion data" refers to data that indicates the user's emotional state as recognized by the emotion analysis engine.
[2769] "Dynamic adjustment" refers to the act of changing an avatar's facial expressions, movements, voice tone, pitch, etc. in real time.
[2770] "Customization options" refers to settings that allow users to change details about their avatar.
[2771] This invention is a system that uses AI to provide a platform that allows users to easily create virtual characters, and by combining it with an emotion analysis engine, it has the ability to dynamically adjust the avatar's expression according to the user's emotions.
[2772] System configuration
[2773] The system consists of the following main components:
[2774] 1. Server: A central processing unit that receives information from users and performs various processes.
[2775] 2. Terminal: A device operated by a user that provides an interface for user input, display, customization, etc.
[2776] 3. Sentiment analysis engine: Software for analyzing user emotions and generating data.
[2777] Technology used
[2778] 1. Image generation AI module (e.g. DeepArt): Generates a 3D or 2D avatar based on basic status information (gender, age, style, etc.) entered by the user.
[2779] 2. Large-scale language models (LLMs) (e.g., GPT-4): Generate utterances based on themes entered by the user.
[2780] 3. Natural language processing engine (e.g., Google Text-to-Speech): Converts the generated speech into audio.
[2781] 4. Emotion analysis engine (e.g., Affectiva SDK): Analyzes the user's facial expressions, tone of voice, and gestures to collect emotional data.
[2782] System Operation
[2783] The system operates in the following steps:
[2784] 1. Create your avatar:
[2785] The user clicks the "Create Avatar" button on the device and enters basic status information (gender, age, style, etc.).
[2786] The terminal transmits this information to the server.
[2787] The server uses an image generation AI module to generate a 3D or 2D avatar and sends a preview to the device.
[2788] The user checks the preview image, performs detailed customization (hairstyle, clothing, etc.), and confirms it.
[2789] The terminal transmits the determined avatar information to the server.
[2790] 2. Speech generation:
[2791] The user inputs topic information and sends it to the server.
[2792] The server uses LLM to generate speech content and sends it to the terminal.
[2793] The user confirms and corrects the content of the comment and sends the finalized content to the server.
[2794] 3. Speech synthesis:
[2795] The server synthesizes the confirmed speech content through a natural language processing engine.
[2796] The server transmits the generated audio file to the terminal.
[2797] The user checks the audio file and requests corrections if necessary.
[2798] 4. Audio and Motion Mapping:
[2799] The server passes the audio file and avatar data to a lip sync engine to generate lip sync and body language.
[2800] The server transmits the generated motion sequence to the terminal.
[2801] The user checks the motion image and requests corrections if necessary.
[2802] 5. Creating subtitles:
[2803] The server automatically generates subtitles based on the confirmed content of the remarks, integrates them with the video, and transmits them to the terminal.
[2804] The user checks the video and subtitles and confirms them after a final check.
[2805] 6. Utilizing sentiment analysis engines:
[2806] The device detects the user's facial expressions and tone of voice and passes them to an emotion analysis engine in real time.
[2807] The terminal transmits the emotion analysis results obtained to the server.
[2808] The server dynamically adjusts the avatar's facial expressions and movements based on the emotional data, and also appropriately changes the tone and pitch of the voice.
[2809] The user checks the final image, makes adjustments and confirms them.
[2810] Specific examples
[2811] Example 1: A user creates a video to express their emotions
[2812] 1. Create your avatar:
[2813] The user clicks the "Create Avatar" button and inputs the avatar's basic status (e.g., male, young man, casual style, etc.).
[2814] The terminal displays an input form and accepts input from the user.
[2815] When the user inputs information and presses the send button, the terminal sends the information to the server.
[2816] The server receives the information, activates the image generation AI module, and generates the avatar.
[2817] The server sends the generated avatar to the device as a preview.
[2818] The device displays a preview, and the user can customize the details (hairstyle, clothing, etc.) and finally confirm the avatar.
[2819] The terminal transmits the determined avatar information to the server.
[2820] 2. Speech generation:
[2821] The user opens the "Comment Content" input form and inputs the topic "Emotional Expressions."
[2822] The terminal transmits the topic information to the server.
[2823] The server generates the message using LLM and sends it to the terminal.
[2824] The user confirms and modifies the generated message and sends the finalized message to the server.
[2825] 3. Speech synthesis:
[2826] The server uses a natural language processing engine to synthesize speech based on the confirmed content of the speech.
[2827] The server transmits the generated audio file to the terminal.
[2828] The user checks the audio file and confirms it if there are no problems.
[2829] 4. Audio and Motion Mapping:
[2830] The server generates lip sync and body language based on the audio file and the determined avatar.
[2831] The server transmits the motion video to the terminal and provides it to the user.
[2832] The user checks the motion image and requests corrections if necessary.
[2833] 5. Creating subtitles:
[2834] The server automatically generates subtitles based on the confirmed remarks, integrates them into the video, and sends them to the terminal.
[2835] The user checks the final video and subtitles, corrects them as necessary, and finalizes them.
[2836] 6. Utilizing sentiment analysis engines:
[2837] The device senses the user's facial expressions and tone of voice in real time and sends the analysis results to the emotion engine.
[2838] The terminal transmits the emotion analysis result to the server.
[2839] The server dynamically adjusts lip sync, facial expressions, and movements based on emotional data, and also appropriately changes the tone and pitch of the voice.
[2840] The user checks the final image, makes adjustments and confirms them.
[2841] By following the above steps, users can easily create high-quality virtual character content. This system also enables real-time expression according to emotions, making it possible to generate a wide variety of content.
[2842] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2843] Step 1: User begins creating an avatar
[2844] The user clicks the "Create Avatar" button on the device.
[2845] Input: User input (basic status information such as gender, age, style, etc.)
[2846] The terminal displays an input form and accepts input from the user.
[2847] How it works: The user enters the required information and presses the submit button.
[2848] Output: The information entered
[2849] Step 2: The server generates the avatar
[2850] The server launches an image generation AI module based on the user's basic status information received from the device.
[2851] Input: Basic status information entered
[2852] How it works: The server uses an image generation AI module to generate a 3D or 2D avatar.
[2853] Output: A preview image of the generated avatar
[2854] Step 3: User customizes avatar
[2855] The terminal displays a preview of the generated avatar to the user.
[2856] Input: A preview image of the generated avatar
[2857] How it works: The user customizes the avatar's details (hairstyle, clothing, etc.) and finally finalizes the avatar.
[2858] Output: A customized confirmed avatar
[2859] Step 4: Enter and submit topic information
[2860] The user opens the "Comment Generation" input form on the terminal and inputs a topic (e.g., "About emotional expression").
[2861] Input: Topic information
[2862] How it works: The user enters a topic and presses the send button.
[2863] Output: Input topic information
[2864] Step 5: Generate speech
[2865] The server receives topic information and generates utterances using a large-scale language model (LLM).
[2866] Input: Topic information
[2867] How it works: The server passes the prompt to the generative AI model to generate the speech.
[2868] Output: Generated speech
[2869] Step 6: Revise and confirm your statement
[2870] The terminal displays the generated comment content to the user.
[2871] Input: Generated speech
[2872] Action: The user confirms and corrects what has been said, and finally confirms it.
[2873] Output: Confirmed statement
[2874] Step 7: Speech synthesis of what is being said
[2875] The server passes the confirmed utterance content to a natural language processing engine (e.g., Google Text-to-Speech) for speech synthesis.
[2876] Input: Confirmed statement
[2877] How it works: The server uses a speech synthesis engine to generate an audio file.
[2878] Output: Generated audio file
[2879] Step 8: Check and correct the audio file
[2880] The terminal provides the generated audio file to the user.
[2881] Input: Generated audio file
[2882] Action: The user reviews the audio file and requests corrections if necessary.
[2883] Output: Finalized audio file
[2884] Step 9: Mapping Audio and Motion
[2885] The server passes the audio file and the determined avatar to a lip sync engine to generate lip sync and body language.
[2886] Input: Audio file, confirmed avatar
[2887] How it works: The server generates lip sync and motion data.
[2888] Output: Generated motion sequence
[2889] Step 10: Check and correct motion footage
[2890] The terminal displays the generated motion image to the user.
[2891] Input: Generated motion sequence
[2892] Action: The user reviews the motion footage and requests corrections if necessary.
[2893] Output: Confirmed motion footage
[2894] Step 11: Creating and merging subtitles
[2895] The server automatically generates subtitles based on the confirmed remarks and integrates them into the video.
[2896] Input: Confirmed speech content, confirmed motion image
[2897] How it works: The server generates a subtitle file and integrates it into the video.
[2898] Output: Integrated video (audio, motion, subtitles)
[2899] Step 12: Reflecting the results of sentiment analysis
[2900] The device passes the user's facial expressions and tone of voice to an emotion analysis engine.
[2901] Input: Real-time facial expressions, tone of voice
[2902] Operation: The device obtains the emotion analysis results and sends them to the server.
[2903] Output: Emotion data
[2904] Step 13: Adjust based on sentiment data
[2905] The server dynamically adjusts the avatar's facial expression...
Claims
1. A platform that uses AI to allow users to easily create virtual characters. means for generating a 3D or Live2D avatar based on information input by a user; A means for generating a speech content based on an input theme; means for synthesizing the generated speech into voice; A means of synchronizing the audio file with the avatar's movements; A method for automatically generating subtitles based on the content of speech and integrating them into the video; A system including:
2. 10. The system of claim 1, further comprising means for providing a plurality of user-selectable customization options in generating the avatar.
3. 10. The system of claim 1, further comprising means for providing an administration tool that allows users to request modifications to the videos they generate.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A