System

The system addresses the challenge of high-quality character development and real-time dialogue in content production by using a generative AI model to create virtual characters with realistic traits and interactions, offering cost-effective and efficient content generation.

JP2026021017APending Publication Date: 2026-02-10SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024122699
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing content production systems using generative AI face challenges in providing high-quality, consistent character development and real-time professional dialogue and animation, which are costly and technically complex.

Method used

A system utilizing a generative AI model to create virtual characters with realistic traits, generating conversations and setting facial expressions and movements based on user requests, and enabling interactive responses in real time.

Benefits of technology

Enables high-quality, cost-effective interactive content generation with natural conversations and animations, meeting user demands through efficient use of generative AI.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021017000001_ABST
    Figure 2026021017000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for generating a virtual character having realistic and consistent character properties according to accumulated data and instructed behavior using a generative artificial intelligence model; means for receiving a content production request from a user; means for generating conversation content by the virtual character based on the received request; means for setting an expression and an action of the virtual character based on the generated conversation content; and means for delivering the set behavior of the virtual character to a user terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, content production using generative AI has expanded, but there is also a growing need for high-quality, consistent character development, rather than simply generating beautiful characters. Furthermore, providing professional dialogue and animation in real time in content production poses high costs and technical hurdles. Therefore, there is a need for a method that can provide high-quality interactive content that meets user demands while reducing the costs and risks associated with casting. [Means for solving the problem]

[0005] The present invention provides a system that uses a generative AI model to generate a virtual character with realistic and consistent character traits based on accumulated data and instructed behavior. The system includes a means for receiving a content production request from a user and generating a conversation for the virtual character based on the received request, and a means for setting the virtual character's facial expressions and movements based on the generated conversation. The system also includes a means for delivering the set virtual character behavior to a user device and generating interactive responses based on user input, thereby achieving a high-quality interactive experience in content creation. This enables the provision of professional conversations and animations at low cost and with low risk.

[0006] A "generative artificial intelligence model" is a model trained using machine learning techniques such as neural networks, and includes algorithms for generating natural, human-like conversations and behaviors.

[0007] A "virtual character" is a fictitious character created using digital technology and displayed through computer graphics or animation.

[0008] A "request" is a request or command sent by a user for particular content or service.

[0009] "Conversational content" refers to the verbal messages and dialogue uttered by the virtual character, and is generated by the generative AI model.

[0010] "Facial expression" refers to a virtual character's facial expression used to indicate emotion or reaction.

[0011] "Movement" refers to the body movement of a virtual character, including specific actions and gestures.

[0012] A "user terminal" refers to a device operated by a user, such as a computer, smartphone, or tablet.

[0013] "Interactive response" refers to the actions and conversations that a virtual character makes in real time in response to input or questions from a user.

[0014] An "external database" is a database that stores information outside the system, providing news data and other resources.

[0015] "API" stands for Application Programming Interface, an interface that enables data exchange and the use of functions between different software applications. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention relates to a system that generates virtual characters using a generative AI model and provides high-quality content desired by users in real time. Below, the program processing of this system is explained in natural language and explained in detail with specific examples.

[0038] 1. Initializing the VTuber generation AI model

[0039] The server loads pre-trained generative AI models, which form the basis for virtual characters to generate natural-sounding speech and behavior.

[0040] 2. Acceptance of Content Creation Requests

[0041] Users submit requests for newscasts and other content production to the system via a human interface, including the desired content description and specific requirements.

[0042] 3. Content-based conversation generation

[0043] The server analyzes the received request, retrieves the latest data from external databases and APIs, and uses the generated AI model to generate the lines and conversations of the virtual character.

[0044] As a specific example, if a user requests a "virtual character who hosts the latest technology news," the server will retrieve the latest technology news data from a news API and generate natural conversation using a generative AI model based on that data.

[0045] 4. Character behavior settings

[0046] The server then sets the virtual character's facial expressions and movements based on the generated dialogue, allowing the character to behave realistically to the viewer.

[0047] For example, when reading the news, a virtual character can show appropriate facial expressions to interesting topics and wave their hands to emphasize key points.

[0048] 5. Content Delivery

[0049] The server transmits the behavior of the virtual character to the user's device, which receives the data and displays it to the user in real time.

[0050] 6. Interactive Responses

[0051] Users can enter questions or comments while watching content. The server analyzes the input and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[0052] For example, if a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using an AI model. This response is then returned to the user by a virtual character.

[0053] In this way, the system provides natural conversations with high-quality virtual characters and can generate and deliver content desired by users in real time.

[0054] The processing flow will be explained below.

[0055] Step 1:

[0056] The server loads pre-trained generative AI models, which are the foundation of the system and allow virtual characters to generate natural-sounding conversations and movements.

[0057] Step 2:

[0058] Through the terminal interface, the user inputs a specific content creation request, which includes a content type, for example "news show host," and a specific topic, for example, the latest technology news.

[0059] Step 3:

[0060] The terminal transmits the request input by the user to the server, including the request content and related details.

[0061] Step 4:

[0062] The server analyzes the received request and extracts the necessary information, including the type and topic of the content.

[0063] Step 5:

[0064] The server retrieves the latest relevant data from external databases and APIs, for example connecting to a news API to provide the latest tech news.

[0065] Step 6:

[0066] The server uses a generative AI model based on the external data it retrieves to generate dialogue for the virtual character, resulting in natural and consistent dialogue.

[0067] Step 7:

[0068] The server then uses the generated dialogue to set the virtual character's facial expressions and movements, using a generative AI model to set the appropriate expressions and movements at the appropriate times.

[0069] Step 8:

[0070] The server structures the data of the virtual character and sends it to the user's device, including information on the content of the conversation, facial expressions, and movements.

[0071] Step 9:

[0072] The device passes the received data to a rendering engine, which displays the virtual character in real time, allowing the user to view the character's movements and conversations on the device.

[0073] Step 10:

[0074] Users can enter questions or comments about the content they are viewing, and the input is sent from the terminal to the server.

[0075] Step 11:

[0076] The server receives and analyzes input from the user, and generates an appropriate response based on the input using a generative AI model.

[0077] Step 12:

[0078] The server then structures the generated response and sends it back to the user's terminal, containing the new conversation content.

[0079] Step 13:

[0080] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[0081] Example 1

[0082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0083] Conventional virtual character generation systems require a significant amount of time and manual setup to enable characters to converse and move naturally. It is also difficult to support real-time data updates and rapid interactive responses from users. This results in a poor user experience and makes the system impractical.

[0084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0085] In this invention, the server includes means for loading a pre-trained generative artificial intelligence model, means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behavior, means for receiving a content creation request from a user, means for acquiring data from an external database or API based on the received request and generating conversation content for the virtual character based on the data, means for setting the facial expressions and movements of the virtual character based on the generated conversation content, and means for delivering the set behavior of the virtual character to a user terminal, thereby enabling the real-time generation and delivery of virtual characters with fast and natural conversation and movements.

[0086] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates natural conversations and actions based on pre-trained data.

[0087] A "virtual character" is a digitally generated fictional character that converses and behaves in a natural, human-like manner.

[0088] A "content creation request" is an instruction from a user to the system to create specific content.

[0089] "Means for setting facial expressions and movements" refers to the methods and technologies for determining and setting the facial expressions and body movements of a virtual character based on the generated conversation content.

[0090] An "external database or API" is a data repository outside the system, or an interface for exchanging information with another program or system.

[0091] "Interactive responses" refer to responses or reactions by a virtual character that are generated in real time based on input from a user.

[0092] "User terminal" refers to a device (e.g., smartphone, tablet, PC, etc.) that a user uses to access and operate the system.

[0093] This invention relates to a system that generates virtual characters using a generative AI model and provides users with high-quality content they desire in real time. Specific embodiments for implementing this system are described in detail below.

[0094] 1. Initializing the VTuber generation AI model

[0095] The server loads a pre-trained generative AI model (e.g., a Transformer model or a natural language processing model like GPT-3) that is required for the virtual character to generate natural-sounding conversations and behaviors. The server loads the model into memory and initializes the necessary pre-processing parameters and settings.

[0096] 2. Acceptance of Content Creation Requests

[0097] Users access the system through a browser or dedicated application and submit content creation requests, which include specific requirements and requirements (e.g., news type or topic segment).

[0098] Examples:

[0099] A user may submit a request saying, "I'd like a virtual character to host the latest technology news."

[0100] 3. Content-based conversation generation

[0101] The server analyzes user requests and retrieves relevant, up-to-date information from external databases and APIs (e.g., news APIs). Based on the retrieved data, the server provides prompts to a generative AI model, which generates lines and dialogue for the virtual character.

[0102] Example prompt sentence:

[0103] "Generate a virtual character to host the latest tech news."

[0104] 4. Character behavior settings

[0105] The server sets the virtual character's facial expressions and movements (e.g., hand gestures and eye movements) based on the generated lines. This is done using a 3D motion engine (e.g., Unity or Unreal Engine). The server analyzes the generated lines, determines facial expressions and movement patterns, and generates the character's movement data.

[0106] 5. Content Delivery

[0107] The server distributes the behavior data of the virtual character to the user's device, which receives the data and displays it to the user in real time using WebRTC and HTTP streaming technologies.

[0108] 6. Interactive Responses

[0109] Users can enter questions or comments while watching. The server receives this input, performs text analysis, and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[0110] Examples:

[0111] If a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using a generative AI model. This response is then returned to the user by a virtual character.

[0112] In this way, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content that users desire in real time.

[0113] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0114] Step 1: Initializing the VTuber generation AI model

[0115] The server loads a pre-trained generative AI model. Loading the model involves reading the model weight data from a specified directory. Specifically, it initializes the parameters and configuration values ​​of the generative AI model (e.g., GPT-3) into memory, making the model immediately available for use.

[0116] Input: Model weight data file

[0117] Output: Initialized generative AI model

[0118] Specific operation: The server loads the model weight data from the specified directory and initializes the necessary parameters.

[0119] Step 2: Accepting a content creation request

[0120] Users submit content creation requests using a browser or dedicated application. The request includes the type of content to be generated and any specific requirements. Once the request is sent, the server receives and analyzes it.

[0121] Input: Content production request from user

[0122] Output: Parsed request content

[0123] Specific operation: A user inputs a request using a web form or application interface and sends it. The server receives the request and analyzes the contents.

[0124] Step 3: Content-based conversation generation

[0125] Based on the analyzed request, the server retrieves the necessary data from external databases or APIs. For example, it retrieves the latest technology news from a news API. The retrieved data is converted into prompts and input into a generative AI model. The model then generates natural-sounding dialogue based on these prompts.

[0126] Input: Parsed request content, external data to be retrieved

[0127] Output: Generated dialogue

[0128] Specific operation: The server sends a data acquisition request to an external API, converts the obtained data into a prompt sentence, and inputs the prompt sentence into the generative AI model to generate dialogue.

[0129] Step 4: Setting character behavior

[0130] The server sets the virtual character's facial expressions and movements based on the generated lines. This is done using a 3D motion engine (such as Unity or Unreal Engine). The server determines facial expressions and body movements and generates them as data so that the lines and movements match.

[0131] Input: Generated dialogue

[0132] Output: Set facial expressions and movement data

[0133] Specific movements: The server analyzes the dialogue and determines facial expressions and movements. A 3D motion engine is used to generate movement data.

[0134] Step 5: Deliver your content

[0135] The server delivers the configured behavior data to the user device using WebRTC or HTTP streaming technology. The device decodes the received data in real time and displays it to the user.

[0136] Input: Set behavior data

[0137] Output: Display on the user's terminal

[0138] Specific operation: The server encodes behavioral data and sends it in streaming format. The device receives the data, decodes it in real time, and displays it.

[0139] Step 6: Interactive Response

[0140] Users can enter questions or comments while watching content. The server analyzes these inputs and uses a generative AI model to generate appropriate responses, which are then sent to the device and displayed to the user.

[0141] Input: User questions and comments

[0142] Output: The generated response

[0143] How it works: The user enters a question or comment into the chat box and sends it. The server analyzes the input and generates a response using the generative AI model. The response is sent to the device and displayed to the user.

[0144] Through these processing steps, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content desired by users in real time.

[0145] (Application example 1)

[0146] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0147] Current content delivery services make it difficult for users to obtain real-time updates and interact with virtual characters based on them. This limits the user experience and lacks dynamic and personalized content delivery. In particular, there is a demand for real-time conversation generation and responses based on news and trend information, but there is a lack of effective methods to achieve this.

[0148] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0149] In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative AI model according to accumulated data and instructed behaviors; means for receiving a content creation request from a user; means for generating conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for acquiring the latest data from an external API and generating conversation content for the virtual character using the generative AI model in response to prompts based on the data; and means for a user to interact with the virtual character via a smart device, thereby enabling a user to enjoy real-time interactive conversations based on the latest information.

[0150] A "generative artificial intelligence model" is an artificial intelligence algorithm that is trained using large datasets and is capable of generating new data and content, especially natural-sounding conversations and behaviors, under certain conditions.

[0151] A "virtual character" is a virtual personality or character that can behave dynamically and realistically on a digital device.

[0152] A "content creation request" is a request that a user sends to the system specifying the specific details and requirements of the content to be created.

[0153] "Conversational content" refers to words and phrases that a virtual character utters to interact with a user or provide information.

[0154] "Expressions and movements" refers to the facial expressions and body movements used by a virtual character to express emotions.

[0155] A "user terminal" is a device such as a smartphone, tablet, PC, or head-mounted display that allows a user to interact with virtual characters or receive content via the Internet.

[0156] An "external API" is a programmatic interface to external databases and services, and is a mechanism used by systems to obtain the latest data in real time.

[0157] A "prompt" is an input text given to a generative AI model, which serves as an instruction for the AI ​​to generate conversations and content based on that text.

[0158] "Smart devices" is a general term for intelligent electronic devices with internet connectivity, such as smartphones, tablets, and smartwatches.

[0159] "Interactive dialogue" refers to two-way communication in which a virtual character generates and responds to questions or comments entered by a user in real time.

[0160] The system for implementing the present invention comprises the following steps: Specific hardware and software usage methods will be described in detail below.

[0161] Program Overview

[0162] First, the server uses a generative AI model to generate the digital characters used on the client device. To do this, it loads a generative AI model that has been trained in advance using a huge dataset. The software used is Python and a generative AI model library (e.g., GPT-3).

[0163] Hardware and Software Details

[0164] Hardware:

[0165] server:

[0166] A server with a powerful processor (e.g., Intel Xeon), lots of memory (e.g., 64GB RAM), and a fast network connection.

[0167] User device:

[0168] Smartphone, head-mounted display (HMD), or computer.

[0169] software:

[0170] Python: Used for program development and operating AI models.

[0171] Generative AI models: Generative AI model libraries (e.g., GPT-3) are used to generate conversation content and character behavior.

[0172] API: Use external APIs (e.g. news API) to get the latest data.

[0173] Data processing and calculation

[0174] The server processes the data and performs calculations using the following procedure.

[0175] 1. Initializing the generative AI model

[0176] The server loads the generative AI model and serves as the foundation for the system, providing the foundation for the virtual characters to generate natural conversations and movements.

[0177] 2. Acceptance of Content Creation Requests

[0178] Users send content creation requests to the server through the human interface of their smart devices, which include specific requests for the virtual character.

[0179] 3. Acquiring external data

[0180] The server retrieves the latest data (e.g., the latest news) in real time through an external API, which is then used to generate content.

[0181] 4. Conversation generation

[0182] Based on the acquired data, a generative AI model is used to generate conversational content for the virtual character. A prompt is set, and the AI ​​model generates an appropriate response based on that prompt.

[0183] 5. Setting the facial expressions and movements of the virtual character

[0184] The facial expressions and movements of the virtual character are set based on the generated conversation content, resulting in more realistic and natural interactions.

[0185] 6. Content Delivery

[0186] The server transmits the behavior of the virtual character that has been set to the user's device, which receives it in real time and displays it to the user.

[0187] Specific examples

[0188] Example 1: News distribution

[0189] The user requests, "What is the latest technology news?"

[0190] The server retrieves the latest technology news data from the news API.

[0191] The generative AI model was given the prompt, "The user wants to know the latest technology news. News content: {latest news}."

[0192] The virtual character responds, "I'll report on the latest tech news."

[0193] The user asked, "Tell me more about this technology."

[0194] A virtual character follows and explains the details of the news.

[0195] Example prompt sentence:

[0196] "Users want to know the latest tech news. Please report on this news in detail."

[0197] This system allows users to enjoy real-time, interactive dialogue based on the latest information.

[0198] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0199] Step 1:

[0200] The server initializes the generative AI model and loads the trained model. As input, it requires the trained generative AI model, which prepares the virtual character to generate natural conversations and actions. As output, it obtains the initialized generative AI model.

[0201] Step 2:

[0202] A user uses a smart device to input a content creation request. The input includes the specific content desired by the user and their requirements. The smart device then sends the request to the server. The output is the received content creation request.

[0203] Step 3:

[0204] The server calls an external API to retrieve the latest data. The input requires the endpoint URL of the external API and authentication information. News data, trend data, etc. are retrieved from the external API. The output is the retrieved latest data.

[0205] Step 4:

[0206] The server generates a prompt sentence based on the acquired data. The latest data acquired from an external API is required as input. The prompt sentence to be given to the generative AI model is formed. The prompt sentence is obtained as output.

[0207] Step 5:

[0208] The server inputs the generated prompt sentence into the AI ​​model to generate the conversation content of the virtual character. The input requires the prompt sentence and the generation AI model. The generation AI model generates natural conversation content based on the prompt sentence. The generated conversation content is obtained as the output.

[0209] Step 6:

[0210] The server sets the facial expressions and movements of the virtual character based on the generated conversation content. The generated conversation content is required as input. Settings are made so that the virtual character will make facial expressions and movements according to the content. The set facial expressions and movements are obtained as output.

[0211] Step 7:

[0212] The server delivers the virtual character's behavior to the user's device. The inputs required are the user's facial expressions, movements, and generated conversations. The smart device receives this information and displays it to the user in real time. The output is the virtual character's behavior displayed on the user's device.

[0213] Step 8:

[0214] The user inputs questions and comments while viewing the content. The user's questions and comments are required as input. The device sends the input content to the server. The user's input content is obtained as output.

[0215] Step 9:

[0216] The server analyzes the input from the user and generates an appropriate response using a generative AI model. The input requires a user's question or comment. The generative AI model generates an appropriate response based on the input. The output is the generated response.

[0217] Step 10:

[0218] The server sends the generated response to the user's terminal, where the virtual character displays the response. The input requires the generated response, which the user terminal displays in real time. The output is the virtual character's response, which is displayed to the user.

[0219] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0220] This invention relates to a system that generates virtual characters using a generative AI model combined with an emotion engine, and provides high-quality, interactive content based on the user's emotions. Below, the program processing of this system is explained in natural language and explained in detail with specific examples.

[0221] Initializing the VTuber generation AI model

[0222] The server first loads a pre-trained generative AI model, which contains the basic algorithms that allow virtual characters to generate natural-sounding speech and movements.

[0223] Emotion engine initialization

[0224] The server also initializes an emotion engine to recognize the user's emotions. This emotion engine analyzes the user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[0225] Accepting content creation requests

[0226] Users input specific content creation requests through the device interface, specifying content types and topics, such as "I would like a host for the latest technology news."

[0227] Content-based conversation generation

[0228] The server receives and analyzes user requests. Based on the analysis results, it retrieves the latest necessary data from external databases and APIs. Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[0229] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[0230] Character behavior settings

[0231] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set the most appropriate facial expressions and movements according to the user's emotional state.

[0232] As a specific example, if the user is excited and happy, the character can be set to smile and move in a lively manner.

[0233] Content Delivery

[0234] The server transmits the behavior data of the configured virtual character to the user's device, which then renders the received data in real time and displays it to the user, allowing the user to watch and listen to the virtual character's conversations and movements.

[0235] Interactive Response

[0236] Users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion engine, the generative AI model generates an appropriate response.

[0237] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion engine determines the user's interest and level of attention, and the generative AI model responds with a positive comment along with appropriate detailed information.

[0238] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[0239] The processing flow will be explained below.

[0240] Step 1:

[0241] The server initializes a generative AI model, which is a pre-trained model that contains the basic algorithms for generating natural-sounding speech and behavior for virtual characters.

[0242] Step 2:

[0243] The server initializes an emotion engine for recognizing the user's emotions. This emotion engine has the function of analyzing the user's facial expressions, tone of voice, and input text information to recognize the user's emotional state in real time.

[0244] Step 3:

[0245] Users input content creation requests through their terminals, making specific requests such as "a virtual character that hosts the latest technology news."

[0246] Step 4:

[0247] The terminal transmits the content creation request input by the user to the server, and the transmitted data includes the type and topic of the content.

[0248] Step 5:

[0249] The server analyzes the request received from the user and retrieves the necessary data from external databases or APIs based on the analysis results. For example, it retrieves the latest technology news data from a news API.

[0250] Step 6:

[0251] The server uses the acquired data to generate conversational content for the virtual character using a generative AI model, ensuring that the conversational content is natural and consistent.

[0252] Step 7:

[0253] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set optimal facial expressions and movements according to the user's real-time emotions.

[0254] Step 8:

[0255] The server structures the data of the virtual character and transmits it to the user's device, including the content of the conversation, facial expressions, and movements.

[0256] Step 9:

[0257] The device passes the received data to a rendering engine to display the virtual character in real time, allowing the user to watch the character's conversations and movements.

[0258] Step 10:

[0259] The user can input questions or comments about the content being viewed. For example, the user can input a question such as "When is the next technical seminar?"

[0260] Step 11:

[0261] The device sends questions and comments from the user to the server, and the sent data includes the user's input.

[0262] Step 12:

[0263] The server receives and analyzes questions and comments from users. Based on the analysis results and information from the emotion engine, an appropriate response is generated using a generative AI model.

[0264] Step 13:

[0265] The server structures the generated response and sends it back to the user terminal, which contains the new conversation content.

[0266] Step 14:

[0267] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[0268] Through these steps, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly.

[0269] Example 2

[0270] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0271] In conventional virtual character systems, the characters' facial expressions and movements are fixed, making it difficult to realize interactive responses and behaviors that correspond to the user's emotions. Furthermore, because the latest information cannot be acquired in real time, the generated content is often out of date. This makes it difficult to maintain user satisfaction and interest, and limits the interactive experience they can provide.

[0272] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0273] In this invention, the server includes means for generating a realistic and consistent video character using a generative computer model in accordance with accumulated information and instructed actions, means for receiving a content creation request from a user, means for generating dialogue content for the video character based on the received request, means for setting facial expressions and actions of the video character based on the generated dialogue content, means for delivering the set behavior of the video character to a user device, means for initializing an emotion analysis engine and analyzing the emotional state of the user, and means for setting optimal facial expressions and actions based on the analysis results. This enables interactive response according to the user's emotions and allows the latest information to be acquired in real time, making it possible to provide high-quality and interesting content.

[0274] A "generative computer model" is a computer system that makes automated decisions and generates for a specific purpose based on stored information and instructed actions.

[0275] A "video character" is a computer-generated digital representation that has human-like appearance and behavior.

[0276] "User" means a person who uses this system to input content creation requests and view interactive content.

[0277] A "content creation request" is a request from a user to create content based on a particular topic or format.

[0278] "Dialogue content" refers to sentences and conversations spoken by a video character, generated based on user requests and external information.

[0279] "Expressions and actions" refer to facial changes and body movements that visually express emotions and intentions of a video character.

[0280] "Behavior data" is information that describes the facial expressions and movements of a video character.

[0281] A "user device" is an electronic device that a user uses to view content.

[0282] An "emotion analysis engine" is a system that analyzes data such as a user's facial expressions, tone of voice, and text to detect their emotional state.

[0283] "Emotional state" refers to the user's current emotional or psychological state.

[0284] This invention relates to a system that uses a generative computer model combined with an emotion analysis engine to generate realistic and consistent video characters and provide high-quality, interactive content based on the user's emotional state. Specific embodiments for implementing this system are described below.

[0285] System Overview

[0286] The system consists of three main components: a server, a terminal, and a user. The server runs the generative computational model and the sentiment analysis engine, and the terminal provides the interface between the user and the server.

[0287] Hardware and Software Configuration

[0288] server:

[0289] The server is a high-performance computer system that runs a generative computational model (e.g., a computational system model) and an emotion analysis engine (e.g., a facial expression analysis program and an emotional tone analysis program).

[0290] The server also includes an interface for connecting with external information sources and APIs (e.g., information retrieval programs).

[0291] Device:

[0292] A terminal is a device (e.g., a computer, smartphone, or tablet) used by a user and provides a user interface.

[0293] The terminal has a network connection for communicating with the server.

[0294] User:

[0295] Using this system, users can input content creation requests and engage in interactive dialogue with virtual characters.

[0296] Specific operation of the system

[0297] Initialization procedure for high performance computer systems:

[0298] The server loads pre-trained generative computer models that allow video characters to generate natural-sounding speech and movements.

[0299] At the same time, the emotion analysis engine is initialized to prepare for analyzing the user's emotional state in real time.

[0300] Steps for users to enter their request:

[0301] Through the interface, users input specific content creation requests, for example, "I would like to host a segment on the latest technology news."

[0302] Other example prompts include "Tell me about the next tech seminar" and "Tell me some interesting technology topics."

[0303] Server data acquisition and analysis procedure:

[0304] The server receives and analyzes requests from users, and based on the analysis results, retrieves the latest data needed from external sources or APIs.

[0305] For example, in the case of technology news, the latest related technology news data is acquired.

[0306] Steps for generating and configuring conversation content:

[0307] The server uses a generative computer model based on the acquired data to generate dialogue content for the video character.

[0308] Based on the generated dialogue, the emotion analysis engine analyzes the user's emotional state and sets the most appropriate facial expressions and actions.

[0309] For example, if the user is excited, the video character is set to have an excited expression and movements.

[0310] To distribute your content:

[0311] The server transmits the behavior data of the set video character to the terminal in real time.

[0312] The device renders the received data in real time and displays it to the user, allowing the user to view interactive content.

[0313] Interactive response steps:

[0314] Users can enter questions or comments about the content they are viewing.

[0315] The terminal sends the input to the server, which again analyzes the user's emotional state using an emotion analysis engine and generates an appropriate response in a generative computer model.

[0316] The server generates a response and sends it back to the terminal, allowing the user to enjoy an interactive dialogue.

[0317] In this way, the system can recognize users' emotions in real time and provide interactive experiences that respond to them. Based on specific scenarios and usage patterns, users can generate and enjoy high-quality content.

[0318] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0319] Step 1:

[0320] Loading a generative AI model

[0321] The server loads a pre-trained generative AI model. Specifically, the server loads the generative AI model (e.g., a computational system model) from storage into memory and initializes it. The input is the model file on storage, and the output is the generative AI model launched in memory.

[0322] Step 2:

[0323] Initializing the sentiment analysis engine

[0324] The server initializes the emotion analysis engine, which prepares it to analyze the user's emotional state in real time. Specifically, the server starts the emotion analysis engine (e.g., a facial expression analysis program or an emotional tone analysis program) and checks its operation using test data. The input is the emotion analysis engine program, and the output is the started emotion analysis engine.

[0325] Step 3:

[0326] Accepting content creation requests

[0327] A user inputs a specific content creation request through the interface. Specifically, the user enters the request text (e.g., "I would like you to host a segment on the latest technology news") into the device's input field and clicks the send button. The input is the request text entered by the user, and the output is the request data sent to the server.

[0328] Step 4:

[0329] Request analysis and data acquisition

[0330] The server analyzes requests received from users and retrieves the necessary data from relevant external information sources or APIs. Specifically, the server performs natural language analysis on the request content, generates an API request, sends it to an external database or API (e.g., an information retrieval program), and receives the retrieved data. The input is the user's request content, and the output is the data retrieved from the external database or API.

[0331] Step 5:

[0332] Conversation generation

[0333] The server uses a generative AI model based on the acquired data to generate dialogue for the video character. Specifically, the server inputs the acquired data into the generative AI model and outputs the generated dialogue in text format. The input is data acquired from an external database or API, and the output is the text data of the generated dialogue.

[0334] Step 6:

[0335] Character facial expressions and movement settings

[0336] Based on the generated conversation content, the server uses an emotion analysis engine to analyze the user's emotional state and set the optimal facial expressions and movements for the video character. Specifically, the server inputs the generated conversation content and the user's real-time emotional data, which the emotion analysis engine analyzes to generate the character's facial expressions and movement data. The input is the generated conversation content and the user's emotional data, and the output is the character's facial expressions and movement data.

[0337] Step 7:

[0338] Sending behavioral data

[0339] The server transmits the behavior data of the configured video character to the terminal in real time. Specifically, it executes a process to transmit the generated behavior data to the terminal via the network. The input is the character's facial expression and movement data, and the output is the behavior data transmitted to the terminal.

[0340] Step 8:

[0341] Rendering and Displaying Data

[0342] The device renders the received behavior data in real time and displays it to the user. Specifically, the device processes the received data using a 3D model and animation engine, and displays a visualized character to the user. The input is the received behavior data, and the output is the video character displayed to the user.

[0343] Step 9:

[0344] Accepting User Input

[0345] Users can enter questions or comments about the content they are viewing. Specifically, the user enters the question or comment in text in the input field on the device and clicks the send button. The input is the question or comment entered by the user, and the output is the input data sent to the server.

[0346] Step 10:

[0347] Interactive response generation and delivery

[0348] The server analyzes the input data received from the user, re-analyzes the user's emotional state using the emotion analysis engine, and generates an appropriate response using the generative AI model. The generated response is sent back to the device, which then displays it to the user. Specifically, the server analyzes the input data, generates a response based on the generative AI model, and the emotion analysis engine complements the response. The input is a question or comment from the user, and the output is an interactive response sent to the device.

[0349] This allows the system to provide a high quality and interactive experience to the user.

[0350] (Application example 2)

[0351] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0352] Conventional virtual character systems provide content in a one-way manner without considering the user's emotional state. This limits the improvement of the user experience and lacks true interactivity. Furthermore, they lack the ability to respond in real time to user input requests, making it difficult to respond flexibly to the user's interests and emotions.

[0353] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative artificial intelligence model in accordance with accumulated data and instructed behavior; means for receiving a content creation request from a user; means for generating a conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for an emotion recognition engine for recognizing the emotional state of the user in real time; and means for optimizing the behavior of the virtual character based on the recognized emotional state. This makes it possible to recognize the emotional state of the user and provide high-quality, interactive content that is customized accordingly.

[0354] A "generative artificial intelligence model" is an algorithm or program that generates the conversation content and behavior of a virtual character based on accumulated data and instructed behavior.

[0355] A "virtual character" is a character that has a realistic and consistent personality in a digital space and can interact with users.

[0356] A "content creation request" is a request or instruction a user makes to a virtual character regarding a particular topic or content.

[0357] An "emotion recognition engine" is a technology that analyzes a user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[0358] An "interactive response" is a response generated by a virtual character to respond in real time to a user's questions or comments and to continue the dialogue with the user.

[0359] "External databases and APIs" means external sources of information that a virtual character uses to obtain up-to-date data for generating dialogue.

[0360] "Behavior" refers to the facial expressions and actions that a virtual character makes toward the user.

[0361] This invention relates to a system that provides high-quality, interactive content based on user emotions using a generative AI model combined with an emotion recognition engine. The program processing of this system is described below.

[0362] The server first loads a pre-trained generative artificial intelligence model, which contains the basic algorithms that allow a virtual character to generate natural-sounding conversations and movements.

[0363] The emotion recognition engine is also initialized at the same time. The emotion recognition engine is a system that analyzes the user's facial expressions, tone of voice, input text, etc. to recognize their emotional state in real time. To identify emotions, it uses technologies such as OpenCV and deep learning models (e.g., TensorFlow, PyTorch).

[0364] Users input specific content creation requests through a smartphone interface, specifying the type of content and the topic, such as "I want you to host the latest technology news."

[0365] The server receives and analyzes the user's request. Based on the analysis results, it retrieves the latest necessary data from an external database, such as a news API (e.g., Google News API). Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[0366] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[0367] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion recognition engine to set the most appropriate facial expressions and movements according to the user's emotional state, and then transmits the set virtual character's behavior data to the user's device.

[0368] The device renders the received data in real time and displays it to the user, allowing the user to see and hear the virtual character's conversations and movements.

[0369] Additionally, users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion recognition engine, the generative AI model generates an appropriate response.

[0370] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion recognition engine determines the user's level of interest, and the generative AI model responds with a positive comment along with appropriate details.

[0371] Prompt Sentence Examples

[0372] User Input: "What's in tech news today?"

[0373] Generated AI prompt:

[0374] A user has requested, "What's in tech news today?" Use the news API to get the latest tech news and generate energetic and positive conversations when users are excited.

[0375] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[0376] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0377] Step 1:

[0378] The server loads and initializes a pre-trained generative artificial intelligence model, taking the parameter data of the pre-trained generative AI model as input and obtaining the prepared state of the AI ​​model as output, which will later generate a virtual character.

[0379] Step 2:

[0380] The server initializes the emotion recognition engine, which takes as input the configuration file of the emotion recognition algorithm and as output is ready to analyze the user's emotional state in real time, so that it can process the emotional data from the user.

[0381] Step 3:

[0382] Users input content creation requests through a smartphone interface. The input includes the type of content and topic desired by the user, and the output provides detailed data about the request, which can then be processed.

[0383] Step 4:

[0384] The server parses the request received from the user, taking the user request data as input and providing the structured data of the parsed request as output, which allows the retrieval of the data needed for the next processing step.

[0385] Step 5:

[0386] The server retrieves the latest data related to the requested content from an external database or API. The input is the structured data of the parsed request, and the output is the latest retrieved data. This allows the server to provide information that meets the latest requests. A specific example is using a news API to retrieve the latest technology news.

[0387] Step 6:

[0388] The server uses a generative AI model based on the latest acquired data to generate the conversation content of the virtual character. The input is news data and other external data, and the output is the generated conversation content data. This gives shape to the dialogue content of the virtual character.

[0389] Step 7:

[0390] The server sets the virtual character's facial expressions and movements based on the generated conversation content. The inputs are conversation content data and user emotion data obtained from an emotion recognition engine, and the output is movement setting data for the virtual character. This allows the character to behave naturally according to the user's emotions.

[0391] Step 8:

[0392] The server sends the behavior data of the configured virtual character to the user's device. The input is the virtual character's action setting data, and the output is the data sent to the user's device. This allows the user to watch the character's conversation and actions.

[0393] Step 9:

[0394] The device renders the received data in real time and displays it to the user. The input is behavior data sent from the server, and the output is a rendered virtual character. This allows the user to enjoy interactive dialogue with the character.

[0395] Step 10:

[0396] The user enters a question or comment about the content they are viewing. The input is the user's text input, and the output is the text data. This is used to generate the next response.

[0397] Step 11:

[0398] The terminal sends the user's input to the server. The input is the user's comment data, and the output is the data transmission to the server, so that the server can prepare an interactive response.

[0399] Step 12:

[0400] The server generates appropriate responses using a generative AI model based on user input and emotion recognition data. The inputs are user comment data and current emotion data, and the output is interactive response data. This allows the virtual character to respond appropriately in real time.

[0401] In this way, by performing each processing step in detail, high-quality interactive virtual character dialogue optimized for the user's emotions is realized.

[0402] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0403] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0404] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0405] [Second embodiment]

[0406] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0407] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0408] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0409] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0410] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0411] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0412] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0413] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0414] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0415] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0416] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0417] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0418] This invention relates to a system that generates virtual characters using a generative AI model and provides high-quality content desired by users in real time. Below, the program processing of this system is explained in natural language and explained in detail with specific examples.

[0419] 1. Initializing the VTuber generation AI model

[0420] The server loads pre-trained generative AI models, which form the basis for virtual characters to generate natural-sounding speech and behavior.

[0421] 2. Acceptance of Content Creation Requests

[0422] Users submit requests for newscasts and other content production to the system via a human interface, including the desired content description and specific requirements.

[0423] 3. Content-based conversation generation

[0424] The server analyzes the received request, retrieves the latest data from external databases and APIs, and uses the generated AI model to generate the lines and conversations of the virtual character.

[0425] As a specific example, if a user requests a "virtual character who hosts the latest technology news," the server will retrieve the latest technology news data from a news API and generate natural conversation using a generative AI model based on that data.

[0426] 4. Character behavior settings

[0427] The server then sets the virtual character's facial expressions and movements based on the generated dialogue, allowing the character to behave realistically to the viewer.

[0428] For example, when reading the news, a virtual character can show appropriate facial expressions to interesting topics and wave their hands to emphasize key points.

[0429] 5. Content Delivery

[0430] The server transmits the behavior of the virtual character to the user's device, which receives the data and displays it to the user in real time.

[0431] 6. Interactive Responses

[0432] Users can enter questions or comments while watching content. The server analyzes the input and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[0433] For example, if a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using an AI model. This response is then returned to the user by a virtual character.

[0434] In this way, the system provides natural conversations with high-quality virtual characters and can generate and deliver content desired by users in real time.

[0435] The processing flow will be explained below.

[0436] Step 1:

[0437] The server loads pre-trained generative AI models, which are the foundation of the system and allow virtual characters to generate natural-sounding conversations and movements.

[0438] Step 2:

[0439] Through the terminal interface, the user inputs a specific content creation request, which includes a content type, for example "news show host," and a specific topic, for example, the latest technology news.

[0440] Step 3:

[0441] The terminal transmits the request input by the user to the server, including the request content and related details.

[0442] Step 4:

[0443] The server analyzes the received request and extracts the necessary information, including the type and topic of the content.

[0444] Step 5:

[0445] The server retrieves the latest relevant data from external databases and APIs, for example connecting to a news API to provide the latest tech news.

[0446] Step 6:

[0447] The server uses a generative AI model based on the external data it retrieves to generate dialogue for the virtual character, resulting in natural and consistent dialogue.

[0448] Step 7:

[0449] The server then uses the generated dialogue to set the virtual character's facial expressions and movements, using a generative AI model to set the appropriate expressions and movements at the appropriate times.

[0450] Step 8:

[0451] The server structures the data of the virtual character and sends it to the user's device, including information on the content of the conversation, facial expressions, and movements.

[0452] Step 9:

[0453] The device passes the received data to a rendering engine, which displays the virtual character in real time, allowing the user to view the character's movements and conversations on the device.

[0454] Step 10:

[0455] Users can enter questions or comments about the content they are viewing, and the input is sent from the terminal to the server.

[0456] Step 11:

[0457] The server receives and analyzes input from the user, and generates an appropriate response based on the input using a generative AI model.

[0458] Step 12:

[0459] The server then structures the generated response and sends it back to the user's terminal, containing the new conversation content.

[0460] Step 13:

[0461] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[0462] Example 1

[0463] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0464] Conventional virtual character generation systems require a significant amount of time and manual setup to enable characters to converse and move naturally. It is also difficult to support real-time data updates and rapid interactive responses from users. This results in a poor user experience and makes the system impractical.

[0465] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0466] In this invention, the server includes means for loading a pre-trained generative artificial intelligence model, means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behavior, means for receiving a content creation request from a user, means for acquiring data from an external database or API based on the received request and generating conversation content for the virtual character based on the data, means for setting the facial expressions and movements of the virtual character based on the generated conversation content, and means for delivering the set behavior of the virtual character to a user terminal, thereby enabling the real-time generation and delivery of virtual characters with fast and natural conversation and movements.

[0467] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates natural conversations and actions based on pre-trained data.

[0468] A "virtual character" is a digitally generated fictional character that converses and behaves in a natural, human-like manner.

[0469] A "content creation request" is an instruction from a user to the system to create specific content.

[0470] "Means for setting facial expressions and movements" refers to the methods and technologies for determining and setting the facial expressions and body movements of a virtual character based on the generated conversation content.

[0471] An "external database or API" is a data repository outside the system, or an interface for exchanging information with another program or system.

[0472] "Interactive responses" refer to responses or reactions by a virtual character that are generated in real time based on input from a user.

[0473] "User terminal" refers to a device (e.g., smartphone, tablet, PC, etc.) that a user uses to access and operate the system.

[0474] This invention relates to a system that generates virtual characters using a generative AI model and provides users with high-quality content they desire in real time. Specific embodiments for implementing this system are described in detail below.

[0475] 1. Initializing the VTuber generation AI model

[0476] The server loads a pre-trained generative AI model (e.g., a Transformer model or a natural language processing model like GPT-3) that is required for the virtual character to generate natural-sounding speech and behavior. The server loads the model into memory and initializes the necessary pre-processing parameters and configuration values.

[0477] 2. Acceptance of Content Creation Requests

[0478] Users access the system through a browser or dedicated application and submit content creation requests, which include specific requirements and requirements (e.g., news type or topic segment).

[0479] Examples:

[0480] A user may submit a request saying, "I'd like a virtual character to host the latest technology news."

[0481] 3. Content-based conversation generation

[0482] The server analyzes user requests and retrieves relevant, up-to-date information from external databases and APIs (e.g., news APIs). Based on the retrieved data, the server provides prompts to a generative AI model, which generates lines and dialogue for the virtual character.

[0483] Example prompt sentence:

[0484] "Generate a virtual character to host the latest tech news."

[0485] 4. Character behavior settings

[0486] The server sets the virtual character's facial expressions and movements (e.g., hand gestures and eye movements) based on the generated lines. This is done using a 3D motion engine (e.g., Unity or Unreal Engine). The server analyzes the generated lines, determines facial expressions and movement patterns, and generates the character's movement data.

[0487] 5. Content Delivery

[0488] The server distributes the behavior data of the virtual character to the user's device, which receives the data and displays it to the user in real time using WebRTC and HTTP streaming technologies.

[0489] 6. Interactive Responses

[0490] Users can enter questions or comments while watching. The server receives this input, performs text analysis, and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[0491] Examples:

[0492] If a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using a generative AI model. This response is then returned to the user by a virtual character.

[0493] In this way, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content that users desire in real time.

[0494] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0495] Step 1: Initializing the VTuber generation AI model

[0496] The server loads a pre-trained generative AI model. Loading the model involves reading the model weight data from a specified directory. Specifically, it initializes the parameters and configuration values ​​of the generative AI model (e.g., GPT-3) into memory, making the model immediately available for use.

[0497] Input: Model weight data file

[0498] Output: Initialized generative AI model

[0499] Specific operation: The server loads the model weight data from the specified directory and initializes the necessary parameters.

[0500] Step 2: Accepting a content creation request

[0501] Users submit content creation requests using a browser or dedicated application. The request includes the type of content to be generated and any specific requirements. Once the request is sent, the server receives and analyzes it.

[0502] Input: Content production request from user

[0503] Output: Parsed request content

[0504] Specific operation: A user inputs a request using a web form or application interface and sends it. The server receives the request and analyzes the contents.

[0505] Step 3: Content-based conversation generation

[0506] Based on the analyzed request, the server retrieves the necessary data from external databases or APIs. For example, it retrieves the latest technology news from a news API. The retrieved data is converted into prompts and input into a generative AI model. The model then generates natural-sounding dialogue based on these prompts.

[0507] Input: Parsed request content, external data to be retrieved

[0508] Output: Generated dialogue

[0509] Specific operation: The server sends a data acquisition request to an external API, converts the obtained data into a prompt sentence, and inputs the prompt sentence into the generative AI model to generate dialogue.

[0510] Step 4: Setting character behavior

[0511] The server sets the virtual character's facial expressions and movements based on the generated lines. This is done using a 3D motion engine (such as Unity or Unreal Engine). The server determines facial expressions and body movements and generates them as data so that the lines and movements match.

[0512] Input: Generated dialogue

[0513] Output: Set facial expressions and movement data

[0514] Specific movements: The server analyzes the dialogue and determines facial expressions and movements. A 3D motion engine is used to generate movement data.

[0515] Step 5: Deliver your content

[0516] The server delivers the configured behavior data to the user device using WebRTC or HTTP streaming technology. The device decodes the received data in real time and displays it to the user.

[0517] Input: Set behavior data

[0518] Output: Display on the user's terminal

[0519] Specific operation: The server encodes behavioral data and sends it in streaming format. The device receives the data, decodes it in real time, and displays it.

[0520] Step 6: Interactive Response

[0521] Users can enter questions or comments while watching content. The server analyzes these inputs and uses a generative AI model to generate appropriate responses, which are then sent to the device and displayed to the user.

[0522] Input: User questions and comments

[0523] Output: The generated response

[0524] How it works: The user enters a question or comment into the chat box and sends it. The server analyzes the input and generates a response using the generative AI model. The response is sent to the device and displayed to the user.

[0525] Through these processing steps, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content desired by users in real time.

[0526] (Application example 1)

[0527] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0528] Current content delivery services make it difficult for users to obtain real-time updates and interact with virtual characters based on them. This limits the user experience and lacks dynamic and personalized content delivery. In particular, there is a demand for real-time conversation generation and responses based on news and trend information, but there is a lack of effective methods to achieve this.

[0529] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0530] In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative AI model according to accumulated data and instructed behaviors; means for receiving a content creation request from a user; means for generating conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for acquiring the latest data from an external API and generating conversation content for the virtual character using the generative AI model in response to prompts based on the data; and means for a user to interact with the virtual character via a smart device, thereby enabling a user to enjoy real-time interactive conversations based on the latest information.

[0531] A "generative artificial intelligence model" is an artificial intelligence algorithm that is trained using large datasets and is capable of generating new data and content, especially natural-sounding conversations and behaviors, under certain conditions.

[0532] A "virtual character" is a virtual personality or character that can behave dynamically and realistically on a digital device.

[0533] A "content creation request" is a request that a user sends to the system specifying the specific details and requirements of the content to be created.

[0534] "Conversational content" refers to words and phrases that a virtual character utters to interact with a user or provide information.

[0535] "Expressions and movements" refers to the facial expressions and body movements used by a virtual character to express emotions.

[0536] A "user terminal" is a device such as a smartphone, tablet, PC, or head-mounted display that allows a user to interact with virtual characters or receive content via the Internet.

[0537] An "external API" is a programmatic interface to external databases and services, and is a mechanism used by systems to obtain the latest data in real time.

[0538] A "prompt" is an input text given to a generative AI model, which serves as an instruction for the AI ​​to generate conversations and content based on that text.

[0539] "Smart devices" is a general term for intelligent electronic devices with internet connectivity, such as smartphones, tablets, and smartwatches.

[0540] "Interactive dialogue" refers to two-way communication in which a virtual character generates and responds to questions or comments entered by a user in real time.

[0541] The system for implementing the present invention comprises the following steps: Specific hardware and software usage methods will be described in detail below.

[0542] Program Overview

[0543] First, the server uses a generative AI model to generate the digital characters used on the client device. To do this, it loads a generative AI model that has been trained in advance using a huge dataset. The software used is Python and a generative AI model library (e.g., GPT-3).

[0544] Hardware and Software Details

[0545] Hardware:

[0546] server:

[0547] A server with a powerful processor (e.g., Intel Xeon), lots of memory (e.g., 64GB RAM), and a fast network connection.

[0548] User device:

[0549] Smartphone, head-mounted display (HMD), or computer.

[0550] software:

[0551] Python: Used for program development and operating AI models.

[0552] Generative AI models: Generative AI model libraries (e.g., GPT-3) are used to generate conversation content and character behavior.

[0553] API: Use external APIs (e.g. news API) to get the latest data.

[0554] Data processing and calculation

[0555] The server processes the data and performs calculations using the following procedure.

[0556] 1. Initializing the generative AI model

[0557] The server loads the generative AI model and serves as the foundation for the system, providing the foundation for the virtual characters to generate natural conversations and movements.

[0558] 2. Acceptance of Content Creation Requests

[0559] Users send content creation requests to the server through the human interface of their smart devices, which include specific requests for the virtual character.

[0560] 3. Acquiring external data

[0561] The server retrieves the latest data (e.g., the latest news) in real time through an external API, which is then used to generate content.

[0562] 4. Conversation generation

[0563] Based on the acquired data, a generative AI model is used to generate conversational content for the virtual character. A prompt is set, and the AI ​​model generates an appropriate response based on that prompt.

[0564] 5. Setting the facial expressions and movements of the virtual character

[0565] The facial expressions and movements of the virtual character are set based on the generated conversation content, resulting in more realistic and natural interactions.

[0566] 6. Content Delivery

[0567] The server transmits the behavior of the virtual character that has been set to the user's device, which receives it in real time and displays it to the user.

[0568] Specific examples

[0569] Example 1: News distribution

[0570] The user requests, "What is the latest technology news?"

[0571] The server retrieves the latest technology news data from the news API.

[0572] The generative AI model was given the prompt, "The user wants to know the latest technology news. News content: {latest news}."

[0573] The virtual character responds, "I'll report on the latest tech news."

[0574] The user asked, "Tell me more about this technology."

[0575] A virtual character follows and explains the details of the news.

[0576] Example prompt sentence:

[0577] "Users want to know the latest tech news. Please report on this news in detail."

[0578] This system allows users to enjoy real-time, interactive dialogue based on the latest information.

[0579] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0580] Step 1:

[0581] The server initializes the generative AI model and loads the trained model. As input, it requires the trained generative AI model, which prepares the virtual character to generate natural conversations and actions. As output, it obtains the initialized generative AI model.

[0582] Step 2:

[0583] A user uses a smart device to input a content creation request. The input includes the specific content desired by the user and their requirements. The smart device then sends the request to the server. The output is the received content creation request.

[0584] Step 3:

[0585] The server calls an external API to retrieve the latest data. The input requires the endpoint URL of the external API and authentication information. News data, trend data, etc. are retrieved from the external API. The output is the retrieved latest data.

[0586] Step 4:

[0587] The server generates a prompt sentence based on the acquired data. The latest data acquired from an external API is required as input. The prompt sentence to be given to the generative AI model is formed. The prompt sentence is obtained as output.

[0588] Step 5:

[0589] The server inputs the generated prompt sentence into the AI ​​model to generate the conversation content of the virtual character. The input requires the prompt sentence and the generation AI model. The generation AI model generates natural conversation content based on the prompt sentence. The generated conversation content is obtained as the output.

[0590] Step 6:

[0591] The server sets the facial expressions and movements of the virtual character based on the generated conversation content. The generated conversation content is required as input. Settings are made so that the virtual character will make facial expressions and movements according to the content. The set facial expressions and movements are obtained as output.

[0592] Step 7:

[0593] The server delivers the virtual character's behavior to the user's device. The inputs required are the user's facial expressions, movements, and generated conversations. The smart device receives this information and displays it to the user in real time. The output is the virtual character's behavior displayed on the user's device.

[0594] Step 8:

[0595] The user inputs questions and comments while viewing the content. The user's questions and comments are required as input. The device sends the input content to the server. The user's input content is obtained as output.

[0596] Step 9:

[0597] The server analyzes the input from the user and generates an appropriate response using a generative AI model. The input requires a user's question or comment. The generative AI model generates an appropriate response based on the input. The output is the generated response.

[0598] Step 10:

[0599] The server sends the generated response to the user's terminal, where the virtual character displays the response. The input requires the generated response, which the user terminal displays in real time. The output is the virtual character's response, which is displayed to the user.

[0600] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0601] This invention relates to a system that generates virtual characters using a generative AI model combined with an emotion engine, and provides high-quality, interactive content based on the user's emotions. Below, the program processing of this system is explained in natural language and explained in detail with specific examples.

[0602] Initializing the VTuber generation AI model

[0603] The server first loads a pre-trained generative AI model, which contains the basic algorithms that allow virtual characters to generate natural-sounding speech and movements.

[0604] Emotion engine initialization

[0605] The server also initializes an emotion engine to recognize the user's emotions. This emotion engine analyzes the user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[0606] Accepting content creation requests

[0607] Users input specific content creation requests through the device interface, specifying content types and topics, such as "I would like a host for the latest technology news."

[0608] Content-based conversation generation

[0609] The server receives and analyzes user requests. Based on the analysis results, it retrieves the latest necessary data from external databases and APIs. Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[0610] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[0611] Character behavior settings

[0612] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set the most appropriate facial expressions and movements according to the user's emotional state.

[0613] As a specific example, if the user is excited and happy, the character can be set to smile and move in a lively manner.

[0614] Content Delivery

[0615] The server transmits the behavior data of the configured virtual character to the user's device, which then renders the received data in real time and displays it to the user, allowing the user to watch and listen to the virtual character's conversations and movements.

[0616] Interactive Response

[0617] Users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion engine, the generative AI model generates an appropriate response.

[0618] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion engine determines the user's interest and level of attention, and the generative AI model responds with a positive comment along with appropriate detailed information.

[0619] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[0620] The processing flow will be explained below.

[0621] Step 1:

[0622] The server initializes a generative AI model, which is a pre-trained model that contains the basic algorithms for generating natural-sounding speech and behavior for virtual characters.

[0623] Step 2:

[0624] The server initializes an emotion engine for recognizing the user's emotions. This emotion engine has the function of analyzing the user's facial expressions, tone of voice, and input text information to recognize the user's emotional state in real time.

[0625] Step 3:

[0626] Users input content creation requests through their terminals, making specific requests such as "a virtual character that hosts the latest technology news."

[0627] Step 4:

[0628] The terminal transmits the content creation request input by the user to the server, and the transmitted data includes the type and topic of the content.

[0629] Step 5:

[0630] The server analyzes the request received from the user and retrieves the necessary data from external databases or APIs based on the analysis results. For example, it retrieves the latest technology news data from a news API.

[0631] Step 6:

[0632] The server uses the acquired data to generate conversational content for the virtual character using a generative AI model, ensuring that the conversational content is natural and consistent.

[0633] Step 7:

[0634] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set optimal facial expressions and movements according to the user's real-time emotions.

[0635] Step 8:

[0636] The server structures the data of the virtual character and transmits it to the user's device, including the content of the conversation, facial expressions, and movements.

[0637] Step 9:

[0638] The device passes the received data to a rendering engine to display the virtual character in real time, allowing the user to watch the character's conversations and movements.

[0639] Step 10:

[0640] The user can input questions or comments about the content being viewed. For example, the user can input a question such as "When is the next technical seminar?"

[0641] Step 11:

[0642] The device sends questions and comments from the user to the server, and the sent data includes the user's input.

[0643] Step 12:

[0644] The server receives and analyzes questions and comments from users. Based on the analysis results and information from the emotion engine, an appropriate response is generated using a generative AI model.

[0645] Step 13:

[0646] The server structures the generated response and sends it back to the user terminal, which contains the new conversation content.

[0647] Step 14:

[0648] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[0649] Through these steps, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly.

[0650] Example 2

[0651] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0652] In conventional virtual character systems, the characters' facial expressions and movements are fixed, making it difficult to realize interactive responses and behaviors that correspond to the user's emotions. Furthermore, because the latest information cannot be acquired in real time, the generated content is often out of date. This makes it difficult to maintain user satisfaction and interest, and limits the interactive experience they can provide.

[0653] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0654] In this invention, the server includes means for generating a realistic and consistent video character using a generative computer model in accordance with accumulated information and instructed actions, means for receiving a content creation request from a user, means for generating dialogue content for the video character based on the received request, means for setting facial expressions and actions of the video character based on the generated dialogue content, means for delivering the set behavior of the video character to a user device, means for initializing an emotion analysis engine and analyzing the emotional state of the user, and means for setting optimal facial expressions and actions based on the analysis results. This enables interactive response according to the user's emotions and allows the latest information to be acquired in real time, making it possible to provide high-quality and interesting content.

[0655] A "generative computer model" is a computer system that makes automated decisions and generates for a specific purpose based on stored information and instructed actions.

[0656] A "video character" is a computer-generated digital representation that has human-like appearance and behavior.

[0657] "User" means a person who uses this system to input content creation requests and view interactive content.

[0658] A "content creation request" is a request from a user to create content based on a particular topic or format.

[0659] "Dialogue content" refers to sentences and conversations spoken by a video character, generated based on user requests and external information.

[0660] "Expressions and actions" refer to facial changes and body movements that visually express emotions and intentions of a video character.

[0661] "Behavior data" is information that describes the facial expressions and movements of a video character.

[0662] A "user device" is an electronic device that a user uses to view content.

[0663] An "emotion analysis engine" is a system that analyzes data such as a user's facial expressions, tone of voice, and text to detect their emotional state.

[0664] "Emotional state" refers to the user's current emotional or psychological state.

[0665] This invention relates to a system that uses a generative computer model combined with an emotion analysis engine to generate realistic and consistent video characters and provide high-quality, interactive content based on the user's emotional state. Specific embodiments for implementing this system are described below.

[0666] System Overview

[0667] The system consists of three main components: a server, a terminal, and a user. The server runs the generative computational model and the sentiment analysis engine, and the terminal provides the interface between the user and the server.

[0668] Hardware and Software Configuration

[0669] server:

[0670] The server is a high-performance computer system that runs a generative computational model (e.g., a computational system model) and an emotion analysis engine (e.g., a facial expression analysis program and an emotional tone analysis program).

[0671] The server also includes an interface for connecting with external information sources and APIs (e.g., information retrieval programs).

[0672] Device:

[0673] A terminal is a device (e.g., a computer, smartphone, or tablet) used by a user and provides a user interface.

[0674] The terminal has a network connection for communicating with the server.

[0675] User:

[0676] Using this system, users can input content creation requests and engage in interactive dialogue with virtual characters.

[0677] Specific operation of the system

[0678] Initialization procedure for high performance computer systems:

[0679] The server loads pre-trained generative computer models that allow video characters to generate natural-sounding speech and movements.

[0680] At the same time, the emotion analysis engine is initialized to prepare for analyzing the user's emotional state in real time.

[0681] Steps for users to enter their request:

[0682] Through the interface, users input specific content creation requests, for example, "I would like to host a segment on the latest technology news."

[0683] Other example prompts include "Tell me about the next tech seminar" and "Tell me some interesting technology topics."

[0684] Server data acquisition and analysis procedure:

[0685] The server receives and analyzes requests from users, and based on the analysis results, retrieves the latest data needed from external sources or APIs.

[0686] For example, in the case of technology news, the latest related technology news data is acquired.

[0687] Steps for generating and configuring conversation content:

[0688] The server uses a generative computer model based on the acquired data to generate dialogue content for the video character.

[0689] Based on the generated dialogue, the emotion analysis engine analyzes the user's emotional state and sets the most appropriate facial expressions and actions.

[0690] For example, if the user is excited, the video character is set to have an excited expression and movements.

[0691] To distribute your content:

[0692] The server transmits the behavior data of the set video character to the terminal in real time.

[0693] The device renders the received data in real time and displays it to the user, allowing the user to view interactive content.

[0694] Interactive response steps:

[0695] Users can enter questions or comments about the content they are viewing.

[0696] The terminal sends the input to the server, which again analyzes the user's emotional state using an emotion analysis engine and generates an appropriate response in a generative computer model.

[0697] The server generates a response and sends it back to the terminal, allowing the user to enjoy an interactive dialogue.

[0698] In this way, the system can recognize users' emotions in real time and provide interactive experiences that respond to them. Based on specific scenarios and usage patterns, users can generate and enjoy high-quality content.

[0699] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0700] Step 1:

[0701] Loading a generative AI model

[0702] The server loads a pre-trained generative AI model. Specifically, the server loads the generative AI model (e.g., a computational system model) from storage into memory and initializes it. The input is the model file on storage, and the output is the generative AI model launched in memory.

[0703] Step 2:

[0704] Initializing the sentiment analysis engine

[0705] The server initializes the emotion analysis engine, which prepares it to analyze the user's emotional state in real time. Specifically, the server starts the emotion analysis engine (e.g., a facial expression analysis program or an emotional tone analysis program) and checks its operation using test data. The input is the emotion analysis engine program, and the output is the started emotion analysis engine.

[0706] Step 3:

[0707] Accepting content creation requests

[0708] A user inputs a specific content creation request through the interface. Specifically, the user enters the request text (e.g., "I would like you to host a segment on the latest technology news") into the device's input field and clicks the send button. The input is the request text entered by the user, and the output is the request data sent to the server.

[0709] Step 4:

[0710] Request analysis and data acquisition

[0711] The server analyzes requests received from users and retrieves the necessary data from relevant external information sources or APIs. Specifically, the server performs natural language analysis on the request content, generates an API request, sends it to an external database or API (e.g., an information retrieval program), and receives the retrieved data. The input is the user's request content, and the output is the data retrieved from the external database or API.

[0712] Step 5:

[0713] Conversation generation

[0714] The server uses a generative AI model based on the acquired data to generate dialogue for the video character. Specifically, the server inputs the acquired data into the generative AI model and outputs the generated dialogue in text format. The input is data acquired from an external database or API, and the output is the text data of the generated dialogue.

[0715] Step 6:

[0716] Character facial expressions and movement settings

[0717] Based on the generated conversation content, the server uses an emotion analysis engine to analyze the user's emotional state and set the optimal facial expressions and movements for the video character. Specifically, the server inputs the generated conversation content and the user's real-time emotional data, which the emotion analysis engine analyzes to generate the character's facial expressions and movement data. The input is the generated conversation content and the user's emotional data, and the output is the character's facial expressions and movement data.

[0718] Step 7:

[0719] Sending behavioral data

[0720] The server transmits the behavior data of the configured video character to the terminal in real time. Specifically, it executes a process to transmit the generated behavior data to the terminal via the network. The input is the character's facial expression and movement data, and the output is the behavior data transmitted to the terminal.

[0721] Step 8:

[0722] Rendering and Displaying Data

[0723] The device renders the received behavior data in real time and displays it to the user. Specifically, the device processes the received data using a 3D model and animation engine, and displays a visualized character to the user. The input is the received behavior data, and the output is the video character displayed to the user.

[0724] Step 9:

[0725] Accepting User Input

[0726] Users can enter questions or comments about the content they are viewing. Specifically, the user enters the question or comment in text in the input field on the device and clicks the send button. The input is the question or comment entered by the user, and the output is the input data sent to the server.

[0727] Step 10:

[0728] Interactive response generation and delivery

[0729] The server analyzes the input data received from the user, re-analyzes the user's emotional state using the emotion analysis engine, and generates an appropriate response using the generative AI model. The generated response is sent back to the device, which then displays it to the user. Specifically, the server analyzes the input data, generates a response based on the generative AI model, and the emotion analysis engine complements the response. The input is a question or comment from the user, and the output is an interactive response sent to the device.

[0730] This allows the system to provide a high quality and interactive experience to the user.

[0731] (Application example 2)

[0732] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0733] Conventional virtual character systems provide content in a one-way manner without considering the user's emotional state. This limits the improvement of the user experience and lacks true interactivity. Furthermore, they lack the ability to respond in real time to user input requests, making it difficult to respond flexibly to the user's interests and emotions.

[0734] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative artificial intelligence model in accordance with accumulated data and instructed behavior; means for receiving a content creation request from a user; means for generating a conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for an emotion recognition engine for recognizing the emotional state of the user in real time; and means for optimizing the behavior of the virtual character based on the recognized emotional state. This makes it possible to recognize the emotional state of the user and provide high-quality, interactive content that is customized accordingly.

[0735] A "generative artificial intelligence model" is an algorithm or program that generates the conversation content and behavior of a virtual character based on accumulated data and instructed behavior.

[0736] A "virtual character" is a character that has a realistic and consistent personality in a digital space and can interact with users.

[0737] A "content creation request" is a request or instruction a user makes to a virtual character regarding a particular topic or content.

[0738] An "emotion recognition engine" is a technology that analyzes a user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[0739] An "interactive response" is a response generated by a virtual character to respond in real time to a user's questions or comments and to continue the dialogue with the user.

[0740] "External databases and APIs" means external sources of information that a virtual character uses to obtain up-to-date data for generating dialogue.

[0741] "Behavior" refers to the facial expressions and actions that a virtual character makes toward the user.

[0742] This invention relates to a system that provides high-quality, interactive content based on user emotions using a generative AI model combined with an emotion recognition engine. The program processing of this system is described below.

[0743] The server first loads a pre-trained generative artificial intelligence model, which contains the basic algorithms that allow a virtual character to generate natural-sounding conversations and movements.

[0744] The emotion recognition engine is also initialized at the same time. The emotion recognition engine is a system that analyzes the user's facial expressions, tone of voice, input text, etc. to recognize their emotional state in real time. To identify emotions, it uses technologies such as OpenCV and deep learning models (e.g., TensorFlow, PyTorch).

[0745] Users input specific content creation requests through a smartphone interface, specifying the type of content and the topic, such as "I want you to host the latest technology news."

[0746] The server receives and analyzes the user's request. Based on the analysis results, it retrieves the latest necessary data from an external database, such as a news API (e.g., Google News API). Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[0747] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[0748] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion recognition engine to set the most appropriate facial expressions and movements according to the user's emotional state, and then transmits the set virtual character's behavior data to the user's device.

[0749] The device renders the received data in real time and displays it to the user, allowing the user to see and hear the virtual character's conversations and movements.

[0750] Additionally, users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion recognition engine, the generative AI model generates an appropriate response.

[0751] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion recognition engine determines the user's level of interest, and the generative AI model responds with a positive comment along with appropriate details.

[0752] Prompt Sentence Examples

[0753] User Input: "What's in tech news today?"

[0754] Generated AI prompt:

[0755] A user has requested, "What's in tech news today?" Use the news API to get the latest tech news and generate energetic and positive conversations when users are excited.

[0756] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[0757] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0758] Step 1:

[0759] The server loads and initializes a pre-trained generative artificial intelligence model, taking the parameter data of the pre-trained generative AI model as input and obtaining the prepared state of the AI ​​model as output, which will later generate a virtual character.

[0760] Step 2:

[0761] The server initializes the emotion recognition engine, which takes as input the configuration file of the emotion recognition algorithm and as output is ready to analyze the user's emotional state in real time, so that it can process the emotional data from the user.

[0762] Step 3:

[0763] Users input content creation requests through a smartphone interface. The input includes the type of content and topic desired by the user, and the output provides detailed data about the request, which can then be processed.

[0764] Step 4:

[0765] The server parses the request received from the user, taking the user request data as input and providing the structured data of the parsed request as output, which allows the retrieval of the data needed for the next processing step.

[0766] Step 5:

[0767] The server retrieves the latest data related to the requested content from an external database or API. The input is the structured data of the parsed request, and the output is the latest retrieved data. This allows the server to provide information that meets the latest requests. A specific example is using a news API to retrieve the latest technology news.

[0768] Step 6:

[0769] The server uses a generative AI model based on the latest acquired data to generate the conversation content of the virtual character. The input is news data and other external data, and the output is the generated conversation content data. This gives shape to the dialogue content of the virtual character.

[0770] Step 7:

[0771] The server sets the virtual character's facial expressions and movements based on the generated conversation content. The inputs are conversation content data and user emotion data obtained from an emotion recognition engine, and the output is movement setting data for the virtual character. This allows the character to behave naturally according to the user's emotions.

[0772] Step 8:

[0773] The server sends the behavior data of the configured virtual character to the user's device. The input is the virtual character's action setting data, and the output is the data sent to the user's device. This allows the user to watch the character's conversation and actions.

[0774] Step 9:

[0775] The device renders the received data in real time and displays it to the user. The input is behavior data sent from the server, and the output is a rendered virtual character. This allows the user to enjoy interactive dialogue with the character.

[0776] Step 10:

[0777] The user enters a question or comment about the content they are viewing. The input is the user's text input, and the output is the text data. This is used to generate the next response.

[0778] Step 11:

[0779] The terminal sends the user's input to the server. The input is the user's comment data, and the output is the data transmission to the server, so that the server can prepare an interactive response.

[0780] Step 12:

[0781] The server generates appropriate responses using a generative AI model based on user input and emotion recognition data. The inputs are user comment data and current emotion data, and the output is interactive response data. This allows the virtual character to respond appropriately in real time.

[0782] In this way, by performing each processing step in detail, high-quality interactive virtual character dialogue optimized for the user's emotions is realized.

[0783] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0784] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0785] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0786] [Third embodiment]

[0787] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0788] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0789] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0790] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0791] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0792] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0793] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0794] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0795] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0796] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0797] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0798] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0799] This invention relates to a system that generates virtual characters using a generative AI model and provides users with high-quality content they desire in real time. Below, we explain the program processing of this system in natural language and provide detailed explanations with specific examples.

[0800] 1. Initializing the VTuber generation AI model

[0801] The server loads pre-trained generative AI models, which serve as the foundation for virtual characters to generate natural-sounding speech and behavior.

[0802] 2. Acceptance of Content Creation Requests

[0803] Users submit requests for newscasts and other content production to the system via a human interface, including the desired content description and specific requirements.

[0804] 3. Content-based conversation generation

[0805] The server analyzes the received request, retrieves the latest data from external databases and APIs, and uses the generated AI model to generate the lines and conversations of the virtual character.

[0806] As a specific example, if a user requests a "virtual character who hosts the latest technology news," the server will retrieve the latest technology news data from a news API and generate natural conversation using a generative AI model based on that data.

[0807] 4. Character behavior settings

[0808] The server then sets the virtual character's facial expressions and movements based on the generated dialogue, allowing the character to behave realistically to the viewer.

[0809] For example, when reading the news, a virtual character can show appropriate facial expressions to interesting topics and wave their hands to emphasize key points.

[0810] 5. Content Delivery

[0811] The server transmits the behavior of the virtual character to the user's device, which receives the data and displays it to the user in real time.

[0812] 6. Interactive Responses

[0813] Users can enter questions or comments while watching content. The server analyzes the input and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[0814] For example, if a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using an AI model. This response is then returned to the user by a virtual character.

[0815] In this way, the system provides natural conversations with high-quality virtual characters and can generate and deliver content desired by users in real time.

[0816] The processing flow will be explained below.

[0817] Step 1:

[0818] The server loads pre-trained generative AI models, which are the foundation of the system and allow virtual characters to generate natural-sounding conversations and movements.

[0819] Step 2:

[0820] Through the terminal interface, the user inputs a specific content creation request, which includes a content type, for example "news show host," and a specific topic, for example, the latest technology news.

[0821] Step 3:

[0822] The terminal transmits the request input by the user to the server, including the request content and related details.

[0823] Step 4:

[0824] The server analyzes the received request and extracts the necessary information, including the type and topic of the content.

[0825] Step 5:

[0826] The server retrieves the latest relevant data from external databases and APIs, for example connecting to a news API to provide the latest tech news.

[0827] Step 6:

[0828] The server uses a generative AI model based on the external data it retrieves to generate dialogue for the virtual character, resulting in natural and consistent dialogue.

[0829] Step 7:

[0830] The server then uses the generated dialogue to set the virtual character's facial expressions and movements, using a generative AI model to set the appropriate expressions and movements at the appropriate times.

[0831] Step 8:

[0832] The server structures the data of the virtual character and sends it to the user's device, including information on the content of the conversation, facial expressions, and movements.

[0833] Step 9:

[0834] The device passes the received data to a rendering engine, which displays the virtual character in real time, allowing the user to view the character's movements and conversations on the device.

[0835] Step 10:

[0836] Users can enter questions or comments about the content they are viewing, and the input is sent from the terminal to the server.

[0837] Step 11:

[0838] The server receives and analyzes input from the user, and generates an appropriate response based on the input using a generative AI model.

[0839] Step 12:

[0840] The server then structures the generated response and sends it back to the user's terminal, containing the new conversation content.

[0841] Step 13:

[0842] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[0843] Example 1

[0844] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0845] Conventional virtual character generation systems require a significant amount of time and manual setup to enable characters to converse and move naturally. It is also difficult to support real-time data updates and rapid interactive responses from users. This results in a poor user experience and makes the system impractical.

[0846] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0847] In this invention, the server includes means for loading a pre-trained generative artificial intelligence model, means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behavior, means for receiving a content creation request from a user, means for acquiring data from an external database or API based on the received request and generating conversation content for the virtual character based on the data, means for setting the facial expressions and movements of the virtual character based on the generated conversation content, and means for delivering the set behavior of the virtual character to a user terminal, thereby enabling the real-time generation and delivery of virtual characters with fast and natural conversation and movements.

[0848] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates natural conversations and actions based on pre-trained data.

[0849] A "virtual character" is a digitally generated fictional character that converses and behaves in a natural, human-like manner.

[0850] A "content creation request" is an instruction from a user to the system to create specific content.

[0851] "Means for setting facial expressions and movements" refers to the methods and technologies for determining and setting the facial expressions and body movements of a virtual character based on the generated conversation content.

[0852] An "external database or API" is a data repository outside the system, or an interface for exchanging information with another program or system.

[0853] "Interactive responses" refer to responses or reactions by a virtual character that are generated in real time based on input from a user.

[0854] "User terminal" refers to a device (e.g., smartphone, tablet, PC, etc.) that a user uses to access and operate the system.

[0855] This invention relates to a system that generates virtual characters using a generative AI model and provides users with high-quality content they desire in real time. Specific embodiments for implementing this system are described in detail below.

[0856] 1. Initializing the VTuber generation AI model

[0857] The server loads a pre-trained generative AI model (e.g., a Transformer model or a natural language processing model like GPT-3) that is required for the virtual character to generate natural-sounding conversations and behaviors. The server loads the model into memory and initializes the necessary pre-processing parameters and settings.

[0858] 2. Acceptance of Content Creation Requests

[0859] Users access the system through a browser or dedicated application and submit content creation requests, which include specific requirements and requirements (e.g., news type or topic segment).

[0860] Examples:

[0861] A user may submit a request saying, "I'd like a virtual character to host the latest technology news."

[0862] 3. Content-based conversation generation

[0863] The server analyzes user requests and retrieves relevant, up-to-date information from external databases and APIs (e.g., news APIs). Based on the retrieved data, the server provides prompts to a generative AI model, which generates lines and dialogue for the virtual character.

[0864] Example prompt sentence:

[0865] "Generate a virtual character to host the latest tech news."

[0866] 4. Character behavior settings

[0867] The server sets the virtual character's facial expressions and movements (e.g., hand gestures and eye movements) based on the generated lines. This is done using a 3D motion engine (e.g., Unity or Unreal Engine). The server analyzes the generated lines, determines facial expressions and movement patterns, and generates the character's movement data.

[0868] 5. Content Delivery

[0869] The server distributes the behavior data of the virtual character to the user's device, which receives the data and displays it to the user in real time using WebRTC and HTTP streaming technologies.

[0870] 6. Interactive Responses

[0871] Users can enter questions or comments while watching. The server receives this input, performs text analysis, and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[0872] Examples:

[0873] If a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using a generative AI model. This response is then returned to the user by a virtual character.

[0874] In this way, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content that users desire in real time.

[0875] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0876] Step 1: Initializing the VTuber generation AI model

[0877] The server loads a pre-trained generative AI model. Loading the model involves reading the model weight data from a specified directory. Specifically, it initializes the parameters and configuration values ​​of the generative AI model (e.g., GPT-3) into memory, making the model immediately available for use.

[0878] Input: Model weight data file

[0879] Output: Initialized generative AI model

[0880] Specific operation: The server loads the model weight data from the specified directory and initializes the necessary parameters.

[0881] Step 2: Accepting a content creation request

[0882] Users submit content creation requests using a browser or dedicated application. The request includes the type of content to be generated and any specific requirements. Once the request is sent, the server receives and analyzes it.

[0883] Input: Content production request from user

[0884] Output: Parsed request content

[0885] Specific operation: A user inputs a request using a web form or application interface and sends it. The server receives the request and analyzes the contents.

[0886] Step 3: Content-based conversation generation

[0887] Based on the analyzed request, the server retrieves the necessary data from external databases or APIs. For example, it retrieves the latest technology news from a news API. The retrieved data is converted into prompts and input into a generative AI model. The model then generates natural-sounding dialogue based on these prompts.

[0888] Input: Parsed request content, external data to be retrieved

[0889] Output: Generated dialogue

[0890] Specific operation: The server sends a data acquisition request to an external API, converts the obtained data into a prompt sentence, and inputs the prompt sentence into the generative AI model to generate dialogue.

[0891] Step 4: Setting character behavior

[0892] The server sets the virtual character's facial expressions and movements based on the generated lines. This is done using a 3D motion engine (such as Unity or Unreal Engine). The server determines facial expressions and body movements and generates them as data so that the lines and movements match.

[0893] Input: Generated dialogue

[0894] Output: Set facial expressions and movement data

[0895] Specific movements: The server analyzes the dialogue and determines facial expressions and movements. A 3D motion engine is used to generate movement data.

[0896] Step 5: Deliver your content

[0897] The server delivers the configured behavior data to the user device using WebRTC or HTTP streaming technology. The device decodes the received data in real time and displays it to the user.

[0898] Input: Set behavior data

[0899] Output: Display on the user's terminal

[0900] Specific operation: The server encodes behavioral data and sends it in streaming format. The device receives the data, decodes it in real time, and displays it.

[0901] Step 6: Interactive Response

[0902] Users can enter questions or comments while watching content. The server analyzes these inputs and uses a generative AI model to generate appropriate responses, which are then sent to the device and displayed to the user.

[0903] Input: User questions and comments

[0904] Output: The generated response

[0905] How it works: The user enters a question or comment into the chat box and sends it. The server analyzes the input and generates a response using the generative AI model. The response is sent to the device and displayed to the user.

[0906] Through these processing steps, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content desired by users in real time.

[0907] (Application example 1)

[0908] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0909] Current content delivery services make it difficult for users to obtain real-time updates and interact with virtual characters based on them. This limits the user experience and lacks dynamic and personalized content delivery. In particular, there is a demand for real-time conversation generation and responses based on news and trend information, but there is a lack of effective methods to achieve this.

[0910] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0911] In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative AI model according to accumulated data and instructed behaviors; means for receiving a content creation request from a user; means for generating conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for acquiring the latest data from an external API and generating conversation content for the virtual character using the generative AI model in response to prompts based on the data; and means for a user to interact with the virtual character via a smart device, thereby enabling a user to enjoy real-time interactive conversations based on the latest information.

[0912] A "generative artificial intelligence model" is an artificial intelligence algorithm that is trained using large datasets and is capable of generating new data and content, especially natural-sounding conversations and behaviors, under certain conditions.

[0913] A "virtual character" is a virtual personality or character that can behave dynamically and realistically on a digital device.

[0914] A "content creation request" is a request that a user sends to the system specifying the specific details and requirements of the content to be created.

[0915] "Conversational content" refers to words and phrases that a virtual character utters to interact with a user or provide information.

[0916] "Expressions and movements" refers to the facial expressions and body movements used by a virtual character to express emotions.

[0917] A "user terminal" is a device such as a smartphone, tablet, PC, or head-mounted display that allows a user to interact with virtual characters or receive content via the Internet.

[0918] An "external API" is a programmatic interface to external databases and services, and is a mechanism used by systems to obtain the latest data in real time.

[0919] A "prompt" is an input text given to a generative AI model, which serves as an instruction for the AI ​​to generate conversations and content based on that text.

[0920] "Smart devices" is a general term for intelligent electronic devices with internet connectivity, such as smartphones, tablets, and smartwatches.

[0921] "Interactive dialogue" refers to two-way communication in which a virtual character generates and responds to questions or comments entered by a user in real time.

[0922] The system for implementing the present invention comprises the following steps: Specific hardware and software usage methods will be described in detail below.

[0923] Program Overview

[0924] First, the server uses a generative AI model to generate the digital characters used on the client device. To do this, it loads a generative AI model that has been trained in advance using a huge dataset. The software used is Python and a generative AI model library (e.g., GPT-3).

[0925] Hardware and Software Details

[0926] Hardware:

[0927] server:

[0928] A server with a powerful processor (e.g., Intel Xeon), lots of memory (e.g., 64GB RAM), and a fast network connection.

[0929] User device:

[0930] Smartphone, head-mounted display (HMD), or computer.

[0931] software:

[0932] Python: Used for program development and operating AI models.

[0933] Generative AI models: Generative AI model libraries (e.g., GPT-3) are used to generate conversation content and character behavior.

[0934] API: Use external APIs (e.g. news API) to get the latest data.

[0935] Data processing and calculation

[0936] The server processes the data and performs calculations using the following procedure.

[0937] 1. Initializing the generative AI model

[0938] The server loads the generative AI model and serves as the foundation for the system, providing the foundation for the virtual characters to generate natural conversations and movements.

[0939] 2. Acceptance of Content Creation Requests

[0940] Users send content creation requests to the server through the human interface of their smart devices, which include specific requests for the virtual character.

[0941] 3. Acquiring external data

[0942] The server retrieves the latest data (e.g., the latest news) in real time through an external API, which is then used to generate content.

[0943] 4. Conversation generation

[0944] Based on the acquired data, a generative AI model is used to generate conversational content for the virtual character. A prompt is set, and the AI ​​model generates an appropriate response based on that prompt.

[0945] 5. Setting the facial expressions and movements of the virtual character

[0946] The facial expressions and movements of the virtual character are set based on the generated conversation content, resulting in more realistic and natural interactions.

[0947] 6. Content Delivery

[0948] The server transmits the behavior of the virtual character that has been set to the user's device, which receives it in real time and displays it to the user.

[0949] Specific examples

[0950] Example 1: News distribution

[0951] The user requests, "What is the latest technology news?"

[0952] The server retrieves the latest technology news data from the news API.

[0953] The generative AI model was given the prompt, "The user wants to know the latest technology news. News content: {latest news}."

[0954] The virtual character responds, "I'll report on the latest tech news."

[0955] The user asked, "Tell me more about this technology."

[0956] A virtual character follows and explains the details of the news.

[0957] Example prompt sentence:

[0958] "Users want to know the latest tech news. Please report on this news in detail."

[0959] This system allows users to enjoy real-time, interactive dialogue based on the latest information.

[0960] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0961] Step 1:

[0962] The server initializes the generative AI model and loads the trained model. As input, it requires the trained generative AI model, which prepares the virtual character to generate natural conversations and actions. As output, it obtains the initialized generative AI model.

[0963] Step 2:

[0964] A user uses a smart device to input a content creation request. The input includes the specific content desired by the user and their requirements. The smart device then sends the request to the server. The output is the received content creation request.

[0965] Step 3:

[0966] The server calls an external API to retrieve the latest data. The input requires the endpoint URL of the external API and authentication information. News data, trend data, etc. are retrieved from the external API. The output is the retrieved latest data.

[0967] Step 4:

[0968] The server generates a prompt sentence based on the acquired data. The latest data acquired from an external API is required as input. The prompt sentence to be given to the generative AI model is formed. The prompt sentence is obtained as output.

[0969] Step 5:

[0970] The server inputs the generated prompt sentence into the AI ​​model to generate the conversation content of the virtual character. The input requires the prompt sentence and the generation AI model. The generation AI model generates natural conversation content based on the prompt sentence. The generated conversation content is obtained as the output.

[0971] Step 6:

[0972] The server sets the facial expressions and movements of the virtual character based on the generated conversation content. The generated conversation content is required as input. Settings are made so that the virtual character will make facial expressions and movements according to the content. The set facial expressions and movements are obtained as output.

[0973] Step 7:

[0974] The server delivers the virtual character's behavior to the user's device. The inputs required are the user's facial expressions, movements, and generated conversations. The smart device receives this information and displays it to the user in real time. The output is the virtual character's behavior displayed on the user's device.

[0975] Step 8:

[0976] The user inputs questions and comments while viewing the content. The user's questions and comments are required as input. The device sends the input content to the server. The user's input content is obtained as output.

[0977] Step 9:

[0978] The server analyzes the input from the user and generates an appropriate response using a generative AI model. The input requires a user's question or comment. The generative AI model generates an appropriate response based on the input. The output is the generated response.

[0979] Step 10:

[0980] The server sends the generated response to the user's terminal, where the virtual character displays the response. The input requires the generated response, which the user terminal displays in real time. The output is the virtual character's response, which is displayed to the user.

[0981] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0982] This invention relates to a system that generates virtual characters using a generative AI model combined with an emotion engine, and provides high-quality, interactive content based on the user's emotions. Below, the program processing of this system is explained in natural language and explained in detail with specific examples.

[0983] Initializing the VTuber generation AI model

[0984] The server first loads a pre-trained generative AI model, which contains the basic algorithms that allow virtual characters to generate natural-sounding speech and movements.

[0985] Emotion engine initialization

[0986] The server also initializes an emotion engine to recognize the user's emotions. This emotion engine analyzes the user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[0987] Accepting content creation requests

[0988] Users input specific content creation requests through the device interface, specifying content types and topics, such as "I would like a host for the latest technology news."

[0989] Content-based conversation generation

[0990] The server receives and analyzes user requests. Based on the analysis results, it retrieves the latest necessary data from external databases and APIs. Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[0991] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[0992] Character behavior settings

[0993] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set the most appropriate facial expressions and movements according to the user's emotional state.

[0994] As a specific example, if the user is excited and happy, the character can be set to smile and move in a lively manner.

[0995] Content Delivery

[0996] The server transmits the behavior data of the configured virtual character to the user's device, which then renders the received data in real time and displays it to the user, allowing the user to watch and listen to the virtual character's conversations and movements.

[0997] Interactive Response

[0998] Users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion engine, the generative AI model generates an appropriate response.

[0999] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion engine determines the user's interest and level of attention, and the generative AI model responds with a positive comment along with appropriate detailed information.

[1000] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[1001] The processing flow will be explained below.

[1002] Step 1:

[1003] The server initializes a generative AI model, which is a pre-trained model that contains the basic algorithms for generating natural-sounding speech and behavior for virtual characters.

[1004] Step 2:

[1005] The server initializes an emotion engine for recognizing the user's emotions. This emotion engine has the function of analyzing the user's facial expressions, tone of voice, and input text information to recognize the user's emotional state in real time.

[1006] Step 3:

[1007] Users input content creation requests through their terminals, making specific requests such as "a virtual character that hosts the latest technology news."

[1008] Step 4:

[1009] The terminal transmits the content creation request input by the user to the server, and the transmitted data includes the type and topic of the content.

[1010] Step 5:

[1011] The server analyzes the request received from the user and retrieves the necessary data from external databases or APIs based on the analysis results. For example, it retrieves the latest technology news data from a news API.

[1012] Step 6:

[1013] The server uses the acquired data to generate conversational content for the virtual character using a generative AI model, ensuring that the conversational content is natural and consistent.

[1014] Step 7:

[1015] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set optimal facial expressions and movements according to the user's real-time emotions.

[1016] Step 8:

[1017] The server structures the data of the virtual character and transmits it to the user's device, including the content of the conversation, facial expressions, and movements.

[1018] Step 9:

[1019] The device passes the received data to a rendering engine to display the virtual character in real time, allowing the user to watch the character's conversations and movements.

[1020] Step 10:

[1021] The user can input questions or comments about the content being viewed. For example, the user can input a question such as "When is the next technical seminar?"

[1022] Step 11:

[1023] The device sends questions and comments from the user to the server, and the sent data includes the user's input.

[1024] Step 12:

[1025] The server receives and analyzes questions and comments from users. Based on the analysis results and information from the emotion engine, an appropriate response is generated using a generative AI model.

[1026] Step 13:

[1027] The server structures the generated response and sends it back to the user terminal, which contains the new conversation content.

[1028] Step 14:

[1029] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[1030] Through these steps, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly.

[1031] Example 2

[1032] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1033] In conventional virtual character systems, the characters' facial expressions and movements are fixed, making it difficult to realize interactive responses and behaviors that correspond to the user's emotions. Furthermore, because the latest information cannot be acquired in real time, the generated content is often out of date. This makes it difficult to maintain user satisfaction and interest, and limits the interactive experience they can provide.

[1034] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1035] In this invention, the server includes means for generating a realistic and consistent video character using a generative computer model in accordance with accumulated information and instructed actions, means for receiving a content creation request from a user, means for generating dialogue content for the video character based on the received request, means for setting facial expressions and actions of the video character based on the generated dialogue content, means for delivering the set behavior of the video character to a user device, means for initializing an emotion analysis engine and analyzing the emotional state of the user, and means for setting optimal facial expressions and actions based on the analysis results. This enables interactive response according to the user's emotions and allows the latest information to be acquired in real time, making it possible to provide high-quality and interesting content.

[1036] A "generative computer model" is a computer system that makes automated decisions and generates for a specific purpose based on stored information and instructed actions.

[1037] A "video character" is a computer-generated digital representation that has human-like appearance and behavior.

[1038] "User" means a person who uses this system to input content creation requests and view interactive content.

[1039] A "content creation request" is a request from a user to create content based on a particular topic or format.

[1040] "Dialogue content" refers to sentences and conversations spoken by a video character, generated based on user requests and external information.

[1041] "Expressions and actions" refer to facial changes and body movements that visually express emotions and intentions of a video character.

[1042] "Behavior data" is information that describes the facial expressions and movements of a video character.

[1043] A "user device" is an electronic device that a user uses to view content.

[1044] An "emotion analysis engine" is a system that analyzes data such as a user's facial expressions, tone of voice, and text to detect their emotional state.

[1045] "Emotional state" refers to the user's current emotional or psychological state.

[1046] This invention relates to a system that uses a generative computer model combined with an emotion analysis engine to generate realistic and consistent video characters and provide high-quality, interactive content based on the user's emotional state. Specific embodiments for implementing this system are described below.

[1047] System Overview

[1048] The system consists of three main components: a server, a terminal, and a user. The server runs the generative computational model and the sentiment analysis engine, and the terminal provides the interface between the user and the server.

[1049] Hardware and Software Configuration

[1050] server:

[1051] The server is a high-performance computer system that runs a generative computational model (e.g., a computational system model) and an emotion analysis engine (e.g., a facial expression analysis program and an emotional tone analysis program).

[1052] The server also includes an interface for connecting with external information sources and APIs (e.g., information retrieval programs).

[1053] Device:

[1054] A terminal is a device (e.g., a computer, smartphone, or tablet) used by a user and provides a user interface.

[1055] The terminal has a network connection for communicating with the server.

[1056] User:

[1057] Using this system, users can input content creation requests and engage in interactive dialogue with virtual characters.

[1058] Specific operation of the system

[1059] Initialization procedure for high performance computer systems:

[1060] The server loads pre-trained generative computer models that allow video characters to generate natural-sounding speech and movements.

[1061] At the same time, the emotion analysis engine is initialized to prepare for analyzing the user's emotional state in real time.

[1062] Steps for users to enter their request:

[1063] Through the interface, users input specific content creation requests, for example, "I would like to host a segment on the latest technology news."

[1064] Other example prompts include "Tell me about the next tech seminar" and "Tell me some interesting technology topics."

[1065] Server data acquisition and analysis procedure:

[1066] The server receives and analyzes requests from users, and based on the analysis results, retrieves the latest data needed from external sources or APIs.

[1067] For example, in the case of technology news, the latest related technology news data is acquired.

[1068] Steps for generating and configuring conversation content:

[1069] The server uses a generative computer model based on the acquired data to generate dialogue content for the video character.

[1070] Based on the generated dialogue, the emotion analysis engine analyzes the user's emotional state and sets the most appropriate facial expressions and actions.

[1071] For example, if the user is excited, the video character is set to have an excited expression and movements.

[1072] To distribute your content:

[1073] The server transmits the behavior data of the set video character to the terminal in real time.

[1074] The device renders the received data in real time and displays it to the user, allowing the user to view interactive content.

[1075] Interactive response steps:

[1076] Users can enter questions or comments about the content they are viewing.

[1077] The terminal sends the input to the server, which again analyzes the user's emotional state using an emotion analysis engine and generates an appropriate response in a generative computer model.

[1078] The server generates a response and sends it back to the terminal, allowing the user to enjoy an interactive dialogue.

[1079] In this way, the system can recognize users' emotions in real time and provide interactive experiences that respond to them. Based on specific scenarios and usage patterns, users can generate and enjoy high-quality content.

[1080] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1081] Step 1:

[1082] Loading a generative AI model

[1083] The server loads a pre-trained generative AI model. Specifically, the server loads the generative AI model (e.g., a computational system model) from storage into memory and initializes it. The input is the model file on storage, and the output is the generative AI model launched in memory.

[1084] Step 2:

[1085] Initializing the sentiment analysis engine

[1086] The server initializes the emotion analysis engine, which prepares it to analyze the user's emotional state in real time. Specifically, the server starts the emotion analysis engine (e.g., a facial expression analysis program or an emotional tone analysis program) and checks its operation using test data. The input is the emotion analysis engine program, and the output is the started emotion analysis engine.

[1087] Step 3:

[1088] Accepting content creation requests

[1089] A user inputs a specific content creation request through the interface. Specifically, the user enters the request text (e.g., "I would like you to host a segment on the latest technology news") into the device's input field and clicks the send button. The input is the request text entered by the user, and the output is the request data sent to the server.

[1090] Step 4:

[1091] Request analysis and data acquisition

[1092] The server analyzes requests received from users and retrieves the necessary data from relevant external information sources or APIs. Specifically, the server performs natural language analysis on the request content, generates an API request, sends it to an external database or API (e.g., an information retrieval program), and receives the retrieved data. The input is the user's request content, and the output is the data retrieved from the external database or API.

[1093] Step 5:

[1094] Conversation generation

[1095] The server uses a generative AI model based on the acquired data to generate dialogue for the video character. Specifically, the server inputs the acquired data into the generative AI model and outputs the generated dialogue in text format. The input is data acquired from an external database or API, and the output is the text data of the generated dialogue.

[1096] Step 6:

[1097] Character facial expressions and movement settings

[1098] Based on the generated conversation content, the server uses an emotion analysis engine to analyze the user's emotional state and set the optimal facial expressions and movements for the video character. Specifically, the server inputs the generated conversation content and the user's real-time emotional data, which the emotion analysis engine analyzes to generate the character's facial expressions and movement data. The input is the generated conversation content and the user's emotional data, and the output is the character's facial expressions and movement data.

[1099] Step 7:

[1100] Sending behavioral data

[1101] The server transmits the behavior data of the configured video character to the terminal in real time. Specifically, it executes a process to transmit the generated behavior data to the terminal via the network. The input is the character's facial expression and movement data, and the output is the behavior data transmitted to the terminal.

[1102] Step 8:

[1103] Rendering and Displaying Data

[1104] The device renders the received behavior data in real time and displays it to the user. Specifically, the device processes the received data using a 3D model and animation engine, and displays a visualized character to the user. The input is the received behavior data, and the output is the video character displayed to the user.

[1105] Step 9:

[1106] Accepting User Input

[1107] Users can enter questions or comments about the content they are viewing. Specifically, the user enters the question or comment in text in the input field on the device and clicks the send button. The input is the question or comment entered by the user, and the output is the input data sent to the server.

[1108] Step 10:

[1109] Interactive response generation and delivery

[1110] The server analyzes the input data received from the user, re-analyzes the user's emotional state using the emotion analysis engine, and generates an appropriate response using the generative AI model. The generated response is sent back to the device, which then displays it to the user. Specifically, the server analyzes the input data, generates a response based on the generative AI model, and the emotion analysis engine complements the response. The input is a question or comment from the user, and the output is an interactive response sent to the device.

[1111] This allows the system to provide a high quality and interactive experience to the user.

[1112] (Application example 2)

[1113] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1114] Conventional virtual character systems provide content in a one-way manner without considering the user's emotional state. This limits the improvement of the user experience and lacks true interactivity. Furthermore, they lack the ability to respond in real time to user input requests, making it difficult to respond flexibly to the user's interests and emotions.

[1115] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative artificial intelligence model in accordance with accumulated data and instructed behavior; means for receiving a content creation request from a user; means for generating a conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for an emotion recognition engine for recognizing the emotional state of the user in real time; and means for optimizing the behavior of the virtual character based on the recognized emotional state. This makes it possible to recognize the emotional state of the user and provide high-quality, interactive content that is customized accordingly.

[1116] A "generative artificial intelligence model" is an algorithm or program that generates the conversation content and behavior of a virtual character based on accumulated data and instructed behavior.

[1117] A "virtual character" is a character that has a realistic and consistent personality in a digital space and can interact with users.

[1118] A "content creation request" is a request or instruction a user makes to a virtual character regarding a particular topic or content.

[1119] An "emotion recognition engine" is a technology that analyzes a user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[1120] An "interactive response" is a response generated by a virtual character to respond in real time to a user's questions or comments and to continue the dialogue with the user.

[1121] "External databases and APIs" means external sources of information that a virtual character uses to obtain up-to-date data for generating dialogue.

[1122] "Behavior" refers to the facial expressions and actions that a virtual character makes toward the user.

[1123] This invention relates to a system that provides high-quality, interactive content based on user emotions using a generative AI model combined with an emotion recognition engine. The program processing of this system is described below.

[1124] The server first loads a pre-trained generative artificial intelligence model, which contains the basic algorithms that allow a virtual character to generate natural-sounding conversations and movements.

[1125] The emotion recognition engine is also initialized at the same time. The emotion recognition engine is a system that analyzes the user's facial expressions, tone of voice, input text, etc. to recognize their emotional state in real time. To identify emotions, it uses technologies such as OpenCV and deep learning models (e.g., TensorFlow, PyTorch).

[1126] Users input specific content creation requests through a smartphone interface, specifying the type of content and the topic, such as "I want you to host the latest technology news."

[1127] The server receives and analyzes the user's request. Based on the analysis results, it retrieves the latest necessary data from an external database, such as a news API (e.g., Google News API). Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[1128] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[1129] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion recognition engine to set the most appropriate facial expressions and movements according to the user's emotional state, and then transmits the set virtual character's behavior data to the user's device.

[1130] The device renders the received data in real time and displays it to the user, allowing the user to see and hear the virtual character's conversations and movements.

[1131] Additionally, users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion recognition engine, the generative AI model generates an appropriate response.

[1132] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion recognition engine determines the user's level of interest, and the generative AI model responds with a positive comment along with appropriate details.

[1133] Prompt Sentence Examples

[1134] User Input: "What's in tech news today?"

[1135] Generated AI prompt:

[1136] A user has requested, "What's in tech news today?" Use the news API to get the latest tech news and generate energetic and positive conversations when users are excited.

[1137] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[1138] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1139] Step 1:

[1140] The server loads and initializes a pre-trained generative artificial intelligence model, taking the parameter data of the pre-trained generative AI model as input and obtaining the prepared state of the AI ​​model as output, which will later generate a virtual character.

[1141] Step 2:

[1142] The server initializes the emotion recognition engine, which takes as input the configuration file of the emotion recognition algorithm and as output is ready to analyze the user's emotional state in real time, so that it can process the emotional data from the user.

[1143] Step 3:

[1144] Users input content creation requests through a smartphone interface. The input includes the type of content and topic desired by the user, and the output provides detailed data about the request, which can then be processed.

[1145] Step 4:

[1146] The server parses the request received from the user, taking the user request data as input and providing the structured data of the parsed request as output, which allows the retrieval of the data needed for the next processing step.

[1147] Step 5:

[1148] The server retrieves the latest data related to the requested content from an external database or API. The input is the structured data of the parsed request, and the output is the latest retrieved data. This allows the server to provide information that meets the latest requests. A specific example is using a news API to retrieve the latest technology news.

[1149] Step 6:

[1150] The server uses a generative AI model based on the latest acquired data to generate the conversation content of the virtual character. The input is news data and other external data, and the output is the generated conversation content data. This gives shape to the dialogue content of the virtual character.

[1151] Step 7:

[1152] The server sets the virtual character's facial expressions and movements based on the generated conversation content. The inputs are conversation content data and user emotion data obtained from an emotion recognition engine, and the output is movement setting data for the virtual character. This allows the character to behave naturally according to the user's emotions.

[1153] Step 8:

[1154] The server sends the behavior data of the configured virtual character to the user's device. The input is the virtual character's action setting data, and the output is the data sent to the user's device. This allows the user to watch the character's conversation and actions.

[1155] Step 9:

[1156] The device renders the received data in real time and displays it to the user. The input is behavior data sent from the server, and the output is a rendered virtual character. This allows the user to enjoy interactive dialogue with the character.

[1157] Step 10:

[1158] The user enters a question or comment about the content they are viewing. The input is the user's text input, and the output is the text data. This is used to generate the next response.

[1159] Step 11:

[1160] The terminal sends the user's input to the server. The input is the user's comment data, and the output is the data transmission to the server, so that the server can prepare an interactive response.

[1161] Step 12:

[1162] The server generates appropriate responses using a generative AI model based on user input and emotion recognition data. The inputs are user comment data and current emotion data, and the output is interactive response data. This allows the virtual character to respond appropriately in real time.

[1163] In this way, by performing each processing step in detail, high-quality interactive virtual character dialogue optimized for the user's emotions is realized.

[1164] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1165] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1166] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1167] [Fourth embodiment]

[1168] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1169] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1170] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1171] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1172] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1173] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1174] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1175] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1176] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1177] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1178] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1179] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1180] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1181] This invention relates to a system that generates virtual characters using a generative AI model and provides high-quality content desired by users in real time. Below, the program processing of this system is explained in natural language and explained in detail with specific examples.

[1182] 1. Initializing the VTuber generation AI model

[1183] The server loads pre-trained generative AI models, which form the basis for virtual characters to generate natural-sounding speech and behavior.

[1184] 2. Acceptance of Content Creation Requests

[1185] Users submit requests for newscasts and other content production to the system via a human interface, including the desired content description and specific requirements.

[1186] 3. Content-based conversation generation

[1187] The server analyzes the received request, retrieves the latest data from external databases and APIs, and uses the generated AI model to generate the lines and conversations of the virtual character.

[1188] As a specific example, if a user requests a "virtual character who hosts the latest technology news," the server will retrieve the latest technology news data from a news API and generate natural conversation using a generative AI model based on that data.

[1189] 4. Character behavior settings

[1190] The server then sets the virtual character's facial expressions and movements based on the generated dialogue, allowing the character to behave realistically to the viewer.

[1191] For example, when reading the news, a virtual character can show appropriate facial expressions to interesting topics and wave their hands to emphasize key points.

[1192] 5. Content Delivery

[1193] The server transmits the behavior of the virtual character to the user's device, which receives the data and displays it to the user in real time.

[1194] 6. Interactive Responses

[1195] Users can enter questions or comments while watching content. The server analyzes the input and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[1196] For example, if a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using an AI model. This response is then returned to the user by a virtual character.

[1197] In this way, the system provides natural conversations with high-quality virtual characters and can generate and deliver content desired by users in real time.

[1198] The processing flow will be explained below.

[1199] Step 1:

[1200] The server loads pre-trained generative AI models, which are the foundation of the system and allow virtual characters to generate natural-sounding conversations and movements.

[1201] Step 2:

[1202] Through the terminal interface, the user inputs a specific content creation request, which includes a content type, for example "news show host," and a specific topic, for example, the latest technology news.

[1203] Step 3:

[1204] The terminal transmits the request input by the user to the server, including the request content and related details.

[1205] Step 4:

[1206] The server analyzes the received request and extracts the necessary information, including the type and topic of the content.

[1207] Step 5:

[1208] The server retrieves the latest relevant data from external databases and APIs, for example connecting to a news API to provide the latest tech news.

[1209] Step 6:

[1210] The server uses a generative AI model based on the external data it retrieves to generate dialogue for the virtual character, resulting in natural and consistent dialogue.

[1211] Step 7:

[1212] The server then uses the generated dialogue to set the virtual character's facial expressions and movements, using a generative AI model to set the appropriate expressions and movements at the appropriate times.

[1213] Step 8:

[1214] The server structures the data of the virtual character and sends it to the user's device, including information on the content of the conversation, facial expressions, and movements.

[1215] Step 9:

[1216] The device passes the received data to a rendering engine, which displays the virtual character in real time, allowing the user to view the character's movements and conversations on the device.

[1217] Step 10:

[1218] Users can enter questions or comments about the content they are viewing, and the input is sent from the terminal to the server.

[1219] Step 11:

[1220] The server receives and analyzes input from the user, and generates an appropriate response based on the input using a generative AI model.

[1221] Step 12:

[1222] The server then structures the generated response and sends it back to the user's terminal, containing the new conversation content.

[1223] Step 13:

[1224] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[1225] Example 1

[1226] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1227] Conventional virtual character generation systems require a significant amount of time and manual setup to enable characters to converse and move naturally. It is also difficult to support real-time data updates and rapid interactive responses from users. This results in a poor user experience and makes the system impractical.

[1228] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1229] In this invention, the server includes means for loading a pre-trained generative artificial intelligence model, means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behavior, means for receiving a content creation request from a user, means for acquiring data from an external database or API based on the received request and generating conversation content for the virtual character based on the data, means for setting the facial expressions and movements of the virtual character based on the generated conversation content, and means for delivering the set behavior of the virtual character to a user terminal, thereby enabling the real-time generation and delivery of virtual characters with fast and natural conversation and movements.

[1230] A "generative artificial intelligence model" is an artificial intelligence algorithm that generates natural conversations and actions based on pre-trained data.

[1231] A "virtual character" is a digitally generated fictional character that converses and behaves in a natural, human-like manner.

[1232] A "content creation request" is an instruction from a user to the system to create specific content.

[1233] "Means for setting facial expressions and movements" refers to the methods and technologies for determining and setting the facial expressions and body movements of a virtual character based on the generated conversation content.

[1234] An "external database or API" is a data repository outside the system, or an interface for exchanging information with another program or system.

[1235] "Interactive responses" refer to responses or reactions by a virtual character that are generated in real time based on input from a user.

[1236] "User terminal" refers to a device (e.g., smartphone, tablet, PC, etc.) that a user uses to access and operate the system.

[1237] This invention relates to a system that generates virtual characters using a generative AI model and provides users with high-quality content they desire in real time. Specific embodiments for implementing this system are described in detail below.

[1238] 1. Initializing the VTuber generation AI model

[1239] The server loads a pre-trained generative AI model (e.g., a Transformer model or a natural language processing model like GPT-3) that is required for the virtual character to generate natural-sounding conversations and behaviors. The server loads the model into memory and initializes the necessary pre-processing parameters and settings.

[1240] 2. Acceptance of Content Creation Requests

[1241] Users access the system through a browser or dedicated application and submit content creation requests, which include specific requirements and requirements (e.g., news type or topic segment).

[1242] Examples:

[1243] A user may submit a request saying, "I'd like a virtual character to host the latest technology news."

[1244] 3. Content-based conversation generation

[1245] The server analyzes user requests and retrieves relevant, up-to-date information from external databases and APIs (e.g., news APIs). Based on the retrieved data, the server provides prompts to a generative AI model, which generates lines and dialogue for the virtual character.

[1246] Example prompt sentence:

[1247] "Generate a virtual character to host the latest tech news."

[1248] 4. Character behavior settings

[1249] The server sets the virtual character's facial expressions and movements (e.g., hand gestures and eye movements) based on the generated lines. This is done using a 3D motion engine (e.g., Unity or Unreal Engine). The server analyzes the generated lines, determines facial expressions and movement patterns, and generates the character's movement data.

[1250] 5. Content Delivery

[1251] The server distributes the behavior data of the virtual character to the user's device, which receives the data and displays it to the user in real time using WebRTC and HTTP streaming technologies.

[1252] 6. Interactive Responses

[1253] Users can enter questions or comments while watching. The server receives this input, performs text analysis, and uses a generative AI model to generate an appropriate response. The generated response is then sent back to the device and displayed to the user.

[1254] Examples:

[1255] If a user types a question like "When is the next technical seminar?", the server analyzes the question and generates an appropriate response using a generative AI model. This response is then returned to the user by a virtual character.

[1256] In this way, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content that users desire in real time.

[1257] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1258] Step 1: Initializing the VTuber generation AI model

[1259] The server loads a pre-trained generative AI model. Loading the model involves reading the model weight data from a specified directory. Specifically, it initializes the parameters and configuration values ​​of the generative AI model (e.g., GPT-3) into memory, making the model immediately available for use.

[1260] Input: Model weight data file

[1261] Output: Initialized generative AI model

[1262] Specific operation: The server loads the model weight data from the specified directory and initializes the necessary parameters.

[1263] Step 2: Accepting a content creation request

[1264] Users submit content creation requests using a browser or dedicated application. The request includes the type of content to be generated and any specific requirements. Once the request is sent, the server receives and analyzes it.

[1265] Input: Content production request from user

[1266] Output: Parsed request content

[1267] Specific operation: A user inputs a request using a web form or application interface and sends it. The server receives the request and analyzes the contents.

[1268] Step 3: Content-based conversation generation

[1269] Based on the analyzed request, the server retrieves the necessary data from external databases or APIs. For example, it retrieves the latest technology news from a news API. The retrieved data is converted into prompts and input into a generative AI model. The model then generates natural-sounding dialogue based on these prompts.

[1270] Input: Parsed request content, external data to be retrieved

[1271] Output: Generated dialogue

[1272] Specific operation: The server sends a data acquisition request to an external API, converts the obtained data into a prompt sentence, and inputs the prompt sentence into the generative AI model to generate dialogue.

[1273] Step 4: Setting character behavior

[1274] The server sets the virtual character's facial expressions and movements based on the generated lines. This is done using a 3D motion engine (such as Unity or Unreal Engine). The server determines facial expressions and body movements and generates them as data so that the lines and movements match.

[1275] Input: Generated dialogue

[1276] Output: Set facial expressions and movement data

[1277] Specific movements: The server analyzes the dialogue and determines facial expressions and movements. A 3D motion engine is used to generate movement data.

[1278] Step 5: Deliver your content

[1279] The server delivers the configured behavior data to the user device using WebRTC or HTTP streaming technology. The device decodes the received data in real time and displays it to the user.

[1280] Input: Set behavior data

[1281] Output: Display on the user's terminal

[1282] Specific operation: The server encodes behavioral data and sends it in streaming format. The device receives the data, decodes it in real time, and displays it.

[1283] Step 6: Interactive Response

[1284] Users can enter questions or comments while watching content. The server analyzes these inputs and uses a generative AI model to generate appropriate responses, which are then sent to the device and displayed to the user.

[1285] Input: User questions and comments

[1286] Output: The generated response

[1287] How it works: The user enters a question or comment into the chat box and sends it. The server analyzes the input and generates a response using the generative AI model. The response is sent to the device and displayed to the user.

[1288] Through these processing steps, the system can provide high-quality virtual characters and natural conversations, and generate and deliver content desired by users in real time.

[1289] (Application example 1)

[1290] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1291] Current content delivery services make it difficult for users to obtain real-time updates and interact with virtual characters based on them. This limits the user experience and lacks dynamic and personalized content delivery. In particular, there is a demand for real-time conversation generation and responses based on news and trend information, but there is a lack of effective methods to achieve this.

[1292] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1293] In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative AI model according to accumulated data and instructed behaviors; means for receiving a content creation request from a user; means for generating conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for acquiring the latest data from an external API and generating conversation content for the virtual character using the generative AI model in response to prompts based on the data; and means for a user to interact with the virtual character via a smart device, thereby enabling a user to enjoy real-time interactive conversations based on the latest information.

[1294] A "generative artificial intelligence model" is an artificial intelligence algorithm that is trained using large datasets and is capable of generating new data and content, especially natural-sounding conversations and behaviors, under certain conditions.

[1295] A "virtual character" is a virtual personality or character that can behave dynamically and realistically on a digital device.

[1296] A "content creation request" is a request that a user sends to the system specifying the specific details and requirements of the content to be created.

[1297] "Conversational content" refers to words and phrases that a virtual character utters to interact with a user or provide information.

[1298] "Expressions and movements" refers to the facial expressions and body movements used by a virtual character to express emotions.

[1299] A "user terminal" is a device such as a smartphone, tablet, PC, or head-mounted display that allows a user to interact with virtual characters or receive content via the Internet.

[1300] An "external API" is a programmatic interface to external databases and services, and is a mechanism used by systems to obtain the latest data in real time.

[1301] A "prompt" is an input text given to a generative AI model, which serves as an instruction for the AI ​​to generate conversations and content based on that text.

[1302] "Smart devices" is a general term for intelligent electronic devices with internet connectivity, such as smartphones, tablets, and smartwatches.

[1303] "Interactive dialogue" refers to two-way communication in which a virtual character generates and responds to questions or comments entered by a user in real time.

[1304] The system for implementing the present invention comprises the following steps: Specific hardware and software usage methods will be described in detail below.

[1305] Program Overview

[1306] First, the server uses a generative AI model to generate the digital characters used on the client device. To do this, it loads a generative AI model that has been trained in advance using a huge dataset. The software used is Python and a generative AI model library (e.g., GPT-3).

[1307] Hardware and Software Details

[1308] Hardware:

[1309] server:

[1310] A server with a powerful processor (e.g., Intel Xeon), lots of memory (e.g., 64GB RAM), and a fast network connection.

[1311] User device:

[1312] Smartphone, head-mounted display (HMD), or computer.

[1313] software:

[1314] Python: Used for program development and operating AI models.

[1315] Generative AI models: Generative AI model libraries (e.g., GPT-3) are used to generate conversation content and character behavior.

[1316] API: Use external APIs (e.g. news API) to get the latest data.

[1317] Data processing and calculation

[1318] The server processes the data and performs calculations using the following procedure.

[1319] 1. Initializing the generative AI model

[1320] The server loads the generative AI model and serves as the foundation for the system, providing the foundation for the virtual characters to generate natural conversations and movements.

[1321] 2. Acceptance of Content Creation Requests

[1322] Users send content creation requests to the server through the human interface of their smart devices, which include specific requests for the virtual character.

[1323] 3. Acquiring external data

[1324] The server retrieves the latest data (e.g., the latest news) in real time through an external API, which is then used to generate content.

[1325] 4. Conversation generation

[1326] Based on the acquired data, a generative AI model is used to generate conversational content for the virtual character. A prompt is set, and the AI ​​model generates an appropriate response based on that prompt.

[1327] 5. Setting the facial expressions and movements of the virtual character

[1328] The facial expressions and movements of the virtual character are set based on the generated conversation content, resulting in more realistic and natural interactions.

[1329] 6. Content Delivery

[1330] The server transmits the behavior of the virtual character that has been set to the user's device, which receives it in real time and displays it to the user.

[1331] Specific examples

[1332] Example 1: News distribution

[1333] The user requests, "What is the latest technology news?"

[1334] The server retrieves the latest technology news data from the news API.

[1335] The generative AI model was given the prompt, "The user wants to know the latest technology news. News content: {latest news}."

[1336] The virtual character responds, "I'll report on the latest tech news."

[1337] The user asked, "Tell me more about this technology."

[1338] A virtual character follows and explains the details of the news.

[1339] Example prompt sentence:

[1340] "Users want to know the latest tech news. Please report on this news in detail."

[1341] This system allows users to enjoy real-time, interactive dialogue based on the latest information.

[1342] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1343] Step 1:

[1344] The server initializes the generative AI model and loads the trained model. As input, it requires the trained generative AI model, which prepares the virtual character to generate natural conversations and actions. As output, it obtains the initialized generative AI model.

[1345] Step 2:

[1346] A user uses a smart device to input a content creation request. The input includes the specific content desired by the user and their requirements. The smart device then sends the request to the server. The output is the received content creation request.

[1347] Step 3:

[1348] The server calls an external API to retrieve the latest data. The input requires the endpoint URL of the external API and authentication information. News data, trend data, etc. are retrieved from the external API. The output is the retrieved latest data.

[1349] Step 4:

[1350] The server generates a prompt sentence based on the acquired data. The latest data acquired from an external API is required as input. The prompt sentence to be given to the generative AI model is formed. The prompt sentence is obtained as output.

[1351] Step 5:

[1352] The server inputs the generated prompt sentence into the AI ​​model to generate the conversation content of the virtual character. The input requires the prompt sentence and the generation AI model. The generation AI model generates natural conversation content based on the prompt sentence. The generated conversation content is obtained as the output.

[1353] Step 6:

[1354] The server sets the facial expressions and movements of the virtual character based on the generated conversation content. The generated conversation content is required as input. Settings are made so that the virtual character will make facial expressions and movements according to the content. The set facial expressions and movements are obtained as output.

[1355] Step 7:

[1356] The server delivers the virtual character's behavior to the user's device. The inputs required are the user's facial expressions, movements, and generated conversations. The smart device receives this information and displays it to the user in real time. The output is the virtual character's behavior displayed on the user's device.

[1357] Step 8:

[1358] The user inputs questions and comments while viewing the content. The user's questions and comments are required as input. The device sends the input content to the server. The user's input content is obtained as output.

[1359] Step 9:

[1360] The server analyzes the input from the user and generates an appropriate response using a generative AI model. The input requires a user's question or comment. The generative AI model generates an appropriate response based on the input. The output is the generated response.

[1361] Step 10:

[1362] The server sends the generated response to the user's terminal, where the virtual character displays the response. The input requires the generated response, which the user terminal displays in real time. The output is the virtual character's response, which is displayed to the user.

[1363] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1364] This invention relates to a system that generates virtual characters using a generative AI model combined with an emotion engine, and provides high-quality, interactive content based on the user's emotions. Below, the program processing of this system is explained in natural language and explained in detail with specific examples.

[1365] Initializing the VTuber generation AI model

[1366] The server first loads a pre-trained generative AI model, which contains the basic algorithms that allow virtual characters to generate natural-sounding speech and movements.

[1367] Emotion engine initialization

[1368] The server also initializes an emotion engine to recognize the user's emotions. This emotion engine analyzes the user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[1369] Accepting content creation requests

[1370] Users input specific content creation requests through the device interface, specifying content types and topics, such as "I would like a host for the latest technology news."

[1371] Content-based conversation generation

[1372] The server receives and analyzes user requests. Based on the analysis results, it retrieves the latest necessary data from external databases and APIs. Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[1373] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[1374] Character behavior settings

[1375] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set the most appropriate facial expressions and movements according to the user's emotional state.

[1376] As a specific example, if the user is excited and happy, the character can be set to smile and move in a lively manner.

[1377] Content Delivery

[1378] The server transmits the behavior data of the configured virtual character to the user's device, which then renders the received data in real time and displays it to the user, allowing the user to watch and listen to the virtual character's conversations and movements.

[1379] Interactive Response

[1380] Users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion engine, the generative AI model generates an appropriate response.

[1381] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion engine determines the user's interest and level of attention, and the generative AI model responds with a positive comment along with appropriate detailed information.

[1382] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[1383] The processing flow will be explained below.

[1384] Step 1:

[1385] The server initializes a generative AI model, which is a pre-trained model that contains the basic algorithms for generating natural-sounding speech and behavior for virtual characters.

[1386] Step 2:

[1387] The server initializes an emotion engine for recognizing the user's emotions. This emotion engine has the function of analyzing the user's facial expressions, tone of voice, and input text information to recognize the user's emotional state in real time.

[1388] Step 3:

[1389] Users input content creation requests through their terminals, making specific requests such as "a virtual character that hosts the latest technology news."

[1390] Step 4:

[1391] The terminal transmits the content creation request input by the user to the server, and the transmitted data includes the type and topic of the content.

[1392] Step 5:

[1393] The server analyzes the request received from the user and retrieves the necessary data from external databases or APIs based on the analysis results. For example, it retrieves the latest technology news data from a news API.

[1394] Step 6:

[1395] The server uses the acquired data to generate conversational content for the virtual character using a generative AI model, ensuring that the conversational content is natural and consistent.

[1396] Step 7:

[1397] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion engine to set optimal facial expressions and movements according to the user's real-time emotions.

[1398] Step 8:

[1399] The server structures the data of the virtual character and transmits it to the user's device, including the content of the conversation, facial expressions, and movements.

[1400] Step 9:

[1401] The device passes the received data to a rendering engine to display the virtual character in real time, allowing the user to watch the character's conversations and movements.

[1402] Step 10:

[1403] The user can input questions or comments about the content being viewed. For example, the user can input a question such as "When is the next technical seminar?"

[1404] Step 11:

[1405] The device sends questions and comments from the user to the server, and the sent data includes the user's input.

[1406] Step 12:

[1407] The server receives and analyzes questions and comments from users. Based on the analysis results and information from the emotion engine, an appropriate response is generated using a generative AI model.

[1408] Step 13:

[1409] The server structures the generated response and sends it back to the user terminal, which contains the new conversation content.

[1410] Step 14:

[1411] The device receives the new conversation data and uses a rendering engine to play it on the virtual character, allowing the user to experience an interactive conversation.

[1412] Through these steps, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly.

[1413] Example 2

[1414] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1415] In conventional virtual character systems, the characters' facial expressions and movements are fixed, making it difficult to realize interactive responses and behaviors that correspond to the user's emotions. Furthermore, because the latest information cannot be acquired in real time, the generated content is often out of date. This makes it difficult to maintain user satisfaction and interest, and limits the interactive experience they can provide.

[1416] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1417] In this invention, the server includes means for generating a realistic and consistent video character using a generative computer model in accordance with accumulated information and instructed actions, means for receiving a content creation request from a user, means for generating dialogue content for the video character based on the received request, means for setting facial expressions and actions of the video character based on the generated dialogue content, means for delivering the set behavior of the video character to a user device, means for initializing an emotion analysis engine and analyzing the emotional state of the user, and means for setting optimal facial expressions and actions based on the analysis results. This enables interactive response according to the user's emotions and allows the latest information to be acquired in real time, making it possible to provide high-quality and interesting content.

[1418] A "generative computer model" is a computer system that makes automated decisions and generates for a specific purpose based on stored information and instructed actions.

[1419] A "video character" is a computer-generated digital representation that has human-like appearance and behavior.

[1420] "User" means a person who uses this system to input content creation requests and view interactive content.

[1421] A "content creation request" is a request from a user to create content based on a particular topic or format.

[1422] "Dialogue content" refers to sentences and conversations spoken by a video character, generated based on user requests and external information.

[1423] "Expressions and actions" refer to facial changes and body movements that visually express emotions and intentions of a video character.

[1424] "Behavior data" is information that describes the facial expressions and movements of a video character.

[1425] A "user device" is an electronic device that a user uses to view content.

[1426] An "emotion analysis engine" is a system that analyzes data such as a user's facial expressions, tone of voice, and text to detect their emotional state.

[1427] "Emotional state" refers to the user's current emotional or psychological state.

[1428] This invention relates to a system that uses a generative computer model combined with an emotion analysis engine to generate realistic and consistent video characters and provide high-quality, interactive content based on the user's emotional state. Specific embodiments for implementing this system are described below.

[1429] System Overview

[1430] The system consists of three main components: a server, a terminal, and a user. The server runs the generative computational model and the sentiment analysis engine, and the terminal provides the interface between the user and the server.

[1431] Hardware and Software Configuration

[1432] server:

[1433] The server is a high-performance computer system that runs a generative computational model (e.g., a computational system model) and an emotion analysis engine (e.g., a facial expression analysis program and an emotional tone analysis program).

[1434] The server also includes an interface for connecting with external information sources and APIs (e.g., information retrieval programs).

[1435] Device:

[1436] A terminal is a device (e.g., a computer, smartphone, or tablet) used by a user and provides a user interface.

[1437] The terminal has a network connection for communicating with the server.

[1438] User:

[1439] Using this system, users can input content creation requests and engage in interactive dialogue with virtual characters.

[1440] Specific operation of the system

[1441] Initialization procedure for high performance computer systems:

[1442] The server loads pre-trained generative computer models that allow video characters to generate natural-sounding speech and movements.

[1443] At the same time, the emotion analysis engine is initialized to prepare for analyzing the user's emotional state in real time.

[1444] Steps for users to enter their request:

[1445] Through the interface, users input specific content creation requests, for example, "I would like to host a segment on the latest technology news."

[1446] Other example prompts include "Tell me about the next tech seminar" and "Tell me some interesting technology topics."

[1447] Server data acquisition and analysis procedure:

[1448] The server receives and analyzes requests from users, and based on the analysis results, retrieves the latest data needed from external sources or APIs.

[1449] For example, in the case of technology news, the latest related technology news data is acquired.

[1450] Steps for generating and configuring conversation content:

[1451] The server uses a generative computer model based on the acquired data to generate dialogue content for the video character.

[1452] Based on the generated dialogue, the emotion analysis engine analyzes the user's emotional state and sets the most appropriate facial expressions and actions.

[1453] For example, if the user is excited, the video character is set to have an excited expression and movements.

[1454] To distribute your content:

[1455] The server transmits the behavior data of the set video character to the terminal in real time.

[1456] The device renders the received data in real time and displays it to the user, allowing the user to view interactive content.

[1457] Interactive response steps:

[1458] Users can enter questions or comments about the content they are viewing.

[1459] The terminal sends the input to the server, which again analyzes the user's emotional state using an emotion analysis engine and generates an appropriate response in a generative computer model.

[1460] The server generates a response and sends it back to the terminal, allowing the user to enjoy an interactive dialogue.

[1461] In this way, the system can recognize users' emotions in real time and provide interactive experiences that respond to them. Based on specific scenarios and usage patterns, users can generate and enjoy high-quality content.

[1462] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1463] Step 1:

[1464] Loading a generative AI model

[1465] The server loads a pre-trained generative AI model. Specifically, the server loads the generative AI model (e.g., a computational system model) from storage into memory and initializes it. The input is the model file on storage, and the output is the generative AI model launched in memory.

[1466] Step 2:

[1467] Initializing the sentiment analysis engine

[1468] The server initializes the emotion analysis engine, which prepares it to analyze the user's emotional state in real time. Specifically, the server starts the emotion analysis engine (e.g., a facial expression analysis program or an emotional tone analysis program) and checks its operation using test data. The input is the emotion analysis engine program, and the output is the started emotion analysis engine.

[1469] Step 3:

[1470] Accepting content creation requests

[1471] A user inputs a specific content creation request through the interface. Specifically, the user enters the request text (e.g., "I would like you to host a segment on the latest technology news") into the device's input field and clicks the send button. The input is the request text entered by the user, and the output is the request data sent to the server.

[1472] Step 4:

[1473] Request analysis and data acquisition

[1474] The server analyzes requests received from users and retrieves the necessary data from relevant external information sources or APIs. Specifically, the server performs natural language analysis on the request content, generates an API request, sends it to an external database or API (e.g., an information retrieval program), and receives the retrieved data. The input is the user's request content, and the output is the data retrieved from the external database or API.

[1475] Step 5:

[1476] Conversation generation

[1477] The server uses a generative AI model based on the acquired data to generate dialogue for the video character. Specifically, the server inputs the acquired data into the generative AI model and outputs the generated dialogue in text format. The input is data acquired from an external database or API, and the output is the text data of the generated dialogue.

[1478] Step 6:

[1479] Character facial expressions and movement settings

[1480] Based on the generated conversation content, the server uses an emotion analysis engine to analyze the user's emotional state and set the optimal facial expressions and movements for the video character. Specifically, the server inputs the generated conversation content and the user's real-time emotional data, which the emotion analysis engine analyzes to generate the character's facial expressions and movement data. The input is the generated conversation content and the user's emotional data, and the output is the character's facial expressions and movement data.

[1481] Step 7:

[1482] Sending behavioral data

[1483] The server transmits the behavior data of the configured video character to the terminal in real time. Specifically, it executes a process to transmit the generated behavior data to the terminal via the network. The input is the character's facial expression and movement data, and the output is the behavior data transmitted to the terminal.

[1484] Step 8:

[1485] Rendering and Displaying Data

[1486] The device renders the received behavior data in real time and displays it to the user. Specifically, the device processes the received data using a 3D model and animation engine, and displays a visualized character to the user. The input is the received behavior data, and the output is the video character displayed to the user.

[1487] Step 9:

[1488] Accepting User Input

[1489] Users can enter questions or comments about the content they are viewing. Specifically, the user enters the question or comment in text in the input field on the device and clicks the send button. The input is the question or comment entered by the user, and the output is the input data sent to the server.

[1490] Step 10:

[1491] Interactive response generation and delivery

[1492] The server analyzes the input data received from the user, re-analyzes the user's emotional state using the emotion analysis engine, and generates an appropriate response using the generative AI model. The generated response is sent back to the device, which then displays it to the user. Specifically, the server analyzes the input data, generates a response based on the generative AI model, and the emotion analysis engine complements the response. The input is a question or comment from the user, and the output is an interactive response sent to the device.

[1493] This allows the system to provide a high quality and interactive experience to the user.

[1494] (Application example 2)

[1495] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1496] Conventional virtual character systems provide content in a one-way manner without considering the user's emotional state. This limits the improvement of the user experience and lacks true interactivity. Furthermore, they lack the ability to respond in real time to user input requests, making it difficult to respond flexibly to the user's interests and emotions.

[1497] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes: means for generating a virtual character with realistic and consistent character characteristics using a generative artificial intelligence model in accordance with accumulated data and instructed behavior; means for receiving a content creation request from a user; means for generating a conversation content for the virtual character based on the received request; means for setting the facial expressions and movements of the virtual character based on the generated conversation content; means for delivering the set behavior of the virtual character to a user terminal; means for an emotion recognition engine for recognizing the emotional state of the user in real time; and means for optimizing the behavior of the virtual character based on the recognized emotional state. This makes it possible to recognize the emotional state of the user and provide high-quality, interactive content that is customized accordingly.

[1498] A "generative artificial intelligence model" is an algorithm or program that generates the conversation content and behavior of a virtual character based on accumulated data and instructed behavior.

[1499] A "virtual character" is a character that has a realistic and consistent personality in a digital space and can interact with users.

[1500] A "content creation request" is a request or instruction a user makes to a virtual character regarding a particular topic or content.

[1501] An "emotion recognition engine" is a technology that analyzes a user's facial expressions, tone of voice, and input text to recognize the user's emotional state in real time.

[1502] An "interactive response" is a response generated by a virtual character to respond in real time to a user's questions or comments and to continue the dialogue with the user.

[1503] "External databases and APIs" are external sources used by virtual characters to obtain up-to-date data for generating dialogue.

[1504] "Behavior" refers to the facial expressions and actions that a virtual character makes toward the user.

[1505] This invention relates to a system that provides high-quality, interactive content based on user emotions using a generative AI model combined with an emotion recognition engine. The program processing of this system is described below.

[1506] The server first loads a pre-trained generative artificial intelligence model, which contains the basic algorithms that allow a virtual character to generate natural-sounding conversations and movements.

[1507] The emotion recognition engine is also initialized at the same time. The emotion recognition engine is a system that analyzes the user's facial expressions, tone of voice, input text, etc. to recognize their emotional state in real time. To identify emotions, it uses technologies such as OpenCV and deep learning models (e.g., TensorFlow, PyTorch).

[1508] Users input specific content creation requests through a smartphone interface, specifying the type of content and the topic, such as "I want you to host the latest technology news."

[1509] The server receives and analyzes the user's request. Based on the analysis results, it retrieves the latest necessary data from an external database, such as a news API (e.g., Google News API). Based on the retrieved data, it uses a generative AI model to generate the conversation content of the virtual character.

[1510] As a concrete example, if a user requests "I want to hear the latest technology news," the server retrieves the latest relevant technology news data from a news API, and then uses the generative AI model to generate summaries and details of the news articles.

[1511] The server sets the virtual character's facial expressions and movements based on the generated conversation content, using an emotion recognition engine to set the most appropriate facial expressions and movements according to the user's emotional state, and then transmits the set virtual character's behavior data to the user's device.

[1512] The device renders the received data in real time and displays it to the user, allowing the user to see and hear the virtual character's conversations and movements.

[1513] Additionally, users can enter questions or comments about the content they are watching. The device sends the input to the server, which analyzes it. Based on the analysis results and information from the emotion recognition engine, the generative AI model generates an appropriate response.

[1514] As a concrete example, if a user types, "Please tell me about the next technical seminar," the server analyzes the details of the seminar, the emotion recognition engine determines the user's level of interest, and the generative AI model responds with a positive comment along with appropriate details.

[1515] Prompt Sentence Examples

[1516] User Input: "What's in tech news today?"

[1517] Generated AI prompt:

[1518] A user has requested, "What's in tech news today?" Use the news API to get the latest tech news and generate energetic and positive conversations when users are excited.

[1519] In this way, the system recognizes the user's emotions and provides a virtual character with interactive conversations and natural movements accordingly, thereby realizing a high-quality interactive experience and improving user satisfaction.

[1520] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1521] Step 1:

[1522] The server loads and initializes a pre-trained generative artificial intelligence model, taking the parameter data of the pre-trained generative AI model as input and obtaining the prepared state of the AI ​​model as output, which will later generate a virtual character.

[1523] Step 2:

[1524] The server initializes the emotion recognition engine, which takes as input the configuration file of the emotion recognition algorithm and as output is ready to analyze the user's emotional state in real time, so that it can process the emotional data from the user.

[1525] Step 3:

[1526] Users input content creation requests through a smartphone interface. The input includes the type of content and topic desired by the user, and the output provides detailed data about the request, which can then be processed.

[1527] Step 4:

[1528] The server parses the request received from the user, taking the user request data as input and providing the structured data of the parsed request as output, which allows the retrieval of the data needed for the next processing step.

[1529] Step 5:

[1530] The server retrieves the latest data related to the requested content from an external database or API. The input is the structured data of the parsed request, and the output is the latest retrieved data. This allows the server to provide information that meets the latest requests. A specific example is using a news API to retrieve the latest technology news.

[1531] Step 6:

[1532] The server uses a generative AI model based on the latest acquired data to generate the conversation content of the virtual character. The input is news data and other external data, and the output is the generated conversation content data. This gives shape to the dialogue content of the virtual character.

[1533] Step 7:

[1534] The server sets the virtual character's facial expressions and movements based on the generated conversation content. The inputs are conversation content data and user emotion data obtained from an emotion recognition engine, and the output is movement setting data for the virtual character. This allows the character to behave naturally according to the user's emotions.

[1535] Step 8:

[1536] The server sends the behavior data of the configured virtual character to the user's device. The input is the virtual character's action setting data, and the output is the data sent to the user's device. This allows the user to watch the character's conversation and actions.

[1537] Step 9:

[1538] The device renders the received data in real time and displays it to the user. The input is behavior data sent from the server, and the output is a rendered virtual character. This allows the user to enjoy interactive dialogue with the character.

[1539] Step 10:

[1540] The user enters a question or comment about the content they are viewing. The input is the user's text input, and the output is the text data. This is used to generate the next response.

[1541] Step 11:

[1542] The terminal sends the user's input to the server. The input is the user's comment data, and the output is the data transmission to the server, so that the server can prepare an interactive response.

[1543] Step 12:

[1544] The server generates appropriate responses using a generative AI model based on user input and emotion recognition data. The inputs are user comment data and current emotion data, and the output is interactive response data. This allows the virtual character to respond appropriately in real time.

[1545] In this way, by performing each processing step in detail, high-quality interactive virtual character dialogue optimized for the user's emotions is realized.

[1546] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1547] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1548] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1549] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1550] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1551] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1552] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1553] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1554] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1555] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1556] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1557] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1558] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1559] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1560] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1561] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1562] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1563] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1564] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1565] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1566] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1567] The following is further disclosed regarding the above embodiment.

[1568] (Claim 1)

[1569] A means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behaviors using a generative artificial intelligence model;

[1570] means for receiving a content creation request from a user;

[1571] means for generating a conversation by a virtual character based on the received request;

[1572] a means for setting facial expressions and movements of a virtual character based on the generated conversation content;

[1573] A means for delivering the behavior of the set virtual character to a user terminal;

[1574] A system including:

[1575] (Claim 2)

[1576] 10. The system of claim 1, further comprising means for generating an interactive response by the virtual character based on input from the user.

[1577] (Claim 3)

[1578] 2. The system according to claim 1, further comprising means for obtaining the latest data from an external database or API and generating conversation content based on the data.

[1579] "Example 1"

[1580] (Claim 1)

[1581] means for loading a pre-trained generative artificial intelligence model;

[1582] A means for generating a virtual character with realistic and consistent character traits according to accumulated data and instructed behaviors;

[1583] means for receiving a content creation request from a user;

[1584] A means for retrieving data from an external database or API based on the received request and generating conversation content for a virtual character based on the data;

[1585] a means for setting facial expressions and movements of a virtual character based on the generated conversation content;

[1586] A means for delivering the behavior of the set virtual character to a user terminal;

[1587] A system including:

[1588] (Claim 2)

[1589] 10. The system of claim 1, further comprising means for generating an interactive response by the virtual character based on input from the user.

[1590] (Claim 3)

[1591] 10. The system of claim 1, further comprising means for generating conversational content in real time using a generative artificial intelligence model based on current data.

[1592] "Application Example 1"

[1593] (Claim 1)

[1594] A means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behaviors using a generative artificial intelligence model;

[1595] means for receiving a content creation request from a user;

[1596] means for generating a conversation by a virtual character based on the received request;

[1597] a means for setting facial expressions and movements of a virtual character based on the generated conversation content;

[1598] A means for delivering the behavior of the set virtual character to a user terminal;

[1599] A means to obtain the latest data from an external API and generate the conversation content of the virtual character using a generative AI model with prompts based on that data;

[1600] a means for a user to interact with a virtual character through a smart device;

[1601] A system including:

[1602] (Claim 2)

[1603] 10. The system of claim 1, further comprising means for generating an interactive response by the virtual character based on input from the user.

[1604] (Claim 3)

[1605] 2. The system of claim 1, further comprising means for obtaining the latest data from an external database or API and using the data to generate conversation content using a generative AI model.

[1606] "Example 2: Combining Emotion Engines"

[1607] (Claim 1)

[1608] A means for generating realistic and consistent visual characters using a generative computer model according to stored information and instructed actions;

[1609] means for receiving content production requests from users;

[1610] means for generating dialogue content for a video character based on the received request;

[1611] a means for setting facial expressions and actions of a video character based on the generated dialogue content;

[1612] means for delivering the set behavior of the video character to a user device;

[1613] means for initializing an emotion analysis engine to analyze the user's emotional state;

[1614] A means to set optimal facial expressions and movements based on the analysis results,

[1615] A system including:

[1616] (Claim 2)

[1617] 10. The system of claim 1, further comprising means for generating an interactive response by the video character based on input from the user.

[1618] (Claim 3)

[1619] 2. The system according to claim 1, further comprising means for obtaining up-to-date information from an external information source or API, and generating dialogue content based on the information.

[1620] "Application example 2 when combining emotion engines"

[1621] (Claim 1)

[1622] A means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behaviors using a generative artificial intelligence model;

[1623] means for receiving a content creation request from a user;

[1624] means for generating a conversation by a virtual character based on the received request;

[1625] a means for setting facial expressions and movements of a virtual character based on the generated conversation content;

[1626] A means for delivering the behavior of the set virtual character to a user terminal;

[1627] an emotion recognition engine for recognizing the emotional state of a user in real time;

[1628] means for optimizing the behavior of the virtual character based on the recognized emotional state;

[1629] A system including:

[1630] (Claim 2)

[1631] 10. The system of claim 1, further comprising means for generating an interactive response to the generated content based on the user's emotional state and request.

[1632] (Claim 3)

[1633] 10. The system of claim 1, further comprising means for obtaining up-to-date data from an external database or API, and generating conversation content based on the data and the user's emotional state. [Explanation of symbols]

[1634] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for generating a virtual character with realistic and consistent character characteristics according to accumulated data and instructed behaviors using a generative artificial intelligence model; means for receiving a content creation request from a user; means for generating a conversation by a virtual character based on the received request; a means for setting facial expressions and movements of a virtual character based on the generated conversation content; A means for delivering the behavior of the set virtual character to a user terminal; A system including:

2. The system of claim 1 , further comprising means for generating an interactive response by the virtual character based on input from the user.

3. The system according to claim 1, further comprising means for acquiring the latest data from an external database or API, and generating conversation content based on the data.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A