System
The system addresses the inefficiencies of conventional drama production by using a generative AI model to create personalized, emotionally tailored short vertical dramas, reducing time and cost while enhancing viewer engagement.
Patent Information
- Application Number
- JP2024130278
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-06
- Publication Date
- 2026-02-19
AI Technical Summary
Conventional drama production methods are costly, time-consuming, and lack the ability to quickly provide diverse storylines tailored to individual viewer preferences and emotional states.
A system that creates a user profile, generates a script using a generative AI model, selects appropriate characters from a database, generates high-resolution 3D models, and delivers the video to users, incorporating an emotion engine to adjust content based on user emotions.
Significantly reduces production time and cost while providing high-quality, personalized short vertical dramas that match viewer preferences and emotional states.
Smart Images

Figure 2026027980000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Current drama production methods require high budgets and time, making them unsuitable for the needs of young people who tend to consume information quickly. Another problem is that actor fees put a strain on production budgets. Furthermore, there is also the issue of not being able to provide diverse storylines in a short amount of time. The objective of this invention is to solve these problems and provide a system that can quickly generate and distribute high-quality short, vertical dramas at low cost. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for creating a user profile, a means for generating a script for a short, vertical drama using a generative AI model, a means for selecting appropriate characters from a database of actor videos and generating high-resolution 3D models, a means for generating videos using the generated script and the selected character models, and a means for delivering the generated videos to users. This system enables the rapid provision of high-quality content while significantly reducing the time and cost required for traditional drama production.
[0006] A "user profile" refers to information that integrates data such as a user's personal information, viewing history, and preferred genres.
[0007] "Generative AI model" refers to an artificial intelligence system that uses machine learning algorithms to automatically generate stories and scripts based on user profiles.
[0008] "Short vertical drama" refers to a drama in a vertical format (usually optimized for smartphones) that can be viewed in a short amount of time.
[0009] A "script" refers to a screenplay that describes the storyline of a drama, the dialogue of characters, and details of scenes.
[0010] An "actor video database" refers to a digital archive that stores video data of various actors.
[0011] "Character" refers to a person who plays a specific role in the generated script.
[0012] A "high-resolution 3D model" refers to a three-dimensional digital model that is generated from video data of an actor and is rendered in great detail.
[0013] "Video editing system" refers to an integrated system of software and hardware for generating actual video using scripts and 3D models.
[0014] An "endpoint" refers to the connection point of a device or application through which a user watches or listens to video.
[0015] "Distribution" refers to the act of sending the generated video to a user's device via the Internet or other means. [Brief explanation of the drawings]
[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0024] [First embodiment]
[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0037] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, and delivers the generated footage to users.
[0038] Program Overview
[0039] The program for this system is composed of, for example, the following steps:
[0040] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile, which then lays the foundation for customization for each user.
[0041] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[0042] The server then selects suitable characters from a database of actor footage based on the generated storyline and generates high-resolution 3D models of the selected characters, which then serve as the basis for the video production.
[0043] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[0044] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[0045] Specific examples
[0046] For example, suppose a user named "Tanaka" likes the action genre. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story.
[0047] Next, the server retrieves data for the actor "Sato" from a database of actor footage and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate a video. Finally, this video is sent to Tanaka's smartphone, where he can watch it on the spot.
[0048] This system can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content. In this way, the present invention can be put into practice.
[0049] The processing flow will be explained below.
[0050] Step 1: Creating a User Profile
[0051] The server retrieves the user's personal information and viewing history from a database.
[0052] The server creates a user profile based on the acquired data.
[0053] For example, it can be determined from the database that user "Tanaka" watches a lot of action movies, and his preferred genre can be recorded as "action" in his profile.
[0054] Step 2: Story Generation
[0055] The server constructs input data for the generative AI model based on the user profile.
[0056] The server provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre" and generates a storyline.
[0057] For example, a generative AI model might generate a story about a brave police officer defeating a villain.
[0058] Step 3: Selecting and modeling your character
[0059] The server extracts the main characters from the generated story.
[0060] The server acquires actor data corresponding to the extracted character from an actor video database.
[0061] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[0062] Step 4: Generate footage
[0063] The server imports the generated storyline and character models into a video editing system.
[0064] The server applies character movements and dialogue based on the script and renders the video.
[0065] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[0066] Step 5: Streaming the video
[0067] The server obtains the terminal endpoint for the user to view.
[0068] The server transmits the generated video to the user's terminal.
[0069] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[0070] Example 1
[0071] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0072] Traditional drama production is time-consuming and costly, making it difficult to provide high-quality content in a short period of time. In addition, there is a lack of methods to quickly provide customized video content tailored to viewer preferences.
[0073] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0074] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting an appropriate character from a character video database and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected character model, and a means for delivering the generated video to the user, thereby enabling the rapid provision of high-quality short vertical drama videos customized to the viewer's preferences.
[0075] A "user profile" is a collection of data created based on a user's personal information, viewing history, preferred genres, etc.
[0076] A "generative AI model" is an algorithm or program that uses artificial intelligence techniques to generate scripts or text in natural language based on a given prompt.
[0077] A "short vertical drama" is a story-based content that can be viewed in a short amount of time, and is primarily intended to be viewed on vertical screens such as smartphones.
[0078] A "script" is a document that describes the dialogue and scene details for video content.
[0079] A "person video database" is a database that stores video data of actors and characters.
[0080] A "character" is a person, animal, or personified being that plays a specific role in the story.
[0081] A "high resolution three-dimensional model" is a detailed three-dimensional object created using computer graphics and displayed at a high resolution.
[0082] "Video" is visual content that expresses movement through a series of images.
[0083] "Distribution" refers to sending the generated video to the user's device so that it can be viewed.
[0084] This invention is a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a human video database, generates high-resolution three-dimensional models, generates video using the generated script and the selected character models, and delivers the generated video to the user.
[0085] Creating a user profile
[0086] First, the server retrieves the user's personal information, viewing history, and preferred genres from a database. This information is used to create a user profile, laying the groundwork for delivering customized content. Using SQL queries, the server extracts the necessary data from the tables where the user information is stored.
[0087] Storyline Generation
[0088] The server then uses the generative AI model to generate a storyline based on the user profile. In this process, the server inputs prompt statements into the generative AI model. The prompt statements can be in the following format:
[0089] "Generating short vertical dramas in the action genre"
[0090] Based on the generated prompts, a generative AI model generates a storyline using a natural language generation algorithm, such as an AI model like GPT-3.
[0091] Character selection and 3D model generation
[0092] After the storyline is generated, the server selects suitable characters from a human video database, filtering data that matches the character attributes (e.g., gender, age, personality, etc.) in the storyline, and uses 3D modeling software such as Blender or Maya to generate high-resolution three-dimensional models of the selected characters.
[0093] Video generation
[0094] The server then combines the generated script with the selected character model to generate a video. Using a video editing system, each line and scene in the script is applied to the character. For example, video editing software such as Adobe Premiere Pro or DaVinci Resolve is used.
[0095] Video distribution
[0096] Finally, the server prepares the resulting video for delivery to the user's device. The video file is encoded into the appropriate format and uploaded to a CDN (Content Delivery Network), where it is delivered to the end user's device using a streaming service (e.g., HLS or DASH).
[0097] Specific examples
[0098] For example, if a user named "Tanaka" likes the action genre, the server creates a profile based on Tanaka's viewing history and preferred genres. The server then provides the generative AI model with a prompt, "Generate a short vertical drama in the action genre," and extracts key characters from the generated storyline. The server then retrieves appropriate actor data from a human video database and generates high-resolution 3D models. These elements are integrated to generate a video, which is ultimately delivered to Tanaka's smartphone. This system allows Tanaka to watch an action-packed short vertical drama on the spot.
[0099] In this way, the system of the present invention can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content that meets the preferences of viewers.
[0100] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0101] Step 1: Creating a User Profile
[0102] Specific behavior:
[0103] The server first obtains the user's personal information, viewing history, and preferred genres from a database.
[0104] Input: User information stored in the database, viewing history, and preferred genres.
[0105] Data processing: Extract data using SQL queries and generate user profile objects.
[0106] Output: A customized user profile object.
[0107] Step 2: Generate a prompt statement
[0108] Specific behavior:
[0109] The server parses information from the user profile and generates prompts to feed into the generative AI model.
[0110] Input: A user profile object.
[0111] Data processing: Create a natural language prompt based on the attributes of the user profile.
[0112] Output: A prompt to input to the generative AI model (e.g., "Generate a short, vertical drama in the action genre").
[0113] Step 3: Generate a storyline
[0114] Specific behavior:
[0115] The server uses a generative AI model to generate a storyline based on the prompt.
[0116] Input: The generated prompt statement.
[0117] Data processing: Use a generative AI model (e.g., GPT-3) to automatically generate storyline text.
[0118] Output: The generated storyline text.
[0119] Step 4: Character Selection
[0120] Specific behavior:
[0121] The server analyzes the character attributes (gender, age, personality, etc.) in the generated storyline and references a video database of characters.
[0122] Input: The generated storyline text.
[0123] Data Processing: Based on character attributes, filter matching database entries to select the best character.
[0124] Output: Data of the selected character.
[0125] Step 5: Generate a high-resolution 3D model
[0126] Specific behavior:
[0127] The server generates a high-resolution three-dimensional model based on the video data of the selected character.
[0128] Input: Selected character's data.
[0129] Data processing: Creating a three-dimensional model using 3D modeling software such as Blender or Maya.
[0130] Output: High resolution 3D character model.
[0131] Step 6: Generate footage
[0132] Specific behavior:
[0133] The server integrates the generated script with the selected character model to generate video content.
[0134] Input: Generated storyline text, high-resolution 3D character models.
[0135] Data processing: Analyze each line and scene in the script and add movement and lines to the characters using a video editing system (Adobe Premiere Pro or DaVinci Resolve).
[0136] Output: The finished video file.
[0137] Step 7: Streaming the video
[0138] Specific behavior:
[0139] The server encodes the finished video into the appropriate format and uploads it to the CDN for delivery to end users.
[0140] Input: Final video file.
[0141] Data processing: Encode the video file, upload it to the CDN, and prepare it for distribution on the streaming service.
[0142] Output: A video stream playable on the end user's device.
[0143] This enables the server to quickly provide high-quality short vertical drama content tailored to the viewer's preferences.
[0144] (Application example 1)
[0145] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0146] Modern content distribution services face challenges in quickly providing content tailored to individual user preferences and interests. Complex video content, such as dramas and animations, requires a high level of customization and rapid generation, but no system currently exists that can achieve this. Therefore, efficiently providing high-quality, personalized video content that satisfies users is a major challenge.
[0147] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0148] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting appropriate characters from an actor video database and generating high-resolution 3D models, a means for generating a video using the generated script and the selected character models, a means for delivering the generated video to a user, a means for generating a story based on the user's genre preferences, and a means for delivering the generated video to a smartphone application. This makes it possible to quickly generate personalized video content based on the user's individual preferences and interests and smoothly deliver it to the user's device.
[0149] A "user profile" is a collection of information such as a user's personal information, viewing history, and preferred genres.
[0150] A "generative AI model" is an artificial intelligence algorithm that uses generative AI to generate sentences or stories based on a specific prompt.
[0151] "Short vertical drama" is a drama-style content in which a story that can be viewed in a short amount of time is displayed on a vertical screen.
[0152] An "actor video database" is a database that stores videos and data of multiple actors.
[0153] A "high-resolution 3D model" is a high-resolution three-dimensional graphic model capable of depicting detailed images.
[0154] "Video generation means" is a function for creating videos using the generated scripts and character models.
[0155] "Distribution means" refers to a system for transmitting the generated video to the user's terminal via the Internet.
[0156] "Genre preferences" refer to the types of dramas and movies that a user is particularly interested in.
[0157] A "smartphone application" is any software program that runs on a smartphone.
[0158] "Story generation means" is a function that allows AI to generate stories based on the user's profile and genre preferences.
[0159] This invention builds a system that generates and delivers customized short, vertical dramas tailored to individual user preferences. This system is realized by creating a user profile, automatically generating drama scripts using a generative AI model based on that information, selecting appropriate characters from a database of actor footage, and generating high-resolution 3D models.
[0160] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile. This user profile is used to gain a detailed understanding of the type of content the user prefers.
[0161] The server then uses a generative AI model to generate a storyline based on the user profile. This generative AI model uses a natural language generation algorithm to automatically generate a story that fits the user's preferred genre. An example of a specific prompt is "Generate a short, vertical drama in the horror genre."
[0162] The server then selects suitable characters from a database of actor footage based on the generated storyline, and generates high-resolution 3D models of the selected characters, which serve as the foundation for the video production.
[0163] The server then generates a video using the generated script and the selected character model. The script is then expanded into lines and scenes using a video editing system (e.g., OpenCV), and the characters move and speak.
[0164] Finally, the server delivers the generated video to a smartphone application. Users can then watch their individually customized, high-quality short vertical dramas on their smartphones. Using a high-performance graphics card (e.g., the NVIDIA RTX series) ensures smooth video playback.
[0165] For example, if a user prefers the action genre, the server creates a profile based on their viewing history and preferred genres and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story. Next, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models. Finally, the generated footage is delivered to the user's smartphone for real-time viewing.
[0166] As described above, by implementing the present invention, it becomes possible to quickly generate personalized video content based on the individual preferences and interests of a user and distribute it to the user's terminal.
[0167] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0168] Step 1:
[0169] The server retrieves the user's personal information, viewing history, preferred genres, etc. from the database and creates a user profile. The input data for the user profile requires the user ID, viewing history, and preferred genres. Based on this, a profile detailing the user's preferences and interests is output.
[0170] Step 2:
[0171] The server uses a generative AI model to generate a storyline based on the user profile. The input data requires the user profile and a prompt, such as "Generate a short, vertical drama in the horror genre." The generative AI model uses these inputs to apply a natural language generation algorithm and output a drama script.
[0172] Step 3:
[0173] The server selects appropriate characters from a video database of actors based on the generated script. In this step, the script content is used as input data, and the IDs and characteristics of the characters that match it are output. This selection determines the main characters of the story.
[0174] Step 4:
[0175] The server generates a high-resolution 3D model of the selected character. The input data for this step is the character's ID and characteristics. Based on this information, a detailed 3D model is output using 3D graphics software (e.g., Blender).
[0176] Step 5:
[0177] The server generates a video using the generated script and the selected character model. Specifically, the script is developed as lines and scenes, and the video is edited so that the characters move and speak the lines. The input data for this step are the script and 3D models, and the final video is output using a video editing system (e.g., OpenCV).
[0178] Step 6:
[0179] The server delivers the generated video to the smartphone application. In this step, the video file is used as input data, and the video is compressed using a high-performance graphics card (e.g., NVIDIA RTX series) and sent to the user's smartphone via the Internet. The final output is video data that can be played on the user's smartphone.
[0180] Step 7:
[0181] The user watches the streamed video on their smartphone. In this step, the smartphone application plays the video file received as input data. Through the application, the user can watch personalized, high-quality short vertical dramas.
[0182] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0183] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, combines it with an emotion engine that recognizes the user's emotions, and delivers the generated footage to the user.
[0184] Program Overview
[0185] The program for this system is composed of, for example, the following steps:
[0186] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, which then forms the basis for customization for each individual user.
[0187] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[0188] The server then uses an emotion engine to recognize the user's current emotional state, which detects emotions by analyzing the user's facial expressions, voice, and text.
[0189] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[0190] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[0191] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[0192] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[0193] Specific examples
[0194] For example, suppose a user named "Tanaka" likes the action genre and is currently in an emotional state of "excitement." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[0195] The server then uses its emotion engine to recognize that Tanaka's emotional state is "excited," and based on this emotional state, dynamically adjusts the input data of the generative AI model to make the generated storyline even more compelling.
[0196] The server extracts the main characters from the generated story, retrieves data on an actor named "Sato" from a database of actor footage, and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate the video.
[0197] Finally, the video is streamed to Tanaka's smartphone, where he can watch it instantly, enjoying an exciting short vertical drama that matches his own emotions.
[0198] This system significantly reduces the time and cost required for conventional drama production, and also makes it possible to quickly provide high-quality content that matches the user's emotions.
[0199] The processing flow will be explained below.
[0200] Step 1: Creating a User Profile
[0201] The server retrieves the user's personal information and viewing history from a database.
[0202] The server creates a user profile based on the acquired data.
[0203] For example, the information that "Tanaka watches a lot of action movies" is found in the database of the user, and the preferred genre is recorded as "action" in the profile.
[0204] Step 2: Recognize emotions
[0205] The server uses an emotion engine to recognize the user's emotional state.
[0206] The server analyzes data such as the user's facial expressions, voice, and text to determine their current emotional state.
[0207] For example, it analyzes Tanaka's facial expressions and voice to recognize that he is in an "excited" state.
[0208] Step 3: Story Generation
[0209] The server builds input data for the generative AI model based on the user profile and recognized emotions.
[0210] The server provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre for excited users," and generates a storyline.
[0211] For example, a generative AI model can generate exciting stories such as "a brave police officer defeats a bad guy."
[0212] Step 4: Character Selection and Modeling
[0213] The server extracts the main characters from the generated story.
[0214] The server acquires actor data corresponding to the extracted character from an actor video database.
[0215] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[0216] Step 5: Generate footage
[0217] The server imports the generated storyline and character models into a video editing system.
[0218] The server applies character movements and dialogue based on the script and renders the video.
[0219] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[0220] Step 6: Streaming the video
[0221] The server obtains the terminal endpoint for the user to view.
[0222] The server transmits the generated video to the user's terminal.
[0223] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[0224] Through the above steps, the present invention can generate storylines tailored to the user's emotions and provide short, personal dramas customized in real time, thereby significantly reducing the time and cost required for traditional drama production and quickly providing high-quality entertainment.
[0225] Example 2
[0226] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0227] Conventional video production systems have had difficulty quickly providing content tailored to a user's preferences and emotional state. Generating high-quality 3D models and videos also requires a huge amount of time and cost. This makes it difficult to provide personalized video content to individual users. Therefore, there is a need for a system that can dynamically recognize emotions based on each user's profile and adjust scripts and videos accordingly.
[0228] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0229] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical video work using a generative AI model, a means for selecting an appropriate person from a video database of actors and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected person model, a means for delivering the generated video to a user, a means for recognizing the user's emotional state, and a means for adjusting the generated script based on the emotional state, thereby enabling the rapid provision of high-quality video content personalized to each user based on the user's profile and real-time emotional state.
[0230] A "user profile" is a profile created based on data such as a user's personal information, viewing history, and preferences, and reflects the user's characteristics and tastes.
[0231] A "generative AI model" is an algorithmic model that uses artificial intelligence to process natural language and generate content, automatically generating scripts for short vertical video works based on user profiles.
[0232] A "short vertical video work" is video content that can generally be viewed in a short period of time, typically from a few minutes to a few tens of minutes, and is intended to be displayed vertically.
[0233] A "script" is a document that describes the lines, scene structure, character actions, etc. in a video work.
[0234] An "actor" is a person who appears in a video work and plays a specific role in it.
[0235] A "video database" is a database that stores video data related to multiple performers and allows for searching and retrieval.
[0236] A "high-resolution 3D model" is a visually detailed and accurate three-dimensional character model generated using computer software.
[0237] "Emotional state" refers to a user's current psychological and emotional state, analyzed using data obtained from facial expressions, voice, text messages, etc.
[0238] "Distribution" refers to the act of transmitting the generated video content to a user's device via the Internet, allowing the user to view it in real time.
[0239] "Adjustment" refers to the process of changing or modifying parts of the script created by the generative AI model based on the emotional state, optimizing it to better match the user's emotions.
[0240] The present invention relates to a system that creates a user profile, generates a script for a short vertical video using a generative AI model, selects suitable actors from a video database of actors, generates high-resolution 3D models, generates a video using the generated script and the selected actor models, and finally delivers the video to the user's device. It also has the ability to recognize the user's emotional state and adjust the generated script based on that.
[0241] Specifically, the system operates in the following steps.
[0242] First, the server retrieves the personal information and viewing history entered by the user from a database and creates a user profile based on information such as the user's name, age, preferred genres, and viewing history.
[0243] Next, the server provides a prompt based on the profile information to the generative AI model. An example of the generative AI model used here is OpenAI's GPT-4. An example of the prompt is "Generate a short, vertical drama in the action genre."
[0244] The generative AI model automatically generates a storyline based on this prompt. For example, it might generate a storyline in which the protagonist infiltrates the enemy's hideout and engages in a spectacular battle.
[0245] The server also uses an emotion engine (e.g., Emotion API or Affectiva) to recognize the user's current emotional state by analyzing emotions from the user's facial expressions, voice, text messages, etc.
[0246] The server dynamically adjusts the input data of the generative AI model based on the recognized emotions, optimizing the storyline to better match the user's emotions. For example, if a user is in an excited state, the server may add more exciting battle scenes.
[0247] Next, the server selects an appropriate person from a video database of actors based on the generated storyline. For example, it retrieves data on an actor named "Sato" from the video database and generates a high-resolution 3D model of him using Unity or Unreal Engine.
[0248] The server uses the generated script and the selected 3D models to generate the video. Video editing software such as Adobe Premiere Pro or Final Cut Pro is used to generate the video. As a result, lines and scenes are developed based on the script, and the video is completed with characters moving and speaking lines.
[0249] Finally, the server delivers the generated video to the user's device using streaming services such as AWS CloudFront or Akamai. The device receives the delivered video in real time, allowing the user to watch it immediately.
[0250] For example, suppose a user named "Tanaka" likes the action genre and is currently in an "excited" emotional state. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt: "Generate a short, vertical drama in the action genre." The generative AI model then generates an exciting storyline that matches Tanaka's "excited state."
[0251] The server then uses its emotion engine to recognize Tanaka's emotional state as "excited" and adjusts the input data of the generative AI model based on that state. The server then selects an actor named "Sato" from the generated storyline, generates a 3D model of him, and finally completes the video and distributes it to Tanaka's smartphone. Tanaka can now enjoy an exciting short vertical drama that matches his emotions at that moment.
[0252] This system can significantly reduce the time and cost required for conventional drama production and quickly provide high-quality content that matches the user's emotions.
[0253] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0254] System program processing flow
[0255] Step 1: Creating a User Profile
[0256] The server obtains the personal information and viewing history entered by the user and creates a user profile. Specifically, the server obtains the user's name, age, preferred genres, viewing history, etc. from a database to generate the profile.
[0257] Input: User information from the database
[0258] Output: User profile
[0259] Specific operation: The server obtains the information "User name: Tanaka, Age: 30, Favorite genre: Action, Viewing history: 5 action movies, 1 comedy movie" and creates a profile for "Tanaka."
[0260] Step 2: Generate a storyline
[0261] The server provides a prompt based on the user profile to the generative AI model to generate a storyline, for example, "Generate a short, vertical drama in the action genre."
[0262] Input: User Profile
[0263] Output: Generated storyline
[0264] Specific operation: Based on the information that "Tanaka's preference is the action genre," the server gives the generative AI model a prompt statement of "Generate a short vertical drama in the action genre." The generative AI model generates a storyline in which "the protagonist infiltrates the enemy's hideout and engages in a spectacular battle."
[0265] Step 3: Recognize emotions
[0266] The device captures the user's facial expressions and voice through a camera and microphone and sends the data to a server, which uses an emotion engine to analyze the user's emotional state.
[0267] Input: User's facial expression data, voice data
[0268] Output: User's emotional state
[0269] Specific operation: The device captures Tanaka's facial expressions with a camera and collects his voice with a microphone. The server analyzes Tanaka's data using an emotion engine and recognizes that Tanaka is in an "excited" state.
[0270] Step 4: Adjusting the storyline
[0271] Based on the recognized emotions, the server dynamically adjusts the input data of the generative AI model, optimizing the storyline to better match the emotions.
[0272] Input: Generated storyline, user's emotional state
[0273] Output: Adjusted storyline
[0274] Specific operation: The server instructs the generative AI model to "add more exciting battle scenes" based on the information "current emotional state: excitement." The generative AI model adds a plot twist to the storyline in which the protagonist uses a secret weapon he found in the basement to wipe out the enemies.
[0275] Step 5: Select and generate a character model
[0276] Based on the generated storyline, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models.
[0277] Enter: the adjusted storyline
[0278] Output: High-resolution 3D character model
[0279] Specific operation: The server retrieves data on "Actor A," who is the best suited to play the main character of the story, from the database and generates a 3D model of Actor A in Unity or Unreal Engine.
[0280] Step 6: Generate footage
[0281] The server uses the generated script and 3D character model to generate video, which is then edited using video editing software such as Adobe Premiere Pro or Final Cut Pro.
[0282] Input: Adjusted storyline, 3D character models
[0283] Output: Generated video
[0284] How it works: Based on the generated script, the server uses a 3D model of Actor A to create a battle scene with an enemy. Adobe Premiere Pro is used to align the timing of lines and scenes and complete the final footage.
[0285] Step 7: Streaming the video
[0286] The server delivers the generated video to the user's device using streaming services such as AWS CloudFront and Akamai.
[0287] Input: Generated video
[0288] Output: Video played on the user's device
[0289] How it works: The server uploads the completed video file to AWS CloudFront and generates a distribution URL. The device receives the URL and immediately plays the video, allowing users to enjoy the short vertical drama on the spot.
[0290] As described above, high-quality video content that matches the user's emotions and preferences is quickly provided through each step.
[0291] (Application example 2)
[0292] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0293] Conventional content delivery services have had difficulty generating and delivering personalized video content in real time based on user preferences and emotions. In particular, providing content that takes into account the user's emotional state is difficult, and there has been a demand for a way to improve the quality of the viewing experience.
[0294] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0295] a means for creating a user profile;
[0296] A means for generating scripts for short vertical dramas using a generative AI model;
[0297] A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models;
[0298] means for generating a video using the generated script and a selected character model;
[0299] a means for delivering the generated video to a user;
[0300] A means for dynamically adjusting input data for a generative AI model using an emotion engine that recognizes user emotions; and
[0301] This will enable the generation and delivery of personalized video content in real time based on the user's emotional state.
[0302] A "user profile" is information created by organizing and analyzing data about a user based on the user's personal information, viewing history, preferred genres, etc.
[0303] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and scripts based on user profiles and specific input data.
[0304] The "actor video database" is a database that stores the video and attribute data of multiple actors, and is used to select appropriate characters.
[0305] "High-Resolution 3D Model" means a highly detailed, high-resolution, three-dimensional computer model used to visually recreate a Selected Character.
[0306] The "emotion engine" is an engine that analyzes and recognizes emotions from a user's facial expressions, voice, text, etc.
[0307] The "video generation means" is a means for developing scenes, lines, etc. and creating video using the generated script and selected character models.
[0308] "Video distribution means" refers to a means for distributing the generated video to end users in real time.
[0309] This invention is a system that creates a user profile, generates personalized short-form dramas for users using a generative AI model, and provides storylines based on the user's emotions using an emotion engine. This system is implemented through the following specific process.
[0310] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, including the user's preferred genres and past viewing history, providing the foundation for customization for each individual user.
[0311] The server then uses a generative AI model to generate a storyline based on the user profile. The generative AI model uses a natural language generation algorithm (e.g., GPT-4) to automatically generate a story that fits the user's preferred genre.
[0312] Furthermore, the server uses an emotion engine to recognize the user's current emotional state. The emotion engine detects emotions by analyzing the user's facial expressions and voice data using a facial recognition camera (e.g., Face API) and voice analysis tools (e.g., Azure Cognitive Services).
[0313] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[0314] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[0315] The server then generates a video using the generated script and the selected character model. The script is then developed into lines and scenes using a video editing system (e.g., Blender), and the characters move and speak.
[0316] Finally, the server delivers the generated video to the user's device. Using a streaming distribution server (e.g., AWS S3 + CloudFront), the video is delivered to the user's device through an endpoint, allowing the user to instantly watch the short vertical drama.
[0317] As a concrete example, suppose a user named "Tanaka" likes the action genre and his current emotional state is "excited." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[0318] An example of a prompt used in this process is:
[0319] "Generate a comedy-action story that is suitable for when the user is in a happy state. The main character should have a sense of humor and the episodes should be thrilling with action scenes."
[0320] As a result, this invention makes it possible to efficiently generate and distribute personalized video content that matches the user's emotions, improving the quality of the viewing experience.
[0321] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0322] Step 1:
[0323] The server retrieves the user's personal information and viewing history from the database and creates a user profile. This profile creation step takes the user ID as input and generates profile data including the user's preferred genres and viewing history as output. Specifically, it uses SQL queries to retrieve user information from the database, organizes the necessary data, and builds the profile.
[0324] Step 2:
[0325] The server uses a generative AI model to generate a storyline based on the user profile. In this step, the user profile data is used as input and the generated storyline is obtained as output. Specifically, a prompt sentence reflecting the user's preferred genre and emotional state is generated, and a story is created using a natural language generation algorithm. The prompt sentence is input into the generative AI model (e.g., GPT-4) to obtain the storyline.
[0326] Step 3:
[0327] The server uses an emotion engine to recognize the user's current emotional state. In this step, the user's real-time facial expressions and voice data are used as input, and the user's emotional state is obtained as output. Specifically, the data is analyzed using a facial recognition camera and voice analysis tools (e.g., Face API, Azure Cognitive Services) to detect the user's emotions.
[0328] Step 4:
[0329] The server dynamically adjusts the input data of the generative AI model based on the recognized emotion. In this step, the emotional state is used as input and a storyline tailored to the emotion is obtained as output. Specifically, the recognized emotion is used to regenerate the prompt sentence, which is then reinput into the generative AI model to generate a story that is appropriate for the emotion. For example, the prompt sentence "Generate an action story that is appropriate for the user's excited state" is input into the generative AI model.
[0330] Step 5:
[0331] The server selects appropriate characters from the actor video database based on the generated storyline and generates high-resolution 3D models. In this step, the storyline is used as input and 3D character models are obtained as output. Specifically, the server extracts the main characters appearing in the story, selects appropriate actors from the video database, and generates their high-resolution 3D models.
[0332] Step 6:
[0333] The server generates a video using the generated script and the selected character model. In this step, the script and character model are used as input, and a short vertical drama video is obtained as output. Specifically, a video editing system (e.g., Blender) is used to create scenes including character movements and dialogue, and the entire video is generated.
[0334] Step 7:
[0335] The server delivers the generated video to the user's device. In this step, the generated video is used as input and the video delivered to the user's device is obtained as output. Specifically, the video is delivered using a streaming distribution server (e.g., AWS S3 + CloudFront), and the user watches this video on their smartphone.
[0336] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0337] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0338] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0339] [Second embodiment]
[0340] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0341] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0342] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0343] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0344] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0345] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0346] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0347] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0348] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0349] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0350] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0351] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0352] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, and delivers the generated footage to users.
[0353] Program Overview
[0354] The program for this system is composed of, for example, the following steps:
[0355] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile, which then lays the foundation for customization for each user.
[0356] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[0357] The server then selects suitable characters from a database of actor footage based on the generated storyline and generates high-resolution 3D models of the selected characters, which then serve as the basis for the video production.
[0358] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[0359] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[0360] Specific examples
[0361] For example, suppose a user named "Tanaka" likes the action genre. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story.
[0362] Next, the server retrieves data for the actor "Sato" from a database of actor footage and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate a video. Finally, this video is sent to Tanaka's smartphone, where he can watch it on the spot.
[0363] This system can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content. In this way, the present invention can be put into practice.
[0364] The processing flow will be explained below.
[0365] Step 1: Creating a User Profile
[0366] The server retrieves the user's personal information and viewing history from a database.
[0367] The server creates a user profile based on the acquired data.
[0368] For example, it can be determined from the database that user "Tanaka" watches a lot of action movies, and his preferred genre can be recorded as "action" in his profile.
[0369] Step 2: Story Generation
[0370] The server constructs input data for the generative AI model based on the user profile.
[0371] The server provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre" and generates a storyline.
[0372] For example, a generative AI model might generate a story about a brave police officer defeating a villain.
[0373] Step 3: Selecting and modeling your character
[0374] The server extracts the main characters from the generated story.
[0375] The server acquires actor data corresponding to the extracted character from an actor video database.
[0376] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[0377] Step 4: Generate footage
[0378] The server imports the generated storyline and character models into a video editing system.
[0379] The server applies character movements and dialogue based on the script and renders the video.
[0380] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[0381] Step 5: Streaming the video
[0382] The server obtains the terminal endpoint for the user to view.
[0383] The server transmits the generated video to the user's terminal.
[0384] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[0385] Example 1
[0386] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0387] Traditional drama production is time-consuming and costly, making it difficult to provide high-quality content in a short period of time. In addition, there is a lack of methods to quickly provide customized video content tailored to viewer preferences.
[0388] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0389] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting an appropriate character from a character video database and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected character model, and a means for delivering the generated video to the user, thereby enabling the rapid provision of high-quality short vertical drama videos customized to the viewer's preferences.
[0390] A "user profile" is a collection of data created based on a user's personal information, viewing history, preferred genres, etc.
[0391] A "generative AI model" is an algorithm or program that uses artificial intelligence techniques to generate scripts or text in natural language based on a given prompt.
[0392] A "short vertical drama" is a story-based content that can be viewed in a short amount of time, and is primarily intended to be viewed on vertical screens such as smartphones.
[0393] A "script" is a document that describes the dialogue and scene details for video content.
[0394] A "person video database" is a database that stores video data of actors and characters.
[0395] A "character" is a person, animal, or personified being that plays a specific role in the story.
[0396] A "high resolution three-dimensional model" is a detailed three-dimensional object created using computer graphics and displayed at a high resolution.
[0397] "Video" is visual content that expresses movement through a series of images.
[0398] "Distribution" refers to sending the generated video to the user's device so that it can be viewed.
[0399] This invention is a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a human video database, generates high-resolution three-dimensional models, generates video using the generated script and the selected character models, and delivers the generated video to the user.
[0400] Creating a user profile
[0401] First, the server retrieves the user's personal information, viewing history, and preferred genres from a database. This information is used to create a user profile, laying the groundwork for delivering customized content. Using SQL queries, the server extracts the necessary data from the tables where the user information is stored.
[0402] Storyline Generation
[0403] The server then uses the generative AI model to generate a storyline based on the user profile. In this process, the server inputs prompt statements into the generative AI model. The prompt statements can be in the following format:
[0404] "Generating short vertical dramas in the action genre"
[0405] Based on the generated prompts, a generative AI model generates a storyline using a natural language generation algorithm, such as an AI model like GPT-3.
[0406] Character selection and 3D model generation
[0407] After the storyline is generated, the server selects suitable characters from a human video database, filtering data that matches the character attributes (e.g., gender, age, personality, etc.) in the storyline, and uses 3D modeling software such as Blender or Maya to generate high-resolution three-dimensional models of the selected characters.
[0408] Video generation
[0409] The server then combines the generated script with the selected character model to generate a video. Using a video editing system, each line and scene in the script is applied to the character. For example, video editing software such as Adobe Premiere Pro or DaVinci Resolve is used.
[0410] Video distribution
[0411] Finally, the server prepares the resulting video for delivery to the user's device. The video file is encoded into the appropriate format and uploaded to a CDN (Content Delivery Network), where it is delivered to the end user's device using a streaming service (e.g., HLS or DASH).
[0412] Specific examples
[0413] For example, if a user named "Tanaka" likes the action genre, the server creates a profile based on Tanaka's viewing history and preferred genres. The server then provides the generative AI model with a prompt, "Generate a short vertical drama in the action genre," and extracts key characters from the generated storyline. The server then retrieves appropriate actor data from a human video database and generates high-resolution 3D models. These elements are integrated to generate a video, which is ultimately delivered to Tanaka's smartphone. This system allows Tanaka to watch an action-packed short vertical drama on the spot.
[0414] In this way, the system of the present invention can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content that meets the preferences of viewers.
[0415] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0416] Step 1: Creating a User Profile
[0417] Specific behavior:
[0418] The server first obtains the user's personal information, viewing history, and preferred genres from a database.
[0419] Input: User information stored in the database, viewing history, and preferred genres.
[0420] Data processing: Extract data using SQL queries and generate user profile objects.
[0421] Output: A customized user profile object.
[0422] Step 2: Generate a prompt statement
[0423] Specific behavior:
[0424] The server parses information from the user profile and generates prompts to feed into the generative AI model.
[0425] Input: A user profile object.
[0426] Data processing: Create a natural language prompt based on the attributes of the user profile.
[0427] Output: A prompt to input to the generative AI model (e.g., "Generate a short, vertical drama in the action genre").
[0428] Step 3: Generate a storyline
[0429] Specific behavior:
[0430] The server uses a generative AI model to generate a storyline based on the prompt.
[0431] Input: The generated prompt statement.
[0432] Data processing: Use a generative AI model (e.g., GPT-3) to automatically generate storyline text.
[0433] Output: The generated storyline text.
[0434] Step 4: Character Selection
[0435] Specific behavior:
[0436] The server analyzes the character attributes (gender, age, personality, etc.) in the generated storyline and references a video database of characters.
[0437] Input: The generated storyline text.
[0438] Data Processing: Based on character attributes, filter matching database entries to select the best character.
[0439] Output: Data of the selected character.
[0440] Step 5: Generate a high-resolution 3D model
[0441] Specific behavior:
[0442] The server generates a high-resolution three-dimensional model based on the video data of the selected character.
[0443] Input: Selected character's data.
[0444] Data processing: Creating a three-dimensional model using 3D modeling software such as Blender or Maya.
[0445] Output: High resolution 3D character model.
[0446] Step 6: Generate footage
[0447] Specific behavior:
[0448] The server integrates the generated script with the selected character model to generate video content.
[0449] Input: Generated storyline text, high-resolution 3D character models.
[0450] Data processing: Analyze each line and scene in the script and add movement and lines to the characters using a video editing system (Adobe Premiere Pro or DaVinci Resolve).
[0451] Output: The finished video file.
[0452] Step 7: Streaming the video
[0453] Specific behavior:
[0454] The server encodes the finished video into the appropriate format and uploads it to the CDN for delivery to end users.
[0455] Input: Final video file.
[0456] Data processing: Encode the video file, upload it to the CDN, and prepare it for distribution on the streaming service.
[0457] Output: A video stream playable on the end user's device.
[0458] This enables the server to quickly provide high-quality short vertical drama content tailored to the viewer's preferences.
[0459] (Application example 1)
[0460] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0461] Modern content distribution services face challenges in quickly providing content tailored to individual user preferences and interests. Complex video content, such as dramas and animations, requires a high level of customization and rapid generation, but no system currently exists that can achieve this. Therefore, efficiently providing high-quality, personalized video content that satisfies users is a major challenge.
[0462] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0463] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting appropriate characters from an actor video database and generating high-resolution 3D models, a means for generating a video using the generated script and the selected character models, a means for delivering the generated video to a user, a means for generating a story based on the user's genre preferences, and a means for delivering the generated video to a smartphone application. This makes it possible to quickly generate personalized video content based on the user's individual preferences and interests and smoothly deliver it to the user's device.
[0464] A "user profile" is a collection of information such as a user's personal information, viewing history, and preferred genres.
[0465] A "generative AI model" is an artificial intelligence algorithm that uses generative AI to generate sentences or stories based on a specific prompt.
[0466] "Short vertical drama" is a drama-style content in which a story that can be viewed in a short amount of time is displayed on a vertical screen.
[0467] An "actor video database" is a database that stores videos and data of multiple actors.
[0468] A "high-resolution 3D model" is a high-resolution three-dimensional graphic model capable of depicting detailed images.
[0469] "Video generation means" is a function for creating videos using the generated scripts and character models.
[0470] "Distribution means" refers to a system for transmitting the generated video to the user's terminal via the Internet.
[0471] "Genre preferences" refer to the types of dramas and movies that a user is particularly interested in.
[0472] A "smartphone application" is any software program that runs on a smartphone.
[0473] "Story generation means" is a function that allows AI to generate stories based on the user's profile and genre preferences.
[0474] This invention builds a system that generates and delivers customized short, vertical dramas tailored to individual user preferences. This system is realized by creating a user profile, automatically generating drama scripts using a generative AI model based on that information, selecting appropriate characters from a database of actor footage, and generating high-resolution 3D models.
[0475] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile. This user profile is used to gain a detailed understanding of the type of content the user prefers.
[0476] The server then uses a generative AI model to generate a storyline based on the user profile. This generative AI model uses a natural language generation algorithm to automatically generate a story that fits the user's preferred genre. An example of a specific prompt is "Generate a short, vertical drama in the horror genre."
[0477] The server then selects suitable characters from a database of actor footage based on the generated storyline, and generates high-resolution 3D models of the selected characters, which serve as the foundation for the video production.
[0478] The server then generates a video using the generated script and the selected character model. The script is then expanded into lines and scenes using a video editing system (e.g., OpenCV), and the characters move and speak.
[0479] Finally, the server delivers the generated video to a smartphone application. Users can then watch their individually customized, high-quality short vertical dramas on their smartphones. Using a high-performance graphics card (e.g., the NVIDIA RTX series) ensures smooth video playback.
[0480] For example, if a user prefers the action genre, the server creates a profile based on their viewing history and preferred genres and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story. Next, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models. Finally, the generated footage is delivered to the user's smartphone for real-time viewing.
[0481] As described above, by implementing the present invention, it becomes possible to quickly generate personalized video content based on the individual preferences and interests of a user and distribute it to the user's terminal.
[0482] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0483] Step 1:
[0484] The server retrieves the user's personal information, viewing history, preferred genres, etc. from the database and creates a user profile. The input data for the user profile requires the user ID, viewing history, and preferred genres. Based on this, a profile detailing the user's preferences and interests is output.
[0485] Step 2:
[0486] The server uses a generative AI model to generate a storyline based on the user profile. The input data requires the user profile and a prompt, such as "Generate a short, vertical drama in the horror genre." The generative AI model uses these inputs to apply a natural language generation algorithm and output a drama script.
[0487] Step 3:
[0488] The server selects appropriate characters from a video database of actors based on the generated script. In this step, the script content is used as input data, and the IDs and characteristics of the characters that match it are output. This selection determines the main characters of the story.
[0489] Step 4:
[0490] The server generates a high-resolution 3D model of the selected character. The input data for this step is the character's ID and characteristics. Based on this information, a detailed 3D model is output using 3D graphics software (e.g., Blender).
[0491] Step 5:
[0492] The server generates a video using the generated script and the selected character model. Specifically, the script is developed as lines and scenes, and the video is edited so that the characters move and speak the lines. The input data for this step are the script and 3D models, and the final video is output using a video editing system (e.g., OpenCV).
[0493] Step 6:
[0494] The server delivers the generated video to the smartphone application. In this step, the video file is used as input data, and the video is compressed using a high-performance graphics card (e.g., NVIDIA RTX series) and sent to the user's smartphone via the Internet. The final output is video data that can be played on the user's smartphone.
[0495] Step 7:
[0496] The user watches the streamed video on their smartphone. In this step, the smartphone application plays the video file received as input data. Through the application, the user can watch personalized, high-quality short vertical dramas.
[0497] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0498] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, combines it with an emotion engine that recognizes the user's emotions, and delivers the generated footage to the user.
[0499] Program Overview
[0500] The program for this system is composed of, for example, the following steps:
[0501] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, which then forms the basis for customization for each individual user.
[0502] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[0503] The server then uses an emotion engine to recognize the user's current emotional state, which detects emotions by analyzing the user's facial expressions, voice, and text.
[0504] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[0505] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[0506] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[0507] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[0508] Specific examples
[0509] For example, suppose a user named "Tanaka" likes the action genre and is currently in an emotional state of "excitement." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[0510] The server then uses its emotion engine to recognize that Tanaka's emotional state is "excited," and based on this emotional state, dynamically adjusts the input data of the generative AI model to make the generated storyline even more compelling.
[0511] The server extracts the main characters from the generated story, retrieves data on an actor named "Sato" from a database of actor footage, and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate the video.
[0512] Finally, the video is streamed to Tanaka's smartphone, where he can watch it instantly, enjoying an exciting short vertical drama that matches his own emotions.
[0513] This system significantly reduces the time and cost required for conventional drama production, and also makes it possible to quickly provide high-quality content that matches the user's emotions.
[0514] The processing flow will be explained below.
[0515] Step 1: Creating a User Profile
[0516] The server retrieves the user's personal information and viewing history from a database.
[0517] The server creates a user profile based on the acquired data.
[0518] For example, the information that "Tanaka watches a lot of action movies" is found in the database of the user, and the preferred genre is recorded as "action" in the profile.
[0519] Step 2: Recognize emotions
[0520] The server uses an emotion engine to recognize the user's emotional state.
[0521] The server analyzes data such as the user's facial expressions, voice, and text to determine their current emotional state.
[0522] For example, it analyzes Tanaka's facial expressions and voice to recognize that he is in an "excited" state.
[0523] Step 3: Story Generation
[0524] The server builds input data for the generative AI model based on the user profile and recognized emotions.
[0525] The server provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre for excited users," and generates a storyline.
[0526] For example, a generative AI model can generate exciting stories such as "a brave police officer defeats a bad guy."
[0527] Step 4: Character Selection and Modeling
[0528] The server extracts the main characters from the generated story.
[0529] The server acquires actor data corresponding to the extracted character from an actor video database.
[0530] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[0531] Step 5: Generate footage
[0532] The server imports the generated storyline and character models into a video editing system.
[0533] The server applies character movements and dialogue based on the script and renders the video.
[0534] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[0535] Step 6: Streaming the video
[0536] The server obtains the terminal endpoint for the user to view.
[0537] The server transmits the generated video to the user's terminal.
[0538] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[0539] Through the above steps, the present invention can generate storylines tailored to the user's emotions and provide short, personal dramas customized in real time, thereby significantly reducing the time and cost required for traditional drama production and quickly providing high-quality entertainment.
[0540] Example 2
[0541] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0542] Conventional video production systems have had difficulty quickly providing content tailored to a user's preferences and emotional state. Generating high-quality 3D models and videos also requires a huge amount of time and cost. This makes it difficult to provide personalized video content to individual users. Therefore, there is a need for a system that can dynamically recognize emotions based on each user's profile and adjust scripts and videos accordingly.
[0543] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0544] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical video work using a generative AI model, a means for selecting an appropriate person from a video database of actors and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected person model, a means for delivering the generated video to a user, a means for recognizing the user's emotional state, and a means for adjusting the generated script based on the emotional state, thereby enabling the rapid provision of high-quality video content personalized to each user based on the user's profile and real-time emotional state.
[0545] A "user profile" is a profile created based on data such as a user's personal information, viewing history, and preferences, and reflects the user's characteristics and tastes.
[0546] A "generative AI model" is an algorithmic model that uses artificial intelligence to process natural language and generate content, automatically generating scripts for short vertical video works based on user profiles.
[0547] A "short vertical video work" is video content that can generally be viewed in a short period of time, typically from a few minutes to a few tens of minutes, and is intended to be displayed vertically.
[0548] A "script" is a document that describes the lines, scene structure, character actions, etc. in a video work.
[0549] An "actor" is a person who appears in a video work and plays a specific role in it.
[0550] A "video database" is a database that stores video data related to multiple performers and allows for searching and retrieval.
[0551] A "high-resolution 3D model" is a visually detailed and accurate three-dimensional character model generated using computer software.
[0552] "Emotional state" refers to a user's current psychological and emotional state, analyzed using data obtained from facial expressions, voice, text messages, etc.
[0553] "Distribution" refers to the act of transmitting the generated video content to a user's device via the Internet, allowing the user to view it in real time.
[0554] "Adjustment" refers to the process of changing or modifying parts of the script created by the generative AI model based on the emotional state, optimizing it to better match the user's emotions.
[0555] The present invention relates to a system that creates a user profile, generates a script for a short vertical video using a generative AI model, selects suitable actors from a video database of actors, generates high-resolution 3D models, generates a video using the generated script and the selected actor models, and finally delivers the video to the user's device. It also has the ability to recognize the user's emotional state and adjust the generated script based on that.
[0556] Specifically, the system operates in the following steps.
[0557] First, the server retrieves the personal information and viewing history entered by the user from a database and creates a user profile based on information such as the user's name, age, preferred genres, and viewing history.
[0558] Next, the server provides a prompt based on the profile information to the generative AI model. An example of the generative AI model used here is OpenAI's GPT-4. An example of the prompt is "Generate a short, vertical drama in the action genre."
[0559] The generative AI model automatically generates a storyline based on this prompt. For example, it might generate a storyline in which the protagonist infiltrates the enemy's hideout and engages in a spectacular battle.
[0560] The server also uses an emotion engine (e.g., Emotion API or Affectiva) to recognize the user's current emotional state by analyzing emotions from the user's facial expressions, voice, text messages, etc.
[0561] The server dynamically adjusts the input data of the generative AI model based on the recognized emotions, optimizing the storyline to better match the user's emotions. For example, if a user is in an excited state, the server may add more exciting battle scenes.
[0562] Next, the server selects an appropriate person from a video database of actors based on the generated storyline. For example, it retrieves data on an actor named "Sato" from the video database and generates a high-resolution 3D model of him using Unity or Unreal Engine.
[0563] The server uses the generated script and the selected 3D models to generate the video. Video editing software such as Adobe Premiere Pro or Final Cut Pro is used to generate the video. As a result, lines and scenes are developed based on the script, and the video is completed with characters moving and speaking lines.
[0564] Finally, the server delivers the generated video to the user's device using streaming services such as AWS CloudFront or Akamai. The device receives the delivered video in real time, allowing the user to watch it immediately.
[0565] For example, suppose a user named "Tanaka" likes the action genre and is currently in an "excited" emotional state. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt: "Generate a short, vertical drama in the action genre." The generative AI model then generates an exciting storyline that matches Tanaka's "excited state."
[0566] The server then uses its emotion engine to recognize Tanaka's emotional state as "excited" and adjusts the input data of the generative AI model based on that state. The server then selects an actor named "Sato" from the generated storyline, generates a 3D model of him, and finally completes the video and distributes it to Tanaka's smartphone. Tanaka can now enjoy an exciting short vertical drama that matches his emotions at that moment.
[0567] This system can significantly reduce the time and cost required for conventional drama production and quickly provide high-quality content that matches the user's emotions.
[0568] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0569] System program processing flow
[0570] Step 1: Creating a User Profile
[0571] The server obtains the personal information and viewing history entered by the user and creates a user profile. Specifically, the server obtains the user's name, age, preferred genres, viewing history, etc. from a database to generate the profile.
[0572] Input: User information from the database
[0573] Output: User profile
[0574] Specific operation: The server obtains the information "User name: Tanaka, Age: 30, Favorite genre: Action, Viewing history: 5 action movies, 1 comedy movie" and creates a profile for "Tanaka."
[0575] Step 2: Generate a storyline
[0576] The server provides a prompt based on the user profile to the generative AI model to generate a storyline, for example, "Generate a short, vertical drama in the action genre."
[0577] Input: User Profile
[0578] Output: Generated storyline
[0579] Specific operation: Based on the information that "Tanaka's preference is the action genre," the server gives the generative AI model a prompt statement of "Generate a short vertical drama in the action genre." The generative AI model generates a storyline in which "the protagonist infiltrates the enemy's hideout and engages in a spectacular battle."
[0580] Step 3: Recognize emotions
[0581] The device captures the user's facial expressions and voice through a camera and microphone and sends the data to a server, which uses an emotion engine to analyze the user's emotional state.
[0582] Input: User's facial expression data, voice data
[0583] Output: User's emotional state
[0584] Specific operation: The device captures Tanaka's facial expressions with a camera and collects his voice with a microphone. The server analyzes Tanaka's data using an emotion engine and recognizes that Tanaka is in an "excited" state.
[0585] Step 4: Adjusting the storyline
[0586] Based on the recognized emotions, the server dynamically adjusts the input data of the generative AI model, optimizing the storyline to better match the emotions.
[0587] Input: Generated storyline, user's emotional state
[0588] Output: Adjusted storyline
[0589] Specific operation: The server instructs the generative AI model to "add more exciting battle scenes" based on the information "current emotional state: excitement." The generative AI model adds a plot twist to the storyline in which the protagonist uses a secret weapon he found in the basement to wipe out the enemies.
[0590] Step 5: Select and generate a character model
[0591] Based on the generated storyline, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models.
[0592] Enter: the adjusted storyline
[0593] Output: High-resolution 3D character model
[0594] Specific operation: The server retrieves data on "Actor A," who is the best suited to play the main character of the story, from the database and generates a 3D model of Actor A in Unity or Unreal Engine.
[0595] Step 6: Generate footage
[0596] The server uses the generated script and 3D character model to generate video, which is then edited using video editing software such as Adobe Premiere Pro or Final Cut Pro.
[0597] Input: Adjusted storyline, 3D character models
[0598] Output: Generated video
[0599] How it works: Based on the generated script, the server uses a 3D model of Actor A to create a battle scene with an enemy. Adobe Premiere Pro is used to align the timing of lines and scenes and complete the final footage.
[0600] Step 7: Streaming the video
[0601] The server delivers the generated video to the user's device using streaming services such as AWS CloudFront and Akamai.
[0602] Input: Generated video
[0603] Output: Video played on the user's device
[0604] How it works: The server uploads the completed video file to AWS CloudFront and generates a distribution URL. The device receives the URL and immediately plays the video, allowing users to enjoy the short vertical drama on the spot.
[0605] As described above, high-quality video content that matches the user's emotions and preferences is quickly provided through each step.
[0606] (Application example 2)
[0607] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0608] Conventional content delivery services have had difficulty generating and delivering personalized video content in real time based on user preferences and emotions. In particular, providing content that takes into account the user's emotional state is difficult, and there has been a demand for a way to improve the quality of the viewing experience.
[0609] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0610] a means for creating a user profile;
[0611] A means for generating scripts for short vertical dramas using a generative AI model;
[0612] A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models;
[0613] means for generating a video using the generated script and a selected character model;
[0614] a means for delivering the generated video to a user;
[0615] A means for dynamically adjusting input data for a generative AI model using an emotion engine that recognizes user emotions; and
[0616] This will enable the generation and delivery of personalized video content in real time based on the user's emotional state.
[0617] A "user profile" is information created by organizing and analyzing data about a user based on the user's personal information, viewing history, preferred genres, etc.
[0618] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and scripts based on user profiles and specific input data.
[0619] The "actor video database" is a database that stores the video and attribute data of multiple actors, and is used to select appropriate characters.
[0620] "High-Resolution 3D Model" means a highly detailed, high-resolution, three-dimensional computer model used to visually recreate a Selected Character.
[0621] The "emotion engine" is an engine that analyzes and recognizes emotions from a user's facial expressions, voice, text, etc.
[0622] The "video generation means" is a means for developing scenes, lines, etc. and creating video using the generated script and selected character models.
[0623] "Video distribution means" refers to a means for distributing the generated video to end users in real time.
[0624] This invention is a system that creates a user profile, generates personalized short-form dramas for users using a generative AI model, and provides storylines based on the user's emotions using an emotion engine. This system is implemented through the following specific process.
[0625] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, including the user's preferred genres and past viewing history, providing the foundation for customization for each individual user.
[0626] The server then uses a generative AI model to generate a storyline based on the user profile. The generative AI model uses a natural language generation algorithm (e.g., GPT-4) to automatically generate a story that fits the user's preferred genre.
[0627] Furthermore, the server uses an emotion engine to recognize the user's current emotional state. The emotion engine detects emotions by analyzing the user's facial expressions and voice data using a facial recognition camera (e.g., Face API) and voice analysis tools (e.g., Azure Cognitive Services).
[0628] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[0629] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[0630] The server then generates a video using the generated script and the selected character model. The script is then developed into lines and scenes using a video editing system (e.g., Blender), and the characters move and speak.
[0631] Finally, the server delivers the generated video to the user's device. Using a streaming distribution server (e.g., AWS S3 + CloudFront), the video is delivered to the user's device through an endpoint, allowing the user to instantly watch the short vertical drama.
[0632] As a concrete example, suppose a user named "Tanaka" likes the action genre and his current emotional state is "excited." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[0633] An example of a prompt used in this process is:
[0634] "Generate a comedy-action story that is suitable for when the user is in a happy state. The main character should have a sense of humor and the episodes should be thrilling with action scenes."
[0635] As a result, this invention makes it possible to efficiently generate and distribute personalized video content that matches the user's emotions, improving the quality of the viewing experience.
[0636] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0637] Step 1:
[0638] The server retrieves the user's personal information and viewing history from the database and creates a user profile. This profile creation step takes the user ID as input and generates profile data including the user's preferred genres and viewing history as output. Specifically, it uses SQL queries to retrieve user information from the database, organizes the necessary data, and builds the profile.
[0639] Step 2:
[0640] The server uses a generative AI model to generate a storyline based on the user profile. In this step, the user profile data is used as input and the generated storyline is obtained as output. Specifically, a prompt sentence reflecting the user's preferred genre and emotional state is generated, and a story is created using a natural language generation algorithm. The prompt sentence is input into the generative AI model (e.g., GPT-4) to obtain the storyline.
[0641] Step 3:
[0642] The server uses an emotion engine to recognize the user's current emotional state. In this step, the user's real-time facial expressions and voice data are used as input, and the user's emotional state is obtained as output. Specifically, the data is analyzed using a facial recognition camera and voice analysis tools (e.g., Face API, Azure Cognitive Services) to detect the user's emotions.
[0643] Step 4:
[0644] The server dynamically adjusts the input data of the generative AI model based on the recognized emotion. In this step, the emotional state is used as input and a storyline tailored to the emotion is obtained as output. Specifically, the recognized emotion is used to regenerate the prompt sentence, which is then reinput into the generative AI model to generate a story that is appropriate for the emotion. For example, the prompt sentence "Generate an action story that is appropriate for the user's excited state" is input into the generative AI model.
[0645] Step 5:
[0646] The server selects appropriate characters from the actor video database based on the generated storyline and generates high-resolution 3D models. In this step, the storyline is used as input and 3D character models are obtained as output. Specifically, the server extracts the main characters appearing in the story, selects appropriate actors from the video database, and generates their high-resolution 3D models.
[0647] Step 6:
[0648] The server generates a video using the generated script and the selected character model. In this step, the script and character model are used as input, and a short vertical drama video is obtained as output. Specifically, a video editing system (e.g., Blender) is used to create scenes including character movements and dialogue, and the entire video is generated.
[0649] Step 7:
[0650] The server delivers the generated video to the user's device. In this step, the generated video is used as input and the video delivered to the user's device is obtained as output. Specifically, the video is delivered using a streaming distribution server (e.g., AWS S3 + CloudFront), and the user watches this video on their smartphone.
[0651] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0652] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0653] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0654] [Third embodiment]
[0655] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0656] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0657] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0658] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0659] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0660] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0661] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0662] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0663] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0664] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0665] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0666] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0667] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, and delivers the generated footage to users.
[0668] Program Overview
[0669] The program for this system is composed of, for example, the following steps:
[0670] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile, which then lays the foundation for customization for each user.
[0671] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[0672] The server then selects suitable characters from a database of actor footage based on the generated storyline and generates high-resolution 3D models of the selected characters, which then serve as the basis for the video production.
[0673] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[0674] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[0675] Specific examples
[0676] For example, suppose a user named "Tanaka" likes the action genre. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story.
[0677] Next, the server retrieves data for the actor "Sato" from a database of actor footage and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate a video. Finally, this video is sent to Tanaka's smartphone, where he can watch it on the spot.
[0678] This system can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content. In this way, the present invention can be put into practice.
[0679] The processing flow will be explained below.
[0680] Step 1: Creating a User Profile
[0681] The server retrieves the user's personal information and viewing history from a database.
[0682] The server creates a user profile based on the acquired data.
[0683] For example, it can be determined from the database that user "Tanaka" watches a lot of action movies, and his preferred genre can be recorded as "action" in his profile.
[0684] Step 2: Story Generation
[0685] The server constructs input data for the generative AI model based on the user profile.
[0686] The server provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre" and generates a storyline.
[0687] For example, a generative AI model might generate a story about a brave police officer defeating a villain.
[0688] Step 3: Selecting and modeling your character
[0689] The server extracts the main characters from the generated story.
[0690] The server acquires actor data corresponding to the extracted character from an actor video database.
[0691] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[0692] Step 4: Generate footage
[0693] The server imports the generated storyline and character models into a video editing system.
[0694] The server applies character movements and dialogue based on the script and renders the video.
[0695] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[0696] Step 5: Streaming the video
[0697] The server obtains the terminal endpoint for the user to view.
[0698] The server transmits the generated video to the user's terminal.
[0699] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[0700] Example 1
[0701] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0702] Traditional drama production is time-consuming and costly, making it difficult to provide high-quality content in a short period of time. In addition, there is a lack of methods to quickly provide customized video content tailored to viewer preferences.
[0703] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0704] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting an appropriate character from a character video database and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected character model, and a means for delivering the generated video to the user, thereby enabling the rapid provision of high-quality short vertical drama videos customized to the viewer's preferences.
[0705] A "user profile" is a collection of data created based on a user's personal information, viewing history, preferred genres, etc.
[0706] A "generative AI model" is an algorithm or program that uses artificial intelligence techniques to generate scripts or text in natural language based on a given prompt.
[0707] A "short vertical drama" is a story-based content that can be viewed in a short amount of time, and is primarily intended to be viewed on vertical screens such as smartphones.
[0708] A "script" is a document that describes the dialogue and scene details for video content.
[0709] A "person video database" is a database that stores video data of actors and characters.
[0710] A "character" is a person, animal, or personified being that plays a specific role in the story.
[0711] A "high resolution three-dimensional model" is a detailed three-dimensional object created using computer graphics and displayed at a high resolution.
[0712] "Video" is visual content that expresses movement through a series of images.
[0713] "Distribution" refers to sending the generated video to the user's device so that it can be viewed.
[0714] This invention is a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a human video database, generates high-resolution three-dimensional models, generates video using the generated script and the selected character models, and delivers the generated video to the user.
[0715] Creating a user profile
[0716] First, the server retrieves the user's personal information, viewing history, and preferred genres from a database. This information is used to create a user profile, laying the groundwork for delivering customized content. Using SQL queries, the server extracts the necessary data from the tables where the user information is stored.
[0717] Storyline Generation
[0718] The server then uses the generative AI model to generate a storyline based on the user profile. In this process, the server inputs prompt statements into the generative AI model. The prompt statements can be in the following format:
[0719] "Generating short vertical dramas in the action genre"
[0720] Based on the generated prompts, a generative AI model generates a storyline using a natural language generation algorithm, such as an AI model like GPT-3.
[0721] Character selection and 3D model generation
[0722] After the storyline is generated, the server selects suitable characters from a human video database, filtering data that matches the character attributes (e.g., gender, age, personality, etc.) in the storyline, and uses 3D modeling software such as Blender or Maya to generate high-resolution three-dimensional models of the selected characters.
[0723] Video generation
[0724] The server then combines the generated script with the selected character model to generate a video. Using a video editing system, each line and scene in the script is applied to the character. For example, video editing software such as Adobe Premiere Pro or DaVinci Resolve is used.
[0725] Video distribution
[0726] Finally, the server prepares the resulting video for delivery to the user's device. The video file is encoded into the appropriate format and uploaded to a CDN (Content Delivery Network), where it is delivered to the end user's device using a streaming service (e.g., HLS or DASH).
[0727] Specific examples
[0728] For example, if a user named "Tanaka" likes the action genre, the server creates a profile based on Tanaka's viewing history and preferred genres. The server then provides the generative AI model with a prompt, "Generate a short vertical drama in the action genre," and extracts key characters from the generated storyline. The server then retrieves appropriate actor data from a human video database and generates high-resolution 3D models. These elements are integrated to generate a video, which is ultimately delivered to Tanaka's smartphone. This system allows Tanaka to watch an action-packed short vertical drama on the spot.
[0729] In this way, the system of the present invention can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content that meets the preferences of viewers.
[0730] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0731] Step 1: Creating a User Profile
[0732] Specific behavior:
[0733] The server first obtains the user's personal information, viewing history, and preferred genres from a database.
[0734] Input: User information stored in the database, viewing history, and preferred genres.
[0735] Data processing: Extract data using SQL queries and generate user profile objects.
[0736] Output: A customized user profile object.
[0737] Step 2: Generate a prompt statement
[0738] Specific behavior:
[0739] The server parses information from the user profile and generates prompts to feed into the generative AI model.
[0740] Input: A user profile object.
[0741] Data processing: Create a natural language prompt based on the attributes of the user profile.
[0742] Output: A prompt to input to the generative AI model (e.g., "Generate a short, vertical drama in the action genre").
[0743] Step 3: Generate a storyline
[0744] Specific behavior:
[0745] The server uses a generative AI model to generate a storyline based on the prompt.
[0746] Input: The generated prompt statement.
[0747] Data processing: Use a generative AI model (e.g., GPT-3) to automatically generate storyline text.
[0748] Output: The generated storyline text.
[0749] Step 4: Character Selection
[0750] Specific behavior:
[0751] The server analyzes the character attributes (gender, age, personality, etc.) in the generated storyline and references a video database of characters.
[0752] Input: The generated storyline text.
[0753] Data Processing: Based on character attributes, filter matching database entries to select the best character.
[0754] Output: Data of the selected character.
[0755] Step 5: Generate a high-resolution 3D model
[0756] Specific behavior:
[0757] The server generates a high-resolution three-dimensional model based on the video data of the selected character.
[0758] Input: Selected character's data.
[0759] Data processing: Creating a three-dimensional model using 3D modeling software such as Blender or Maya.
[0760] Output: High resolution 3D character model.
[0761] Step 6: Generate footage
[0762] Specific behavior:
[0763] The server integrates the generated script with the selected character model to generate video content.
[0764] Input: Generated storyline text, high-resolution 3D character models.
[0765] Data processing: Analyze each line and scene in the script and add movement and lines to the characters using a video editing system (Adobe Premiere Pro or DaVinci Resolve).
[0766] Output: The finished video file.
[0767] Step 7: Streaming the video
[0768] Specific behavior:
[0769] The server encodes the finished video into the appropriate format and uploads it to the CDN for delivery to end users.
[0770] Input: Final video file.
[0771] Data processing: Encode the video file, upload it to the CDN, and prepare it for distribution on the streaming service.
[0772] Output: A video stream playable on the end user's device.
[0773] This enables the server to quickly provide high-quality short vertical drama content tailored to the viewer's preferences.
[0774] (Application example 1)
[0775] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0776] Modern content distribution services face challenges in quickly providing content tailored to individual user preferences and interests. Complex video content, such as dramas and animations, requires a high level of customization and rapid generation, but no system currently exists that can achieve this. Therefore, efficiently providing high-quality, personalized video content that satisfies users is a major challenge.
[0777] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0778] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting appropriate characters from an actor video database and generating high-resolution 3D models, a means for generating a video using the generated script and the selected character models, a means for delivering the generated video to a user, a means for generating a story based on the user's genre preferences, and a means for delivering the generated video to a smartphone application. This makes it possible to quickly generate personalized video content based on the user's individual preferences and interests and smoothly deliver it to the user's device.
[0779] A "user profile" is a collection of information such as a user's personal information, viewing history, and preferred genres.
[0780] A "generative AI model" is an artificial intelligence algorithm that uses generative AI to generate sentences or stories based on a specific prompt.
[0781] "Short vertical drama" is a drama-style content in which a story that can be viewed in a short amount of time is displayed on a vertical screen.
[0782] An "actor video database" is a database that stores videos and data of multiple actors.
[0783] A "high-resolution 3D model" is a high-resolution three-dimensional graphic model capable of depicting detailed images.
[0784] "Video generation means" is a function for creating videos using the generated scripts and character models.
[0785] "Distribution means" refers to a system for transmitting the generated video to the user's terminal via the Internet.
[0786] "Genre preferences" refer to the types of dramas and movies that a user is particularly interested in.
[0787] A "smartphone application" is any software program that runs on a smartphone.
[0788] "Story generation means" is a function that allows AI to generate stories based on the user's profile and genre preferences.
[0789] This invention builds a system that generates and delivers customized short, vertical dramas tailored to individual user preferences. This system is realized by creating a user profile, automatically generating drama scripts using a generative AI model based on that information, selecting appropriate characters from a database of actor footage, and generating high-resolution 3D models.
[0790] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile. This user profile is used to gain a detailed understanding of the type of content the user prefers.
[0791] The server then uses a generative AI model to generate a storyline based on the user profile. This generative AI model uses a natural language generation algorithm to automatically generate a story that fits the user's preferred genre. An example of a specific prompt is "Generate a short, vertical drama in the horror genre."
[0792] The server then selects suitable characters from a database of actor footage based on the generated storyline, and generates high-resolution 3D models of the selected characters, which serve as the foundation for the video production.
[0793] The server then generates a video using the generated script and the selected character model. The script is then expanded into lines and scenes using a video editing system (e.g., OpenCV), and the characters move and speak.
[0794] Finally, the server delivers the generated video to a smartphone application. Users can then watch their individually customized, high-quality short vertical dramas on their smartphones. Using a high-performance graphics card (e.g., the NVIDIA RTX series) ensures smooth video playback.
[0795] For example, if a user prefers the action genre, the server creates a profile based on their viewing history and preferred genres and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story. Next, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models. Finally, the generated footage is delivered to the user's smartphone for real-time viewing.
[0796] As described above, by implementing the present invention, it becomes possible to quickly generate personalized video content based on the individual preferences and interests of a user and distribute it to the user's terminal.
[0797] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0798] Step 1:
[0799] The server retrieves the user's personal information, viewing history, preferred genres, etc. from the database and creates a user profile. The input data for the user profile requires the user ID, viewing history, and preferred genres. Based on this, a profile detailing the user's preferences and interests is output.
[0800] Step 2:
[0801] The server uses a generative AI model to generate a storyline based on the user profile. The input data requires the user profile and a prompt, such as "Generate a short, vertical drama in the horror genre." The generative AI model uses these inputs to apply a natural language generation algorithm and output a drama script.
[0802] Step 3:
[0803] The server selects appropriate characters from a video database of actors based on the generated script. In this step, the script content is used as input data, and the IDs and characteristics of the characters that match it are output. This selection determines the main characters of the story.
[0804] Step 4:
[0805] The server generates a high-resolution 3D model of the selected character. The input data for this step is the character's ID and characteristics. Based on this information, a detailed 3D model is output using 3D graphics software (e.g., Blender).
[0806] Step 5:
[0807] The server generates a video using the generated script and the selected character model. Specifically, the script is developed as lines and scenes, and the video is edited so that the characters move and speak the lines. The input data for this step are the script and 3D models, and the final video is output using a video editing system (e.g., OpenCV).
[0808] Step 6:
[0809] The server delivers the generated video to the smartphone application. In this step, the video file is used as input data, and the video is compressed using a high-performance graphics card (e.g., NVIDIA RTX series) and sent to the user's smartphone via the Internet. The final output is video data that can be played on the user's smartphone.
[0810] Step 7:
[0811] The user watches the streamed video on their smartphone. In this step, the smartphone application plays the video file received as input data. Through the application, the user can watch personalized, high-quality short vertical dramas.
[0812] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0813] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, combines it with an emotion engine that recognizes the user's emotions, and delivers the generated footage to the user.
[0814] Program Overview
[0815] The program for this system is composed of, for example, the following steps:
[0816] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, which then forms the basis for customization for each individual user.
[0817] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[0818] The server then uses an emotion engine to recognize the user's current emotional state, which detects emotions by analyzing the user's facial expressions, voice, and text.
[0819] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[0820] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[0821] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[0822] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[0823] Specific examples
[0824] For example, suppose a user named "Tanaka" likes the action genre and is currently in an emotional state of "excitement." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[0825] The server then uses its emotion engine to recognize that Tanaka's emotional state is "excited," and based on this emotional state, dynamically adjusts the input data of the generative AI model to make the generated storyline even more compelling.
[0826] The server extracts the main characters from the generated story, retrieves data on an actor named "Sato" from a database of actor footage, and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate the video.
[0827] Finally, the video is streamed to Tanaka's smartphone, where he can watch it instantly, enjoying an exciting short vertical drama that matches his own emotions.
[0828] This system significantly reduces the time and cost required for conventional drama production, and also makes it possible to quickly provide high-quality content that matches the user's emotions.
[0829] The processing flow will be explained below.
[0830] Step 1: Creating a User Profile
[0831] The server retrieves the user's personal information and viewing history from a database.
[0832] The server creates a user profile based on the acquired data.
[0833] For example, the information that "Tanaka watches a lot of action movies" is found in the database of the user, and the preferred genre is recorded as "action" in the profile.
[0834] Step 2: Recognize emotions
[0835] The server uses an emotion engine to recognize the user's emotional state.
[0836] The server analyzes data such as the user's facial expressions, voice, and text to determine their current emotional state.
[0837] For example, it analyzes Tanaka's facial expressions and voice to recognize that he is in an "excited" state.
[0838] Step 3: Story Generation
[0839] The server builds input data for the generative AI model based on the user profile and recognized emotions.
[0840] The server provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre for excited users," and generates a storyline.
[0841] For example, a generative AI model can generate exciting stories such as "a brave police officer defeats a bad guy."
[0842] Step 4: Character Selection and Modeling
[0843] The server extracts the main characters from the generated story.
[0844] The server acquires actor data corresponding to the extracted character from an actor video database.
[0845] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[0846] Step 5: Generate footage
[0847] The server imports the generated storyline and character models into a video editing system.
[0848] The server applies character movements and dialogue based on the script and renders the video.
[0849] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[0850] Step 6: Streaming the video
[0851] The server obtains the terminal endpoint for the user to view.
[0852] The server transmits the generated video to the user's terminal.
[0853] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[0854] Through the above steps, the present invention can generate storylines tailored to the user's emotions and provide short, personal dramas customized in real time, thereby significantly reducing the time and cost required for traditional drama production and quickly providing high-quality entertainment.
[0855] Example 2
[0856] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0857] Conventional video production systems have had difficulty quickly providing content tailored to a user's preferences and emotional state. Generating high-quality 3D models and videos also requires a huge amount of time and cost. This makes it difficult to provide personalized video content to individual users. Therefore, there is a need for a system that can dynamically recognize emotions based on each user's profile and adjust scripts and videos accordingly.
[0858] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0859] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical video work using a generative AI model, a means for selecting an appropriate person from a video database of actors and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected person model, a means for delivering the generated video to a user, a means for recognizing the user's emotional state, and a means for adjusting the generated script based on the emotional state, thereby enabling the rapid provision of high-quality video content personalized to each user based on the user's profile and real-time emotional state.
[0860] A "user profile" is a profile created based on data such as a user's personal information, viewing history, and preferences, and reflects the user's characteristics and tastes.
[0861] A "generative AI model" is an algorithmic model that uses artificial intelligence to process natural language and generate content, automatically generating scripts for short vertical video works based on user profiles.
[0862] A "short vertical video work" is video content that can generally be viewed in a short period of time, typically from a few minutes to a few tens of minutes, and is intended to be displayed vertically.
[0863] A "script" is a document that describes the lines, scene structure, character actions, etc. in a video work.
[0864] An "actor" is a person who appears in a video work and plays a specific role in it.
[0865] A "video database" is a database that stores video data related to multiple performers and allows for searching and retrieval.
[0866] A "high-resolution 3D model" is a visually detailed and accurate three-dimensional character model generated using computer software.
[0867] "Emotional state" refers to a user's current psychological and emotional state, analyzed using data obtained from facial expressions, voice, text messages, etc.
[0868] "Distribution" refers to the act of transmitting the generated video content to a user's device via the Internet, allowing the user to view it in real time.
[0869] "Adjustment" refers to the process of changing or modifying parts of the script created by the generative AI model based on the emotional state, optimizing it to better match the user's emotions.
[0870] The present invention relates to a system that creates a user profile, generates a script for a short vertical video using a generative AI model, selects suitable actors from a video database of actors, generates high-resolution 3D models, generates a video using the generated script and the selected actor models, and finally delivers the video to the user's device. It also has the ability to recognize the user's emotional state and adjust the generated script based on that.
[0871] Specifically, the system operates in the following steps.
[0872] First, the server retrieves the personal information and viewing history entered by the user from a database and creates a user profile based on information such as the user's name, age, preferred genres, and viewing history.
[0873] Next, the server provides a prompt based on the profile information to the generative AI model. An example of the generative AI model used here is OpenAI's GPT-4. An example of the prompt is "Generate a short, vertical drama in the action genre."
[0874] The generative AI model automatically generates a storyline based on this prompt. For example, it might generate a storyline in which the protagonist infiltrates the enemy's hideout and engages in a spectacular battle.
[0875] The server also uses an emotion engine (e.g., Emotion API or Affectiva) to recognize the user's current emotional state by analyzing emotions from the user's facial expressions, voice, text messages, etc.
[0876] The server dynamically adjusts the input data of the generative AI model based on the recognized emotions, optimizing the storyline to better match the user's emotions. For example, if a user is in an excited state, the server may add more exciting battle scenes.
[0877] Next, the server selects an appropriate person from a video database of actors based on the generated storyline. For example, it retrieves data on an actor named "Sato" from the video database and generates a high-resolution 3D model of him using Unity or Unreal Engine.
[0878] The server uses the generated script and the selected 3D models to generate the video. Video editing software such as Adobe Premiere Pro or Final Cut Pro is used to generate the video. As a result, lines and scenes are developed based on the script, and the video is completed with characters moving and speaking lines.
[0879] Finally, the server delivers the generated video to the user's device using streaming services such as AWS CloudFront or Akamai. The device receives the delivered video in real time, allowing the user to watch it immediately.
[0880] For example, suppose a user named "Tanaka" likes the action genre and is currently in an "excited" emotional state. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt: "Generate a short, vertical drama in the action genre." The generative AI model then generates an exciting storyline that matches Tanaka's "excited state."
[0881] The server then uses its emotion engine to recognize Tanaka's emotional state as "excited" and adjusts the input data of the generative AI model based on that state. The server then selects an actor named "Sato" from the generated storyline, generates a 3D model of him, and finally completes the video and distributes it to Tanaka's smartphone. Tanaka can now enjoy an exciting short vertical drama that matches his emotions at that moment.
[0882] This system can significantly reduce the time and cost required for conventional drama production and quickly provide high-quality content that matches the user's emotions.
[0883] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0884] System program processing flow
[0885] Step 1: Creating a User Profile
[0886] The server obtains the personal information and viewing history entered by the user and creates a user profile. Specifically, the server obtains the user's name, age, preferred genres, viewing history, etc. from a database to generate the profile.
[0887] Input: User information from the database
[0888] Output: User profile
[0889] Specific operation: The server obtains the information "User name: Tanaka, Age: 30, Favorite genre: Action, Viewing history: 5 action movies, 1 comedy movie" and creates a profile for "Tanaka."
[0890] Step 2: Generate a storyline
[0891] The server provides a prompt based on the user profile to the generative AI model to generate a storyline, for example, "Generate a short, vertical drama in the action genre."
[0892] Input: User Profile
[0893] Output: Generated storyline
[0894] Specific operation: Based on the information that "Tanaka's preference is the action genre," the server gives the generative AI model a prompt statement of "Generate a short vertical drama in the action genre." The generative AI model generates a storyline in which "the protagonist infiltrates the enemy's hideout and engages in a spectacular battle."
[0895] Step 3: Recognize emotions
[0896] The device captures the user's facial expressions and voice through a camera and microphone and sends the data to a server, which uses an emotion engine to analyze the user's emotional state.
[0897] Input: User's facial expression data, voice data
[0898] Output: User's emotional state
[0899] Specific operation: The device captures Tanaka's facial expressions with a camera and collects his voice with a microphone. The server analyzes Tanaka's data using an emotion engine and recognizes that Tanaka is in an "excited" state.
[0900] Step 4: Adjusting the storyline
[0901] Based on the recognized emotions, the server dynamically adjusts the input data of the generative AI model, optimizing the storyline to better match the emotions.
[0902] Input: Generated storyline, user's emotional state
[0903] Output: Adjusted storyline
[0904] Specific operation: The server instructs the generative AI model to "add more exciting battle scenes" based on the information "current emotional state: excitement." The generative AI model adds a plot twist to the storyline in which the protagonist uses a secret weapon he found in the basement to wipe out the enemies.
[0905] Step 5: Select and generate a character model
[0906] Based on the generated storyline, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models.
[0907] Enter: the adjusted storyline
[0908] Output: High-resolution 3D character model
[0909] Specific operation: The server retrieves data on "Actor A," who is the best suited to play the main character of the story, from the database and generates a 3D model of Actor A in Unity or Unreal Engine.
[0910] Step 6: Generate footage
[0911] The server uses the generated script and 3D character model to generate video, which is then edited using video editing software such as Adobe Premiere Pro or Final Cut Pro.
[0912] Input: Adjusted storyline, 3D character models
[0913] Output: Generated video
[0914] How it works: Based on the generated script, the server uses a 3D model of Actor A to create a battle scene with an enemy. Adobe Premiere Pro is used to align the timing of lines and scenes and complete the final footage.
[0915] Step 7: Streaming the video
[0916] The server delivers the generated video to the user's device using streaming services such as AWS CloudFront and Akamai.
[0917] Input: Generated video
[0918] Output: Video played on the user's device
[0919] How it works: The server uploads the completed video file to AWS CloudFront and generates a distribution URL. The device receives the URL and immediately plays the video, allowing users to enjoy the short vertical drama on the spot.
[0920] As described above, high-quality video content that matches the user's emotions and preferences is quickly provided through each step.
[0921] (Application example 2)
[0922] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0923] Conventional content delivery services have had difficulty generating and delivering personalized video content in real time based on user preferences and emotions. In particular, providing content that takes into account the user's emotional state is difficult, and there has been a demand for a way to improve the quality of the viewing experience.
[0924] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0925] a means for creating a user profile;
[0926] A means for generating scripts for short vertical dramas using a generative AI model;
[0927] A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models;
[0928] means for generating a video using the generated script and a selected character model;
[0929] a means for delivering the generated video to a user;
[0930] A means for dynamically adjusting input data for a generative AI model using an emotion engine that recognizes user emotions; and
[0931] This will enable the generation and delivery of personalized video content in real time based on the user's emotional state.
[0932] A "user profile" is information created by organizing and analyzing data about a user based on the user's personal information, viewing history, preferred genres, etc.
[0933] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and scripts based on user profiles and specific input data.
[0934] The "actor video database" is a database that stores the video and attribute data of multiple actors, and is used to select appropriate characters.
[0935] "High-Resolution 3D Model" means a highly detailed, high-resolution, three-dimensional computer model used to visually recreate a Selected Character.
[0936] The "emotion engine" is an engine that analyzes and recognizes emotions from a user's facial expressions, voice, text, etc.
[0937] The "video generation means" is a means for developing scenes, lines, etc. and creating video using the generated script and selected character models.
[0938] "Video distribution means" refers to a means for distributing the generated video to end users in real time.
[0939] This invention is a system that creates a user profile, generates personalized short-form dramas for users using a generative AI model, and provides storylines based on the user's emotions using an emotion engine. This system is implemented through the following specific process.
[0940] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, including the user's preferred genres and past viewing history, providing the foundation for customization for each individual user.
[0941] The server then uses a generative AI model to generate a storyline based on the user profile. The generative AI model uses a natural language generation algorithm (e.g., GPT-4) to automatically generate a story that fits the user's preferred genre.
[0942] Furthermore, the server uses an emotion engine to recognize the user's current emotional state. The emotion engine detects emotions by analyzing the user's facial expressions and voice data using a facial recognition camera (e.g., Face API) and voice analysis tools (e.g., Azure Cognitive Services).
[0943] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[0944] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[0945] The server then generates a video using the generated script and the selected character model. The script is then developed into lines and scenes using a video editing system (e.g., Blender), and the characters move and speak.
[0946] Finally, the server delivers the generated video to the user's device. Using a streaming distribution server (e.g., AWS S3 + CloudFront), the video is delivered to the user's device through an endpoint, allowing the user to instantly watch the short vertical drama.
[0947] As a concrete example, suppose a user named "Tanaka" likes the action genre and his current emotional state is "excited." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[0948] An example of a prompt used in this process is:
[0949] "Generate a comedy-action story that is suitable for when the user is in a happy state. The main character should have a sense of humor and the episodes should be thrilling with action scenes."
[0950] As a result, this invention makes it possible to efficiently generate and distribute personalized video content that matches the user's emotions, improving the quality of the viewing experience.
[0951] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0952] Step 1:
[0953] The server retrieves the user's personal information and viewing history from the database and creates a user profile. This profile creation step takes the user ID as input and generates profile data including the user's preferred genres and viewing history as output. Specifically, it uses SQL queries to retrieve user information from the database, organizes the necessary data, and builds the profile.
[0954] Step 2:
[0955] The server uses a generative AI model to generate a storyline based on the user profile. In this step, the user profile data is used as input and the generated storyline is obtained as output. Specifically, a prompt sentence reflecting the user's preferred genre and emotional state is generated, and a story is created using a natural language generation algorithm. The prompt sentence is input into the generative AI model (e.g., GPT-4) to obtain the storyline.
[0956] Step 3:
[0957] The server uses an emotion engine to recognize the user's current emotional state. In this step, the user's real-time facial expressions and voice data are used as input, and the user's emotional state is obtained as output. Specifically, the data is analyzed using a facial recognition camera and voice analysis tools (e.g., Face API, Azure Cognitive Services) to detect the user's emotions.
[0958] Step 4:
[0959] The server dynamically adjusts the input data of the generative AI model based on the recognized emotion. In this step, the emotional state is used as input and a storyline tailored to the emotion is obtained as output. Specifically, the recognized emotion is used to regenerate the prompt sentence, which is then reinput into the generative AI model to generate a story that is appropriate for the emotion. For example, the prompt sentence "Generate an action story that is appropriate for the user's excited state" is input into the generative AI model.
[0960] Step 5:
[0961] The server selects appropriate characters from the actor video database based on the generated storyline and generates high-resolution 3D models. In this step, the storyline is used as input and 3D character models are obtained as output. Specifically, the server extracts the main characters appearing in the story, selects appropriate actors from the video database, and generates their high-resolution 3D models.
[0962] Step 6:
[0963] The server generates a video using the generated script and the selected character model. In this step, the script and character model are used as input, and a short vertical drama video is obtained as output. Specifically, a video editing system (e.g., Blender) is used to create scenes including character movements and dialogue, and the entire video is generated.
[0964] Step 7:
[0965] The server delivers the generated video to the user's device. In this step, the generated video is used as input and the video delivered to the user's device is obtained as output. Specifically, the video is delivered using a streaming distribution server (e.g., AWS S3 + CloudFront), and the user watches this video on their smartphone.
[0966] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0967] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0968] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[0969] [Fourth embodiment]
[0970] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0971] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0972] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0973] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0974] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0975] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0976] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0977] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0978] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0979] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0980] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0981] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0982] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[0983] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, and delivers the generated footage to users.
[0984] Program Overview
[0985] The program for this system is composed of, for example, the following steps:
[0986] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile, which then lays the foundation for customization for each user.
[0987] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[0988] The server then selects suitable characters from a database of actor footage based on the generated storyline and generates high-resolution 3D models of the selected characters, which then serve as the basis for the video production.
[0989] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[0990] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[0991] Specific examples
[0992] For example, suppose a user named "Tanaka" likes the action genre. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story.
[0993] Next, the server retrieves data for the actor "Sato" from a database of actor footage and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate a video. Finally, this video is sent to Tanaka's smartphone, where he can watch it on the spot.
[0994] This system can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content. In this way, the present invention can be put into practice.
[0995] The processing flow will be explained below.
[0996] Step 1: Creating a User Profile
[0997] The server retrieves the user's personal information and viewing history from a database.
[0998] The server creates a user profile based on the acquired data.
[0999] For example, it can be determined from the database that user "Tanaka" watches a lot of action movies, and his preferred genre can be recorded as "action" in his profile.
[1000] Step 2: Story Generation
[1001] The server constructs input data for the generative AI model based on the user profile.
[1002] The server provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre" and generates a storyline.
[1003] For example, a generative AI model might generate a story about a brave police officer defeating a villain.
[1004] Step 3: Selecting and modeling your character
[1005] The server extracts the main characters from the generated story.
[1006] The server acquires actor data corresponding to the extracted character from an actor video database.
[1007] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[1008] Step 4: Generate footage
[1009] The server imports the generated storyline and character models into a video editing system.
[1010] The server applies character movements and dialogue based on the script and renders the video.
[1011] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[1012] Step 5: Streaming the video
[1013] The server obtains the terminal endpoint for the user to view.
[1014] The server transmits the generated video to the user's terminal.
[1015] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[1016] Example 1
[1017] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1018] Traditional drama production is time-consuming and costly, making it difficult to provide high-quality content in a short period of time. In addition, there is a lack of methods to quickly provide customized video content tailored to viewer preferences.
[1019] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1020] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting an appropriate character from a character video database and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected character model, and a means for delivering the generated video to the user, thereby enabling the rapid provision of high-quality short vertical drama videos customized to the viewer's preferences.
[1021] A "user profile" is a collection of data created based on a user's personal information, viewing history, preferred genres, etc.
[1022] A "generative AI model" is an algorithm or program that uses artificial intelligence techniques to generate scripts or text in natural language based on a given prompt.
[1023] A "short vertical drama" is a story-based content that can be viewed in a short amount of time, and is primarily intended to be viewed on vertical screens such as smartphones.
[1024] A "script" is a document that describes the dialogue and scene details for video content.
[1025] A "person video database" is a database that stores video data of actors and characters.
[1026] A "character" is a person, animal, or personified being that plays a specific role in the story.
[1027] A "high resolution three-dimensional model" is a detailed three-dimensional object created using computer graphics and displayed at a high resolution.
[1028] "Video" is visual content that expresses movement through a series of images.
[1029] "Distribution" refers to sending the generated video to the user's device so that it can be viewed.
[1030] This invention is a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a human video database, generates high-resolution three-dimensional models, generates video using the generated script and the selected character models, and delivers the generated video to the user.
[1031] Creating a user profile
[1032] First, the server retrieves the user's personal information, viewing history, and preferred genres from a database. This information is used to create a user profile, laying the groundwork for delivering customized content. Using SQL queries, the server extracts the necessary data from the tables where the user information is stored.
[1033] Storyline Generation
[1034] The server then uses the generative AI model to generate a storyline based on the user profile. In this process, the server inputs prompt statements into the generative AI model. The prompt statements can be in the following format:
[1035] "Generating short vertical dramas in the action genre"
[1036] Based on the generated prompts, a generative AI model generates a storyline using a natural language generation algorithm, such as an AI model like GPT-3.
[1037] Character selection and 3D model generation
[1038] After the storyline is generated, the server selects suitable characters from a human video database, filtering data that matches the character attributes (e.g., gender, age, personality, etc.) in the storyline, and uses 3D modeling software such as Blender or Maya to generate high-resolution three-dimensional models of the selected characters.
[1039] Video generation
[1040] The server then combines the generated script with the selected character model to generate a video. Using a video editing system, each line and scene in the script is applied to the character. For example, video editing software such as Adobe Premiere Pro or DaVinci Resolve is used.
[1041] Video distribution
[1042] Finally, the server prepares the resulting video for delivery to the user's device. The video file is encoded into the appropriate format and uploaded to a CDN (Content Delivery Network), where it is delivered to the end user's device using a streaming service (e.g., HLS or DASH).
[1043] Specific examples
[1044] For example, if a user named "Tanaka" likes the action genre, the server creates a profile based on Tanaka's viewing history and preferred genres. The server then provides the generative AI model with a prompt, "Generate a short vertical drama in the action genre," and extracts key characters from the generated storyline. The server then retrieves appropriate actor data from a human video database and generates high-resolution 3D models. These elements are integrated to generate a video, which is ultimately delivered to Tanaka's smartphone. This system allows Tanaka to watch an action-packed short vertical drama on the spot.
[1045] In this way, the system of the present invention can significantly reduce the time and cost required for conventional drama production, and can quickly provide high-quality content that meets the preferences of viewers.
[1046] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1047] Step 1: Creating a User Profile
[1048] Specific behavior:
[1049] The server first obtains the user's personal information, viewing history, and preferred genres from a database.
[1050] Input: User information stored in the database, viewing history, and preferred genres.
[1051] Data processing: Extract data using SQL queries and generate user profile objects.
[1052] Output: A customized user profile object.
[1053] Step 2: Generate a prompt statement
[1054] Specific behavior:
[1055] The server parses information from the user profile and generates prompts to feed into the generative AI model.
[1056] Input: A user profile object.
[1057] Data processing: Create a natural language prompt based on the attributes of the user profile.
[1058] Output: A prompt to input to the generative AI model (e.g., "Generate a short, vertical drama in the action genre").
[1059] Step 3: Generate a storyline
[1060] Specific behavior:
[1061] The server uses a generative AI model to generate a storyline based on the prompt.
[1062] Input: The generated prompt statement.
[1063] Data processing: Use a generative AI model (e.g., GPT-3) to automatically generate storyline text.
[1064] Output: The generated storyline text.
[1065] Step 4: Character Selection
[1066] Specific behavior:
[1067] The server analyzes the character attributes (gender, age, personality, etc.) in the generated storyline and references a video database of characters.
[1068] Input: The generated storyline text.
[1069] Data Processing: Based on character attributes, filter matching database entries to select the best character.
[1070] Output: Data of the selected character.
[1071] Step 5: Generate a high-resolution 3D model
[1072] Specific behavior:
[1073] The server generates a high-resolution three-dimensional model based on the video data of the selected character.
[1074] Input: Selected character's data.
[1075] Data processing: Creating a three-dimensional model using 3D modeling software such as Blender or Maya.
[1076] Output: High resolution 3D character model.
[1077] Step 6: Generate footage
[1078] Specific behavior:
[1079] The server integrates the generated script with the selected character model to generate video content.
[1080] Input: Generated storyline text, high-resolution 3D character models.
[1081] Data processing: Analyze each line and scene in the script and add movement and lines to the characters using a video editing system (Adobe Premiere Pro or DaVinci Resolve).
[1082] Output: The finished video file.
[1083] Step 7: Streaming the video
[1084] Specific behavior:
[1085] The server encodes the finished video into the appropriate format and uploads it to the CDN for delivery to end users.
[1086] Input: Final video file.
[1087] Data processing: Encode the video file, upload it to the CDN, and prepare it for distribution on the streaming service.
[1088] Output: A video stream playable on the end user's device.
[1089] This enables the server to quickly provide high-quality short vertical drama content tailored to the viewer's preferences.
[1090] (Application example 1)
[1091] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1092] Modern content distribution services face challenges in quickly providing content tailored to individual user preferences and interests. Complex video content, such as dramas and animations, requires a high level of customization and rapid generation, but no system currently exists that can achieve this. Therefore, efficiently providing high-quality, personalized video content that satisfies users is a major challenge.
[1093] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1094] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical drama using a generative AI model, a means for selecting appropriate characters from an actor video database and generating high-resolution 3D models, a means for generating a video using the generated script and the selected character models, a means for delivering the generated video to a user, a means for generating a story based on the user's genre preferences, and a means for delivering the generated video to a smartphone application. This makes it possible to quickly generate personalized video content based on the user's individual preferences and interests and smoothly deliver it to the user's device.
[1095] A "user profile" is a collection of information such as a user's personal information, viewing history, and preferred genres.
[1096] A "generative AI model" is an artificial intelligence algorithm that uses generative AI to generate sentences or stories based on a specific prompt.
[1097] "Short vertical drama" is a drama-style content in which a story that can be viewed in a short amount of time is displayed on a vertical screen.
[1098] An "actor video database" is a database that stores videos and data of multiple actors.
[1099] A "high-resolution 3D model" is a high-resolution three-dimensional graphic model capable of depicting detailed images.
[1100] "Video generation means" is a function for creating videos using the generated scripts and character models.
[1101] "Distribution means" refers to a system for transmitting the generated video to the user's terminal via the Internet.
[1102] "Genre preferences" refer to the types of dramas and movies that a user is particularly interested in.
[1103] A "smartphone application" is any software program that runs on a smartphone.
[1104] "Story generation means" is a function that allows AI to generate stories based on the user's profile and genre preferences.
[1105] This invention builds a system that generates and delivers customized short, vertical dramas tailored to individual user preferences. This system is realized by creating a user profile, automatically generating drama scripts using a generative AI model based on that information, selecting appropriate characters from a database of actor footage, and generating high-resolution 3D models.
[1106] First, the server retrieves the user's personal information, viewing history, favorite genres, etc. from a database to create a user profile. This user profile is used to gain a detailed understanding of the type of content the user prefers.
[1107] The server then uses a generative AI model to generate a storyline based on the user profile. This generative AI model uses a natural language generation algorithm to automatically generate a story that fits the user's preferred genre. An example of a specific prompt is "Generate a short, vertical drama in the horror genre."
[1108] The server then selects suitable characters from a database of actor footage based on the generated storyline, and generates high-resolution 3D models of the selected characters, which serve as the foundation for the video production.
[1109] The server then generates a video using the generated script and the selected character model. The script is then expanded into lines and scenes using a video editing system (e.g., OpenCV), and the characters move and speak.
[1110] Finally, the server delivers the generated video to a smartphone application. Users can then watch their individually customized, high-quality short vertical dramas on their smartphones. Using a high-performance graphics card (e.g., the NVIDIA RTX series) ensures smooth video playback.
[1111] For example, if a user prefers the action genre, the server creates a profile based on their viewing history and preferred genres and provides the generative AI model with a prompt such as "Generate a short, vertical drama in the action genre." The generative AI model generates an action-packed storyline, and the server extracts key characters from the story. Next, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models. Finally, the generated footage is delivered to the user's smartphone for real-time viewing.
[1112] As described above, by implementing the present invention, it becomes possible to quickly generate personalized video content based on the individual preferences and interests of a user and distribute it to the user's terminal.
[1113] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1114] Step 1:
[1115] The server retrieves the user's personal information, viewing history, preferred genres, etc. from the database and creates a user profile. The input data for the user profile requires the user ID, viewing history, and preferred genres. Based on this, a profile detailing the user's preferences and interests is output.
[1116] Step 2:
[1117] The server uses a generative AI model to generate a storyline based on the user profile. The input data requires the user profile and a prompt, such as "Generate a short, vertical drama in the horror genre." The generative AI model uses these inputs to apply a natural language generation algorithm and output a drama script.
[1118] Step 3:
[1119] The server selects appropriate characters from a video database of actors based on the generated script. In this step, the script content is used as input data, and the IDs and characteristics of the characters that match it are output. This selection determines the main characters of the story.
[1120] Step 4:
[1121] The server generates a high-resolution 3D model of the selected character. The input data for this step is the character's ID and characteristics. Based on this information, a detailed 3D model is output using 3D graphics software (e.g., Blender).
[1122] Step 5:
[1123] The server generates a video using the generated script and the selected character model. Specifically, the script is developed as lines and scenes, and the video is edited so that the characters move and speak the lines. The input data for this step are the script and 3D models, and the final video is output using a video editing system (e.g., OpenCV).
[1124] Step 6:
[1125] The server delivers the generated video to the smartphone application. In this step, the video file is used as input data, and the video is compressed using a high-performance graphics card (e.g., NVIDIA RTX series) and sent to the user's smartphone via the Internet. The final output is video data that can be played on the user's smartphone.
[1126] Step 7:
[1127] The user watches the streamed video on their smartphone. In this step, the smartphone application plays the video file received as input data. Through the application, the user can watch personalized, high-quality short vertical dramas.
[1128] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1129] The present invention provides a system that creates a user profile, generates a script for a short vertical drama using a generative AI model, selects appropriate characters from a database of actor footage, generates high-resolution 3D models, generates footage using the generated script and the selected character models, combines it with an emotion engine that recognizes the user's emotions, and delivers the generated footage to the user.
[1130] Program Overview
[1131] The program for this system is composed of, for example, the following steps:
[1132] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, which then forms the basis for customization for each individual user.
[1133] The server then uses a generative AI model to generate a storyline based on the user profile, using, for example, a natural language generation algorithm to automatically generate a story that fits the user's genre preferences.
[1134] The server then uses an emotion engine to recognize the user's current emotional state, which detects emotions by analyzing the user's facial expressions, voice, and text.
[1135] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[1136] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[1137] The server then generates a video using the generated script and the selected character model. The video editing system then develops the script into lines and scenes, and the characters move and speak.
[1138] Finally, the server delivers the generated video to the user's device via the endpoint, allowing the user to instantly watch the short vertical drama.
[1139] Specific examples
[1140] For example, suppose a user named "Tanaka" likes the action genre and is currently in an emotional state of "excitement." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[1141] The server then uses its emotion engine to recognize that Tanaka's emotional state is "excited," and based on this emotional state, dynamically adjusts the input data of the generative AI model to make the generated storyline even more compelling.
[1142] The server extracts the main characters from the generated story, retrieves data on an actor named "Sato" from a database of actor footage, and generates a high-resolution 3D model. The generated 3D model is then combined with the script to generate the video.
[1143] Finally, the video is streamed to Tanaka's smartphone, where he can watch it instantly, enjoying an exciting short vertical drama that matches his own emotions.
[1144] This system significantly reduces the time and cost required for conventional drama production, and also makes it possible to quickly provide high-quality content that matches the user's emotions.
[1145] The processing flow will be explained below.
[1146] Step 1: Creating a User Profile
[1147] The server retrieves the user's personal information and viewing history from a database.
[1148] The server creates a user profile based on the acquired data.
[1149] For example, the information that "Tanaka watches a lot of action movies" is found in the database of the user, and the preferred genre is recorded as "action" in the profile.
[1150] Step 2: Recognize emotions
[1151] The server uses an emotion engine to recognize the user's emotional state.
[1152] The server analyzes data such as the user's facial expressions, voice, and text to determine their current emotional state.
[1153] For example, it analyzes Tanaka's facial expressions and voice to recognize that he is in an "excited" state.
[1154] Step 3: Story Generation
[1155] The server builds input data for the generative AI model based on the user profile and recognized emotions.
[1156] The server provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre for excited users," and generates a storyline.
[1157] For example, a generative AI model can generate exciting stories such as "a brave police officer defeats a bad guy."
[1158] Step 4: Character Selection and Modeling
[1159] The server extracts the main characters from the generated story.
[1160] The server acquires actor data corresponding to the extracted character from an actor video database.
[1161] For example, data on the actor "Sato" who corresponds to the "brave police officer" is acquired and a high-resolution 3D model is generated.
[1162] Step 5: Generate footage
[1163] The server imports the generated storyline and character models into a video editing system.
[1164] The server applies character movements and dialogue based on the script and renders the video.
[1165] For example, a 3D model of "Sato" will be generated as a "brave police officer" speaking lines and moving to unfold the story.
[1166] Step 6: Streaming the video
[1167] The server obtains the terminal endpoint for the user to view.
[1168] The server transmits the generated video to the user's terminal.
[1169] For example, the generated short vertical drama will be delivered to the smartphone of user "Tanaka" and can be viewed on the spot.
[1170] Through the above steps, the present invention can generate storylines tailored to the user's emotions and provide short, personal dramas customized in real time, thereby significantly reducing the time and cost required for traditional drama production and quickly providing high-quality entertainment.
[1171] Example 2
[1172] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1173] Conventional video production systems have had difficulty quickly providing content tailored to a user's preferences and emotional state. Generating high-quality 3D models and videos also requires a huge amount of time and cost. This makes it difficult to provide personalized video content to individual users. Therefore, there is a need for a system that can dynamically recognize emotions based on each user's profile and adjust scripts and videos accordingly.
[1174] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1175] In this invention, the server includes a means for creating a user profile, a means for generating a script for a short vertical video work using a generative AI model, a means for selecting an appropriate person from a video database of actors and generating a high-resolution 3D model, a means for generating a video using the generated script and the selected person model, a means for delivering the generated video to a user, a means for recognizing the user's emotional state, and a means for adjusting the generated script based on the emotional state, thereby enabling the rapid provision of high-quality video content personalized to each user based on the user's profile and real-time emotional state.
[1176] A "user profile" is a profile created based on data such as a user's personal information, viewing history, and preferences, and reflects the user's characteristics and tastes.
[1177] A "generative AI model" is an algorithmic model that uses artificial intelligence to process natural language and generate content, automatically generating scripts for short vertical video works based on user profiles.
[1178] A "short vertical video work" is video content that can generally be viewed in a short period of time, typically from a few minutes to a few tens of minutes, and is intended to be displayed vertically.
[1179] A "script" is a document that describes the lines, scene structure, character actions, etc. in a video work.
[1180] An "actor" is a person who appears in a video work and plays a specific role in it.
[1181] A "video database" is a database that stores video data related to multiple performers and allows for searching and retrieval.
[1182] A "high-resolution 3D model" is a visually detailed and accurate three-dimensional character model generated using computer software.
[1183] "Emotional state" refers to a user's current psychological and emotional state, analyzed using data obtained from facial expressions, voice, text messages, etc.
[1184] "Distribution" refers to the act of transmitting the generated video content to a user's device via the Internet, allowing the user to view it in real time.
[1185] "Adjustment" refers to the process of changing or modifying parts of the script created by the generative AI model based on the emotional state, optimizing it to better match the user's emotions.
[1186] The present invention relates to a system that creates a user profile, generates a script for a short vertical video using a generative AI model, selects suitable actors from a video database of actors, generates high-resolution 3D models, generates a video using the generated script and the selected actor models, and finally delivers the video to the user's device. It also has the ability to recognize the user's emotional state and adjust the generated script based on that.
[1187] Specifically, the system operates in the following steps.
[1188] First, the server retrieves the personal information and viewing history entered by the user from a database and creates a user profile based on information such as the user's name, age, preferred genres, and viewing history.
[1189] Next, the server provides a prompt based on the profile information to the generative AI model. An example of the generative AI model used here is OpenAI's GPT-4. An example of the prompt is "Generate a short, vertical drama in the action genre."
[1190] The generative AI model automatically generates a storyline based on this prompt. For example, it might generate a storyline in which the protagonist infiltrates the enemy's hideout and engages in a spectacular battle.
[1191] The server also uses an emotion engine (e.g., Emotion API or Affectiva) to recognize the user's current emotional state by analyzing emotions from the user's facial expressions, voice, text messages, etc.
[1192] The server dynamically adjusts the input data of the generative AI model based on the recognized emotions, optimizing the storyline to better match the user's emotions. For example, if a user is in an excited state, the server may add more exciting battle scenes.
[1193] Next, the server selects an appropriate person from a video database of actors based on the generated storyline. For example, it retrieves data on an actor named "Sato" from the video database and generates a high-resolution 3D model of him using Unity or Unreal Engine.
[1194] The server uses the generated script and the selected 3D models to generate the video. Video editing software such as Adobe Premiere Pro or Final Cut Pro is used to generate the video. As a result, lines and scenes are developed based on the script, and the video is completed with characters moving and speaking lines.
[1195] Finally, the server delivers the generated video to the user's device using streaming services such as AWS CloudFront or Akamai. The device receives the delivered video in real time, allowing the user to watch it immediately.
[1196] For example, suppose a user named "Tanaka" likes the action genre and is currently in an "excited" emotional state. The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt: "Generate a short, vertical drama in the action genre." The generative AI model then generates an exciting storyline that matches Tanaka's "excited state."
[1197] The server then uses its emotion engine to recognize Tanaka's emotional state as "excited" and adjusts the input data of the generative AI model based on that state. The server then selects an actor named "Sato" from the generated storyline, generates a 3D model of him, and finally completes the video and distributes it to Tanaka's smartphone. Tanaka can now enjoy an exciting short vertical drama that matches his emotions at that moment.
[1198] This system can significantly reduce the time and cost required for conventional drama production and quickly provide high-quality content that matches the user's emotions.
[1199] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1200] System program processing flow
[1201] Step 1: Creating a User Profile
[1202] The server obtains the personal information and viewing history entered by the user and creates a user profile. Specifically, the server obtains the user's name, age, preferred genres, viewing history, etc. from a database to generate the profile.
[1203] Input: User information from the database
[1204] Output: User profile
[1205] Specific operation: The server obtains the information "User name: Tanaka, Age: 30, Favorite genre: Action, Viewing history: 5 action movies, 1 comedy movie" and creates a profile for "Tanaka."
[1206] Step 2: Generate a storyline
[1207] The server provides a prompt based on the user profile to the generative AI model to generate a storyline, for example, "Generate a short, vertical drama in the action genre."
[1208] Input: User Profile
[1209] Output: Generated storyline
[1210] Specific operation: Based on the information that "Tanaka's preference is the action genre," the server gives the generative AI model a prompt statement of "Generate a short vertical drama in the action genre." The generative AI model generates a storyline in which "the protagonist infiltrates the enemy's hideout and engages in a spectacular battle."
[1211] Step 3: Recognize emotions
[1212] The device captures the user's facial expressions and voice through a camera and microphone and sends the data to a server, which uses an emotion engine to analyze the user's emotional state.
[1213] Input: User's facial expression data, voice data
[1214] Output: User's emotional state
[1215] Specific operation: The device captures Tanaka's facial expressions with a camera and collects his voice with a microphone. The server analyzes Tanaka's data using an emotion engine and recognizes that Tanaka is in an "excited" state.
[1216] Step 4: Adjusting the storyline
[1217] Based on the recognized emotions, the server dynamically adjusts the input data of the generative AI model, optimizing the storyline to better match the emotions.
[1218] Input: Generated storyline, user's emotional state
[1219] Output: Adjusted storyline
[1220] Specific operation: The server instructs the generative AI model to "add more exciting battle scenes" based on the information "current emotional state: excitement." The generative AI model adds a plot twist to the storyline in which the protagonist uses a secret weapon he found in the basement to wipe out the enemies.
[1221] Step 5: Select and generate a character model
[1222] Based on the generated storyline, the server selects appropriate characters from a database of actor footage and generates high-resolution 3D models.
[1223] Enter: the adjusted storyline
[1224] Output: High-resolution 3D character model
[1225] Specific operation: The server retrieves data on "Actor A," who is the best suited to play the main character of the story, from the database and generates a 3D model of Actor A in Unity or Unreal Engine.
[1226] Step 6: Generate footage
[1227] The server uses the generated script and 3D character model to generate video, which is then edited using video editing software such as Adobe Premiere Pro or Final Cut Pro.
[1228] Input: Adjusted storyline, 3D character models
[1229] Output: Generated video
[1230] How it works: Based on the generated script, the server uses a 3D model of Actor A to create a battle scene with an enemy. Adobe Premiere Pro is used to align the timing of lines and scenes and complete the final footage.
[1231] Step 7: Streaming the video
[1232] The server delivers the generated video to the user's device using streaming services such as AWS CloudFront and Akamai.
[1233] Input: Generated video
[1234] Output: Video played on the user's device
[1235] How it works: The server uploads the completed video file to AWS CloudFront and generates a distribution URL. The device receives the URL and immediately plays the video, allowing users to enjoy the short vertical drama on the spot.
[1236] As described above, high-quality video content that matches the user's emotions and preferences is quickly provided through each step.
[1237] (Application example 2)
[1238] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1239] Conventional content delivery services have had difficulty generating and delivering personalized video content in real time based on user preferences and emotions. In particular, providing content that takes into account the user's emotional state is difficult, and there has been a demand for a way to improve the quality of the viewing experience.
[1240] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1241] a means for creating a user profile;
[1242] A means for generating scripts for short vertical dramas using a generative AI model;
[1243] A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models;
[1244] means for generating a video using the generated script and a selected character model;
[1245] a means for delivering the generated video to a user;
[1246] A means for dynamically adjusting input data for a generative AI model using an emotion engine that recognizes user emotions; and
[1247] This will enable the generation and delivery of personalized video content in real time based on the user's emotional state.
[1248] A "user profile" is information created by organizing and analyzing data about a user based on the user's personal information, viewing history, preferred genres, etc.
[1249] A "generative AI model" is an artificial intelligence algorithm that automatically generates stories and scripts based on user profiles and specific input data.
[1250] The "actor video database" is a database that stores the video and attribute data of multiple actors, and is used to select appropriate characters.
[1251] "High-Resolution 3D Model" means a highly detailed, high-resolution, three-dimensional computer model used to visually recreate a Selected Character.
[1252] The "emotion engine" is an engine that analyzes and recognizes emotions from a user's facial expressions, voice, text, etc.
[1253] The "video generation means" is a means for developing scenes, lines, etc. and creating video using the generated script and selected character models.
[1254] "Video distribution means" refers to a means for distributing the generated video to end users in real time.
[1255] This invention is a system that creates a user profile, generates personalized short-form dramas for users using a generative AI model, and provides storylines based on the user's emotions using an emotion engine. This system is implemented through the following specific process.
[1256] First, the server retrieves the user's personal information and viewing history from a database to create a user profile, including the user's preferred genres and past viewing history, providing the foundation for customization for each individual user.
[1257] The server then uses a generative AI model to generate a storyline based on the user profile. The generative AI model uses a natural language generation algorithm (e.g., GPT-4) to automatically generate a story that fits the user's preferred genre.
[1258] Furthermore, the server uses an emotion engine to recognize the user's current emotional state. The emotion engine detects emotions by analyzing the user's facial expressions and voice data using a facial recognition camera (e.g., Face API) and voice analysis tools (e.g., Azure Cognitive Services).
[1259] The server then dynamically adjusts the input data of the generative AI model based on the recognized emotions, generating a storyline that matches the user's emotions, and also altering the pre-generated script storyline accordingly.
[1260] Based on the generated storyline, the server selects suitable characters from a database of actor footage and generates high-resolution 3D models of the selected characters, which then serve as the basis for video production.
[1261] The server then generates a video using the generated script and the selected character model. The script is then developed into lines and scenes using a video editing system (e.g., Blender), and the characters move and speak.
[1262] Finally, the server delivers the generated video to the user's device. Using a streaming distribution server (e.g., AWS S3 + CloudFront), the video is delivered to the user's device through an endpoint, allowing the user to instantly watch the short vertical drama.
[1263] As a concrete example, suppose a user named "Tanaka" likes the action genre and his current emotional state is "excited." The server creates a profile based on Tanaka's viewing history and preferred genres, and provides the generative AI model with a prompt to "generate a short, vertical drama in the action genre." The generative AI model then generates an even more exciting storyline that matches Tanaka's excitement level.
[1264] An example of a prompt used in this process is:
[1265] "Generate a comedy-action story that is suitable for when the user is in a happy state. The main character should have a sense of humor and the episodes should be thrilling with action scenes."
[1266] As a result, this invention makes it possible to efficiently generate and distribute personalized video content that matches the user's emotions, improving the quality of the viewing experience.
[1267] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1268] Step 1:
[1269] The server retrieves the user's personal information and viewing history from the database and creates a user profile. This profile creation step takes the user ID as input and generates profile data including the user's preferred genres and viewing history as output. Specifically, it uses SQL queries to retrieve user information from the database, organizes the necessary data, and builds the profile.
[1270] Step 2:
[1271] The server uses a generative AI model to generate a storyline based on the user profile. In this step, the user profile data is used as input and the generated storyline is obtained as output. Specifically, a prompt sentence reflecting the user's preferred genre and emotional state is generated, and a story is created using a natural language generation algorithm. The prompt sentence is input into the generative AI model (e.g., GPT-4) to obtain the storyline.
[1272] Step 3:
[1273] The server uses an emotion engine to recognize the user's current emotional state. In this step, the user's real-time facial expressions and voice data are used as input, and the user's emotional state is obtained as output. Specifically, the data is analyzed using a facial recognition camera and voice analysis tools (e.g., Face API, Azure Cognitive Services) to detect the user's emotions.
[1274] Step 4:
[1275] The server dynamically adjusts the input data of the generative AI model based on the recognized emotion. In this step, the emotional state is used as input and a storyline tailored to the emotion is obtained as output. Specifically, the recognized emotion is used to regenerate the prompt sentence, which is then reinput into the generative AI model to generate a story that is appropriate for the emotion. For example, the prompt sentence "Generate an action story that is appropriate for the user's excited state" is input into the generative AI model.
[1276] Step 5:
[1277] The server selects appropriate characters from the actor video database based on the generated storyline and generates high-resolution 3D models. In this step, the storyline is used as input and 3D character models are obtained as output. Specifically, the server extracts the main characters appearing in the story, selects appropriate actors from the video database, and generates their high-resolution 3D models.
[1278] Step 6:
[1279] The server generates a video using the generated script and the selected character model. In this step, the script and character model are used as input, and a short vertical drama video is obtained as output. Specifically, a video editing system (e.g., Blender) is used to create scenes including character movements and dialogue, and the entire video is generated.
[1280] Step 7:
[1281] The server delivers the generated video to the user's device. In this step, the generated video is used as input and the video delivered to the user's device is obtained as output. Specifically, the video is delivered using a streaming distribution server (e.g., AWS S3 + CloudFront), and the user watches this video on their smartphone.
[1282] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1283] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1284] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1285] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1286] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1287] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1288] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1289] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1290] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1291] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1292] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1293] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1294] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1295] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1296] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1297] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1298] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1299] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1300] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1301] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1302] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1303] The following is further disclosed regarding the above embodiment.
[1304] (Claim 1)
[1305] a means for creating a user profile;
[1306] A means for generating scripts for short vertical dramas using a generative AI model;
[1307] A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models;
[1308] means for generating a video using the generated script and a selected character model;
[1309] a means for delivering the generated video to a user;
[1310] A system including:
[1311] (Claim 2)
[1312] 10. The system of claim 1, further comprising: means for constructing input data to the generative AI model based on a user profile.
[1313] (Claim 3)
[1314] The system of claim 1 , further comprising: means for extracting a main character from the generated script.
[1315] "Example 1"
[1316] (Claim 1)
[1317] a means for creating a user profile;
[1318] A means for generating scripts for short vertical dramas using a generative AI model;
[1319] a means for selecting an appropriate character from a human image database and generating a high-resolution three-dimensional model;
[1320] means for generating a video using the generated script and a selected character model;
[1321] a means for delivering the generated video to a user;
[1322] A system including:
[1323] (Claim 2)
[1324] 10. The system of claim 1, further comprising: means for constructing input data to the generative AI model based on a user profile.
[1325] (Claim 3)
[1326] The system of claim 1 , further comprising: means for extracting a main character from the generated script.
[1327] "Application Example 1"
[1328] (Claim 1)
[1329] a means for creating a user profile;
[1330] A means for generating scripts for short vertical dramas using a generative AI model;
[1331] A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models;
[1332] means for generating a video using the generated script and a selected character model;
[1333] a means for delivering the generated video to a user;
[1334] a means for generating stories based on a user's genre preferences;
[1335] A means for delivering the generated video to a smartphone application;
[1336] A system including:
[1337] (Claim 2)
[1338] 10. The system of claim 1, further comprising: means for constructing input data to the generative AI model based on a user profile.
[1339] (Claim 3)
[1340] The system of claim 1 , further comprising: means for extracting a main character from the generated script.
[1341] "Example 2: Combining Emotion Engines"
[1342] (Claim 1)
[1343] a means for creating a user profile;
[1344] A means for generating a script for a short vertical video work using a generative AI model;
[1345] A means to select suitable actors from a video database and generate high-resolution 3D models;
[1346] means for generating a video using the generated script and a selected character model;
[1347] a means for delivering the generated video to a user;
[1348] a means for recognizing the emotional state of a user;
[1349] means for adjusting the generated script based on said emotional state;
[1350] A system including:
[1351] (Claim 2)
[1352] 10. The system of claim 1, further comprising: means for constructing input data to the generative AI model based on a user profile.
[1353] (Claim 3)
[1354] The system of claim 1 , further comprising: means for extracting key characters from the generated script.
[1355] "Application example 2 when combining emotion engines"
[1356] (Claim 1)
[1357] a means for creating a user profile;
[1358] A means for generating scripts for short vertical dramas using a generative AI model;
[1359] A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models;
[1360] means for generating a video using the generated script and a selected character model;
[1361] a means for delivering the generated video to a user;
[1362] A means for dynamically adjusting input data for a generative AI model using an emotion engine that recognizes user emotions; and
[1363] A system including:
[1364] (Claim 2)
[1365] 10. The system of claim 1, further comprising: means for constructing input data to the generative AI model based on a user profile.
[1366] (Claim 3)
[1367] The system of claim 1 , further comprising: means for extracting a main character from the generated script. [Explanation of symbols]
[1368] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for creating a user profile; A means for generating scripts for short vertical dramas using a generative AI model; A means to select appropriate characters from a database of actor footage and generate high-resolution 3D models; means for generating a video using the generated script and a selected character model; a means for delivering the generated video to a user; A system including:
2. The system of claim 1 , further comprising: means for constructing input data to the generative AI model based on a user profile.
3. The system of claim 1 further comprising: means for extracting a main character from the generated script.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A