System
The system automates filmmaking by using generative AI to analyze user ideas, generate film elements, and integrate them, allowing users to produce high-quality films efficiently.
Patent Information
- Application Number
- JP2024131502
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-02-20
AI Technical Summary
Traditional filmmaking is costly and time-consuming, requiring significant resources and specialized knowledge, making it difficult for individuals to create high-quality films independently.
A system that utilizes generative AI to analyze user ideas and stories, generate scene and character information, create an appropriate cast, produce video and audio data, and integrate these elements to produce a film, allowing users to preview and make revisions.
Enables anyone to easily create high-quality films by automating the filmmaking process, reducing the effort and expertise required.
Smart Images

Figure 2026028885000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Traditional filmmaking requires a huge amount of money and time, making it extremely difficult for individuals to create large-scale films, as it requires fundraising, arranging casts, and specialized knowledge. This has resulted in many potential creators losing opportunities to express their creativity. Furthermore, even when producing a short film independently, the enormous effort required for shooting, editing, and directing makes it difficult to produce a high-quality work. The present invention aims to solve these problems by automating the filmmaking process, making it possible for anyone to easily create high-quality films. [Means for solving the problem]
[0005] The present invention provides a system that includes a means for analyzing an idea or story input by a user and generating scene and character information based on the results. The system further includes a means for automatically generating an appropriate cast based on the generated character information, and a means for generating video data based on the scene information and cast information. The system also includes a means for generating character voices and background music, and a means for integrating and editing the generated video data, voices, and music. Users can preview the completed film and collect feedback to make revisions to the work. Integrating these elements greatly simplifies the filmmaking process, allowing anyone to easily create high-quality films.
[0006] "User" refers to an individual or organization that intends to use this system to create a film.
[0007] "Idea" refers to the user's basic concept for the plot, theme, or scenario of a movie.
[0008] "Story" refers to the flow of the story, including the specific content and progression of the film, and the development of each scene.
[0009] "Analysis" refers to the process of extracting scene and character information based on the idea and story entered by the user.
[0010] A "scene" is a sequence of images, sounds, and actions that form a particular situation or setting in a film.
[0011] "Character" refers to people, animals, and other entities with a specific role that appear in the film.
[0012] "Cast" refers to the model data of the actors and voice actors who play the characters in the film.
[0013] "Video data" refers to digital video information including the background, character movements, camera work, etc. corresponding to a scene.
[0014] "Audio" refers to auditory elements such as character dialogue, sound effects, and natural environmental sounds.
[0015] "Background music" refers to music or background music that is appropriate for each scene in a movie.
[0016] "Editing" refers to the process of integrating the generated video data, audio, and music to create an overall flow for the film.
[0017] "Preview" refers to the act of a user viewing a completed movie and checking its contents.
[0018] "Feedback" refers to information that conveys to the system any improvements or corrections that the user feels through the preview.
[0019] "Modifications" refers to changes made to the visual data, audio, and music of a movie based on user feedback. [Brief explanation of the drawings]
[0020] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0021] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0022] First, the terms used in the following description will be explained.
[0023] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0024] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0025] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0026] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0027] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0028] [First embodiment]
[0029] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0030] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0032] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0033] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0034] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0036] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0037] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0038] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0039] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0040] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0041] This invention relates to the "Master's Egg Platform," a platform that uses generative AI to automatically create the elements necessary for filmmaking. The system begins by analyzing the idea and story entered by the user and generating scene and character information. The system then generates a cast suited to the characters, generates and edits video data and audio, and completes the final film.
[0042] Program processing overview
[0043] 1. Story input and analysis
[0044] The user inputs the plot or scenario of their own movie into the device. The device receives the user's input and sends it to the server as story data. The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. For example, if the user inputs "a mystery that takes place in a rural town," the server will identify the following scenes and characters:
[0045] Scene 1: The protagonist arrives in a rural town.
[0046] Scene 2: A mysterious incident occurs.
[0047] Scene 3: The protagonist works with the detective.
[0048] Characters: Protagonist, Detective, Villager.
[0049] 2. Cast Selection
[0050] Based on the analyzed character information, the server uses generative AI to generate appropriate cast members, such as a young male protagonist or a middle-aged male detective, dynamically generating model data for the appearance and voice of each character.
[0051] 3. Image Generation
[0052] Next, the server generates video data based on the scene information. The scene background, character movements, camera work, and other aspects are automatically generated. For example, a "rural town" is generated as the background for Scene 1 to look natural, and the scene in which the main character gets off at the train station is depicted.
[0053] 4. Speech and Music Generation
[0054] The server requests the AI to generate lines and sound effects to generate character voices and background music. For example, the main character's line, "Is this a rural town?", can be generated in a calm voice while unsettling background music plays.
[0055] 5. Direction and Editing
[0056] The server integrates the video, audio, and music to create the overall structure of the film, including smooth transitions between scenes and applying appropriate production effects to create a complete film.
[0057] 6. Final review and feedback
[0058] The user previews the completed movie on their device and checks the content. If necessary, corrections can be sent as feedback to the server. The server receives the feedback and makes the necessary changes by requesting corrections from the generation AI. For example, if the user feels that "the background in Scene 1 is too light," the server will regenerate the background.
[0059] By automatically generating the elements necessary for film production, the present invention provides an environment in which users can easily create high-quality films. This invention allows users to focus on the creative aspects, making it easy for anyone to realize their filmmaking dreams.
[0060] The processing flow will be explained below.
[0061] Step 1:
[0062] Users input the plot and scenario of their own movie into the terminal, which then collects the ideas and story data entered by the user based on a form and formats it into story data.
[0063] Step 2:
[0064] The device sends the formatted story data to the server, which receives the story data and stores it in a database.
[0065] Step 3:
[0066] The server passes the received story data to a natural language processing engine, which then analyzes each element of the plot and extracts scene and character information.
[0067] Step 4:
[0068] The server creates a scene list and a character list from the analysis results. For example, it compiles information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[0069] Step 5:
[0070] The server passes the character list to the generation AI module and asks it to generate an appropriate cast. The generation AI dynamically generates face, body, and voice models for the characters.
[0071] Step 6:
[0072] The server receives the cast data generated by the AI and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[0073] Step 7:
[0074] The server starts generating video data based on the scene information. It requests the AI to generate the background, character movements, and camera work for each scene, and generates the video data.
[0075] Step 8:
[0076] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[0077] Step 9:
[0078] The server requests character dialogue, necessary sound effects, and background music from the AI generator, which then generates the appropriate audio data and background music for the scene.
[0079] Step 10:
[0080] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[0081] Step 11:
[0082] The server calls the editing module to integrate the video data, audio, and background music, and the editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects.
[0083] Step 12:
[0084] The server generates the final edited movie data and provides it to the user, who can then preview the completed movie through their terminal and check the content.
[0085] Step 13:
[0086] Users can provide feedback from their devices to the server about improvements and corrections they feel are necessary through the preview. The feedback is sent to the server as specific instructions for improvement.
[0087] Step 14:
[0088] The server analyzes the received feedback and requests the generation AI to make any necessary corrections, such as regenerating the background of Scene 1.
[0089] Step 15:
[0090] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[0091] Example 1
[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0093] The traditional filmmaking process requires a significant amount of time and expertise, making it difficult for ordinary people to easily create films. Furthermore, the creation and integration of each element (scenes, cast, video, audio, music, editing) requires a lot of manual work, making it inefficient. A method was needed to solve these issues and enable users to easily create high-quality films.
[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0095] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing a user to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, and means for operating based on information including specific names of software or hardware to be used, thereby enabling a user to easily and quickly create high-quality movies.
[0096] "User" means any person or entity that uses the Filmmaking Platform.
[0097] "Server" refers to a computer system that receives, analyzes, and processes input data from users and generates and integrates each element required for film production.
[0098] "Terminal" refers to a computing device through which a user inputs a film plot or scenario and which provides an interface with the film production platform.
[0099] "Story" refers to the plot or scenario of a movie entered by the user.
[0100] "Character information" refers to information about characters extracted through story analysis.
[0101] "Cast" refers to model data of the appearance and voice of characters generated based on character information.
[0102] "Scene information" refers to information about the background and content of a scene extracted through story analysis.
[0103] "Video data" refers to visual data generated based on scene information and cast information.
[0104] "Audio" refers to data related to a character's lines and vocalizations.
[0105] "Background music" refers to music created to enhance the atmosphere of a film scene.
[0106] "Integration" refers to the process of assembling the generated video data, audio, and background music into a single film.
[0107] "Editing" refers to the process of arranging the overall structure of a film by adding transitions between scenes and dramatic effects.
[0108] "Feedback" means suggestions for corrections or improvements provided by Users through Previews.
[0109] "Generative AI Model" refers to the artificial intelligence model used to generate video, audio, and background music.
[0110] A "prompt" refers to text that a user enters to give specific instructions to a generative AI model.
[0111] This invention relates to a platform that uses generative AI to automatically create elements necessary for user-generated filmmaking. The platform begins by analyzing the idea and story input by the user and generating scene and character information.
[0112] The user inputs the plot or scenario of his / her own movie using the terminal. For example, if the plot is "a mystery that takes place in a rural town," the terminal sends the input to the server as story data.
[0113] The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. A typical example of this software is OpenAI's GPT series. Through this process, the plot of "Mystery Set in a Country Town" is analyzed as follows:
[0114] Scene 1: The protagonist arrives in a rural town.
[0115] Scene 2: A mysterious incident occurs.
[0116] Scene 3: The protagonist works with the detective.
[0117] Characters: Protagonist, Detective, Villager.
[0118] Next, the server uses generative AI to generate an appropriate cast based on the analyzed character information, such as Stable Diffusion. In this step, model data for the appearance and voice of characters such as a young male protagonist or a middle-aged male detective is generated.
[0119] The server then generates video data based on the scene information. At this stage, image generation AI such as DALL-E 2 is used to automatically create the scene's background, character movements, and camerawork. Specifically, a "rural town" is generated as the background for Scene 1, depicting the scene in which the protagonist gets off at the train station.
[0120] The server then generates the character voices and background music, using Amazon Polly for voice generation and Jukedeck for music generation. This allows, for example, the protagonist's line, "Is this a rural town?" to be generated in a calm voice while unsettling background music plays.
[0121] The server uses editing software such as Adobe Premiere Pro to combine and edit the generated video, audio, and music, adding transitions between scenes and applying appropriate effects to create a complete film.
[0122] Finally, the user previews the completed movie on their device and checks the content. If necessary, the user can send feedback to the server. For example, if the user feels that the background in Scene 1 is too weak, the server will use generative AI to regenerate the background.
[0123] For example, the following prompts can be used:
[0124] "I want to make a mystery movie in which the protagonist works with a detective to solve a mystery in a rural town. Scene 1 is when the protagonist arrives in the town, Scene 2 is when a mysterious incident occurs, and Scene 3 is when the protagonist and detective work together. The characters are the protagonist, the detective, and the villagers."
[0125] As a result, the present invention significantly improves the efficiency of the movie production process and provides an environment in which users can easily produce high-quality movies.
[0126] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0127] Step 1:
[0128] The user inputs the basic plot or scenario of their movie into the terminal. The input story data is sent from the terminal to the server. For example, the user inputs a prompt such as "A mystery that takes place in a rural town." The input data is in text format and is sent to the server.
[0129] Step 2:
[0130] The server analyzes the story data received from the device. This analysis uses natural language processing technology. Specifically, it extracts scene and character information from the story data. For example, it uses OpenAI's natural language processing model to extract scenes such as "Scene 1: The protagonist arrives in a rural town," "Scene 2: A mysterious incident occurs," and "Scene 3: The protagonist cooperates with a detective." The output data is JSON data containing scene and character information.
[0131] Step 3:
[0132] The server uses a generative AI model to generate appropriate cast members based on the analyzed character information. Specifically, it creates model data for the characters' appearances and voices. Generative AI models such as "Stable Diffusion" are used for this step. For example, based on information that the protagonist is a young man and the detective is a middle-aged man, it generates appearance and voice models for each character. The output data is the character's 3D model data and voice data.
[0133] Step 4:
[0134] The server generates video data based on the scene and cast information obtained in the previous step. At this stage, image generation AI such as "DALL-E 2" is used to automatically create the scene's background, character movements, and camerawork. For example, it generates and outputs a scene in which the protagonist gets off at a train station in a rural town. The output data is a continuous image data of the scene, in other words, a video clip.
[0135] Step 5:
[0136] The server then generates the character's voice and background music. It uses a "voice synthesis API" for voice generation and "music generation software" for music generation. For example, it generates the protagonist's line, "Is this a rural town?", and adds background music that creates an unsettling feeling. The output data is a voice file and a music file.
[0137] Step 6:
[0138] The server then integrates and edits the generated video, audio, and music. This is done using video editing software. Specifically, it applies transitions between scenes and special effects to create a continuous video that functions as a single movie. For example, it applies a fade-in and fade-out effect from scene 1 to scene 2. The output data is a completed movie file.
[0139] Step 7:
[0140] The user previews the completed movie on the device and checks its content. The user then sends feedback from the device to the server. For example, the user may give feedback such as "The background in Scene 1 is too light." The input data is in the form of text feedback and is sent to the server.
[0141] Step 8:
[0142] The server makes the necessary corrections based on the feedback received from the user. The server then uses the generative AI model again to reflect the corrections. In this step, corrections such as "regenerating the background" are made. The output data is the corrected movie file.
[0143] This series of processes allows users to easily and quickly create high-quality movies.
[0144] (Application example 1)
[0145] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0146] In conventional film production, the creation and editing of scripts, images, and audio is done manually, requiring a great deal of time and effort. Furthermore, film production in a virtual reality environment requires specialized knowledge, making it difficult for general users to easily use. Furthermore, there has been no system that can automatically and accurately create films in a virtual reality environment, from story ideas to detailed image construction and audio integration. Therefore, there has been a demand for an environment that allows users to easily create high-quality films in a virtual reality environment.
[0147] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0148] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for integrating and editing the generated video data, audio, and music, means for a user to input a movie idea in a virtual reality environment and perform story analysis and scene generation using a generative AI model, and means for reproducing the generated scenes and characters in the virtual reality environment. This allows users to easily produce movies in a virtual reality environment, enabling the creation of high-quality movies.
[0149] A "user" is an entity that uses this system to input movie ideas and stories and create the final movie.
[0150] "Ideas and stories" are information about the plot and scenario of a movie entered by the user.
[0151] The "analysis means" is a device or program that has the function of analyzing the idea or story input by the user and generating scene and character information.
[0152] "Character information" is information about the characters in the movie generated by the analysis means.
[0153] The "cast generation means" is a device or program that has the function of generating an appropriate cast based on the character information.
[0154] "Scene information" is information about the scene and setting of the movie generated by the analysis means.
[0155] The "video generation means" is a device or program that has the function of generating video data based on the scene information and cast information.
[0156] "Voice and music generating means" refers to a device or program that has the function of generating the voice and background music of the character.
[0157] The "integrated editing means" is a device or program that has the function of integrating and editing the generated video data, audio, and music.
[0158] A "previewer" is a device or program that has the ability to allow users to preview the completed movie and gather feedback.
[0159] The "feedback correction means" is a device or program that has the function of correcting the video data, audio, and music based on the feedback.
[0160] A "virtual reality environment" is an environment in which users can create and watch movies in a virtual cinematic space using dedicated hardware.
[0161] A "generative AI model" is a model that uses artificial intelligence to analyze stories and generate scenes based on user input.
[0162] "Scene playback means" refers to a device or program that has the function of playing back the generated scenes and characters in a virtual reality environment.
[0163] The present invention is a system that allows users to create movies in a virtual reality environment. By inputting ideas and stories, the system automatically generates scenes and characters, enabling the creation and playback of movies in the virtual reality environment.
[0164] The system consists of the following main components:
[0165] 1. Input and Analysis
[0166] Users input ideas and stories using a virtual reality environment, such as a head-mounted display. This input is done through voice input or gesture input. The server receives this and analyzes it using a generative AI model. Natural language processing technology is used for analysis. For example, if a user inputs "a story about a young farmer fighting a dragon in a medieval fantasy world," the server will generate scene and character information based on this.
[0167] 2. Scene and character generation
[0168] The server generates scene and character information based on the analysis results. Based on the character information, it then generates appropriate cast members. The generative AI model dynamically generates character appearance and voice model data, as well as background information for the scene, based on user input.
[0169] 3. Video and audio generation
[0170] The server then generates video data based on the scene and character information. The scene's background, character movements, and camerawork are automatically generated. Character voices and background music are also generated using generative AI models. For example, if a character says something like, "Is this a rural town?", the server generates the voice and adds appropriate background music.
[0171] 4. Merging and Editing
[0172] The server then integrates and edits the generated video data, audio, and music, including smooth transitions between scenes and appropriate stage effects, and the integrated editing process completes the entire film.
[0173] 5. Preview and Feedback
[0174] The completed movie can be previewed by the user in a virtual reality environment. The user watches the movie and sends feedback to the server if necessary. A feedback corrector corrects the video data, audio, and music based on the feedback to improve the quality.
[0175] 6. Playback in a Virtual Reality Environment
[0176] Finally, users can play the completed movie in a virtual reality environment and enjoy the movie-making experience. For example, the following scene is generated based on the user's input: "A story about a young farmer fighting a dragon in a medieval fantasy world."
[0177] Example prompt sentence:
[0178] Generate scenes and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world.
[0179] Generated AI output example:
[0180] Scene 1: A farmer lives an ordinary life in the village.
[0181] Scene 2: The dragon attacks the village.
[0182] Scene 3: The farmer takes up his sword and confronts the dragon.
[0183] Characters: Farmer, dragon, villagers.
[0184] In this way, this invention allows users to intuitively create movies in a virtual reality environment. By utilizing generative AI models and enabling the automatic generation of scenes and characters, the effort required for movie production is significantly reduced, allowing anyone to enjoy high-quality movie production.
[0185] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0186] Step 1:
[0187] The user wears a head-mounted display in a virtual reality environment and launches a dedicated application. Story ideas and plots are input using voice or gestures. The input data is sent to the server as text data on the device. The server begins analysis based on the user's input (prompt): "A story about a young farmer fighting a dragon in a medieval fantasy world."
[0188] Step 2:
[0189] The server analyzes the received text data. A generative AI model is used for this analysis, and natural language processing techniques are used to extract scene and character information for the story. The specific prompt is: "Generate a scene and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world." Based on this, the server generates the following scene and character information. The output data is as follows:
[0190] Scene 1: A farmer lives an ordinary life in the village.
[0191] Scene 2: The dragon attacks the village.
[0192] Scene 3: The farmer takes up his sword and confronts the dragon.
[0193] Characters: Farmer, dragon, villagers.
[0194] Step 3:
[0195] The server generates an appropriate cast based on the generated scene information and character information. This cast generation includes model data for the character's appearance and voice. This process also uses a generative AI model. The output is the following cast information:
[0196] Cast 1: Young farmer character
[0197] Cast 2: Dragon character
[0198] Cast 3: Villager characters
[0199] Step 4:
[0200] The server generates video data based on scene and cast information. This video generation process includes the scene background, character movements, and camera work. The server generates a "village landscape" as the background for Scene 1 and constructs the scene by combining the character movements. The output is video data.
[0201] Step 5:
[0202] The server generates character voices and background music for the generated scenes. Specifically, it uses a generative AI model to generate character lines as audio data and inserts appropriate music into the background. For example, in Scene 1, the "farmer" might say, "Is this a village?" and idyllic rural music would play in the background. The output is audio data and music data.
[0203] Step 6:
[0204] The server then integrates and edits the generated video, audio, and music data. This process involves smooth transitions between scenes and applying appropriate special effects. The integrated editing process assembles the data into a single film. The output is the integrated film data.
[0205] Step 7:
[0206] The user previews the integrated movie in the virtual reality environment. After the preview, the user provides feedback as needed. The feedback is sent back to the server, and the server modifies the video data, audio data, and music data based on the feedback. For example, if the user feels that "the background of Scene 1 is too light," the server regenerates the background. The output is the modified movie data.
[0207] Step 8:
[0208] The modified movie data is finally played in the virtual reality environment, allowing users to enjoy their own high-quality movie in a virtual reality environment. The user's final experience is the output.
[0209] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0210] This invention relates to the "Master's Egg Platform," a movie creation platform that combines an emotion engine that recognizes the user's emotions. This system is based on analyzing the ideas and stories entered by the user and generating scene and character information based on them. Furthermore, it aims to reflect the user's emotions in the process of generating a cast based on the character information, generating and editing video data and audio, and completing the final movie.
[0211] Program processing overview
[0212] 1. Story Input and Emotion Recognition
[0213] When a user inputs a movie plot or scenario into the device, the emotion engine simultaneously collects facial expressions and voice tones via cameras and sensors. The device processes this data and analyzes the user's emotional state (e.g., excited, sad, surprised, etc.). The results are then sent to the server along with the story data.
[0214] 2. Story analysis and emotional reflection
[0215] The server analyzes the received story data and emotion data to generate scene and character information. This analysis uses natural language processing and emotion analysis technologies. For example, if a user excitedly talks about a "mystery that takes place in a rural town," the server will consider settings that will give the scene and characters a sense of urgency:
[0216] Scene 1: The protagonist arrives in a rural town.
[0217] Scene 2: A mysterious incident occurs.
[0218] Scene 3: The protagonist works with the detective.
[0219] Characters: Protagonist, Detective, Villager.
[0220] 3. Cast Generation
[0221] The server then passes the character information, reflecting the emotional data, to the AI generator to generate an appropriate cast. For example, the main character is generated with facial and vocal characteristics that express excitement.
[0222] 4. Image Generation
[0223] Next, the server generates video data that reflects the emotional data based on the scene information, for example by using camera work and lighting that create a sense of tension.
[0224] 5. Speech and Music Generation
[0225] The server generates character voices and background music based on the user's emotional data. For example, exciting scenes will have fast-paced background music.
[0226] 6. Direction and Editing
[0227] The server integrates video data, audio, and music to edit the entire movie. The emotional data is used to adjust the production effects. Specifically, scene transitions and sound effects are set according to the emotions.
[0228] 7. Final review and feedback
[0229] The user can preview the completed movie on their device. During the preview, the user's emotions are recognized again and their reactions to the movie's content are recorded. Based on this, improvements and corrections are sent to the server as feedback.
[0230] 8. Modify and Regenerate
[0231] The server receives the feedback and requests the AI to make any necessary corrections. For example, if the user feels that the background in Scene 1 is too light, the AI will regenerate that background. The AI will also readjust the video and audio based on the emotional feedback.
[0232] This invention allows the user's emotions to be reflected throughout the entire system, enabling more intuitive and emotional movie production. Users can easily express their own emotions, and as a result, high-quality movies can be easily produced.
[0233] The processing flow will be explained below.
[0234] Step 1:
[0235] The user logs in to the device's input interface and inputs the plot and scenario of their movie. The camera and microphone collect the user's facial expressions and voice tone in real time, and the emotion engine analyzes the user's emotional state. For example, if the user is speaking excitedly, that emotional data (excitement state) is obtained.
[0236] Step 2:
[0237] The device sends the collected story data and emotion data to the server. For example, the plot of a "mystery that takes place in a rural town" entered by the user is sent along with the emotional excitement data at the time.
[0238] Step 3:
[0239] The server then passes the received story data to the natural language processing engine and begins analyzing the story. At the same time, it also incorporates the results of the emotion data analysis by the emotion engine. For example, the server takes into account settings that reflect the user's excitement level when analyzing the plot and extracting scene and character information.
[0240] Step 4:
[0241] The server creates a scene list and a character list based on the analysis results. For example, it creates a list containing information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[0242] Step 5:
[0243] The server passes the character list to a generation AI module, which generates an appropriate cast. For example, it generates a face model for the main character with bright eyes and a powerful voice profile, reflecting the user's excitement level.
[0244] Step 6:
[0245] The server receives the generated cast data and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[0246] Step 7:
[0247] The server begins generating video data based on the scene information. It requests the AI to create the background, character movements, and camerawork for each scene, and generates the video data. Emotional data is also reflected, and camerawork and lighting that create a sense of tension are set, for example.
[0248] Step 8:
[0249] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[0250] Step 9:
[0251] The server requests character dialogue, necessary sound effects, and background music from the AI generator. The AI then generates audio data and background music appropriate for the scene. The user's emotions are reflected in the sound, so for example, an exciting scene will have fast-paced background music.
[0252] Step 10:
[0253] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[0254] Step 11:
[0255] The server calls the editing module to integrate the video data, audio, and music. The editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects. Effects based on emotional data can also be added.
[0256] Step 12:
[0257] The server generates the final edited movie data and provides it to the user. The user can preview the finished movie through their device and check the content. The camera and microphone are also active during the preview to monitor the user's emotions.
[0258] Step 13:
[0259] The user can then provide feedback from their device to the server regarding improvements or corrections they feel are necessary through the preview. This feedback, along with emotional data, is then sent to the server as specific instructions. For example, if the user feels that the background in Scene 1 is too light, this data is sent as a correction instruction.
[0260] Step 14:
[0261] The server analyzes the received feedback and requests the AI to make any necessary corrections. Taking emotional feedback into consideration, the AI may, for example, regenerate a clearer background. At the same time, the video and audio are also adjusted in response to changes in emotions.
[0262] Step 15:
[0263] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[0264] Example 2
[0265] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0266] Conventional movie production platforms face the problem of difficulty in creating sophisticated scene compositions and character settings that reflect user emotions. Furthermore, they lack the technology to analyze user emotions in real time and reflect them in movie production. This makes it difficult to create personalized movies that directly reflect user emotions.
[0267] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0268] In this invention, the server includes a means for recognizing the emotional state of the user, a means for reflecting the recognized emotional state in the scene and character information, and a means for generating and adjusting video data and audio based on the user's emotions, thereby enabling advanced scene composition and character settings that reflect the user's emotions, and realizing the production of personalized movies.
[0269] "Means for recognizing the user's emotional state" refers to technology that uses sensors such as cameras and microphones to analyze facial expressions and voice tones based on plots and scenarios entered by the user through a terminal, thereby identifying the user's emotions.
[0270] "Means for reflecting the recognized emotional state in the scene and character information" refers to a technology for incorporating emotional characteristics into the generated scene and character information based on the user's emotional data identified by emotion analysis technology.
[0271] "Means for generating and adjusting video data and audio based on the user's emotions" refers to technology that utilizes the user's emotional data to generate and adjust video data, character audio, and background music using a video generation engine and audio synthesis system.
[0272] "Means for generating scene and character information" refers to technology for analyzing the idea and story input by the user and generating detailed information about each scene and character in the movie.
[0273] "Means for generating a cast" refers to technology for generating appropriate character appearances and voice characteristics based on the generated character information.
[0274] "Means for generating video data" refers to a technology that automatically generates specific video based on scene information and character information.
[0275] "Means for generating character voices and background music" refers to technology that automatically generates character voices and background music to match movie scenes.
[0276] "Means for integrating and editing the generated video data, audio, and music" refers to a technology for integrating and editing the various generated video data, audio, and music into a single film work.
[0277] "Means for collecting feedback" refers to technology that collects opinions and impressions from users when they preview the completed film and transmits them to the system.
[0278] "Means for modifying video data, audio, and music" refers to techniques for readjusting and modifying generated video data, audio, and music based on feedback collected from users.
[0279] "Natural language processing technology" refers to technology for analyzing text data entered by a user, understanding its meaning, and processing the information.
[0280] "Automatically" means that the system generates and edits data on its own, with little or no human intervention required.
[0281] This invention relates to a movie creation platform called "Master's Egg Platform" that combines an emotion engine that recognizes user emotions. The system aims to automatically create a movie by generating scene and character information based on the idea and story input by the user, and reflecting emotion data.
[0282] Hardware and Software Configuration
[0283] Hardware
[0284] Camera: Used to capture the user's facial expressions.
[0285] Microphone: Used to collect the user's voice tones.
[0286] Sensors: Can also be used to collect other biometric information (e.g., heart rate).
[0287] Terminal: The device (e.g., PC, tablet) through which the user enters the scenario.
[0288] Server: A high-performance computer for analyzing data and generating images.
[0289] software
[0290] Emotion recognition engine: Analyzes facial expressions and vocal tone to identify the user's emotions.
[0291] Natural language processing engine: Analyzes user-entered scenarios and extracts information.
[0292] Generative AI model: Generates video and audio based on character and scene information.
[0293] Image generation engine: Automatically generates specific images based on scene information.
[0294] Speech synthesis system: Generates character voices and background music.
[0295] Video editing software: Used to integrate and edit the generated data (e.g. Adobe Premiere Pro, Final Cut Pro).
[0296] Specific examples
[0297] 1. Story Input and Emotion Recognition
[0298] The user inputs the plot or scenario of a movie into a dedicated application on the device. At the same time, the device's built-in camera and microphone are activated to record the user's facial expressions and voice tone. The emotion recognition engine analyzes this data in real time and identifies the user's emotional state, such as "excited" or "sad." For example, the emotion engine can analyze the "user's excitement" while the user is writing the plot of a movie.
[0299] 2. Story analysis and emotional reflection
[0300] The server receives the scenario data and emotion data sent by the user and analyzes the scenario using a natural language processing engine. Based on this analysis, scene and character information is generated to reflect the identified emotion. For example, if a user excitedly enters "a mystery set in a rural town," the scenes and characters will reflect a sense of tension.
[0301] "Scene 1: The protagonist arrives in a rural town.
[0302] Scene 2: A mysterious incident occurs.
[0303] Characters: Protagonist, Detective, Villager.
[0304] 3. Cast Generation
[0305] The server then requests the generative AI model to design an appropriate cast based on the generated character information. The generative AI model then generates cast data taking into account emotional data. For example, a cast with facial expressions and voices that express the protagonist's excitement is generated.
[0306] 4. Image Generation
[0307] The server uses a video generation engine to render specific images based on the scene information. At this time, camera work and lighting are also set to reflect the emotional data. As an example of a prompt, the user can specify, "Generate a video scene that emphasizes the sense of tension."
[0308] 5. Speech and Music Generation
[0309] The server uses a speech synthesis system to generate character voices and background music. A music generation engine works to create background music that matches the emotion of the scene. An example prompt could be, "Generate fast-paced background music that matches an exciting scene."
[0310] 6. Direction and Editing
[0311] The server integrates the generated video data, audio, and music using video editing software. Each element is placed on a timeline, and transitions between scenes and sound effects are added based on the emotional data. Once editing is complete, the movie data is exported and saved as the final file.
[0312] 7. Final review and feedback
[0313] The user plays and watches the completed movie on their device. During the preview, cameras and microphones again recognize the user's emotions and collect feedback data. Even subtle elements and parts that are easily overlooked are recorded. An example prompt is, "Identify scenes where the user is nervous."
[0314] 8. Modify and Regenerate
[0315] The server analyzes the feedback data and identifies corrections based on its content. The necessary correction information is input into the generation AI, which then regenerates the film. The entire film is then re-edited and the final corrected data is provided to the user.
[0316] This system allows for more personal and sophisticated filmmaking that reflects the user's emotions. The combination of generative AI models and emotion analysis technology allows users to easily create more emotionally appealing films.
[0317] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0318] Step 1:
[0319] The user inputs the plot and story of the movie into the device. The device's built-in camera and microphone simultaneously operate to record the user's facial expressions and voice tone. Using this data as input, the device analyzes the emotional data using an emotion recognition engine. This outputs an emotional state such as "excited" or "sad." This emotional data and story data are then sent to the server.
[0320] Specific behavior:
[0321] Launch the dedicated application and enter the plot and scenario.
[0322] Real-time analysis of video and audio is performed.
[0323] Step 2:
[0324] The server receives the received story data and emotion data as input. It analyzes the story data using a natural language processing engine and generates scene and character information. It then reflects the emotion data in the scene and character information obtained from the analysis and outputs the final scene and character information.
[0325] Specific behavior:
[0326] The scenario is analyzed using a natural language processing engine.
[0327] Add emotional elements to identified scenes and characters.
[0328] Step 3:
[0329] The server passes scene and character information to a generative AI model to generate an appropriate cast. The input is scene and character information, and the output is cast data that reflects emotions.
[0330] Specific behavior:
[0331] Input the character's emotional characteristics into the generative AI model.
[0332] Receive and save the generated cast data.
[0333] Step 4:
[0334] The server inputs scene information reflecting the emotion data into the image generation engine, and generates specific image data. The output is image data.
[0335] Specific behavior:
[0336] The image generation engine renders images based on scene information.
[0337] Save the video data in storage.
[0338] Step 5:
[0339] The server uses a voice synthesis system and a music generation engine to generate character voices and background music. The input is scene information and character information, and the output is voice data and music data.
[0340] Specific behavior:
[0341] Generate character voices using a voice synthesis system.
[0342] Create background music with a music generation engine.
[0343] Save the audio and music data in a database.
[0344] Step 6:
[0345] The server integrates and edits the video data, audio, and music, and the output is integrated movie data.
[0346] Specific behavior:
[0347] Using video editing software, arrange all the elements on the timeline.
[0348] Edit the film together, adding transitions between scenes and sound effects.
[0349] Export and save the edited movie data.
[0350] Step 7:
[0351] The user reviews the completed movie on the device. The device's built-in camera and microphone record the user's reactions, which are then used as input and sent back to the server as feedback data. The output is the user's feedback data.
[0352] Specific behavior:
[0353] Watch the finished film.
[0354] Collect users' emotional responses in real time.
[0355] Send the feedback data to the server.
[0356] Step 8:
[0357] The server receives the feedback data as input, analyzes it, identifies corrections, requests the necessary corrections to the generation AI, and regenerates and re-edits the movie. The output is the corrected movie data.
[0358] Specific behavior:
[0359] Analyze feedback data to identify areas that need correction.
[0360] Request corrections from the generation AI and receive regenerated data.
[0361] The movie is re-edited based on the regenerated data, and the final corrected data is saved.
[0362] This allows for seamless movie production that reflects the user's emotions.
[0363] (Application example 2)
[0364] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0365] Conventional filmmaking platforms have difficulty incorporating user emotions into the filmmaking process, and lack a means to easily create a film that reflects the user's emotions. Furthermore, there is no adequate system for re-editing a film after it is completed, incorporating user feedback. Therefore, a system that enables more intuitive and high-quality filmmaking is needed.
[0366] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0367] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing users to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, means for recognizing users' emotions, analyzing the emotion data, and reflecting it in movie production, and means for measuring users' emotions in real time using a smart device, thereby enabling intuitive, high-quality movie production that reflects users' emotions.
[0368] An "idea" is a concept or plot that a user comes up with for film production.
[0369] "Story" refers to the plot or scenario of a movie, and is input by the user.
[0370] A "scene" is a sequence of images that unfolds consecutively in a film, showing an event at a specific place or time.
[0371] "Character information" refers to detailed data about the people and characters appearing in the film, including their personalities, appearances, roles, etc.
[0372] "Cast" refers to the actors and voice actors who play the characters in a film.
[0373] "Video data" means the digital data that constitutes the visual content of a film, including filmed footage and computer graphics.
[0374] "Audio" refers to character dialogue and other audible elements (e.g., sound effects, narration).
[0375] "Background music" is music used to enhance the emotion or atmosphere of a film scene.
[0376] "Editing" is the process of integrating and adjusting individual video data, audio, and background music.
[0377] A "preview" is when a user watches and checks the completed movie in advance.
[0378] "Feedback" refers to opinions and suggestions for improvement provided by users who have previewed the product.
[0379] "Emotion" refers to the internal state, such as excitement, sadness, or surprise, that a user experiences while making a movie.
[0380] A "smart device" is an electronic device equipped with a camera, microphone, sensor, etc. that can collect and measure emotional data.
[0381] "Natural language processing technology" is a technology for analyzing text data entered by a user and understanding its meaning.
[0382] A "generative AI model" is an artificial intelligence technology that automatically generates movie scenes, characters, and cast information based on input data.
[0383] A "prompt sentence" is an input sentence that instructs a generative AI model on the specific content to be generated.
[0384] This invention relates to a movie creation platform that combines an emotion engine that recognizes user emotions. The entire system is configured as follows.
[0385] Hardware and software used
[0386] 1. Hardware:
[0387] Smartphone (camera, microphone, touch screen)
[0388] Head-mounted display (HMD)
[0389] 2. Software:
[0390] Emotion engine (e.g. Microsoft Azure Emotion API)
[0391] Natural language processing engine (e.g. Google Cloud Natural Language API)
[0392] Generative AI models (e.g., OpenAI GPT-4)
[0393] Image generation tools (e.g. Unreal Engine, Blender)
[0394] Audio and music generation tools (e.g. Adobe Audition)
[0395] Program processing overview
[0396] 1. Story Input and Emotion Recognition
[0397] Users input the plot and scenario of a movie using their smartphone. During the input process, the smartphone's camera and microphone collect the user's facial expressions and tone of voice, which are then analyzed by the emotion engine. The analysis results are sent to the server along with the story data.
[0398] 2. Story analysis and emotional reflection
[0399] The server analyzes the received story data and emotion data to generate scene and character information. A natural language processing engine supports this, and a generative AI model creates detailed scene and character settings.
[0400] 3. Cast Generation
[0401] The server generates an appropriate cast based on the character information and emotion data, creating a character with facial and vocal characteristics that match the emotion.
[0402] 4. Image Generation
[0403] The server generates video data based on scene information and emotion data. Video generation tools are used to automatically adjust camera work, lighting settings, and other aspects.
[0404] 5. Speech and Music Generation
[0405] The server reflects the emotional data when generating the character's voice and background music. The appropriate voice tone and background music are set using a voice and music generation tool.
[0406] 6. Direction and Editing
[0407] The server integrates the generated video data, audio, and music for editing, where the dramatic effects are adjusted based on the emotional data.
[0408] 7. Final review and feedback
[0409] The user previews the completed movie, during which the user's emotions are recognized again and their reactions to the movie are recorded and sent to the server.
[0410] 8. Modify and Regenerate
[0411] The server receives the feedback and makes any necessary corrections, which are then passed on to a generative AI model that readjusts the video and audio based on the emotional feedback.
[0412] Specific examples
[0413] Example prompt sentence:
[0414] Prompt: "A mystery that takes place in a rural town."
[0415] User Sentiment: Excited
[0416] Generated scene:
[0417] The protagonist arrives in a rural town.
[0418] A mysterious incident occurs.
[0419] The protagonist works with a detective.
[0420] Generated cast:
[0421] Protagonist (with an excited look on his face)
[0422] Detective (with a serious look)
[0423] Villager (with a frightened look on his face)
[0424] Generated music:
[0425] Fast-paced background music
[0426] Tense sound effects
[0427] Produced video:
[0428] Camera work that creates a sense of tension
[0429] Dim lighting creates a mysterious atmosphere
[0430] Expected feedback:
[0431] "Scene 1 background is light" => Regenerate background
[0432] "The music is a little loud" => Readjust the music
[0433] Through this process, users can create high-quality films that reflect their emotions intuitively. Furthermore, by utilizing smart devices, emotions can be measured and reflected in real time, and continuous feedback is expected to improve the quality of the film.
[0434] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0435] Step 1:
[0436] A user inputs a movie plot or scenario using a smartphone. The input text data is sent to the emotion engine along with the user's facial expressions and voice tone, which are collected in real time by the smartphone's camera and microphone. The emotion engine analyzes the collected data and detects the user's emotional state. The detected emotion data and story data are sent to the server. The input is the plot or scenario, and the output is the detected emotion data and story data.
[0437] Step 2:
[0438] The server analyzes the received story data and emotion data. It uses a natural language processing engine to analyze the story data and generate scene and character information. The generated scene and character information is provided to a generative AI model to determine more detailed scene settings and character details. The input is story and emotion data, and the output is scene and character information.
[0439] Step 3:
[0440] The server generates an appropriate cast based on the generated character information and emotion data. It uses a generative AI model to match the character's facial and vocal characteristics to the emotion. The cast information is integrated with scene and character data. The input is character information and emotion data, and the output is cast information.
[0441] Step 4:
[0442] The server generates video data based on scene information and cast information. Using a video generation tool (e.g., Unreal Engine, Blender), camera work and lighting settings are automatically adjusted to match the emotion. The generated video data is sent to the subsequent process. The input is scene information and cast information, and the output is video data.
[0443] Step 5:
[0444] The server generates the character's voice and background music. Using a voice and music generation tool (e.g., Adobe Audition), it sets the appropriate voice tone and background music based on the emotional data. The generated voice and music data are integrated with the video data. The input is character information and emotional data, and the output is voice and music data.
[0445] Step 6:
[0446] The server integrates and edits the generated video data, audio, and music. The production effects are adjusted based on the emotional data. For example, scene transitions and sound effects are set according to emotions. The edited movie data is ready to be previewed by the user. The input is video data, audio data, and music data, and the output is the edited movie.
[0447] Step 7:
[0448] The user previews the completed movie using a smartphone or a head-mounted display. During the preview, the user's emotions are recognized again and their reactions to the movie are recorded. This reaction data is sent to the server and used as feedback. The input is the movie being previewed, and the output is the reaction data.
[0449] Step 8:
[0450] The server receives user feedback and requests the generative AI model to make any necessary modifications. Based on the emotional feedback, the video and audio are adjusted and a new version of the movie is generated. The input is the user feedback and the output is the modified movie.
[0451] Through these steps, it becomes possible to create high-quality movies that reflect the user's emotions.
[0452] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0453] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0454] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0455] [Second embodiment]
[0456] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0457] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0458] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0459] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0460] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0461] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0462] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0463] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0464] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0465] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0466] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0467] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0468] This invention relates to the "Master's Egg Platform," a platform that uses generative AI to automatically create the elements necessary for filmmaking. The system begins by analyzing the idea and story entered by the user and generating scene and character information. The system then generates a cast suited to the characters, generates and edits video data and audio, and completes the final film.
[0469] Program processing overview
[0470] 1. Story input and analysis
[0471] The user inputs the plot or scenario of their own movie into the device. The device receives the user's input and sends it to the server as story data. The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. For example, if the user inputs "a mystery that takes place in a rural town," the server will identify the following scenes and characters:
[0472] Scene 1: The protagonist arrives in a rural town.
[0473] Scene 2: A mysterious incident occurs.
[0474] Scene 3: The protagonist works with the detective.
[0475] Characters: Protagonist, Detective, Villager.
[0476] 2. Cast Selection
[0477] Based on the analyzed character information, the server uses generative AI to generate appropriate cast members, such as a young male protagonist or a middle-aged male detective, dynamically generating model data for the appearance and voice of each character.
[0478] 3. Image Generation
[0479] Next, the server generates video data based on the scene information. The scene background, character movements, camera work, and other aspects are automatically generated. For example, a "rural town" is generated as the background for Scene 1 to look natural, and the scene in which the main character gets off at the train station is depicted.
[0480] 4. Speech and Music Generation
[0481] The server requests the AI to generate lines and sound effects to generate character voices and background music. For example, the main character's line, "Is this a rural town?", can be generated in a calm voice while unsettling background music plays.
[0482] 5. Direction and Editing
[0483] The server integrates the video, audio, and music to create the overall structure of the film, including smooth transitions between scenes and applying appropriate production effects to create a complete film.
[0484] 6. Final review and feedback
[0485] The user previews the completed movie on their device and checks the content. If necessary, corrections can be sent as feedback to the server. The server receives the feedback and makes the necessary changes by requesting corrections from the generation AI. For example, if the user feels that "the background in Scene 1 is too light," the server will regenerate the background.
[0486] By automatically generating the elements necessary for film production, the present invention provides an environment in which users can easily create high-quality films. This invention allows users to focus on the creative aspects, making it easy for anyone to realize their filmmaking dreams.
[0487] The processing flow will be explained below.
[0488] Step 1:
[0489] Users input the plot and scenario of their own movie into the terminal, which then collects the ideas and story data entered by the user based on a form and formats it into story data.
[0490] Step 2:
[0491] The device sends the formatted story data to the server, which receives the story data and stores it in a database.
[0492] Step 3:
[0493] The server passes the received story data to a natural language processing engine, which then analyzes each element of the plot and extracts scene and character information.
[0494] Step 4:
[0495] The server creates a scene list and a character list from the analysis results. For example, it compiles information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[0496] Step 5:
[0497] The server passes the character list to the generation AI module and asks it to generate an appropriate cast. The generation AI dynamically generates face, body, and voice models for the characters.
[0498] Step 6:
[0499] The server receives the cast data generated by the AI and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[0500] Step 7:
[0501] The server starts generating video data based on the scene information. It requests the AI to generate the background, character movements, and camera work for each scene, and generates the video data.
[0502] Step 8:
[0503] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[0504] Step 9:
[0505] The server requests character dialogue, necessary sound effects, and background music from the AI generator, which then generates the appropriate audio data and background music for the scene.
[0506] Step 10:
[0507] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[0508] Step 11:
[0509] The server calls the editing module to integrate the video data, audio, and background music, and the editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects.
[0510] Step 12:
[0511] The server generates the final edited movie data and provides it to the user, who can then preview the completed movie through their terminal and check the content.
[0512] Step 13:
[0513] Users can provide feedback from their devices to the server about improvements and corrections they feel are necessary through the preview. The feedback is sent to the server as specific instructions for improvement.
[0514] Step 14:
[0515] The server analyzes the received feedback and requests the generation AI to make any necessary corrections, such as regenerating the background of Scene 1.
[0516] Step 15:
[0517] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[0518] Example 1
[0519] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0520] The traditional filmmaking process requires a significant amount of time and expertise, making it difficult for ordinary people to easily create films. Furthermore, the creation and integration of each element (scenes, cast, video, audio, music, editing) requires a lot of manual work, making it inefficient. A method was needed to solve these issues and enable users to easily create high-quality films.
[0521] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0522] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing a user to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, and means for operating based on information including specific names of software or hardware to be used, thereby enabling a user to easily and quickly create high-quality movies.
[0523] "User" means any person or entity that uses the Filmmaking Platform.
[0524] "Server" refers to a computer system that receives, analyzes, and processes input data from users and generates and integrates each element required for film production.
[0525] "Terminal" refers to a computing device through which a user inputs a film plot or scenario and which provides an interface with the film production platform.
[0526] "Story" refers to the plot or scenario of a movie entered by the user.
[0527] "Character information" refers to information about characters extracted through story analysis.
[0528] "Cast" refers to model data of the appearance and voice of characters generated based on character information.
[0529] "Scene information" refers to information about the background and content of a scene extracted through story analysis.
[0530] "Video data" refers to visual data generated based on scene information and cast information.
[0531] "Audio" refers to data related to a character's lines and vocalizations.
[0532] "Background music" refers to music created to enhance the atmosphere of a film scene.
[0533] "Integration" refers to the process of assembling the generated video data, audio, and background music into a single film.
[0534] "Editing" refers to the process of arranging the overall structure of a film by adding transitions between scenes and dramatic effects.
[0535] "Feedback" means suggestions for corrections or improvements provided by Users through Previews.
[0536] "Generative AI Model" refers to the artificial intelligence model used to generate video, audio, and background music.
[0537] A "prompt" refers to text that a user enters to give specific instructions to a generative AI model.
[0538] This invention relates to a platform that uses generative AI to automatically create elements necessary for user-generated filmmaking. The platform begins by analyzing the idea and story input by the user and generating scene and character information.
[0539] The user inputs the plot or scenario of his / her own movie using the terminal. For example, if the plot is "a mystery that takes place in a rural town," the terminal sends the input to the server as story data.
[0540] The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. A typical example of this software is OpenAI's GPT series. Through this process, the plot of "Mystery Set in a Country Town" is analyzed as follows:
[0541] Scene 1: The protagonist arrives in a rural town.
[0542] Scene 2: A mysterious incident occurs.
[0543] Scene 3: The protagonist works with the detective.
[0544] Characters: Protagonist, Detective, Villager.
[0545] Next, the server uses generative AI to generate an appropriate cast based on the analyzed character information, such as Stable Diffusion. In this step, model data for the appearance and voice of characters such as a young male protagonist or a middle-aged male detective is generated.
[0546] The server then generates video data based on the scene information. At this stage, image generation AI such as DALL-E 2 is used to automatically create the scene's background, character movements, and camerawork. Specifically, a "rural town" is generated as the background for Scene 1, depicting the scene in which the protagonist gets off at the train station.
[0547] The server then generates the character voices and background music, using Amazon Polly for voice generation and Jukedeck for music generation. This allows, for example, the protagonist's line, "Is this a rural town?" to be generated in a calm voice while unsettling background music plays.
[0548] The server uses editing software such as Adobe Premiere Pro to combine and edit the generated video, audio, and music, adding transitions between scenes and applying appropriate effects to create a complete film.
[0549] Finally, the user previews the completed movie on their device and checks the content. If necessary, the user can send feedback to the server. For example, if the user feels that the background in Scene 1 is too weak, the server will use generative AI to regenerate the background.
[0550] For example, the following prompts can be used:
[0551] "I want to make a mystery movie in which the protagonist works with a detective to solve a mystery in a rural town. Scene 1 is when the protagonist arrives in the town, Scene 2 is when a mysterious incident occurs, and Scene 3 is when the protagonist and detective work together. The characters are the protagonist, the detective, and the villagers."
[0552] As a result, the present invention significantly improves the efficiency of the movie production process and provides an environment in which users can easily produce high-quality movies.
[0553] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0554] Step 1:
[0555] The user inputs the basic plot or scenario of their movie into the terminal. The input story data is sent from the terminal to the server. For example, the user inputs a prompt such as "A mystery that takes place in a rural town." The input data is in text format and is sent to the server.
[0556] Step 2:
[0557] The server analyzes the story data received from the device. This analysis uses natural language processing technology. Specifically, it extracts scene and character information from the story data. For example, it uses OpenAI's natural language processing model to extract scenes such as "Scene 1: The protagonist arrives in a rural town," "Scene 2: A mysterious incident occurs," and "Scene 3: The protagonist cooperates with a detective." The output data is JSON data containing scene and character information.
[0558] Step 3:
[0559] The server uses a generative AI model to generate appropriate cast members based on the analyzed character information. Specifically, it creates model data for the characters' appearances and voices. Generative AI models such as "Stable Diffusion" are used for this step. For example, based on information that the protagonist is a young man and the detective is a middle-aged man, it generates appearance and voice models for each character. The output data is the character's 3D model data and voice data.
[0560] Step 4:
[0561] The server generates video data based on the scene and cast information obtained in the previous step. At this stage, image generation AI such as "DALL-E 2" is used to automatically create the scene's background, character movements, and camerawork. For example, it generates and outputs a scene in which the protagonist gets off at a train station in a rural town. The output data is a continuous image data of the scene, in other words, a video clip.
[0562] Step 5:
[0563] The server then generates the character's voice and background music. It uses a "voice synthesis API" for voice generation and "music generation software" for music generation. For example, it generates the protagonist's line, "Is this a rural town?", and adds background music that creates an unsettling feeling. The output data is a voice file and a music file.
[0564] Step 6:
[0565] The server then integrates and edits the generated video, audio, and music. This is done using video editing software. Specifically, it applies transitions between scenes and special effects to create a continuous video that functions as a single movie. For example, it applies a fade-in and fade-out effect from scene 1 to scene 2. The output data is a completed movie file.
[0566] Step 7:
[0567] The user previews the completed movie on the device and checks its content. The user then sends feedback from the device to the server. For example, the user may give feedback such as "The background in Scene 1 is too light." The input data is in the form of text feedback and is sent to the server.
[0568] Step 8:
[0569] The server makes the necessary corrections based on the feedback received from the user. The server then uses the generative AI model again to reflect the corrections. In this step, corrections such as "regenerating the background" are made. The output data is the corrected movie file.
[0570] This series of processes allows users to easily and quickly create high-quality movies.
[0571] (Application example 1)
[0572] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0573] In conventional film production, the creation and editing of scripts, images, and audio is done manually, requiring a great deal of time and effort. Furthermore, film production in a virtual reality environment requires specialized knowledge, making it difficult for general users to easily use. Furthermore, there has been no system that can automatically and accurately create films in a virtual reality environment, from story ideas to detailed image construction and audio integration. Therefore, there has been a demand for an environment that allows users to easily create high-quality films in a virtual reality environment.
[0574] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0575] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for integrating and editing the generated video data, audio, and music, means for a user to input a movie idea in a virtual reality environment and perform story analysis and scene generation using a generative AI model, and means for reproducing the generated scenes and characters in the virtual reality environment. This allows users to easily produce movies in a virtual reality environment, enabling the creation of high-quality movies.
[0576] A "user" is an entity that uses this system to input movie ideas and stories and create the final movie.
[0577] "Ideas and stories" are information about the plot and scenario of a movie entered by the user.
[0578] The "analysis means" is a device or program that has the function of analyzing the idea or story input by the user and generating scene and character information.
[0579] "Character information" is information about the characters in the movie generated by the analysis means.
[0580] The "cast generation means" is a device or program that has the function of generating an appropriate cast based on the character information.
[0581] "Scene information" is information about the scene and setting of the movie generated by the analysis means.
[0582] The "video generation means" is a device or program that has the function of generating video data based on the scene information and cast information.
[0583] "Voice and music generating means" refers to a device or program that has the function of generating the voice and background music of the character.
[0584] The "integrated editing means" is a device or program that has the function of integrating and editing the generated video data, audio, and music.
[0585] A "previewer" is a device or program that has the ability to allow users to preview the completed movie and gather feedback.
[0586] The "feedback correction means" is a device or program that has the function of correcting the video data, audio, and music based on the feedback.
[0587] A "virtual reality environment" is an environment in which users can create and watch movies in a virtual cinematic space using dedicated hardware.
[0588] A "generative AI model" is a model that uses artificial intelligence to analyze stories and generate scenes based on user input.
[0589] "Scene playback means" refers to a device or program that has the function of playing back the generated scenes and characters in a virtual reality environment.
[0590] The present invention is a system that allows users to create movies in a virtual reality environment. By inputting ideas and stories, the system automatically generates scenes and characters, enabling the creation and playback of movies in the virtual reality environment.
[0591] The system consists of the following main components:
[0592] 1. Input and Analysis
[0593] Users input ideas and stories using a virtual reality environment, such as a head-mounted display. This input is done through voice input or gesture input. The server receives this and analyzes it using a generative AI model. Natural language processing technology is used for analysis. For example, if a user inputs "a story about a young farmer fighting a dragon in a medieval fantasy world," the server will generate scene and character information based on this.
[0594] 2. Scene and character generation
[0595] The server generates scene and character information based on the analysis results. Based on the character information, it then generates appropriate cast members. The generative AI model dynamically generates character appearance and voice model data, as well as background information for the scene, based on user input.
[0596] 3. Video and audio generation
[0597] The server then generates video data based on the scene and character information. The scene's background, character movements, and camerawork are automatically generated. Character voices and background music are also generated using generative AI models. For example, if a character says something like, "Is this a rural town?", the server generates the voice and adds appropriate background music.
[0598] 4. Merging and Editing
[0599] The server then integrates and edits the generated video data, audio, and music, including smooth transitions between scenes and appropriate stage effects, and the integrated editing process completes the entire film.
[0600] 5. Preview and Feedback
[0601] The completed movie can be previewed by the user in a virtual reality environment. The user watches the movie and sends feedback to the server if necessary. A feedback corrector corrects the video data, audio, and music based on the feedback to improve the quality.
[0602] 6. Playback in a Virtual Reality Environment
[0603] Finally, users can play the completed movie in a virtual reality environment and enjoy the movie-making experience. For example, the following scene is generated based on the user's input: "A story about a young farmer fighting a dragon in a medieval fantasy world."
[0604] Example prompt sentence:
[0605] Generate scenes and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world.
[0606] Generated AI output example:
[0607] Scene 1: A farmer lives an ordinary life in the village.
[0608] Scene 2: The dragon attacks the village.
[0609] Scene 3: The farmer takes up his sword and confronts the dragon.
[0610] Characters: Farmer, dragon, villagers.
[0611] In this way, this invention allows users to intuitively create movies in a virtual reality environment. By utilizing generative AI models and enabling the automatic generation of scenes and characters, the effort required for movie production is significantly reduced, allowing anyone to enjoy high-quality movie production.
[0612] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0613] Step 1:
[0614] The user wears a head-mounted display in a virtual reality environment and launches a dedicated application. Story ideas and plots are input using voice or gestures. The input data is sent to the server as text data on the device. The server begins analysis based on the user's input (prompt): "A story about a young farmer fighting a dragon in a medieval fantasy world."
[0615] Step 2:
[0616] The server analyzes the received text data. A generative AI model is used for this analysis, and natural language processing techniques are used to extract scene and character information for the story. The specific prompt is: "Generate a scene and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world." Based on this, the server generates the following scene and character information. The output data is as follows:
[0617] Scene 1: A farmer lives an ordinary life in the village.
[0618] Scene 2: The dragon attacks the village.
[0619] Scene 3: The farmer takes up his sword and confronts the dragon.
[0620] Characters: Farmer, dragon, villagers.
[0621] Step 3:
[0622] The server generates an appropriate cast based on the generated scene information and character information. This cast generation includes model data for the character's appearance and voice. This process also uses a generative AI model. The output is the following cast information:
[0623] Cast 1: Young farmer character
[0624] Cast 2: Dragon character
[0625] Cast 3: Villager characters
[0626] Step 4:
[0627] The server generates video data based on scene and cast information. This video generation process includes the scene background, character movements, and camera work. The server generates a "village landscape" as the background for Scene 1 and constructs the scene by combining the character movements. The output is video data.
[0628] Step 5:
[0629] The server generates character voices and background music for the generated scenes. Specifically, it uses a generative AI model to generate character lines as audio data and inserts appropriate music into the background. For example, in Scene 1, the "farmer" might say, "Is this a village?" and idyllic rural music would play in the background. The output is audio data and music data.
[0630] Step 6:
[0631] The server then integrates and edits the generated video, audio, and music data. This process involves smooth transitions between scenes and applying appropriate special effects. The integrated editing process assembles the data into a single film. The output is the integrated film data.
[0632] Step 7:
[0633] The user previews the integrated movie in the virtual reality environment. After the preview, the user provides feedback as needed. The feedback is sent back to the server, and the server modifies the video data, audio data, and music data based on the feedback. For example, if the user feels that "the background of Scene 1 is too light," the server regenerates the background. The output is the modified movie data.
[0634] Step 8:
[0635] The modified movie data is finally played in the virtual reality environment, allowing users to enjoy their own high-quality movie in a virtual reality environment. The user's final experience is the output.
[0636] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0637] This invention relates to the "Master's Egg Platform," a movie creation platform that combines an emotion engine that recognizes the user's emotions. This system is based on analyzing the ideas and stories entered by the user and generating scene and character information based on them. Furthermore, it aims to reflect the user's emotions in the process of generating a cast based on the character information, generating and editing video data and audio, and completing the final movie.
[0638] Program processing overview
[0639] 1. Story Input and Emotion Recognition
[0640] When a user inputs a movie plot or scenario into the device, the emotion engine simultaneously collects facial expressions and voice tones via cameras and sensors. The device processes this data and analyzes the user's emotional state (e.g., excited, sad, surprised, etc.). The results are then sent to the server along with the story data.
[0641] 2. Story analysis and emotional reflection
[0642] The server analyzes the received story data and emotion data to generate scene and character information. This analysis uses natural language processing and emotion analysis technologies. For example, if a user excitedly talks about a "mystery that takes place in a rural town," the server will consider settings that will give the scene and characters a sense of urgency:
[0643] Scene 1: The protagonist arrives in a rural town.
[0644] Scene 2: A mysterious incident occurs.
[0645] Scene 3: The protagonist works with the detective.
[0646] Characters: Protagonist, Detective, Villager.
[0647] 3. Cast Generation
[0648] The server then passes the character information, reflecting the emotional data, to the AI generator to generate an appropriate cast. For example, the main character is generated with facial and vocal characteristics that express excitement.
[0649] 4. Image Generation
[0650] Next, the server generates video data that reflects the emotional data based on the scene information, for example by using camera work and lighting that create a sense of tension.
[0651] 5. Speech and Music Generation
[0652] The server generates character voices and background music based on the user's emotional data. For example, exciting scenes will have fast-paced background music.
[0653] 6. Direction and Editing
[0654] The server integrates video data, audio, and music to edit the entire movie. The emotional data is used to adjust the production effects. Specifically, scene transitions and sound effects are set according to the emotions.
[0655] 7. Final review and feedback
[0656] The user can preview the completed movie on their device. During the preview, the user's emotions are recognized again and their reactions to the movie's content are recorded. Based on this, improvements and corrections are sent to the server as feedback.
[0657] 8. Modify and Regenerate
[0658] The server receives the feedback and requests the AI to make any necessary corrections. For example, if the user feels that the background in Scene 1 is too light, the AI will regenerate that background. The AI will also readjust the video and audio based on the emotional feedback.
[0659] This invention allows the user's emotions to be reflected throughout the entire system, enabling more intuitive and emotional movie production. Users can easily express their own emotions, and as a result, high-quality movies can be easily produced.
[0660] The processing flow will be explained below.
[0661] Step 1:
[0662] The user logs in to the device's input interface and inputs the plot and scenario of their movie. The camera and microphone collect the user's facial expressions and voice tone in real time, and the emotion engine analyzes the user's emotional state. For example, if the user is speaking excitedly, that emotional data (excitement state) is obtained.
[0663] Step 2:
[0664] The device sends the collected story data and emotion data to the server. For example, the plot of a "mystery that takes place in a rural town" entered by the user is sent along with the emotional excitement data at the time.
[0665] Step 3:
[0666] The server then passes the received story data to the natural language processing engine and begins analyzing the story. At the same time, it also incorporates the results of the emotion data analysis by the emotion engine. For example, the server takes into account settings that reflect the user's excitement level when analyzing the plot and extracting scene and character information.
[0667] Step 4:
[0668] The server creates a scene list and a character list based on the analysis results. For example, it creates a list containing information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[0669] Step 5:
[0670] The server passes the character list to a generation AI module, which generates an appropriate cast. For example, it generates a face model for the main character with bright eyes and a powerful voice profile, reflecting the user's excitement level.
[0671] Step 6:
[0672] The server receives the generated cast data and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[0673] Step 7:
[0674] The server begins generating video data based on the scene information. It requests the AI to create the background, character movements, and camerawork for each scene, and generates the video data. Emotional data is also reflected, and camerawork and lighting that create a sense of tension are set, for example.
[0675] Step 8:
[0676] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[0677] Step 9:
[0678] The server requests character dialogue, necessary sound effects, and background music from the AI generator. The AI then generates audio data and background music appropriate for the scene. The user's emotions are reflected in the sound, so for example, an exciting scene will have fast-paced background music.
[0679] Step 10:
[0680] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[0681] Step 11:
[0682] The server calls the editing module to integrate the video data, audio, and music. The editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects. Effects based on emotional data can also be added.
[0683] Step 12:
[0684] The server generates the final edited movie data and provides it to the user. The user can preview the finished movie through their device and check the content. The camera and microphone are also active during the preview to monitor the user's emotions.
[0685] Step 13:
[0686] The user can then provide feedback from their device to the server regarding improvements or corrections they feel are necessary through the preview. This feedback, along with emotional data, is then sent to the server as specific instructions. For example, if the user feels that the background in Scene 1 is too light, this data is sent as a correction instruction.
[0687] Step 14:
[0688] The server analyzes the received feedback and requests the AI to make any necessary corrections. Taking emotional feedback into consideration, the AI may, for example, regenerate a clearer background. At the same time, the video and audio are also adjusted in response to changes in emotions.
[0689] Step 15:
[0690] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[0691] Example 2
[0692] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0693] Conventional movie production platforms face the problem of difficulty in creating sophisticated scene compositions and character settings that reflect user emotions. Furthermore, they lack the technology to analyze user emotions in real time and reflect them in movie production. This makes it difficult to create personalized movies that directly reflect user emotions.
[0694] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0695] In this invention, the server includes a means for recognizing the emotional state of the user, a means for reflecting the recognized emotional state in the scene and character information, and a means for generating and adjusting video data and audio based on the user's emotions, thereby enabling advanced scene composition and character settings that reflect the user's emotions, and realizing the production of personalized movies.
[0696] "Means for recognizing the user's emotional state" refers to technology that uses sensors such as cameras and microphones to analyze facial expressions and voice tones based on plots and scenarios entered by the user through a terminal, thereby identifying the user's emotions.
[0697] "Means for reflecting the recognized emotional state in the scene and character information" refers to a technology for incorporating emotional characteristics into the generated scene and character information based on the user's emotional data identified by emotion analysis technology.
[0698] "Means for generating and adjusting video data and audio based on the user's emotions" refers to technology that utilizes the user's emotional data to generate and adjust video data, character audio, and background music using a video generation engine and audio synthesis system.
[0699] "Means for generating scene and character information" refers to technology for analyzing the idea and story input by the user and generating detailed information about each scene and character in the movie.
[0700] "Means for generating a cast" refers to technology for generating appropriate character appearances and voice characteristics based on the generated character information.
[0701] "Means for generating video data" refers to a technology that automatically generates specific video based on scene information and character information.
[0702] "Means for generating character voices and background music" refers to technology that automatically generates character voices and background music to match movie scenes.
[0703] "Means for integrating and editing the generated video data, audio, and music" refers to a technology for integrating and editing the various generated video data, audio, and music into a single film work.
[0704] "Means for collecting feedback" refers to technology that collects opinions and impressions from users when they preview the completed film and transmits them to the system.
[0705] "Means for modifying video data, audio, and music" refers to techniques for readjusting and modifying generated video data, audio, and music based on feedback collected from users.
[0706] "Natural language processing technology" refers to technology for analyzing text data entered by a user, understanding its meaning, and processing the information.
[0707] "Automatically" means that the system generates and edits data on its own, with little or no human intervention required.
[0708] This invention relates to a movie creation platform called "Master's Egg Platform" that combines an emotion engine that recognizes user emotions. The system aims to automatically create a movie by generating scene and character information based on the idea and story input by the user, and reflecting emotion data.
[0709] Hardware and Software Configuration
[0710] Hardware
[0711] Camera: Used to capture the user's facial expressions.
[0712] Microphone: Used to collect the user's voice tones.
[0713] Sensors: Can also be used to collect other biometric information (e.g., heart rate).
[0714] Terminal: The device (e.g., PC, tablet) through which the user enters the scenario.
[0715] Server: A high-performance computer for analyzing data and generating images.
[0716] software
[0717] Emotion recognition engine: Analyzes facial expressions and vocal tone to identify the user's emotions.
[0718] Natural language processing engine: Analyzes user-entered scenarios and extracts information.
[0719] Generative AI model: Generates video and audio based on character and scene information.
[0720] Image generation engine: Automatically generates specific images based on scene information.
[0721] Speech synthesis system: Generates character voices and background music.
[0722] Video editing software: Used to integrate and edit the generated data (e.g. Adobe Premiere Pro, Final Cut Pro).
[0723] Specific examples
[0724] 1. Story Input and Emotion Recognition
[0725] The user inputs the plot or scenario of a movie into a dedicated application on the device. At the same time, the device's built-in camera and microphone are activated to record the user's facial expressions and voice tone. The emotion recognition engine analyzes this data in real time and identifies the user's emotional state, such as "excited" or "sad." For example, the emotion engine can analyze the "user's excitement" while the user is writing the plot of a movie.
[0726] 2. Story analysis and emotional reflection
[0727] The server receives the scenario data and emotion data sent by the user and analyzes the scenario using a natural language processing engine. Based on this analysis, scene and character information is generated to reflect the identified emotion. For example, if a user excitedly enters "a mystery set in a rural town," the scenes and characters will reflect a sense of tension.
[0728] "Scene 1: The protagonist arrives in a rural town.
[0729] Scene 2: A mysterious incident occurs.
[0730] Characters: Protagonist, Detective, Villager.
[0731] 3. Cast Generation
[0732] The server then requests the generative AI model to design an appropriate cast based on the generated character information. The generative AI model then generates cast data taking into account emotional data. For example, a cast with facial expressions and voices that express the protagonist's excitement is generated.
[0733] 4. Image Generation
[0734] The server uses a video generation engine to render specific images based on the scene information. At this time, camera work and lighting are also set to reflect the emotional data. As an example of a prompt, the user can specify, "Generate a video scene that emphasizes the sense of tension."
[0735] 5. Speech and Music Generation
[0736] The server uses a speech synthesis system to generate character voices and background music. A music generation engine works to create background music that matches the emotion of the scene. An example prompt could be, "Generate fast-paced background music that matches an exciting scene."
[0737] 6. Direction and Editing
[0738] The server integrates the generated video data, audio, and music using video editing software. Each element is placed on a timeline, and transitions between scenes and sound effects are added based on the emotional data. Once editing is complete, the movie data is exported and saved as the final file.
[0739] 7. Final review and feedback
[0740] The user plays and watches the completed movie on their device. During the preview, cameras and microphones again recognize the user's emotions and collect feedback data. Even subtle elements and parts that are easily overlooked are recorded. An example prompt is, "Identify scenes where the user is nervous."
[0741] 8. Modify and Regenerate
[0742] The server analyzes the feedback data and identifies corrections based on its content. The necessary correction information is input into the generation AI, which then regenerates the film. The entire film is then re-edited and the final corrected data is provided to the user.
[0743] This system allows for more personal and sophisticated filmmaking that reflects the user's emotions. The combination of generative AI models and emotion analysis technology allows users to easily create more emotionally appealing films.
[0744] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0745] Step 1:
[0746] The user inputs the plot and story of the movie into the device. The device's built-in camera and microphone simultaneously operate to record the user's facial expressions and voice tone. Using this data as input, the device analyzes the emotional data using an emotion recognition engine. This outputs an emotional state such as "excited" or "sad." This emotional data and story data are then sent to the server.
[0747] Specific behavior:
[0748] Launch the dedicated application and enter the plot and scenario.
[0749] Real-time analysis of video and audio is performed.
[0750] Step 2:
[0751] The server receives the received story data and emotion data as input. It analyzes the story data using a natural language processing engine and generates scene and character information. It then reflects the emotion data in the scene and character information obtained from the analysis and outputs the final scene and character information.
[0752] Specific behavior:
[0753] The scenario is analyzed using a natural language processing engine.
[0754] Add emotional elements to identified scenes and characters.
[0755] Step 3:
[0756] The server passes scene and character information to a generative AI model to generate an appropriate cast. The input is scene and character information, and the output is cast data that reflects emotions.
[0757] Specific behavior:
[0758] Input the character's emotional characteristics into the generative AI model.
[0759] Receive and save the generated cast data.
[0760] Step 4:
[0761] The server inputs scene information reflecting the emotion data into the image generation engine, and generates specific image data. The output is image data.
[0762] Specific behavior:
[0763] The image generation engine renders images based on scene information.
[0764] Save the video data in storage.
[0765] Step 5:
[0766] The server uses a voice synthesis system and a music generation engine to generate character voices and background music. The input is scene information and character information, and the output is voice data and music data.
[0767] Specific behavior:
[0768] Generate character voices using a voice synthesis system.
[0769] Create background music with a music generation engine.
[0770] Save the audio and music data in a database.
[0771] Step 6:
[0772] The server integrates and edits the video data, audio, and music, and the output is integrated movie data.
[0773] Specific behavior:
[0774] Using video editing software, arrange all the elements on the timeline.
[0775] Edit the film together, adding transitions between scenes and sound effects.
[0776] Export and save the edited movie data.
[0777] Step 7:
[0778] The user reviews the completed movie on the device. The device's built-in camera and microphone record the user's reactions, which are then used as input and sent back to the server as feedback data. The output is the user's feedback data.
[0779] Specific behavior:
[0780] Watch the finished film.
[0781] Collect users' emotional responses in real time.
[0782] Send the feedback data to the server.
[0783] Step 8:
[0784] The server receives the feedback data as input, analyzes it, identifies corrections, requests the necessary corrections to the generation AI, and regenerates and re-edits the movie. The output is the corrected movie data.
[0785] Specific behavior:
[0786] Analyze feedback data to identify areas that need correction.
[0787] Request corrections from the generation AI and receive regenerated data.
[0788] The movie is re-edited based on the regenerated data, and the final corrected data is saved.
[0789] This allows for seamless movie production that reflects the user's emotions.
[0790] (Application example 2)
[0791] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0792] Conventional filmmaking platforms have difficulty incorporating user emotions into the filmmaking process, and lack a means to easily create a film that reflects the user's emotions. Furthermore, there is no adequate system for re-editing a film after it is completed, incorporating user feedback. Therefore, a system that enables more intuitive and high-quality filmmaking is needed.
[0793] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0794] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing users to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, means for recognizing users' emotions, analyzing the emotion data, and reflecting it in movie production, and means for measuring users' emotions in real time using a smart device, thereby enabling intuitive, high-quality movie production that reflects users' emotions.
[0795] An "idea" is a concept or plot that a user comes up with for film production.
[0796] "Story" refers to the plot or scenario of a movie, and is input by the user.
[0797] A "scene" is a sequence of images that unfolds consecutively in a film, showing an event at a specific place or time.
[0798] "Character information" refers to detailed data about the people and characters appearing in the film, including their personalities, appearances, roles, etc.
[0799] "Cast" refers to the actors and voice actors who play the characters in a film.
[0800] "Video data" means the digital data that constitutes the visual content of a film, including filmed footage and computer graphics.
[0801] "Audio" refers to character dialogue and other audible elements (e.g., sound effects, narration).
[0802] "Background music" is music used to enhance the emotion or atmosphere of a film scene.
[0803] "Editing" is the process of integrating and adjusting individual video data, audio, and background music.
[0804] A "preview" is when a user watches and checks the completed movie in advance.
[0805] "Feedback" refers to opinions and suggestions for improvement provided by users who have previewed the product.
[0806] "Emotion" refers to the internal state, such as excitement, sadness, or surprise, that a user experiences while making a movie.
[0807] A "smart device" is an electronic device equipped with a camera, microphone, sensor, etc. that can collect and measure emotional data.
[0808] "Natural language processing technology" is a technology for analyzing text data entered by a user and understanding its meaning.
[0809] A "generative AI model" is an artificial intelligence technology that automatically generates movie scenes, characters, and cast information based on input data.
[0810] A "prompt sentence" is an input sentence that instructs a generative AI model on the specific content to be generated.
[0811] This invention relates to a movie creation platform that combines an emotion engine that recognizes user emotions. The entire system is configured as follows.
[0812] Hardware and software used
[0813] 1. Hardware:
[0814] Smartphone (camera, microphone, touch screen)
[0815] Head-mounted display (HMD)
[0816] 2. Software:
[0817] Emotion engine (e.g. Microsoft Azure Emotion API)
[0818] Natural language processing engine (e.g. Google Cloud Natural Language API)
[0819] Generative AI models (e.g., OpenAI GPT-4)
[0820] Image generation tools (e.g. Unreal Engine, Blender)
[0821] Audio and music generation tools (e.g. Adobe Audition)
[0822] Program processing overview
[0823] 1. Story Input and Emotion Recognition
[0824] Users input the plot and scenario of a movie using their smartphone. During the input process, the smartphone's camera and microphone collect the user's facial expressions and tone of voice, which are then analyzed by the emotion engine. The analysis results are sent to the server along with the story data.
[0825] 2. Story analysis and emotional reflection
[0826] The server analyzes the received story data and emotion data to generate scene and character information. A natural language processing engine supports this, and a generative AI model creates detailed scene and character settings.
[0827] 3. Cast Generation
[0828] The server generates an appropriate cast based on the character information and emotion data, creating a character with facial and vocal characteristics that match the emotion.
[0829] 4. Image Generation
[0830] The server generates video data based on scene information and emotion data. Video generation tools are used to automatically adjust camera work, lighting settings, and other aspects.
[0831] 5. Speech and Music Generation
[0832] The server reflects the emotional data when generating the character's voice and background music. The appropriate voice tone and background music are set using a voice and music generation tool.
[0833] 6. Direction and Editing
[0834] The server integrates the generated video data, audio, and music for editing, where the dramatic effects are adjusted based on the emotional data.
[0835] 7. Final review and feedback
[0836] The user previews the completed movie, during which the user's emotions are recognized again and their reactions to the movie are recorded and sent to the server.
[0837] 8. Modify and Regenerate
[0838] The server receives the feedback and makes any necessary corrections, which are then passed on to a generative AI model that readjusts the video and audio based on the emotional feedback.
[0839] Specific examples
[0840] Example prompt sentence:
[0841] Prompt: "A mystery that takes place in a rural town."
[0842] User Sentiment: Excited
[0843] Generated scene:
[0844] The protagonist arrives in a rural town.
[0845] A mysterious incident occurs.
[0846] The protagonist works with a detective.
[0847] Generated cast:
[0848] Protagonist (with an excited look on his face)
[0849] Detective (with a serious look)
[0850] Villager (with a frightened look on his face)
[0851] Generated music:
[0852] Fast-paced background music
[0853] Tense sound effects
[0854] Produced video:
[0855] Camera work that creates a sense of tension
[0856] Dim lighting creates a mysterious atmosphere
[0857] Expected feedback:
[0858] "Scene 1 background is light" => Regenerate background
[0859] "The music is a little loud" => Readjust the music
[0860] Through this process, users can create high-quality films that reflect their emotions intuitively. Furthermore, by utilizing smart devices, emotions can be measured and reflected in real time, and continuous feedback is expected to improve the quality of the film.
[0861] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0862] Step 1:
[0863] A user inputs a movie plot or scenario using a smartphone. The input text data is sent to the emotion engine along with the user's facial expressions and voice tone, which are collected in real time by the smartphone's camera and microphone. The emotion engine analyzes the collected data and detects the user's emotional state. The detected emotion data and story data are sent to the server. The input is the plot or scenario, and the output is the detected emotion data and story data.
[0864] Step 2:
[0865] The server analyzes the received story data and emotion data. It uses a natural language processing engine to analyze the story data and generate scene and character information. The generated scene and character information is provided to a generative AI model to determine more detailed scene settings and character details. The input is story and emotion data, and the output is scene and character information.
[0866] Step 3:
[0867] The server generates an appropriate cast based on the generated character information and emotion data. It uses a generative AI model to match the character's facial and vocal characteristics to the emotion. The cast information is integrated with scene and character data. The input is character information and emotion data, and the output is cast information.
[0868] Step 4:
[0869] The server generates video data based on scene information and cast information. Using a video generation tool (e.g., Unreal Engine, Blender), camera work and lighting settings are automatically adjusted to match the emotion. The generated video data is sent to the subsequent process. The input is scene information and cast information, and the output is video data.
[0870] Step 5:
[0871] The server generates the character's voice and background music. Using a voice and music generation tool (e.g., Adobe Audition), it sets the appropriate voice tone and background music based on the emotional data. The generated voice and music data are integrated with the video data. The input is character information and emotional data, and the output is voice and music data.
[0872] Step 6:
[0873] The server integrates and edits the generated video data, audio, and music. The production effects are adjusted based on the emotional data. For example, scene transitions and sound effects are set according to emotions. The edited movie data is ready to be previewed by the user. The input is video data, audio data, and music data, and the output is the edited movie.
[0874] Step 7:
[0875] The user previews the completed movie using a smartphone or a head-mounted display. During the preview, the user's emotions are recognized again and their reactions to the movie are recorded. This reaction data is sent to the server and used as feedback. The input is the movie being previewed, and the output is the reaction data.
[0876] Step 8:
[0877] The server receives user feedback and requests the generative AI model to make any necessary modifications. Based on the emotional feedback, the video and audio are adjusted and a new version of the movie is generated. The input is the user feedback and the output is the modified movie.
[0878] Through these steps, it becomes possible to create high-quality movies that reflect the user's emotions.
[0879] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0880] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0881] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0882] [Third embodiment]
[0883] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0884] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0885] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0886] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0887] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0888] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0889] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0890] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0891] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0892] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0893] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0894] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0895] This invention relates to the "Master's Egg Platform," a platform that uses generative AI to automatically create the elements necessary for filmmaking. The system begins by analyzing the idea and story entered by the user and generating scene and character information. The system then generates a cast suited to the characters, generates and edits video data and audio, and completes the final film.
[0896] Program processing overview
[0897] 1. Story input and analysis
[0898] The user inputs the plot or scenario of their own movie into the device. The device receives the user's input and sends it to the server as story data. The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. For example, if the user inputs "a mystery that takes place in a rural town," the server will identify the following scenes and characters:
[0899] Scene 1: The protagonist arrives in a rural town.
[0900] Scene 2: A mysterious incident occurs.
[0901] Scene 3: The protagonist works with the detective.
[0902] Characters: Protagonist, Detective, Villager.
[0903] 2. Cast Selection
[0904] Based on the analyzed character information, the server uses generative AI to generate appropriate cast members, such as a young male protagonist or a middle-aged male detective, dynamically generating model data for the appearance and voice of each character.
[0905] 3. Image Generation
[0906] Next, the server generates video data based on the scene information. The scene background, character movements, camera work, and other aspects are automatically generated. For example, a "rural town" is generated as the background for Scene 1 to look natural, and the scene in which the main character gets off at the train station is depicted.
[0907] 4. Speech and Music Generation
[0908] The server requests the AI to generate lines and sound effects to generate character voices and background music. For example, the main character's line, "Is this a rural town?", can be generated in a calm voice while unsettling background music plays.
[0909] 5. Direction and Editing
[0910] The server integrates the video, audio, and music to create the overall structure of the film, including smooth transitions between scenes and applying appropriate production effects to create a complete film.
[0911] 6. Final review and feedback
[0912] The user previews the completed movie on their device and checks the content. If necessary, corrections can be sent as feedback to the server. The server receives the feedback and makes the necessary changes by requesting corrections from the generation AI. For example, if the user feels that "the background in Scene 1 is too light," the server will regenerate the background.
[0913] By automatically generating the elements necessary for film production, the present invention provides an environment in which users can easily create high-quality films. This invention allows users to focus on the creative aspects, making it easy for anyone to realize their filmmaking dreams.
[0914] The processing flow will be explained below.
[0915] Step 1:
[0916] Users input the plot and scenario of their own movie into the terminal, which then collects the ideas and story data entered by the user based on a form and formats it into story data.
[0917] Step 2:
[0918] The device sends the formatted story data to the server, which receives the story data and stores it in a database.
[0919] Step 3:
[0920] The server passes the received story data to a natural language processing engine, which then analyzes each element of the plot and extracts scene and character information.
[0921] Step 4:
[0922] The server creates a scene list and a character list from the analysis results. For example, it compiles information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[0923] Step 5:
[0924] The server passes the character list to the generation AI module and asks it to generate an appropriate cast. The generation AI dynamically generates face, body, and voice models for the characters.
[0925] Step 6:
[0926] The server receives the cast data generated by the AI and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[0927] Step 7:
[0928] The server starts generating video data based on the scene information. It requests the AI to generate the background, character movements, and camera work for each scene, and generates the video data.
[0929] Step 8:
[0930] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[0931] Step 9:
[0932] The server requests character dialogue, necessary sound effects, and background music from the AI generator, which then generates the appropriate audio data and background music for the scene.
[0933] Step 10:
[0934] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[0935] Step 11:
[0936] The server calls the editing module to integrate the video data, audio, and background music, and the editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects.
[0937] Step 12:
[0938] The server generates the final edited movie data and provides it to the user, who can then preview the completed movie through their terminal and check the content.
[0939] Step 13:
[0940] Users can provide feedback from their devices to the server about improvements and corrections they feel are necessary through the preview. The feedback is sent to the server as specific instructions for improvement.
[0941] Step 14:
[0942] The server analyzes the received feedback and requests the generation AI to make any necessary corrections, such as regenerating the background of Scene 1.
[0943] Step 15:
[0944] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[0945] Example 1
[0946] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0947] The traditional filmmaking process requires a significant amount of time and expertise, making it difficult for ordinary people to easily create films. Furthermore, the creation and integration of each element (scenes, cast, video, audio, music, editing) requires a lot of manual work, making it inefficient. A method was needed to solve these issues and enable users to easily create high-quality films.
[0948] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0949] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing a user to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, and means for operating based on information including specific names of software or hardware to be used, thereby enabling a user to easily and quickly create high-quality movies.
[0950] "User" means any person or entity that uses the Filmmaking Platform.
[0951] "Server" refers to a computer system that receives, analyzes, and processes input data from users and generates and integrates each element required for film production.
[0952] "Terminal" refers to a computing device through which a user inputs a film plot or scenario and which provides an interface with the film production platform.
[0953] "Story" refers to the plot or scenario of a movie entered by the user.
[0954] "Character information" refers to information about characters extracted through story analysis.
[0955] "Cast" refers to model data of the appearance and voice of characters generated based on character information.
[0956] "Scene information" refers to information about the background and content of a scene extracted through story analysis.
[0957] "Video data" refers to visual data generated based on scene information and cast information.
[0958] "Audio" refers to data related to a character's lines and vocalizations.
[0959] "Background music" refers to music created to enhance the atmosphere of a film scene.
[0960] "Integration" refers to the process of assembling the generated video data, audio, and background music into a single film.
[0961] "Editing" refers to the process of arranging the overall structure of a film by adding transitions between scenes and dramatic effects.
[0962] "Feedback" means suggestions for corrections or improvements provided by Users through Previews.
[0963] "Generative AI Model" refers to the artificial intelligence model used to generate video, audio, and background music.
[0964] A "prompt" refers to text that a user enters to give specific instructions to a generative AI model.
[0965] This invention relates to a platform that uses generative AI to automatically create elements necessary for user-generated filmmaking. The platform begins by analyzing the idea and story input by the user and generating scene and character information.
[0966] The user inputs the plot or scenario of his / her own movie using the terminal. For example, if the plot is "a mystery that takes place in a rural town," the terminal sends the input to the server as story data.
[0967] The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. A typical example of this software is OpenAI's GPT series. Through this process, the plot of "Mystery Set in a Country Town" is analyzed as follows:
[0968] Scene 1: The protagonist arrives in a rural town.
[0969] Scene 2: A mysterious incident occurs.
[0970] Scene 3: The protagonist works with the detective.
[0971] Characters: Protagonist, Detective, Villager.
[0972] Next, the server uses generative AI to generate an appropriate cast based on the analyzed character information, such as Stable Diffusion. In this step, model data for the appearance and voice of characters such as a young male protagonist or a middle-aged male detective is generated.
[0973] The server then generates video data based on the scene information. At this stage, image generation AI such as DALL-E 2 is used to automatically create the scene's background, character movements, and camerawork. Specifically, a "rural town" is generated as the background for Scene 1, depicting the scene in which the protagonist gets off at the train station.
[0974] The server then generates the character voices and background music, using Amazon Polly for voice generation and Jukedeck for music generation. This allows, for example, the protagonist's line, "Is this a rural town?" to be generated in a calm voice while unsettling background music plays.
[0975] The server uses editing software such as Adobe Premiere Pro to combine and edit the generated video, audio, and music, adding transitions between scenes and applying appropriate effects to create a complete film.
[0976] Finally, the user previews the completed movie on their device and checks the content. If necessary, the user can send feedback to the server. For example, if the user feels that the background in Scene 1 is too weak, the server will use generative AI to regenerate the background.
[0977] For example, the following prompts can be used:
[0978] "I want to make a mystery movie in which the protagonist works with a detective to solve a mystery in a rural town. Scene 1 is when the protagonist arrives in the town, Scene 2 is when a mysterious incident occurs, and Scene 3 is when the protagonist and detective work together. The characters are the protagonist, the detective, and the villagers."
[0979] As a result, the present invention significantly improves the efficiency of the movie production process and provides an environment in which users can easily produce high-quality movies.
[0980] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0981] Step 1:
[0982] The user inputs the basic plot or scenario of their movie into the terminal. The input story data is sent from the terminal to the server. For example, the user inputs a prompt such as "A mystery that takes place in a rural town." The input data is in text format and is sent to the server.
[0983] Step 2:
[0984] The server analyzes the story data received from the device. This analysis uses natural language processing technology. Specifically, it extracts scene and character information from the story data. For example, it uses OpenAI's natural language processing model to extract scenes such as "Scene 1: The protagonist arrives in a rural town," "Scene 2: A mysterious incident occurs," and "Scene 3: The protagonist cooperates with a detective." The output data is JSON data containing scene and character information.
[0985] Step 3:
[0986] The server uses a generative AI model to generate appropriate cast members based on the analyzed character information. Specifically, it creates model data for the characters' appearances and voices. Generative AI models such as "Stable Diffusion" are used for this step. For example, based on information that the protagonist is a young man and the detective is a middle-aged man, it generates appearance and voice models for each character. The output data is the character's 3D model data and voice data.
[0987] Step 4:
[0988] The server generates video data based on the scene and cast information obtained in the previous step. At this stage, image generation AI such as "DALL-E 2" is used to automatically create the scene's background, character movements, and camerawork. For example, it generates and outputs a scene in which the protagonist gets off at a train station in a rural town. The output data is a continuous image data of the scene, in other words, a video clip.
[0989] Step 5:
[0990] The server then generates the character's voice and background music. It uses a "voice synthesis API" for voice generation and "music generation software" for music generation. For example, it generates the protagonist's line, "Is this a rural town?", and adds background music that creates an unsettling feeling. The output data is a voice file and a music file.
[0991] Step 6:
[0992] The server then integrates and edits the generated video, audio, and music. This is done using video editing software. Specifically, it applies transitions between scenes and special effects to create a continuous video that functions as a single movie. For example, it applies a fade-in and fade-out effect from scene 1 to scene 2. The output data is a completed movie file.
[0993] Step 7:
[0994] The user previews the completed movie on the device and checks its content. The user then sends feedback from the device to the server. For example, the user may give feedback such as "The background in Scene 1 is too light." The input data is in the form of text feedback and is sent to the server.
[0995] Step 8:
[0996] The server makes the necessary corrections based on the feedback received from the user. The server then uses the generative AI model again to reflect the corrections. In this step, corrections such as "regenerating the background" are made. The output data is the corrected movie file.
[0997] This series of processes allows users to easily and quickly create high-quality movies.
[0998] (Application example 1)
[0999] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1000] In conventional film production, the creation and editing of scripts, images, and audio is done manually, requiring a great deal of time and effort. Furthermore, film production in a virtual reality environment requires specialized knowledge, making it difficult for general users to easily use. Furthermore, there has been no system that can automatically and accurately create films in a virtual reality environment, from story ideas to detailed image construction and audio integration. Therefore, there has been a demand for an environment that allows users to easily create high-quality films in a virtual reality environment.
[1001] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1002] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for integrating and editing the generated video data, audio, and music, means for a user to input a movie idea in a virtual reality environment and perform story analysis and scene generation using a generative AI model, and means for reproducing the generated scenes and characters in the virtual reality environment. This allows users to easily produce movies in a virtual reality environment, enabling the creation of high-quality movies.
[1003] A "user" is an entity that uses this system to input movie ideas and stories and create the final movie.
[1004] "Ideas and stories" are information about the plot and scenario of a movie entered by the user.
[1005] The "analysis means" is a device or program that has the function of analyzing the idea or story input by the user and generating scene and character information.
[1006] "Character information" is information about the characters in the movie generated by the analysis means.
[1007] The "cast generation means" is a device or program that has the function of generating an appropriate cast based on the character information.
[1008] "Scene information" is information about the scene and setting of the movie generated by the analysis means.
[1009] The "video generation means" is a device or program that has the function of generating video data based on the scene information and cast information.
[1010] "Voice and music generating means" refers to a device or program that has the function of generating the voice and background music of the character.
[1011] The "integrated editing means" is a device or program that has the function of integrating and editing the generated video data, audio, and music.
[1012] A "previewer" is a device or program that has the ability to allow users to preview the completed movie and gather feedback.
[1013] The "feedback correction means" is a device or program that has the function of correcting the video data, audio, and music based on the feedback.
[1014] A "virtual reality environment" is an environment in which users can create and watch movies in a virtual cinematic space using dedicated hardware.
[1015] A "generative AI model" is a model that uses artificial intelligence to analyze stories and generate scenes based on user input.
[1016] "Scene playback means" refers to a device or program that has the function of playing back the generated scenes and characters in a virtual reality environment.
[1017] The present invention is a system that allows users to create movies in a virtual reality environment. By inputting ideas and stories, the system automatically generates scenes and characters, enabling the creation and playback of movies in the virtual reality environment.
[1018] The system consists of the following main components:
[1019] 1. Input and Analysis
[1020] Users input ideas and stories using a virtual reality environment, such as a head-mounted display. This input is done through voice input or gesture input. The server receives this and analyzes it using a generative AI model. Natural language processing technology is used for analysis. For example, if a user inputs "a story about a young farmer fighting a dragon in a medieval fantasy world," the server will generate scene and character information based on this.
[1021] 2. Scene and character generation
[1022] The server generates scene and character information based on the analysis results. Based on the character information, it then generates appropriate cast members. The generative AI model dynamically generates character appearance and voice model data, as well as background information for the scene, based on user input.
[1023] 3. Video and audio generation
[1024] The server then generates video data based on the scene and character information. The scene's background, character movements, and camerawork are automatically generated. Character voices and background music are also generated using generative AI models. For example, if a character says something like, "Is this a rural town?", the server generates the voice and adds appropriate background music.
[1025] 4. Merging and Editing
[1026] The server then integrates and edits the generated video data, audio, and music, including smooth transitions between scenes and appropriate stage effects, and the integrated editing process completes the entire film.
[1027] 5. Preview and Feedback
[1028] The completed movie can be previewed by the user in a virtual reality environment. The user watches the movie and sends feedback to the server if necessary. A feedback corrector corrects the video data, audio, and music based on the feedback to improve the quality.
[1029] 6. Playback in a Virtual Reality Environment
[1030] Finally, users can play the completed movie in a virtual reality environment and enjoy the movie-making experience. For example, the following scene is generated based on the user's input: "A story about a young farmer fighting a dragon in a medieval fantasy world."
[1031] Example prompt sentence:
[1032] Generate scenes and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world.
[1033] Generated AI output example:
[1034] Scene 1: A farmer lives an ordinary life in the village.
[1035] Scene 2: The dragon attacks the village.
[1036] Scene 3: The farmer takes up his sword and confronts the dragon.
[1037] Characters: Farmer, dragon, villagers.
[1038] In this way, this invention allows users to intuitively create movies in a virtual reality environment. By utilizing generative AI models and enabling the automatic generation of scenes and characters, the effort required for movie production is significantly reduced, allowing anyone to enjoy high-quality movie production.
[1039] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1040] Step 1:
[1041] The user wears a head-mounted display in a virtual reality environment and launches a dedicated application. Story ideas and plots are input using voice or gestures. The input data is sent to the server as text data on the device. The server begins analysis based on the user's input (prompt): "A story about a young farmer fighting a dragon in a medieval fantasy world."
[1042] Step 2:
[1043] The server analyzes the received text data. A generative AI model is used for this analysis, and natural language processing techniques are used to extract scene and character information for the story. The specific prompt is: "Generate a scene and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world." Based on this, the server generates the following scene and character information. The output data is as follows:
[1044] Scene 1: A farmer lives an ordinary life in the village.
[1045] Scene 2: The dragon attacks the village.
[1046] Scene 3: The farmer takes up his sword and confronts the dragon.
[1047] Characters: Farmer, dragon, villagers.
[1048] Step 3:
[1049] The server generates an appropriate cast based on the generated scene information and character information. This cast generation includes model data for the character's appearance and voice. This process also uses a generative AI model. The output is the following cast information:
[1050] Cast 1: Young farmer character
[1051] Cast 2: Dragon character
[1052] Cast 3: Villager characters
[1053] Step 4:
[1054] The server generates video data based on scene and cast information. This video generation process includes the scene background, character movements, and camera work. The server generates a "village landscape" as the background for Scene 1 and constructs the scene by combining the character movements. The output is video data.
[1055] Step 5:
[1056] The server generates character voices and background music for the generated scenes. Specifically, it uses a generative AI model to generate character lines as audio data and inserts appropriate music into the background. For example, in Scene 1, the "farmer" might say, "Is this a village?" and idyllic rural music would play in the background. The output is audio data and music data.
[1057] Step 6:
[1058] The server then integrates and edits the generated video, audio, and music data. This process involves smooth transitions between scenes and applying appropriate special effects. The integrated editing process assembles the data into a single film. The output is the integrated film data.
[1059] Step 7:
[1060] The user previews the integrated movie in the virtual reality environment. After the preview, the user provides feedback as needed. The feedback is sent back to the server, and the server modifies the video data, audio data, and music data based on the feedback. For example, if the user feels that "the background of Scene 1 is too light," the server regenerates the background. The output is the modified movie data.
[1061] Step 8:
[1062] The modified movie data is finally played in the virtual reality environment, allowing users to enjoy their own high-quality movie in a virtual reality environment. The user's final experience is the output.
[1063] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1064] This invention relates to the "Master's Egg Platform," a movie creation platform that combines an emotion engine that recognizes the user's emotions. This system is based on analyzing the ideas and stories entered by the user and generating scene and character information based on them. Furthermore, it aims to reflect the user's emotions in the process of generating a cast based on the character information, generating and editing video data and audio, and completing the final movie.
[1065] Program processing overview
[1066] 1. Story Input and Emotion Recognition
[1067] When a user inputs a movie plot or scenario into the device, the emotion engine simultaneously collects facial expressions and voice tones via cameras and sensors. The device processes this data and analyzes the user's emotional state (e.g., excited, sad, surprised, etc.). The results are then sent to the server along with the story data.
[1068] 2. Story analysis and emotional reflection
[1069] The server analyzes the received story data and emotion data to generate scene and character information. This analysis uses natural language processing and emotion analysis technologies. For example, if a user excitedly talks about a "mystery that takes place in a rural town," the server will consider settings that will give the scene and characters a sense of urgency:
[1070] Scene 1: The protagonist arrives in a rural town.
[1071] Scene 2: A mysterious incident occurs.
[1072] Scene 3: The protagonist works with the detective.
[1073] Characters: Protagonist, Detective, Villager.
[1074] 3. Cast Generation
[1075] The server then passes the character information, reflecting the emotional data, to the AI generator to generate an appropriate cast. For example, the main character is generated with facial and vocal characteristics that express excitement.
[1076] 4. Image Generation
[1077] Next, the server generates video data that reflects the emotional data based on the scene information, for example by using camera work and lighting that create a sense of tension.
[1078] 5. Speech and Music Generation
[1079] The server generates character voices and background music based on the user's emotional data. For example, exciting scenes will have fast-paced background music.
[1080] 6. Direction and Editing
[1081] The server integrates video data, audio, and music to edit the entire movie. The emotional data is used to adjust the production effects. Specifically, scene transitions and sound effects are set according to the emotions.
[1082] 7. Final review and feedback
[1083] The user can preview the completed movie on their device. During the preview, the user's emotions are recognized again and their reactions to the movie's content are recorded. Based on this, improvements and corrections are sent to the server as feedback.
[1084] 8. Modify and Regenerate
[1085] The server receives the feedback and requests the AI to make any necessary corrections. For example, if the user feels that the background in Scene 1 is too light, the AI will regenerate that background. The AI will also readjust the video and audio based on the emotional feedback.
[1086] This invention allows the user's emotions to be reflected throughout the entire system, enabling more intuitive and emotional movie production. Users can easily express their own emotions, and as a result, high-quality movies can be easily produced.
[1087] The processing flow will be explained below.
[1088] Step 1:
[1089] The user logs in to the device's input interface and inputs the plot and scenario of their movie. The camera and microphone collect the user's facial expressions and voice tone in real time, and the emotion engine analyzes the user's emotional state. For example, if the user is speaking excitedly, that emotional data (excitement state) is obtained.
[1090] Step 2:
[1091] The device sends the collected story data and emotion data to the server. For example, the plot of a "mystery that takes place in a rural town" entered by the user is sent along with the emotional excitement data at the time.
[1092] Step 3:
[1093] The server then passes the received story data to the natural language processing engine and begins analyzing the story. At the same time, it also incorporates the results of the emotion data analysis by the emotion engine. For example, the server takes into account settings that reflect the user's excitement level when analyzing the plot and extracting scene and character information.
[1094] Step 4:
[1095] The server creates a scene list and a character list based on the analysis results. For example, it creates a list containing information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[1096] Step 5:
[1097] The server passes the character list to a generation AI module, which generates an appropriate cast. For example, it generates a face model for the main character with bright eyes and a powerful voice profile, reflecting the user's excitement level.
[1098] Step 6:
[1099] The server receives the generated cast data and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[1100] Step 7:
[1101] The server begins generating video data based on the scene information. It requests the AI to create the background, character movements, and camerawork for each scene, and generates the video data. Emotional data is also reflected, and camerawork and lighting that create a sense of tension are set, for example.
[1102] Step 8:
[1103] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[1104] Step 9:
[1105] The server requests character dialogue, necessary sound effects, and background music from the AI generator. The AI then generates audio data and background music appropriate for the scene. The user's emotions are reflected in the sound, so for example, an exciting scene will have fast-paced background music.
[1106] Step 10:
[1107] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[1108] Step 11:
[1109] The server calls the editing module to integrate the video data, audio, and music. The editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects. Effects based on emotional data can also be added.
[1110] Step 12:
[1111] The server generates the final edited movie data and provides it to the user. The user can preview the finished movie through their device and check the content. The camera and microphone are also active during the preview to monitor the user's emotions.
[1112] Step 13:
[1113] The user can then provide feedback from their device to the server regarding improvements or corrections they feel are necessary through the preview. This feedback, along with emotional data, is then sent to the server as specific instructions. For example, if the user feels that the background in Scene 1 is too light, this data is sent as a correction instruction.
[1114] Step 14:
[1115] The server analyzes the received feedback and requests the AI to make any necessary corrections. Taking emotional feedback into consideration, the AI may, for example, regenerate a clearer background. At the same time, the video and audio are also adjusted in response to changes in emotions.
[1116] Step 15:
[1117] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[1118] Example 2
[1119] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1120] Conventional movie production platforms face the problem of difficulty in creating sophisticated scene compositions and character settings that reflect user emotions. Furthermore, they lack the technology to analyze user emotions in real time and reflect them in movie production. This makes it difficult to create personalized movies that directly reflect user emotions.
[1121] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1122] In this invention, the server includes a means for recognizing the emotional state of the user, a means for reflecting the recognized emotional state in the scene and character information, and a means for generating and adjusting video data and audio based on the user's emotions, thereby enabling advanced scene composition and character settings that reflect the user's emotions, and realizing the production of personalized movies.
[1123] "Means for recognizing the user's emotional state" refers to technology that uses sensors such as cameras and microphones to analyze facial expressions and voice tones based on plots and scenarios entered by the user through a terminal, thereby identifying the user's emotions.
[1124] "Means for reflecting the recognized emotional state in the scene and character information" refers to a technology for incorporating emotional characteristics into the generated scene and character information based on the user's emotional data identified by emotion analysis technology.
[1125] "Means for generating and adjusting video data and audio based on the user's emotions" refers to technology that utilizes the user's emotional data to generate and adjust video data, character audio, and background music using a video generation engine and audio synthesis system.
[1126] "Means for generating scene and character information" refers to technology for analyzing the idea and story input by the user and generating detailed information about each scene and character in the movie.
[1127] "Means for generating a cast" refers to technology for generating appropriate character appearances and voice characteristics based on the generated character information.
[1128] "Means for generating video data" refers to a technology that automatically generates specific video based on scene information and character information.
[1129] "Means for generating character voices and background music" refers to technology that automatically generates character voices and background music to match movie scenes.
[1130] "Means for integrating and editing the generated video data, audio, and music" refers to a technology for integrating and editing the various generated video data, audio, and music into a single film work.
[1131] "Means for collecting feedback" refers to technology that collects opinions and impressions from users when they preview the completed film and transmits them to the system.
[1132] "Means for modifying video data, audio, and music" refers to techniques for readjusting and modifying generated video data, audio, and music based on feedback collected from users.
[1133] "Natural language processing technology" refers to technology for analyzing text data entered by a user, understanding its meaning, and processing the information.
[1134] "Automatically" means that the system generates and edits data on its own, with little or no human intervention required.
[1135] This invention relates to a movie creation platform called "Master's Egg Platform" that combines an emotion engine that recognizes user emotions. The system aims to automatically create a movie by generating scene and character information based on the idea and story input by the user, and reflecting emotion data.
[1136] Hardware and Software Configuration
[1137] Hardware
[1138] Camera: Used to capture the user's facial expressions.
[1139] Microphone: Used to collect the user's voice tones.
[1140] Sensors: Can also be used to collect other biometric information (e.g., heart rate).
[1141] Terminal: The device (e.g., PC, tablet) through which the user enters the scenario.
[1142] Server: A high-performance computer for analyzing data and generating images.
[1143] software
[1144] Emotion recognition engine: Analyzes facial expressions and vocal tone to identify the user's emotions.
[1145] Natural language processing engine: Analyzes user-entered scenarios and extracts information.
[1146] Generative AI model: Generates video and audio based on character and scene information.
[1147] Image generation engine: Automatically generates specific images based on scene information.
[1148] Speech synthesis system: Generates character voices and background music.
[1149] Video editing software: Used to integrate and edit the generated data (e.g. Adobe Premiere Pro, Final Cut Pro).
[1150] Specific examples
[1151] 1. Story Input and Emotion Recognition
[1152] The user inputs the plot or scenario of a movie into a dedicated application on the device. At the same time, the device's built-in camera and microphone are activated to record the user's facial expressions and voice tone. The emotion recognition engine analyzes this data in real time and identifies the user's emotional state, such as "excited" or "sad." For example, the emotion engine can analyze the "user's excitement" while the user is writing the plot of a movie.
[1153] 2. Story analysis and emotional reflection
[1154] The server receives the scenario data and emotion data sent by the user and analyzes the scenario using a natural language processing engine. Based on this analysis, scene and character information is generated to reflect the identified emotion. For example, if a user excitedly enters "a mystery set in a rural town," the scenes and characters will reflect a sense of tension.
[1155] "Scene 1: The protagonist arrives in a rural town.
[1156] Scene 2: A mysterious incident occurs.
[1157] Characters: Protagonist, Detective, Villager.
[1158] 3. Cast Generation
[1159] The server then requests the generative AI model to design an appropriate cast based on the generated character information. The generative AI model then generates cast data taking into account emotional data. For example, a cast with facial expressions and voices that express the protagonist's excitement is generated.
[1160] 4. Image Generation
[1161] The server uses a video generation engine to render specific images based on the scene information. At this time, camera work and lighting are also set to reflect the emotional data. As an example of a prompt, the user can specify, "Generate a video scene that emphasizes the sense of tension."
[1162] 5. Speech and Music Generation
[1163] The server uses a speech synthesis system to generate character voices and background music. A music generation engine works to create background music that matches the emotion of the scene. An example prompt could be, "Generate fast-paced background music that matches an exciting scene."
[1164] 6. Direction and Editing
[1165] The server integrates the generated video data, audio, and music using video editing software. Each element is placed on a timeline, and transitions between scenes and sound effects are added based on the emotional data. Once editing is complete, the movie data is exported and saved as the final file.
[1166] 7. Final review and feedback
[1167] The user plays and watches the completed movie on their device. During the preview, cameras and microphones again recognize the user's emotions and collect feedback data. Even subtle elements and parts that are easily overlooked are recorded. An example prompt is, "Identify scenes where the user is nervous."
[1168] 8. Modify and Regenerate
[1169] The server analyzes the feedback data and identifies corrections based on its content. The necessary correction information is input into the generation AI, which then regenerates the film. The entire film is then re-edited and the final corrected data is provided to the user.
[1170] This system allows for more personal and sophisticated filmmaking that reflects the user's emotions. The combination of generative AI models and emotion analysis technology allows users to easily create more emotionally appealing films.
[1171] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1172] Step 1:
[1173] The user inputs the plot and story of the movie into the device. The device's built-in camera and microphone simultaneously operate to record the user's facial expressions and voice tone. Using this data as input, the device analyzes the emotional data using an emotion recognition engine. This outputs an emotional state such as "excited" or "sad." This emotional data and story data are then sent to the server.
[1174] Specific behavior:
[1175] Launch the dedicated application and enter the plot and scenario.
[1176] Real-time analysis of video and audio is performed.
[1177] Step 2:
[1178] The server receives the received story data and emotion data as input. It analyzes the story data using a natural language processing engine and generates scene and character information. It then reflects the emotion data in the scene and character information obtained from the analysis and outputs the final scene and character information.
[1179] Specific behavior:
[1180] The scenario is analyzed using a natural language processing engine.
[1181] Add emotional elements to identified scenes and characters.
[1182] Step 3:
[1183] The server passes scene and character information to a generative AI model to generate an appropriate cast. The input is scene and character information, and the output is cast data that reflects emotions.
[1184] Specific behavior:
[1185] Input the character's emotional characteristics into the generative AI model.
[1186] Receive and save the generated cast data.
[1187] Step 4:
[1188] The server inputs scene information reflecting the emotion data into the image generation engine, and generates specific image data. The output is image data.
[1189] Specific behavior:
[1190] The image generation engine renders images based on scene information.
[1191] Save the video data in storage.
[1192] Step 5:
[1193] The server uses a voice synthesis system and a music generation engine to generate character voices and background music. The input is scene information and character information, and the output is voice data and music data.
[1194] Specific behavior:
[1195] Generate character voices using a voice synthesis system.
[1196] Create background music with a music generation engine.
[1197] Save the audio and music data in a database.
[1198] Step 6:
[1199] The server integrates and edits the video data, audio, and music, and the output is integrated movie data.
[1200] Specific behavior:
[1201] Using video editing software, arrange all the elements on the timeline.
[1202] Edit the film together, adding transitions between scenes and sound effects.
[1203] Export and save the edited movie data.
[1204] Step 7:
[1205] The user reviews the completed movie on the device. The device's built-in camera and microphone record the user's reactions, which are then used as input and sent back to the server as feedback data. The output is the user's feedback data.
[1206] Specific behavior:
[1207] Watch the finished film.
[1208] Collect users' emotional responses in real time.
[1209] Send the feedback data to the server.
[1210] Step 8:
[1211] The server receives the feedback data as input, analyzes it, identifies corrections, requests the necessary corrections to the generation AI, and regenerates and re-edits the movie. The output is the corrected movie data.
[1212] Specific behavior:
[1213] Analyze feedback data to identify areas that need correction.
[1214] Request corrections from the generation AI and receive regenerated data.
[1215] The movie is re-edited based on the regenerated data, and the final corrected data is saved.
[1216] This allows for seamless movie production that reflects the user's emotions.
[1217] (Application example 2)
[1218] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1219] Conventional filmmaking platforms have difficulty incorporating user emotions into the filmmaking process, and lack a means to easily create a film that reflects the user's emotions. Furthermore, there is no adequate system for re-editing a film after it is completed, incorporating user feedback. Therefore, a system that enables more intuitive and high-quality filmmaking is needed.
[1220] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1221] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing users to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, means for recognizing users' emotions, analyzing the emotion data, and reflecting it in movie production, and means for measuring users' emotions in real time using a smart device, thereby enabling intuitive, high-quality movie production that reflects users' emotions.
[1222] An "idea" is a concept or plot that a user comes up with for film production.
[1223] "Story" refers to the plot or scenario of a movie, and is input by the user.
[1224] A "scene" is a sequence of images that unfolds consecutively in a film, showing an event at a specific place or time.
[1225] "Character information" refers to detailed data about the people and characters appearing in the film, including their personalities, appearances, roles, etc.
[1226] "Cast" refers to the actors and voice actors who play the characters in a film.
[1227] "Video data" means the digital data that constitutes the visual content of a film, including filmed footage and computer graphics.
[1228] "Audio" refers to character dialogue and other audible elements (e.g., sound effects, narration).
[1229] "Background music" is music used to enhance the emotion or atmosphere of a film scene.
[1230] "Editing" is the process of integrating and adjusting individual video data, audio, and background music.
[1231] A "preview" is when a user watches and checks the completed movie in advance.
[1232] "Feedback" refers to opinions and suggestions for improvement provided by users who have previewed the product.
[1233] "Emotion" refers to the internal state, such as excitement, sadness, or surprise, that a user experiences while making a movie.
[1234] A "smart device" is an electronic device equipped with a camera, microphone, sensor, etc. that can collect and measure emotional data.
[1235] "Natural language processing technology" is a technology for analyzing text data entered by a user and understanding its meaning.
[1236] A "generative AI model" is an artificial intelligence technology that automatically generates movie scenes, characters, and cast information based on input data.
[1237] A "prompt sentence" is an input sentence that instructs a generative AI model on the specific content to be generated.
[1238] This invention relates to a movie creation platform that combines an emotion engine that recognizes user emotions. The entire system is configured as follows.
[1239] Hardware and software used
[1240] 1. Hardware:
[1241] Smartphone (camera, microphone, touch screen)
[1242] Head-mounted display (HMD)
[1243] 2. Software:
[1244] Emotion engine (e.g. Microsoft Azure Emotion API)
[1245] Natural language processing engine (e.g. Google Cloud Natural Language API)
[1246] Generative AI models (e.g., OpenAI GPT-4)
[1247] Image generation tools (e.g. Unreal Engine, Blender)
[1248] Audio and music generation tools (e.g. Adobe Audition)
[1249] Program processing overview
[1250] 1. Story Input and Emotion Recognition
[1251] Users input the plot and scenario of a movie using their smartphone. During the input process, the smartphone's camera and microphone collect the user's facial expressions and tone of voice, which are then analyzed by the emotion engine. The analysis results are sent to the server along with the story data.
[1252] 2. Story analysis and emotional reflection
[1253] The server analyzes the received story data and emotion data to generate scene and character information. A natural language processing engine supports this, and a generative AI model creates detailed scene and character settings.
[1254] 3. Cast Generation
[1255] The server generates an appropriate cast based on the character information and emotion data, creating a character with facial and vocal characteristics that match the emotion.
[1256] 4. Image Generation
[1257] The server generates video data based on scene information and emotion data. Video generation tools are used to automatically adjust camera work, lighting settings, and other aspects.
[1258] 5. Speech and Music Generation
[1259] The server reflects the emotional data when generating the character's voice and background music. The appropriate voice tone and background music are set using a voice and music generation tool.
[1260] 6. Direction and Editing
[1261] The server integrates the generated video data, audio, and music for editing, where the dramatic effects are adjusted based on the emotional data.
[1262] 7. Final review and feedback
[1263] The user previews the completed movie, during which the user's emotions are recognized again and their reactions to the movie are recorded and sent to the server.
[1264] 8. Modify and Regenerate
[1265] The server receives the feedback and makes any necessary corrections, which are then passed on to a generative AI model that readjusts the video and audio based on the emotional feedback.
[1266] Specific examples
[1267] Example prompt sentence:
[1268] Prompt: "A mystery that takes place in a rural town."
[1269] User Sentiment: Excited
[1270] Generated scene:
[1271] The protagonist arrives in a rural town.
[1272] A mysterious incident occurs.
[1273] The protagonist works with a detective.
[1274] Generated cast:
[1275] Protagonist (with an excited look on his face)
[1276] Detective (with a serious look)
[1277] Villager (with a frightened look on his face)
[1278] Generated music:
[1279] Fast-paced background music
[1280] Tense sound effects
[1281] Produced video:
[1282] Camera work that creates a sense of tension
[1283] Dim lighting creates a mysterious atmosphere
[1284] Expected feedback:
[1285] "Scene 1 background is light" => Regenerate background
[1286] "The music is a little loud" => Readjust the music
[1287] Through this process, users can create high-quality films that reflect their emotions intuitively. Furthermore, by utilizing smart devices, emotions can be measured and reflected in real time, and continuous feedback is expected to improve the quality of the film.
[1288] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1289] Step 1:
[1290] A user inputs a movie plot or scenario using a smartphone. The input text data is sent to the emotion engine along with the user's facial expressions and voice tone, which are collected in real time by the smartphone's camera and microphone. The emotion engine analyzes the collected data and detects the user's emotional state. The detected emotion data and story data are sent to the server. The input is the plot or scenario, and the output is the detected emotion data and story data.
[1291] Step 2:
[1292] The server analyzes the received story data and emotion data. It uses a natural language processing engine to analyze the story data and generate scene and character information. The generated scene and character information is provided to a generative AI model to determine more detailed scene settings and character details. The input is story and emotion data, and the output is scene and character information.
[1293] Step 3:
[1294] The server generates an appropriate cast based on the generated character information and emotion data. It uses a generative AI model to match the character's facial and vocal characteristics to the emotion. The cast information is integrated with scene and character data. The input is character information and emotion data, and the output is cast information.
[1295] Step 4:
[1296] The server generates video data based on scene information and cast information. Using a video generation tool (e.g., Unreal Engine, Blender), camera work and lighting settings are automatically adjusted to match the emotion. The generated video data is sent to the subsequent process. The input is scene information and cast information, and the output is video data.
[1297] Step 5:
[1298] The server generates the character's voice and background music. Using a voice and music generation tool (e.g., Adobe Audition), it sets the appropriate voice tone and background music based on the emotional data. The generated voice and music data are integrated with the video data. The input is character information and emotional data, and the output is voice and music data.
[1299] Step 6:
[1300] The server integrates and edits the generated video data, audio, and music. The production effects are adjusted based on the emotional data. For example, scene transitions and sound effects are set according to emotions. The edited movie data is ready to be previewed by the user. The input is video data, audio data, and music data, and the output is the edited movie.
[1301] Step 7:
[1302] The user previews the completed movie using a smartphone or a head-mounted display. During the preview, the user's emotions are recognized again and their reactions to the movie are recorded. This reaction data is sent to the server and used as feedback. The input is the movie being previewed, and the output is the reaction data.
[1303] Step 8:
[1304] The server receives user feedback and requests the generative AI model to make any necessary modifications. Based on the emotional feedback, the video and audio are adjusted and a new version of the movie is generated. The input is the user feedback and the output is the modified movie.
[1305] Through these steps, it becomes possible to create high-quality movies that reflect the user's emotions.
[1306] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1307] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1308] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1309] [Fourth embodiment]
[1310] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1311] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1312] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1313] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1314] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1315] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1316] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1317] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1318] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1319] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1320] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1321] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1322] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1323] This invention relates to the "Master's Egg Platform," a platform that uses generative AI to automatically create the elements necessary for filmmaking. The system begins by analyzing the idea and story entered by the user and generating scene and character information. The system then generates a cast suited to the characters, generates and edits video data and audio, and completes the final film.
[1324] Program processing overview
[1325] 1. Story input and analysis
[1326] The user inputs the plot or scenario of their own movie into the device. The device receives the user's input and sends it to the server as story data. The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. For example, if the user inputs "a mystery that takes place in a rural town," the server will identify the following scenes and characters:
[1327] Scene 1: The protagonist arrives in a rural town.
[1328] Scene 2: A mysterious incident occurs.
[1329] Scene 3: The protagonist works with the detective.
[1330] Characters: Protagonist, Detective, Villager.
[1331] 2. Cast Selection
[1332] Based on the analyzed character information, the server uses generative AI to generate appropriate cast members, such as a young male protagonist or a middle-aged male detective, dynamically generating model data for the appearance and voice of each character.
[1333] 3. Image Generation
[1334] Next, the server generates video data based on the scene information. The scene background, character movements, camera work, and other aspects are automatically generated. For example, a "rural town" is generated as the background for Scene 1 to look natural, and the scene in which the main character gets off at the train station is depicted.
[1335] 4. Speech and Music Generation
[1336] The server requests the AI to generate lines and sound effects to generate character voices and background music. For example, the main character's line, "Is this a rural town?", can be generated in a calm voice while unsettling background music plays.
[1337] 5. Direction and Editing
[1338] The server integrates the video, audio, and music to create the overall structure of the film, including smooth transitions between scenes and applying appropriate production effects to create a complete film.
[1339] 6. Final review and feedback
[1340] The user previews the completed movie on their device and checks the content. If necessary, corrections can be sent as feedback to the server. The server receives the feedback and makes the necessary changes by requesting corrections from the generation AI. For example, if the user feels that "the background in Scene 1 is too light," the server will regenerate the background.
[1341] By automatically generating the elements necessary for film production, the present invention provides an environment in which users can easily create high-quality films. This invention allows users to focus on the creative aspects, making it easy for anyone to realize their filmmaking dreams.
[1342] The processing flow will be explained below.
[1343] Step 1:
[1344] Users input the plot and scenario of their own movie into the terminal, which then collects the ideas and story data entered by the user based on a form and formats it into story data.
[1345] Step 2:
[1346] The device sends the formatted story data to the server, which receives the story data and stores it in a database.
[1347] Step 3:
[1348] The server passes the received story data to a natural language processing engine, which then analyzes each element of the plot and extracts scene and character information.
[1349] Step 4:
[1350] The server creates a scene list and a character list from the analysis results. For example, it compiles information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[1351] Step 5:
[1352] The server passes the character list to the generation AI module and asks it to generate an appropriate cast. The generation AI dynamically generates face, body, and voice models for the characters.
[1353] Step 6:
[1354] The server receives the cast data generated by the AI and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[1355] Step 7:
[1356] The server starts generating video data based on the scene information. It requests the AI to generate the background, character movements, and camera work for each scene, and generates the video data.
[1357] Step 8:
[1358] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[1359] Step 9:
[1360] The server requests character dialogue, necessary sound effects, and background music from the AI generator, which then generates the appropriate audio data and background music for the scene.
[1361] Step 10:
[1362] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[1363] Step 11:
[1364] The server calls the editing module to integrate the video data, audio, and background music, and the editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects.
[1365] Step 12:
[1366] The server generates the final edited movie data and provides it to the user, who can then preview the completed movie through their terminal and check the content.
[1367] Step 13:
[1368] Users can provide feedback from their devices to the server about improvements and corrections they feel are necessary through the preview. The feedback is sent to the server as specific instructions for improvement.
[1369] Step 14:
[1370] The server analyzes the received feedback and requests the generation AI to make any necessary corrections, such as regenerating the background of Scene 1.
[1371] Step 15:
[1372] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[1373] Example 1
[1374] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1375] The traditional filmmaking process requires a significant amount of time and expertise, making it difficult for ordinary people to easily create films. Furthermore, the creation and integration of each element (scenes, cast, video, audio, music, editing) requires a lot of manual work, making it inefficient. A method was needed to solve these issues and enable users to easily create high-quality films.
[1376] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1377] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing a user to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, and means for operating based on information including specific names of software or hardware to be used, thereby enabling a user to easily and quickly create high-quality movies.
[1378] "User" means any person or entity that uses the Filmmaking Platform.
[1379] "Server" refers to a computer system that receives, analyzes, and processes input data from users and generates and integrates each element required for film production.
[1380] "Terminal" refers to a computing device through which a user inputs a film plot or scenario and which provides an interface with the film production platform.
[1381] "Story" refers to the plot or scenario of a movie entered by the user.
[1382] "Character information" refers to information about characters extracted through story analysis.
[1383] "Cast" refers to model data of the appearance and voice of characters generated based on character information.
[1384] "Scene information" refers to information about the background and content of a scene extracted through story analysis.
[1385] "Video data" refers to visual data generated based on scene information and cast information.
[1386] "Audio" refers to data related to a character's lines and vocalizations.
[1387] "Background music" refers to music created to enhance the atmosphere of a film scene.
[1388] "Integration" refers to the process of assembling the generated video data, audio, and background music into a single film.
[1389] "Editing" refers to the process of arranging the overall structure of a film by adding transitions between scenes and dramatic effects.
[1390] "Feedback" means suggestions for corrections or improvements provided by Users through Previews.
[1391] "Generative AI Model" refers to the artificial intelligence model used to generate video, audio, and background music.
[1392] A "prompt" refers to text that a user enters to give specific instructions to a generative AI model.
[1393] This invention relates to a platform that uses generative AI to automatically create elements necessary for user-generated filmmaking. The platform begins by analyzing the idea and story input by the user and generating scene and character information.
[1394] The user inputs the plot or scenario of his / her own movie using the terminal. For example, if the plot is "a mystery that takes place in a rural town," the terminal sends the input to the server as story data.
[1395] The server analyzes the received story data and extracts scene and character information. This analysis uses natural language processing technology. A typical example of this software is OpenAI's GPT series. Through this process, the plot of "Mystery Set in a Country Town" is analyzed as follows:
[1396] Scene 1: The protagonist arrives in a rural town.
[1397] Scene 2: A mysterious incident occurs.
[1398] Scene 3: The protagonist works with the detective.
[1399] Characters: Protagonist, Detective, Villager.
[1400] Next, the server uses generative AI to generate an appropriate cast based on the analyzed character information, such as Stable Diffusion. In this step, model data for the appearance and voice of characters such as a young male protagonist or a middle-aged male detective is generated.
[1401] The server then generates video data based on the scene information. At this stage, image generation AI such as DALL-E 2 is used to automatically create the scene's background, character movements, and camerawork. Specifically, a "rural town" is generated as the background for Scene 1, depicting the scene in which the protagonist gets off at the train station.
[1402] The server then generates the character voices and background music, using Amazon Polly for voice generation and Jukedeck for music generation. This allows, for example, the protagonist's line, "Is this a rural town?" to be generated in a calm voice while unsettling background music plays.
[1403] The server uses editing software such as Adobe Premiere Pro to combine and edit the generated video, audio, and music, adding transitions between scenes and applying appropriate effects to create a complete film.
[1404] Finally, the user previews the completed movie on their device and checks the content. If necessary, the user can send feedback to the server. For example, if the user feels that the background in Scene 1 is too weak, the server will use generative AI to regenerate the background.
[1405] For example, the following prompts can be used:
[1406] "I want to make a mystery movie in which the protagonist works with a detective to solve a mystery in a rural town. Scene 1 is when the protagonist arrives in the town, Scene 2 is when a mysterious incident occurs, and Scene 3 is when the protagonist and detective work together. The characters are the protagonist, the detective, and the villagers."
[1407] As a result, the present invention significantly improves the efficiency of the movie production process and provides an environment in which users can easily produce high-quality movies.
[1408] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1409] Step 1:
[1410] The user inputs the basic plot or scenario of their movie into the terminal. The input story data is sent from the terminal to the server. For example, the user inputs a prompt such as "A mystery that takes place in a rural town." The input data is in text format and is sent to the server.
[1411] Step 2:
[1412] The server analyzes the story data received from the device. This analysis uses natural language processing technology. Specifically, it extracts scene and character information from the story data. For example, it uses OpenAI's natural language processing model to extract scenes such as "Scene 1: The protagonist arrives in a rural town," "Scene 2: A mysterious incident occurs," and "Scene 3: The protagonist cooperates with a detective." The output data is JSON data containing scene and character information.
[1413] Step 3:
[1414] The server uses a generative AI model to generate appropriate cast members based on the analyzed character information. Specifically, it creates model data for the characters' appearances and voices. Generative AI models such as "Stable Diffusion" are used for this step. For example, based on information that the protagonist is a young man and the detective is a middle-aged man, it generates appearance and voice models for each character. The output data is the character's 3D model data and voice data.
[1415] Step 4:
[1416] The server generates video data based on the scene and cast information obtained in the previous step. At this stage, image generation AI such as "DALL-E 2" is used to automatically create the scene's background, character movements, and camerawork. For example, it generates and outputs a scene in which the protagonist gets off at a train station in a rural town. The output data is a continuous image data of the scene, in other words, a video clip.
[1417] Step 5:
[1418] The server then generates the character's voice and background music. It uses a "voice synthesis API" for voice generation and "music generation software" for music generation. For example, it generates the protagonist's line, "Is this a rural town?", and adds background music that creates an unsettling feeling. The output data is a voice file and a music file.
[1419] Step 6:
[1420] The server then integrates and edits the generated video, audio, and music. This is done using video editing software. Specifically, it applies transitions between scenes and special effects to create a continuous video that functions as a single movie. For example, it applies a fade-in and fade-out effect from scene 1 to scene 2. The output data is a completed movie file.
[1421] Step 7:
[1422] The user previews the completed movie on the device and checks its content. The user then sends feedback from the device to the server. For example, the user may give feedback such as "The background in Scene 1 is too light." The input data is in the form of text feedback and is sent to the server.
[1423] Step 8:
[1424] The server makes the necessary corrections based on the feedback received from the user. The server then uses the generative AI model again to reflect the corrections. In this step, corrections such as "regenerating the background" are made. The output data is the corrected movie file.
[1425] This series of processes allows users to easily and quickly create high-quality movies.
[1426] (Application example 1)
[1427] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1428] In conventional film production, the creation and editing of scripts, images, and audio is done manually, requiring a great deal of time and effort. Furthermore, film production in a virtual reality environment requires specialized knowledge, making it difficult for general users to easily use. Furthermore, there has been no system that can automatically and accurately create films in a virtual reality environment, from story ideas to detailed image construction and audio integration. Therefore, there has been a demand for an environment that allows users to easily create high-quality films in a virtual reality environment.
[1429] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1430] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for integrating and editing the generated video data, audio, and music, means for a user to input a movie idea in a virtual reality environment and perform story analysis and scene generation using a generative AI model, and means for reproducing the generated scenes and characters in the virtual reality environment. This allows users to easily produce movies in a virtual reality environment, enabling the creation of high-quality movies.
[1431] A "user" is an entity that uses this system to input movie ideas and stories and create the final movie.
[1432] "Ideas and stories" are information about the plot and scenario of a movie entered by the user.
[1433] The "analysis means" is a device or program that has the function of analyzing the idea or story input by the user and generating scene and character information.
[1434] "Character information" is information about the characters in the movie generated by the analysis means.
[1435] The "cast generation means" is a device or program that has the function of generating an appropriate cast based on the character information.
[1436] "Scene information" is information about the scene and setting of the movie generated by the analysis means.
[1437] The "video generation means" is a device or program that has the function of generating video data based on the scene information and cast information.
[1438] "Voice and music generating means" refers to a device or program that has the function of generating the voice and background music of the character.
[1439] The "integrated editing means" is a device or program that has the function of integrating and editing the generated video data, audio, and music.
[1440] A "previewer" is a device or program that has the ability to allow users to preview the completed movie and gather feedback.
[1441] The "feedback correction means" is a device or program that has the function of correcting the video data, audio, and music based on the feedback.
[1442] A "virtual reality environment" is an environment in which users can create and watch movies in a virtual cinematic space using dedicated hardware.
[1443] A "generative AI model" is a model that uses artificial intelligence to analyze stories and generate scenes based on user input.
[1444] "Scene playback means" refers to a device or program that has the function of playing back the generated scenes and characters in a virtual reality environment.
[1445] The present invention is a system that allows users to create movies in a virtual reality environment. By inputting ideas and stories, the system automatically generates scenes and characters, enabling the creation and playback of movies in the virtual reality environment.
[1446] The system consists of the following main components:
[1447] 1. Input and Analysis
[1448] Users input ideas and stories using a virtual reality environment, such as a head-mounted display. This input is done through voice input or gesture input. The server receives this and analyzes it using a generative AI model. Natural language processing technology is used for analysis. For example, if a user inputs "a story about a young farmer fighting a dragon in a medieval fantasy world," the server will generate scene and character information based on this.
[1449] 2. Scene and character generation
[1450] The server generates scene and character information based on the analysis results. Based on the character information, it then generates appropriate cast members. The generative AI model dynamically generates character appearance and voice model data, as well as background information for the scene, based on user input.
[1451] 3. Video and audio generation
[1452] The server then generates video data based on the scene and character information. The scene's background, character movements, and camerawork are automatically generated. Character voices and background music are also generated using generative AI models. For example, if a character says something like, "Is this a rural town?", the server generates the voice and adds appropriate background music.
[1453] 4. Merging and Editing
[1454] The server then integrates and edits the generated video data, audio, and music, including smooth transitions between scenes and appropriate stage effects, and the integrated editing process completes the entire film.
[1455] 5. Preview and Feedback
[1456] The completed movie can be previewed by the user in a virtual reality environment. The user watches the movie and sends feedback to the server if necessary. A feedback corrector corrects the video data, audio, and music based on the feedback to improve the quality.
[1457] 6. Playback in a Virtual Reality Environment
[1458] Finally, users can play the completed movie in a virtual reality environment and enjoy the movie-making experience. For example, the following scene is generated based on the user's input: "A story about a young farmer fighting a dragon in a medieval fantasy world."
[1459] Example prompt sentence:
[1460] Generate scenes and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world.
[1461] Generated AI output example:
[1462] Scene 1: A farmer lives an ordinary life in the village.
[1463] Scene 2: The dragon attacks the village.
[1464] Scene 3: The farmer takes up his sword and confronts the dragon.
[1465] Characters: Farmer, dragon, villagers.
[1466] In this way, this invention allows users to intuitively create movies in a virtual reality environment. By utilizing generative AI models and enabling the automatic generation of scenes and characters, the effort required for movie production is significantly reduced, allowing anyone to enjoy high-quality movie production.
[1467] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1468] Step 1:
[1469] The user wears a head-mounted display in a virtual reality environment and launches a dedicated application. Story ideas and plots are input using voice or gestures. The input data is sent to the server as text data on the device. The server begins analysis based on the user's input (prompt): "A story about a young farmer fighting a dragon in a medieval fantasy world."
[1470] Step 2:
[1471] The server analyzes the received text data. A generative AI model is used for this analysis, and natural language processing techniques are used to extract scene and character information for the story. The specific prompt is: "Generate a scene and characters based on a story about a young farmer fighting a dragon in a medieval fantasy world." Based on this, the server generates the following scene and character information. The output data is as follows:
[1472] Scene 1: A farmer lives an ordinary life in the village.
[1473] Scene 2: The dragon attacks the village.
[1474] Scene 3: The farmer takes up his sword and confronts the dragon.
[1475] Characters: Farmer, dragon, villagers.
[1476] Step 3:
[1477] The server generates an appropriate cast based on the generated scene information and character information. This cast generation includes model data for the character's appearance and voice. This process also uses a generative AI model. The output is the following cast information:
[1478] Cast 1: Young farmer character
[1479] Cast 2: Dragon character
[1480] Cast 3: Villager characters
[1481] Step 4:
[1482] The server generates video data based on scene and cast information. This video generation process includes the scene background, character movements, and camera work. The server generates a "village landscape" as the background for Scene 1 and constructs the scene by combining the character movements. The output is video data.
[1483] Step 5:
[1484] The server generates character voices and background music for the generated scenes. Specifically, it uses a generative AI model to generate character lines as audio data and inserts appropriate music into the background. For example, in Scene 1, the "farmer" might say, "Is this a village?" and idyllic rural music would play in the background. The output is audio data and music data.
[1485] Step 6:
[1486] The server then integrates and edits the generated video, audio, and music data. This process involves smooth transitions between scenes and applying appropriate special effects. The integrated editing process assembles the data into a single film. The output is the integrated film data.
[1487] Step 7:
[1488] The user previews the integrated movie in the virtual reality environment. After the preview, the user provides feedback as needed. The feedback is sent back to the server, and the server modifies the video data, audio data, and music data based on the feedback. For example, if the user feels that "the background of Scene 1 is too light," the server regenerates the background. The output is the modified movie data.
[1489] Step 8:
[1490] The modified movie data is finally played in the virtual reality environment, allowing users to enjoy their own high-quality movie in a virtual reality environment. The user's final experience is the output.
[1491] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1492] This invention relates to the "Master's Egg Platform," a movie creation platform that combines an emotion engine that recognizes the user's emotions. This system is based on analyzing the ideas and stories entered by the user and generating scene and character information based on them. Furthermore, it aims to reflect the user's emotions in the process of generating a cast based on the character information, generating and editing video data and audio, and completing the final movie.
[1493] Program processing overview
[1494] 1. Story Input and Emotion Recognition
[1495] When a user inputs a movie plot or scenario into the device, the emotion engine simultaneously collects facial expressions and voice tones via cameras and sensors. The device processes this data and analyzes the user's emotional state (e.g., excited, sad, surprised, etc.). The results are then sent to the server along with the story data.
[1496] 2. Story analysis and emotional reflection
[1497] The server analyzes the received story data and emotion data to generate scene and character information. This analysis uses natural language processing and emotion analysis technologies. For example, if a user excitedly talks about a "mystery that takes place in a rural town," the server will consider settings that will give the scene and characters a sense of urgency:
[1498] Scene 1: The protagonist arrives in a rural town.
[1499] Scene 2: A mysterious incident occurs.
[1500] Scene 3: The protagonist works with the detective.
[1501] Characters: Protagonist, Detective, Villager.
[1502] 3. Cast Generation
[1503] The server then passes the character information, reflecting the emotional data, to the AI generator to generate an appropriate cast. For example, the main character is generated with facial and vocal characteristics that express excitement.
[1504] 4. Image Generation
[1505] Next, the server generates video data that reflects the emotional data based on the scene information, for example by using camera work and lighting that create a sense of tension.
[1506] 5. Speech and Music Generation
[1507] The server generates character voices and background music based on the user's emotional data. For example, exciting scenes will have fast-paced background music.
[1508] 6. Direction and Editing
[1509] The server integrates video data, audio, and music to edit the entire movie. The emotional data is used to adjust the production effects. Specifically, scene transitions and sound effects are set according to the emotions.
[1510] 7. Final review and feedback
[1511] The user can preview the completed movie on their device. During the preview, the user's emotions are recognized again and their reactions to the movie's content are recorded. Based on this, improvements and corrections are sent to the server as feedback.
[1512] 8. Modify and Regenerate
[1513] The server receives the feedback and requests the AI to make any necessary corrections. For example, if the user feels that the background in Scene 1 is too light, the AI will regenerate that background. The AI will also readjust the video and audio based on the emotional feedback.
[1514] This invention allows the user's emotions to be reflected throughout the entire system, enabling more intuitive and emotional movie production. Users can easily express their own emotions, and as a result, high-quality movies can be easily produced.
[1515] The processing flow will be explained below.
[1516] Step 1:
[1517] The user logs in to the device's input interface and inputs the plot and scenario of their movie. The camera and microphone collect the user's facial expressions and voice tone in real time, and the emotion engine analyzes the user's emotional state. For example, if the user is speaking excitedly, that emotional data (excitement state) is obtained.
[1518] Step 2:
[1519] The device sends the collected story data and emotion data to the server. For example, the plot of a "mystery that takes place in a rural town" entered by the user is sent along with the emotional excitement data at the time.
[1520] Step 3:
[1521] The server then passes the received story data to the natural language processing engine and begins analyzing the story. At the same time, it also incorporates the results of the emotion data analysis by the emotion engine. For example, the server takes into account settings that reflect the user's excitement level when analyzing the plot and extracting scene and character information.
[1522] Step 4:
[1523] The server creates a scene list and a character list based on the analysis results. For example, it creates a list containing information such as Scene 1 "The protagonist arrives in a rural town" and Characters "The protagonist, the detective, and the villagers."
[1524] Step 5:
[1525] The server passes the character list to a generation AI module, which generates an appropriate cast. For example, it generates a face model for the main character with bright eyes and a powerful voice profile, reflecting the user's excitement level.
[1526] Step 6:
[1527] The server receives the generated cast data and adds it to the character list. For example, the main character's face model and voice profile are linked to the character.
[1528] Step 7:
[1529] The server begins generating video data based on the scene information. It requests the AI to create the background, character movements, and camerawork for each scene, and generates the video data. Emotional data is also reflected, and camerawork and lighting that create a sense of tension are set, for example.
[1530] Step 8:
[1531] The server adds the generated video data to a scene list and associates the video data corresponding to each scene. For example, scene 1 includes background video and character movements.
[1532] Step 9:
[1533] The server requests character dialogue, necessary sound effects, and background music from the AI generator. The AI then generates audio data and background music appropriate for the scene. The user's emotions are reflected in the sound, so for example, an exciting scene will have fast-paced background music.
[1534] Step 10:
[1535] The server adds the generated audio data and background music to the scene list and associates the audio and music for each scene. For example, Scene 1 contains the dialogue and background music.
[1536] Step 11:
[1537] The server calls the editing module to integrate the video data, audio, and music. The editing module edits the entire movie, applying smooth transitions between scenes and appropriate effects. Effects based on emotional data can also be added.
[1538] Step 12:
[1539] The server generates the final edited movie data and provides it to the user. The user can preview the finished movie through their device and check the content. The camera and microphone are also active during the preview to monitor the user's emotions.
[1540] Step 13:
[1541] The user can then provide feedback from their device to the server regarding improvements or corrections they feel are necessary through the preview. This feedback, along with emotional data, is then sent to the server as specific instructions. For example, if the user feels that the background in Scene 1 is too light, this data is sent as a correction instruction.
[1542] Step 14:
[1543] The server analyzes the received feedback and requests the AI to make any necessary corrections. Taking emotional feedback into consideration, the AI may, for example, regenerate a clearer background. At the same time, the video and audio are also adjusted in response to changes in emotions.
[1544] Step 15:
[1545] The server updates the corrected data and regenerates the movie data for final confirmation. The user previews it again and confirms the final content.
[1546] Example 2
[1547] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1548] Conventional movie production platforms face the problem of difficulty in creating sophisticated scene compositions and character settings that reflect user emotions. Furthermore, they lack the technology to analyze user emotions in real time and reflect them in movie production. This makes it difficult to create personalized movies that directly reflect user emotions.
[1549] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1550] In this invention, the server includes a means for recognizing the emotional state of the user, a means for reflecting the recognized emotional state in the scene and character information, and a means for generating and adjusting video data and audio based on the user's emotions, thereby enabling advanced scene composition and character settings that reflect the user's emotions, and realizing the production of personalized movies.
[1551] "Means for recognizing the user's emotional state" refers to technology that uses sensors such as cameras and microphones to analyze facial expressions and voice tones based on plots and scenarios entered by the user through a terminal, thereby identifying the user's emotions.
[1552] "Means for reflecting the recognized emotional state in the scene and character information" refers to a technology for incorporating emotional characteristics into the generated scene and character information based on the user's emotional data identified by emotion analysis technology.
[1553] "Means for generating and adjusting video data and audio based on the user's emotions" refers to technology that utilizes the user's emotional data to generate and adjust video data, character audio, and background music using a video generation engine and audio synthesis system.
[1554] "Means for generating scene and character information" refers to technology for analyzing the idea and story input by the user and generating detailed information about each scene and character in the movie.
[1555] "Means for generating a cast" refers to technology for generating appropriate character appearances and voice characteristics based on the generated character information.
[1556] "Means for generating video data" refers to a technology that automatically generates specific video based on scene information and character information.
[1557] "Means for generating character voices and background music" refers to technology that automatically generates character voices and background music to match movie scenes.
[1558] "Means for integrating and editing the generated video data, audio, and music" refers to a technology for integrating and editing the various generated video data, audio, and music into a single film work.
[1559] "Means for collecting feedback" refers to technology that collects opinions and impressions from users when they preview the completed film and transmits them to the system.
[1560] "Means for modifying video data, audio, and music" refers to techniques for readjusting and modifying generated video data, audio, and music based on feedback collected from users.
[1561] "Natural language processing technology" refers to technology for analyzing text data entered by a user, understanding its meaning, and processing the information.
[1562] "Automatically" means that the system generates and edits data on its own, with little or no human intervention required.
[1563] This invention relates to a movie creation platform called "Master's Egg Platform" that combines an emotion engine that recognizes user emotions. The system aims to automatically create a movie by generating scene and character information based on the idea and story input by the user, and reflecting emotion data.
[1564] Hardware and Software Configuration
[1565] Hardware
[1566] Camera: Used to capture the user's facial expressions.
[1567] Microphone: Used to collect the user's voice tones.
[1568] Sensors: Can also be used to collect other biometric information (e.g., heart rate).
[1569] Terminal: The device (e.g., PC, tablet) through which the user enters the scenario.
[1570] Server: A high-performance computer for analyzing data and generating images.
[1571] software
[1572] Emotion recognition engine: Analyzes facial expressions and vocal tone to identify the user's emotions.
[1573] Natural language processing engine: Analyzes user-entered scenarios and extracts information.
[1574] Generative AI model: Generates video and audio based on character and scene information.
[1575] Image generation engine: Automatically generates specific images based on scene information.
[1576] Speech synthesis system: Generates character voices and background music.
[1577] Video editing software: Used to integrate and edit the generated data (e.g. Adobe Premiere Pro, Final Cut Pro).
[1578] Specific examples
[1579] 1. Story Input and Emotion Recognition
[1580] The user inputs the plot or scenario of a movie into a dedicated application on the device. At the same time, the device's built-in camera and microphone are activated to record the user's facial expressions and voice tone. The emotion recognition engine analyzes this data in real time and identifies the user's emotional state, such as "excited" or "sad." For example, the emotion engine can analyze the "user's excitement" while the user is writing the plot of a movie.
[1581] 2. Story analysis and emotional reflection
[1582] The server receives the scenario data and emotion data sent by the user and analyzes the scenario using a natural language processing engine. Based on this analysis, scene and character information is generated to reflect the identified emotion. For example, if a user excitedly enters "a mystery set in a rural town," the scenes and characters will reflect a sense of tension.
[1583] "Scene 1: The protagonist arrives in a rural town.
[1584] Scene 2: A mysterious incident occurs.
[1585] Characters: Protagonist, Detective, Villager.
[1586] 3. Cast Generation
[1587] The server then requests the generative AI model to design an appropriate cast based on the generated character information. The generative AI model then generates cast data taking into account emotional data. For example, a cast with facial expressions and voices that express the protagonist's excitement is generated.
[1588] 4. Image Generation
[1589] The server uses a video generation engine to render specific images based on the scene information. At this time, camera work and lighting are also set to reflect the emotional data. As an example of a prompt, the user can specify, "Generate a video scene that emphasizes the sense of tension."
[1590] 5. Speech and Music Generation
[1591] The server uses a speech synthesis system to generate character voices and background music. A music generation engine works to create background music that matches the emotion of the scene. An example prompt could be, "Generate fast-paced background music that matches an exciting scene."
[1592] 6. Direction and Editing
[1593] The server integrates the generated video data, audio, and music using video editing software. Each element is placed on a timeline, and transitions between scenes and sound effects are added based on the emotional data. Once editing is complete, the movie data is exported and saved as the final file.
[1594] 7. Final review and feedback
[1595] The user plays and watches the completed movie on their device. During the preview, cameras and microphones again recognize the user's emotions and collect feedback data. Even subtle elements and parts that are easily overlooked are recorded. An example prompt is, "Identify scenes where the user is nervous."
[1596] 8. Modify and Regenerate
[1597] The server analyzes the feedback data and identifies corrections based on its content. The necessary correction information is input into the generation AI, which then regenerates the film. The entire film is then re-edited and the final corrected data is provided to the user.
[1598] This system allows for more personal and sophisticated filmmaking that reflects the user's emotions. The combination of generative AI models and emotion analysis technology allows users to easily create more emotionally appealing films.
[1599] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1600] Step 1:
[1601] The user inputs the plot and story of the movie into the device. The device's built-in camera and microphone simultaneously operate to record the user's facial expressions and voice tone. Using this data as input, the device analyzes the emotional data using an emotion recognition engine. This outputs an emotional state such as "excited" or "sad." This emotional data and story data are then sent to the server.
[1602] Specific behavior:
[1603] Launch the dedicated application and enter the plot and scenario.
[1604] Real-time analysis of video and audio is performed.
[1605] Step 2:
[1606] The server receives the received story data and emotion data as input. It analyzes the story data using a natural language processing engine and generates scene and character information. It then reflects the emotion data in the scene and character information obtained from the analysis and outputs the final scene and character information.
[1607] Specific behavior:
[1608] The scenario is analyzed using a natural language processing engine.
[1609] Add emotional elements to identified scenes and characters.
[1610] Step 3:
[1611] The server passes scene and character information to a generative AI model to generate an appropriate cast. The input is scene and character information, and the output is cast data that reflects emotions.
[1612] Specific behavior:
[1613] Input the character's emotional characteristics into the generative AI model.
[1614] Receive and save the generated cast data.
[1615] Step 4:
[1616] The server inputs scene information reflecting the emotion data into the image generation engine, and generates specific image data. The output is image data.
[1617] Specific behavior:
[1618] The image generation engine renders images based on scene information.
[1619] Save the video data in storage.
[1620] Step 5:
[1621] The server uses a voice synthesis system and a music generation engine to generate character voices and background music. The input is scene information and character information, and the output is voice data and music data.
[1622] Specific behavior:
[1623] Generate character voices using a voice synthesis system.
[1624] Create background music with a music generation engine.
[1625] Save the audio and music data in a database.
[1626] Step 6:
[1627] The server integrates and edits the video data, audio, and music, and the output is integrated movie data.
[1628] Specific behavior:
[1629] Using video editing software, arrange all the elements on the timeline.
[1630] Edit the film together, adding transitions between scenes and sound effects.
[1631] Export and save the edited movie data.
[1632] Step 7:
[1633] The user reviews the completed movie on the device. The device's built-in camera and microphone record the user's reactions, which are then used as input and sent back to the server as feedback data. The output is the user's feedback data.
[1634] Specific behavior:
[1635] Watch the finished film.
[1636] Collect users' emotional responses in real time.
[1637] Send the feedback data to the server.
[1638] Step 8:
[1639] The server receives the feedback data as input, analyzes it, identifies corrections, requests the necessary corrections to the generation AI, and regenerates and re-edits the movie. The output is the corrected movie data.
[1640] Specific behavior:
[1641] Analyze feedback data to identify areas that need correction.
[1642] Request corrections from the generation AI and receive regenerated data.
[1643] The movie is re-edited based on the regenerated data, and the final corrected data is saved.
[1644] This allows for seamless movie production that reflects the user's emotions.
[1645] (Application example 2)
[1646] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1647] Conventional filmmaking platforms have difficulty incorporating user emotions into the filmmaking process, and lack a means to easily create a film that reflects the user's emotions. Furthermore, there is no adequate system for re-editing a film after it is completed, incorporating user feedback. Therefore, a system that enables more intuitive and high-quality filmmaking is needed.
[1648] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1649] In this invention, the server includes means for analyzing ideas and stories input by a user and generating scene and character information, means for generating an appropriate cast based on the character information, means for generating video data based on the scene information and cast information, means for generating voices and background music for the characters, means for integrating and editing the generated video data, voices, and music, means for allowing users to preview the completed movie and collect feedback, means for modifying the video data, voices, and music based on the feedback, means for recognizing users' emotions, analyzing the emotion data, and reflecting it in movie production, and means for measuring users' emotions in real time using a smart device, thereby enabling intuitive, high-quality movie production that reflects users' emotions.
[1650] An "idea" is a concept or plot that a user comes up with for film production.
[1651] "Story" refers to the plot or scenario of a movie, and is input by the user.
[1652] A "scene" is a sequence of images that unfolds consecutively in a film, showing an event at a specific place or time.
[1653] "Character information" refers to detailed data about the people and characters appearing in the film, including their personalities, appearances, roles, etc.
[1654] "Cast" refers to the actors and voice actors who play the characters in a film.
[1655] "Video data" means the digital data that constitutes the visual content of a film, including filmed footage and computer graphics.
[1656] "Audio" refers to character dialogue and other audible elements (e.g., sound effects, narration).
[1657] "Background music" is music used to enhance the emotion or atmosphere of a film scene.
[1658] "Editing" is the process of integrating and adjusting individual video data, audio, and background music.
[1659] A "preview" is when a user watches and checks the completed movie in advance.
[1660] "Feedback" refers to opinions and suggestions for improvement provided by users who have previewed the product.
[1661] "Emotion" refers to the internal state, such as excitement, sadness, or surprise, that a user experiences while making a movie.
[1662] A "smart device" is an electronic device equipped with a camera, microphone, sensor, etc. that can collect and measure emotional data.
[1663] "Natural language processing technology" is a technology for analyzing text data entered by a user and understanding its meaning.
[1664] A "generative AI model" is an artificial intelligence technology that automatically generates movie scenes, characters, and cast information based on input data.
[1665] A "prompt sentence" is an input sentence that instructs a generative AI model on the specific content to be generated.
[1666] This invention relates to a movie creation platform that combines an emotion engine that recognizes user emotions. The entire system is configured as follows.
[1667] Hardware and software used
[1668] 1. Hardware:
[1669] Smartphone (camera, microphone, touch screen)
[1670] Head-mounted display (HMD)
[1671] 2. Software:
[1672] Emotion engine (e.g. Microsoft Azure Emotion API)
[1673] Natural language processing engine (e.g. Google Cloud Natural Language API)
[1674] Generative AI models (e.g., OpenAI GPT-4)
[1675] Image generation tools (e.g. Unreal Engine, Blender)
[1676] Audio and music generation tools (e.g. Adobe Audition)
[1677] Program processing overview
[1678] 1. Story Input and Emotion Recognition
[1679] Users input the plot and scenario of a movie using their smartphone. During the input process, the smartphone's camera and microphone collect the user's facial expressions and tone of voice, which are then analyzed by the emotion engine. The analysis results are sent to the server along with the story data.
[1680] 2. Story analysis and emotional reflection
[1681] The server analyzes the received story data and emotion data to generate scene and character information. A natural language processing engine supports this, and a generative AI model creates detailed scene and character settings.
[1682] 3. Cast Generation
[1683] The server generates an appropriate cast based on the character information and emotion data, creating a character with facial and vocal characteristics that match the emotion.
[1684] 4. Image Generation
[1685] The server generates video data based on scene information and emotion data. Video generation tools are used to automatically adjust camera work, lighting settings, and other aspects.
[1686] 5. Speech and Music Generation
[1687] The server reflects the emotional data when generating the character's voice and background music. The appropriate voice tone and background music are set using a voice and music generation tool.
[1688] 6. Direction and Editing
[1689] The server integrates the generated video data, audio, and music for editing, where the dramatic effects are adjusted based on the emotional data.
[1690] 7. Final review and feedback
[1691] The user previews the completed movie, during which the user's emotions are recognized again and their reactions to the movie are recorded and sent to the server.
[1692] 8. Modify and Regenerate
[1693] The server receives the feedback and makes any necessary corrections, which are then passed on to a generative AI model that readjusts the video and audio based on the emotional feedback.
[1694] Specific examples
[1695] Example prompt sentence:
[1696] Prompt: "A mystery that takes place in a rural town."
[1697] User Sentiment: Excited
[1698] Generated scene:
[1699] The protagonist arrives in a rural town.
[1700] A mysterious incident occurs.
[1701] The protagonist works with a detective.
[1702] Generated cast:
[1703] Protagonist (with an excited look on his face)
[1704] Detective (with a serious look)
[1705] Villager (with a frightened look on his face)
[1706] Generated music:
[1707] Fast-paced background music
[1708] Tense sound effects
[1709] Produced video:
[1710] Camera work that creates a sense of tension
[1711] Dim lighting creates a mysterious atmosphere
[1712] Expected feedback:
[1713] "Scene 1 background is light" => Regenerate background
[1714] "The music is a little loud" => Readjust the music
[1715] Through this process, users can create high-quality films that reflect their emotions intuitively. Furthermore, by utilizing smart devices, emotions can be measured and reflected in real time, and continuous feedback is expected to improve the quality of the film.
[1716] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1717] Step 1:
[1718] A user inputs a movie plot or scenario using a smartphone. The input text data is sent to the emotion engine along with the user's facial expressions and voice tone, which are collected in real time by the smartphone's camera and microphone. The emotion engine analyzes the collected data and detects the user's emotional state. The detected emotion data and story data are sent to the server. The input is the plot or scenario, and the output is the detected emotion data and story data.
[1719] Step 2:
[1720] The server analyzes the received story data and emotion data. It uses a natural language processing engine to analyze the story data and generate scene and character information. The generated scene and character information is provided to a generative AI model to determine more detailed scene settings and character details. The input is story and emotion data, and the output is scene and character information.
[1721] Step 3:
[1722] The server generates an appropriate cast based on the generated character information and emotion data. It uses a generative AI model to match the character's facial and vocal characteristics to the emotion. The cast information is integrated with scene and character data. The input is character information and emotion data, and the output is cast information.
[1723] Step 4:
[1724] The server generates video data based on scene information and cast information. Using a video generation tool (e.g., Unreal Engine, Blender), camera work and lighting settings are automatically adjusted to match the emotion. The generated video data is sent to the subsequent process. The input is scene information and cast information, and the output is video data.
[1725] Step 5:
[1726] The server generates the character's voice and background music. Using a voice and music generation tool (e.g., Adobe Audition), it sets the appropriate voice tone and background music based on the emotional data. The generated voice and music data are integrated with the video data. The input is character information and emotional data, and the output is voice and music data.
[1727] Step 6:
[1728] The server integrates and edits the generated video data, audio, and music. The production effects are adjusted based on the emotional data. For example, scene transitions and sound effects are set according to emotions. The edited movie data is ready to be previewed by the user. The input is video data, audio data, and music data, and the output is the edited movie.
[1729] Step 7:
[1730] The user previews the completed movie using a smartphone or a head-mounted display. During the preview, the user's emotions are recognized again and their reactions to the movie are recorded. This reaction data is sent to the server and used as feedback. The input is the movie being previewed, and the output is the reaction data.
[1731] Step 8:
[1732] The server receives user feedback and requests the generative AI model to make any necessary modifications. Based on the emotional feedback, the video and audio are adjusted and a new version of the movie is generated. The input is the user feedback and the output is the modified movie.
[1733] Through these steps, it becomes possible to create high-quality movies that reflect the user's emotions.
[1734] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1735] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1736] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1737] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1738] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1739] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1740] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1741] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1742] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1743] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1744] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1745] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1746] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1747] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1748] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1749] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1750] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1751] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1752] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1753] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1754] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1755] The following is further disclosed regarding the above embodiment.
[1756] (Claim 1)
[1757] means for analyzing an idea or story input by a user and generating scene and character information;
[1758] means for generating an appropriate cast based on the character information;
[1759] means for generating video data based on the scene information and cast information;
[1760] means for generating the voice and background music of said character;
[1761] a means for integrating and editing the generated video data, audio, and music;
[1762] a means for allowing users to preview the completed movie and gather feedback;
[1763] means for modifying said video data, audio and music based on said feedback;
[1764] A system including:
[1765] (Claim 2)
[1766] 2. The system according to claim 1, wherein the analyzing means uses natural language processing techniques to analyze the user's ideas and stories.
[1767] (Claim 3)
[1768] 2. The system according to claim 1, wherein said generating means automatically performs integrated editing of the generated video data, audio, and music.
[1769] "Example 1"
[1770] (Claim 1)
[1771] means for analyzing an idea or story input by a user and generating scene and character information;
[1772] means for generating an appropriate cast based on the character information;
[1773] means for generating video data based on the scene information and cast information;
[1774] means for generating the voice and background music of said character;
[1775] a means for integrating and editing the generated video data, audio, and music;
[1776] a means for allowing users to preview the completed movie and gather feedback;
[1777] means for modifying said video data, audio and music based on said feedback;
[1778] means that act on information including the specific name of the software or hardware used;
[1779] A system including:
[1780] (Claim 2)
[1781] 2. The system according to claim 1, wherein the analyzing means uses natural language processing techniques to analyze the user's ideas and stories.
[1782] (Claim 3)
[1783] 2. The system according to claim 1, wherein said generating means automatically performs integrated editing of the generated video data, audio, and music.
[1784] (Claim 4)
[1785] 10. The system of claim 1, wherein the user previews the completed movie and provides feedback on corrections.
[1786] (Claim 5)
[1787] The system according to claim 1, wherein the correction means automatically reflects the correction points using a generative AI model.
[1788] "Application Example 1"
[1789] (Claim 1)
[1790] means for analyzing an idea or story input by a user and generating scene and character information;
[1791] means for generating an appropriate cast based on the character information;
[1792] means for generating video data based on the scene information and cast information;
[1793] means for generating the voice and background music of said character;
[1794] a means for integrating and editing the generated video data, audio, and music;
[1795] a means for allowing users to preview the completed movie and gather feedback;
[1796] means for modifying said video data, audio and music based on said feedback;
[1797] A means for users to input movie ideas in a virtual reality environment and use a generative AI model to analyze the story and generate scenes;
[1798] a means for playing the generated scenes and characters in a virtual reality environment; and
[1799] A system including:
[1800] (Claim 2)
[1801] 2. The system according to claim 1, wherein the analyzing means uses natural language processing techniques to analyze the user's ideas and stories.
[1802] (Claim 3)
[1803] 2. The system according to claim 1, wherein said generating means automatically performs integrated editing of the generated video data, audio, and music.
[1804] "Example 2: Combining Emotion Engines"
[1805] (Claim 1)
[1806] means for analyzing an idea or story input by a user and generating scene and character information;
[1807] means for generating an appropriate cast based on the character information;
[1808] means for generating video data based on the scene information and cast information;
[1809] means for generating the voice and background music of said character;
[1810] a means for integrating and editing the generated video data, audio, and music;
[1811] a means for allowing users to preview the completed movie and gather feedback;
[1812] means for modifying said video data, audio and music based on said feedback;
[1813] means for recognizing the emotional state of a user;
[1814] means for reflecting the recognized emotional state in the scene and character information;
[1815] means for generating and adjusting video data and audio based on a user's emotions;
[1816] A system including:
[1817] (Claim 2)
[1818] 2. The system according to claim 1, wherein the analyzing means uses natural language processing techniques to analyze the user's ideas and stories.
[1819] (Claim 3)
[1820] 2. The system according to claim 1, wherein said generating means automatically performs integrated editing of the generated video data, audio, and music.
[1821] "Application example 2 when combining emotion engines"
[1822] (Claim 1)
[1823] means for analyzing an idea or story input by a user and generating scene and character information;
[1824] means for generating an appropriate cast based on the character information;
[1825] means for generating video data based on the scene information and cast information;
[1826] means for generating the voice and background music of said character;
[1827] a means for integrating and editing the generated video data, audio, and music;
[1828] a means for allowing users to preview the completed movie and gather feedback;
[1829] means for modifying said video data, audio and music based on said feedback;
[1830] A means of recognizing the user's emotions, analyzing the emotional data, and reflecting it in film production;
[1831] A means of measuring user emotions in real time using a smart device;
[1832] A system including:
[1833] (Claim 2)
[1834] 2. The system according to claim 1, wherein the analysis means uses natural language processing technology to analyze the user's ideas and stories, and further uses the user's emotional data to reflect this in film production.
[1835] (Claim 3)
[1836] The system according to claim 1, characterized in that the generating means automatically performs integrated editing of the generated video data, audio, and music, and further re-edits the data taking into account the user's real-time emotional feedback. [Explanation of symbols]
[1837] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for analyzing an idea or story input by a user and generating scene and character information; means for generating an appropriate cast based on the character information; means for generating video data based on the scene information and cast information; means for generating the voice and background music of said character; a means for integrating and editing the generated video data, audio, and music; a means for allowing users to preview the completed movie and gather feedback; means for modifying said video data, audio and music based on said feedback; A system including:
2. 2. The system of claim 1, wherein the analyzing means uses natural language processing techniques to analyze the user's ideas and stories.
3. 2. The system according to claim 1, wherein said generating means automatically performs integrated editing of the generated video data, audio, and music.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A