System
The system addresses the misalignment between screenwriters and directors by allowing for efficient filmmaking through a generative model that reflects user feedback, ensuring the final footage aligns with the screenwriter's vision.
Patent Information
- Application Number
- JP2024119080
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
The traditional filmmaking process often fails to accurately reflect the screenwriter's intentions due to discrepancies between the screenwriter and director, and the feedback process is time-consuming and inefficient.
A system that allows users to input a movie script, convert it into movie footage using a generative model, check and provide feedback, reflect the feedback in the model to regenerate footage, and store the final footage, enabling efficient creation of movies that align with the screenwriter's vision.
Enables the efficient generation of movie footage that faithfully represents the screenwriter's intentions by incorporating user feedback and adjusting the generative process, reducing time and effort in the filmmaking process.
Smart Images

Figure 2026018019000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In the traditional filmmaking process, the screenwriter's intentions are not always reflected in the film, and there are often discrepancies between the intentions of the screenwriter and the director. This makes it difficult for the screenwriter to create a film that faithfully reflects his or her intentions. Furthermore, the current feedback process requires a lot of time and effort to incorporate corrections, so there is a need for more efficient filmmaking. [Means for solving the problem]
[0005] To solve this problem, we provide the following means: a system including a means for inputting a movie script, a means for converting the input script into movie footage using a generative model, a means for checking the generated footage and inputting feedback, a means for reflecting the input feedback in the generative model and regenerating the footage, and a means for saving and providing the final movie footage. Furthermore, by providing a means for adjusting the generation process when the generative model regenerates the footage based on the feedback, and a means for providing an interface that allows a user to input the script and feedback via a terminal, it is possible to efficiently create a movie that faithfully reflects the screenwriter's intentions.
[0006] A "script" is a document that describes the overall story, dialogue, and scene details of a work such as a movie or play.
[0007] A "generative model" is an algorithm or program that automatically generates images or videos based on input text data.
[0008] "Cinematographic footage" means a sequence presented in a visually viewable format as a cinematographic film.
[0009] "Feedback" means comments or suggestions for correction or improvement provided by a User.
[0010] "Regeneration" is the process of generating a new image based on feedback.
[0011] "Storage" refers to recording data or information in a way that allows it to be accessed at a later time.
[0012] "Providing" means supplying something in a form accessible to the user.
[0013] A "terminal" is a device such as a computer, smartphone, or tablet that is operated by a user.
[0014] "Interface" refers to the input and output means and screen that allow a user to operate a system. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0023] [First embodiment]
[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0036] The present invention relates to a system that allows a user to input a movie script and automatically generates movie images based on the script using a generative model. An embodiment of this system will be described below.
[0037] 1. Entering the script
[0038] The user uses the device to input the screenplay in text format, and the device is provided with a dedicated application or a web interface, allowing the user to write detailed descriptions and dialogue for each scene.
[0039] 2. Submit your script
[0040] The terminal sends the input script to the server, where it is converted into a standard data format such as JSON and sent to the server using an HTTP POST request.
[0041] 3. Initial video generation
[0042] The server inputs the received script data into the generative model to generate the first footage of the movie. The generative model automatically creates images and video sequences for each scene based on the text data of the script.
[0043] 4. First video review and feedback
[0044] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. This feedback is entered in text format, detailing specific corrections and improvements.
[0045] 5. Submitting Feedback
[0046] The device sends the user's feedback to the server, which is also converted to JSON format and sent to the server using an HTTP POST request.
[0047] 6. Brush up the video
[0048] The server then applies the feedback data to the generative model and regenerates the new video. In this process, the parameters of the generative model are adjusted and details are modified according to the user's requests.
[0049] 7. Storage and provision of final footage
[0050] Once the user is satisfied with the final video, it is stored on the server and provided in a format that the user can download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link.
[0051] Specific examples
[0052] For example, suppose a user (scriptwriter) inputs a scene such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. If the user provides feedback after viewing the initial footage, such as "I would like the rain to be a little stronger," the server reflects that feedback in the generative model and generates the footage again. If the user reviews the regenerated footage again and is satisfied, it is saved as the final footage and a download link is provided.
[0053] In this way, a system can be realized that enables a user (screenwriter) to efficiently generate movie images that faithfully reflect his or her intentions.
[0054] The processing flow will be explained below.
[0055] Step 1:
[0056] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[0057] Step 2:
[0058] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[0059] Step 3:
[0060] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[0061] Step 4:
[0062] The server encodes the generated initial video data and sends it to the device, where it is provided in the form of a URL that can be accessed by the user.
[0063] Step 5:
[0064] The user accesses the provided URL using a device and watches the initial video. While checking the video, the user can input any corrections or improvements they would like to see as feedback.
[0065] Step 6:
[0066] The device converts the user's feedback into JSON format and sends it to the server again using an HTTP POST request.
[0067] Step 7:
[0068] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback.
[0069] Step 8:
[0070] The server then encodes the regenerated video and sends it back to the device, and this process can be repeated as many times as necessary.
[0071] Step 9:
[0072] The user checks the regenerated image and repeats steps 5 to 8 until satisfied.
[0073] Step 10:
[0074] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[0075] Example 1
[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0077] Conventional methods require a great deal of time and effort to produce movie footage. Furthermore, the process of creating footage that accurately reflects the intent of the script is difficult to correct or adjust. This makes it difficult to efficiently create footage while maintaining quality during movie production.
[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0079] In this invention, the server includes means for a user to input a movie script in text format using a computer terminal, means for converting the input script into a standard data format and transmitting it to the server via a network, means for the server to generate an initial video of the movie based on the script data using a generative model, means for transmitting the generated initial video to the user terminal and for the user to check the video and input feedback, means for transmitting the feedback input by the user to the server and reflecting the feedback in the generative model to regenerate the video, and means for saving the final generated video and providing it in a format that the user can download. This makes it possible to efficiently visualize a movie script and flexibly modify and adjust it according to user requests.
[0080] A "user" is someone who uses the system to input a movie script, review the generated footage, and provide feedback.
[0081] A "computer terminal" is an electronic device used by a user to input a movie script and feedback.
[0082] "Text format" is a data format that handles character information and allows detailed descriptions of each scene and dialogue.
[0083] A "standard data format" is a format defined to maintain data compatibility, and the JSON format is an example of this.
[0084] A "network" is a communications infrastructure for transmitting and receiving data between computer terminals and servers.
[0085] A "server" is a central computer system that processes and stores data and provides services to users over a network.
[0086] A "generative model" is a machine learning model for automatically generating movie footage based on input script data.
[0087] The "first footage" is the first footage of a movie generated by the generative model based on the script data.
[0088] "Feedback" refers to complaints and requests for corrections that users input after viewing the initial footage, and describes in detail specific corrections and improvements.
[0089] "Regeneration" is the process by which a generative model creates new images based on user feedback.
[0090] "Storage" refers to the act of storing the final generated video on a server.
[0091] "Providing" refers to the act of making the saved final video accessible in the form of a link or the like so that the user can download it.
[0092] This invention relates to a system that automatically generates movie footage based on a movie script input by a user using a generative AI model. Specific embodiments of this system are described below.
[0093] Hardware and Software Configuration
[0094] Hardware
[0095] The system includes the following hardware components:
[0096] Terminal: A device where the user can enter the script for the film, review the resulting footage, and provide feedback. This can typically be a computer, tablet, or smartphone.
[0097] Server: A central computer system for processing and storing data, with powerful processors and large amounts of storage.
[0098] software
[0099] Typical software used in each step of the program is as follows:
[0100] Dedicated application or web interface: An interface for users to enter scripts and provide feedback.
[0101] Standard Data Format Conversion Library: A Python "json" library for converting script data and feedback data into JSON format.
[0102] HTTP communication library: The Python "requests" library used to send data from the device to the server.
[0103] Generative AI model: A model used to generate video based on script data. Examples include OpenAI's GPT-4 and Facebook's Fairseq.
[0104] Program processing flow
[0105] 1. Entering the script
[0106] The user uses a terminal to input a movie script in text format. Detailed descriptions and dialogue for each scene can be added via a dedicated application or a web interface. For example, the user might input a scene in which the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining.
[0107] 2. Submit your script
[0108] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, specifically using the Python "requests" library.
[0109] 3. Initial video generation
[0110] The server then inputs the received script data into a generative AI model, such as OpenAI's GPT-4 or Facebook's Fairseq, to generate the first footage. This automatically creates images and video sequences for each scene based on the script's text data.
[0111] 4. First video review and feedback
[0112] The generated initial video is sent from the server to the device. The user reviews the video and enters feedback about any complaints or requests for corrections. This feedback is also converted to JSON format and sent to the server using an HTTP POST request.
[0113] 5. Submitting Feedback
[0114] The device converts the user's feedback into JSON format and sends it to the server, again using an HTTP POST request.
[0115] 6. Brush up the video
[0116] The server then applies the feedback data to the generative AI model and regenerates a new image. The parameters of the generative model are adjusted, and the image is modified according to the user's request, for example, to make the rain stronger.
[0117] 7. Storage and provision of final footage
[0118] Once the user is satisfied with the final video, the server stores it and provides it in a format that the user can download. The user can obtain the video file via a download link.
[0119] By applying this invention, users can efficiently generate movie images that faithfully reflect their intentions. Furthermore, by incorporating user feedback into the images, it becomes possible to improve the quality of movie production.
[0120] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0121] Step 1:
[0122] A user uses a computer terminal to input a movie script in text format. The user then uses a dedicated application or a web interface to write detailed descriptions and dialogue for each scene. The input data is a text-format script. For example, the user might input "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." Subsequent processing begins based on this input data.
[0123] Step 2:
[0124] The terminal converts the input script into a standard data format (JSON format). The terminal uses Python's "json" library to convert text data into JSON. This conversion generates a script in a structured data format. This JSON data is sent to the server using an HTTP POST request. Specifically, the terminal sends the data to the specified URL on the server via the network.
[0125] Step 3:
[0126] The server parses the received JSON-formatted script data and inputs it into a generative AI model. The server then uses a generative AI model (such as OpenAI's GPT-4 or Facebook's Fairseq) to generate the first footage of the movie based on the script data. In this process, images and video sequences of each scene are automatically created from the text data. The output is a first footage file. The server then compiles this into a single video file.
[0127] Step 4:
[0128] The generated initial video is sent from the server to the device. The device displays this video to the user. The user reviews the video and inputs feedback about any complaints or requests for corrections. For example, a user might input feedback such as "I'd like the rain to be a little stronger." This input data is also in text format and is later converted to JSON format.
[0129] Step 5:
[0130] The device converts the feedback entered by the user into JSON format. The device again uses the Python "json" library to structure the feedback data. This feedback data is then sent to the server using an HTTP POST request. Specifically, the device again sends the feedback data over the network to the specified URL on the server.
[0131] Step 6:
[0132] The server applies the feedback data to the generative model and regenerates a new video. The parameters of the generative model are adjusted, and the video is modified according to the user's requests. This regeneration process produces a video that is an improvement over the initial video. The output is a modified video file.
[0133] Step 7:
[0134] The server saves the final video. The server stores the video file in a secure storage and generates a download link that the user can access. The download link is provided to the user, who clicks on the link to obtain the video file. Specifically, the server generates a link based on the path of the saved video file and sends it to the user's device.
[0135] Through these steps, the system can efficiently visualize a movie script and flexibly make corrections and adjustments according to the user's requests.
[0136] (Application example 1)
[0137] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0138] Conventional systems that automatically generate movie footage based on a movie script input in text format often fail to fully reflect the user's intentions. Another problem is the time and effort required to review the generated footage and provide feedback. Furthermore, they do not provide a realistic viewing environment to enhance the viewing experience.
[0139] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0140] In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating the images, means for saving and providing the final movie images, means for viewing the images using a terminal, and means for providing a realistic viewing experience using a head-mounted display. This enables users to efficiently generate images that reflect their intentions while enjoying a realistic viewing experience.
[0141] A "movie script" is a textual document that details the scenes, characters, dialogue, and action of a movie.
[0142] A "generative model" is a machine learning model for automatically generating images and video sequences based on input text data.
[0143] "Video" refers to visual data in which objects or scenes move over time, and in the case of film, includes scripted sequences.
[0144] "Feedback" refers to opinions and requests for corrections that a user provides regarding a generated video.
[0145] A "terminal" is an electronic device that a user operates, and includes a smartphone, tablet, computer, etc.
[0146] A "head-mounted display" is a display device that is worn by the user on the head, and is used to provide VR environments, etc.
[0147] An "immersive viewing experience" is a visual and auditory experience that immerses the viewer as if they are inside the movie.
[0148] This invention is a system that inputs a movie script in text format and automatically generates movie images based on it. An embodiment for realizing this system is described below.
[0149] First, the user uses the device to input the script for the movie. The device is provided with a dedicated application or web interface, and the user can write detailed descriptions and dialogue for each scene. For example, if the user wants to input a scene in which "the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the user can input it in the following text format:
[0150] "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[0151] The input script is converted into a standard data format (e.g., JSON) and sent to the server using an HTTP POST request. The server then inputs the received script data into a generative model to generate the first footage of the movie. This generative model automatically generates images and video sequences based on the text data entered by the user. For example, a machine learning library such as TensorFlow or PyTorch in Python can be used to build the generative model.
[0152] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. The feedback is again converted into a standard data format and sent to the server using an HTTP POST request. If the user wants to provide feedback such as "I'd like the rain to be a little stronger," they can enter it in the following text format:
[0153] "I want the rain to be a little stronger."
[0154] The server then applies the feedback data to the generative model and regenerates a new image. During this process, the parameters of the generative model are adjusted and details are modified according to the user's requests. The regenerated image is then sent back to the device for the user to review.
[0155] Once the user is satisfied with the final video, it is stored on the server and made available for download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link. It is also possible to provide an immersive viewing experience using a smartphone or head-mounted display. For example, head-mounted displays such as Oculus Rift and HTC Vive can be used.
[0156] This system allows users to efficiently generate video that reflects their own intentions while enjoying a realistic viewing experience.
[0157] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0158] Step 1:
[0159] The user uses the device to input a movie script. The input is in text format, including detailed descriptions and dialogue for each scene. The device receives this input and converts it into a standard data format (such as JSON).
[0160] Step 2:
[0161] The device sends the converted script data to the server using an HTTP POST request. Specifically, a JSON object containing the script data is sent, and the server receives it.
[0162] Step 3:
[0163] The server inputs the received script data into a generative model to generate the first footage of the movie. This generation process uses machine learning libraries (such as TensorFlow and PyTorch) to generate scene images and video sequences from the input text data. The generated footage is stored on the server.
[0164] Step 4:
[0165] The server sends the generated initial video as an HTTP response to the terminal, which receives the video data and plays it on the device to provide it to the user.
[0166] Step 5:
[0167] The user can then use the device to view the generated video and enter feedback. This feedback is in the form of text, including suggestions for improvement or correction, such as "I'd like the rain to be a little stronger." The device then converts this feedback back into JSON format.
[0168] Step 6:
[0169] The terminal sends the converted feedback data to the server using an HTTP POST request, and the server retrieves the received feedback data.
[0170] Step 7:
[0171] The server then applies the feedback data to the generative model and regenerates a new video. At this time, the parameters of the generative model are adjusted based on the user's feedback, and details are corrected. The regenerated video is then stored back on the server.
[0172] Step 8:
[0173] The server sends the regenerated video as an HTTP response to the device, which receives the video data and plays it on the device for presentation to the user. This process is repeated until the user is finally satisfied.
[0174] Step 9:
[0175] When the final video is generated to the user's satisfaction, the server saves the final video and generates a downloadable link, which the device provides to the user, who clicks on the link to obtain the video file.
[0176] Step 10:
[0177] Users can view the final video using a device or a head-mounted display, creating an immersive viewing experience. For example, they can enjoy movies in a 3D environment using devices such as Oculus Rift or HTC Vive.
[0178] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0179] This invention relates to a system that automatically generates movie images based on a movie script input by a user using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described below.
[0180] 1. Entering the script
[0181] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[0182] 2. Submit your script
[0183] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[0184] 3. Initial video generation
[0185] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[0186] 4. First footage review and emotion analysis
[0187] The server encodes the generated initial video data and sends it to the device. When the user watches the video, the emotion engine analyzes the user's facial expressions and voice to recognize the user's emotions.
[0188] 5. Automatic feedback generation
[0189] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "The rain intensity in scene 3 needs to be increased" is automatically generated.
[0190] 6. Submitting Feedback
[0191] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again using an HTTP POST request.
[0192] 7. Brush up the video
[0193] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[0194] 8. Checking the regenerated video and making final adjustments
[0195] The regenerated image is sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[0196] 9. Storage and provision of final footage
[0197] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[0198] Specific examples
[0199] For example, suppose a user (scriptwriter) inputs a scene such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user inputs feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it automatically generates feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[0200] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[0201] The processing flow will be explained below.
[0202] Step 1:
[0203] The user uses a device to input a movie script into a dedicated application or web interface. Specifically, the user writes detailed descriptions and dialogue for each scene in text format and saves it as script data.
[0204] Step 2:
[0205] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, at which point the script data is transferred to the server via the network.
[0206] Step 3:
[0207] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data. This process utilizes natural language processing and image generation technology.
[0208] Step 4:
[0209] The server encodes the generated initial video data and sends it to the device, which is provided with a link to the video file, and the user clicks on the link to watch the video.
[0210] Step 5:
[0211] Users access the provided link using their device and watch the initial video. While reviewing the video, they can enter feedback on corrections and improvements they would like to see. Feedback is written in a specific format, such as "Make the rain stronger in scene 3."
[0212] Step 6:
[0213] When users provide feedback, the emotion engine analyzes their facial expressions and voice to recognize their emotions. The emotion engine collects emotion data in real time while the user is watching the video.
[0214] Step 7:
[0215] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "You need to increase the intensity of the rain in scene 3" is automatically generated.
[0216] Step 8:
[0217] The device converts the user feedback and feedback automatically generated by the emotion engine into JSON format and sends it back to the server using an HTTP POST request.
[0218] Step 9:
[0219] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[0220] Step 10:
[0221] The regenerated image is then sent to the device. The user checks the regenerated image and repeats steps 5 to 9 until they are satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[0222] Step 11:
[0223] The final video that the user is satisfied with is saved by the server and a downloadable link is generated. The server then provides this link to the user's device, allowing the user to retrieve the completed video file. This process effectively generates an optimized movie video that takes the user's emotions into account.
[0224] Example 2
[0225] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0226] The traditional film production process, from scriptwriting to video generation and editing, is complex and time-consuming, often requiring large resources and advanced technical skills. It is also difficult to reflect user emotions in real time, making it impossible to optimize video content based on viewer emotions. Furthermore, there are limited ways to efficiently incorporate feedback, making the process of improving video quality cumbersome. There was a need for a system that could solve these problems and efficiently generate high-quality video.
[0227] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input a movie script using a terminal, a means for converting the input script into JSON format and sending it to the server, a means for the server to convert the script into movie footage using a generative model, a means for the user to check the generated footage and analyze the user's emotions using an emotion engine, a means for automatically generating feedback based on the emotion analysis, a means for converting user or automatically generated feedback into JSON format and sending it to the server, a means for reflecting the input feedback in the generative model and regenerating the footage, and a means for saving and providing the final movie footage. This enables optimization of the footage based on the user's emotions and an efficient feedback process.
[0228] A "user" is a user of the system who inputs a movie script, reviews the generated footage, and provides feedback.
[0229] A "terminal" is an information processing device that a user uses to input a movie script and feedback, and specifically includes a PC, smartphone, tablet, etc.
[0230] A "screenplay" is a document that describes in text form the details and dialogue of each scene of a movie.
[0231] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that represents data as key-value pairs.
[0232] An "HTTP POST request" is a type of HyperText Transfer Protocol used to send data to a server.
[0233] A "server" is a computer system that receives and processes scripts and feedback submitted by users.
[0234] A "generative model" is a machine learning model for automatically generating movie footage from input script data, and includes, for example, generative AI models (e.g., DALL-E and GPT-4).
[0235] "Movie footage" is visual content that is automatically generated by a generative model based on a script.
[0236] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice and recognizes their emotions.
[0237] "Feedback" refers to evaluations of videos and suggestions for improvement provided by the user or the emotion engine.
[0238] "Encoding" is the process of compressing and converting video data into a particular format.
[0239] The "final footage" is the completed film footage that the user has provided feedback on repeatedly and is finally satisfied with.
[0240] "Storage" refers to storing the final video generated in a data storage system.
[0241] A "download link" is a web address provided to users to obtain the final video.
[0242] This invention relates to a system that automatically generates images using a generative AI model based on a movie script entered by the user, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[0243] 1. Entering the script
[0244] The user uses a device to input a movie script into a dedicated application or web interface. The user writes detailed descriptions of each scene and lines in text format. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining."
[0245] 2. Submit your script
[0246] The terminal converts the input script into JSON format using a script parsing module that converts text input into structured data, and then sends the converted JSON data to the server using an HTTP POST request.
[0247] 3. Initial video generation
[0248] The server analyzes the received JSON data and inputs it into a generative AI model (e.g., DALL-E or GPT-4). The generative AI model generates movie images based on the text data of the script. The generated image data is first temporarily saved.
[0249] 4. First footage review and emotion analysis
[0250] The server encodes the generated video and sends it to the device. When the user watches the video, the device's camera and microphone record the user's facial expressions and voice, and the emotion engine analyzes the data. The emotion engine recognizes the user's emotions in real time and determines emotions such as joy, sadness, and surprise.
[0251] 5. Automatic feedback generation
[0252] The server automatically generates feedback based on the analysis results of the emotion engine. For example, if the user looks dissatisfied while watching a video, the server generates specific feedback such as, "The rain intensity in scene 3 needs to be increased."
[0253] 6. Submitting Feedback
[0254] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server as an HTTP POST request.
[0255] 7. Brush up the video
[0256] The server re-inputs the received feedback into the generative model to regenerate the video. The generative model then modifies the video based on the feedback and outputs the refined video data.
[0257] 8. Checking the regenerated video and making final adjustments
[0258] The regenerated image is then sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine continues to monitor the user's emotions during this process and generates further feedback as needed.
[0259] 9. Storage and provision of final footage
[0260] Once the user is satisfied with the final video, it is stored on the server and a downloadable link is generated. The user can obtain this link via their device and download the final video file.
[0261] For example, if a user inputs a prompt such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the server uses the generative AI model to sequentially generate each scene. If the user reviews the footage and provides feedback such as "I would like the rain to be stronger in Scene 3," or if the emotion engine detects dissatisfaction from the user's facial expression, the server will automatically generate feedback such as "The atmosphere in Scene 3 needs to be more dramatic." In this way, the user can generate movie footage that faithfully reflects their intentions in an efficient and emotion-optimized manner.
[0262] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0263] Step 1:
[0264] A user uses a device to input a movie script into a dedicated application or web interface. The input is text data including details and dialogue for each scene of the movie. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The input is a text script sent to the device. The output is the input script stored in the device's memory.
[0265] Step 2:
[0266] The terminal converts the input script into JSON format. Specifically, the script analysis module analyzes the text data and generates structured JSON data. This JSON data is then sent to the server using an HTTP POST request. The text script is used as input, and the data converted to JSON format is generated as output and sent to the server.
[0267] Step 3:
[0268] The server analyzes the received JSON data and inputs it into a generative AI model. The generative model (e.g., DALL-E or GPT-4) initially generates movie footage based on the text data of the script. The JSON data received by the server is used as input, and the initially generated video data is generated as output and temporarily stored on the server.
[0269] Step 4:
[0270] The server encodes the generated video data and sends it to the device via HTTP. Video encoding technologies such as H.264 are used for encoding. When the user watches the video on the device, the device's camera and microphone record the user's facial expressions and voice in real time, and this data is analyzed by the emotion engine. The video data sent from the server is received as input, and the user's emotion data is generated as output.
[0271] Step 5:
[0272] The emotion engine automatically generates feedback based on the analyzed emotional data. For example, specific feedback is created based on the emotions inferred from the user's facial expressions and voice, such as "The rain intensity in scene 3 needs to be reduced." The user's emotional data is used as input, and the text feedback data is generated as output.
[0273] Step 6:
[0274] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again via an HTTP POST request. The user feedback text is used as input, and the data converted into JSON format is generated as output and sent to the server.
[0275] Step 7:
[0276] The server re-inputs the received feedback data into the generative model and regenerates the video. The generative model outputs a modified video based on the previous video and the new feedback data. The feedback data and the initial video data are used as input, and the modified video data is generated as output and temporarily stored on the server.
[0277] Step 8:
[0278] The regenerated video is sent back to the device for the user to review. The emotion engine continues to analyze the user's reactions and generates new feedback as needed. The modified video data is used as input, and the user's emotion data and new feedback are generated as output.
[0279] Step 9:
[0280] The video data that the user is finally satisfied with is stored on the server. The server generates and provides a download link for the final video to the user. The final video data is used as input, and a download link URL is generated as output and sent to the device.
[0281] (Application example 2)
[0282] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0283] Not only does filmmaking require a significant amount of time and effort, but it is also difficult to create footage that reflects the user's intentions and emotions. In particular, generating a satisfactory video based on a user-provided script requires a lot of trial and error, which hinders the creative process. This challenge stems from the lack of a means to quickly and efficiently generate high-quality footage and achieve results that satisfy the user.
[0284] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating images, means for saving and providing the final movie images, means for analyzing the user's facial expressions and voice to acquire emotional data, and means for automatically generating feedback based on the acquired emotional data. This makes it possible to quickly and efficiently generate high-quality images that reflect the user's intentions and emotions.
[0285] A "movie script" is a text document that describes the content of a movie, the dialogue of the characters, and details of the scenes.
[0286] A "generative model" is an artificial intelligence (AI) system that automatically generates images based on input text data.
[0287] "Feedback" refers to the user's evaluation of the generated video and requests for corrections.
[0288] "Regeneration" is the process of regenerating a previously generated image based on feedback and emotional data.
[0289] "Emotional data" is emotional information obtained by analyzing the user's facial expressions and voice.
[0290] "Automatic generation" refers to the process by which the system automatically generates feedback.
[0291] "Terminal" means a user's device used to input scripts and manage feedback.
[0292] "Interface" refers to the operating screen and input form that allows users to input scripts and feedback via a terminal.
[0293] "Storage" refers to the act of storing the final generated video as data.
[0294] This invention relates to a system that allows a user to input a movie script and automatically generates movie images based on that script using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[0295] Hardware and Software
[0296] First, a user inputs a movie script using a device such as a smartphone. The device has a dedicated application or web interface through which the user can write each scene and line in detail.
[0297] The terminal converts the input script into JSON format and sends it using an HTTP POST request to a server, which is built using a server-side framework such as Flask or Django and has a generative model for analyzing the received script data.
[0298] Image Generation
[0299] The server inputs the received script data into a generative model (e.g., GPT-4) to automatically generate the first movie footage. This video generation process uses the model's deep learning technology. The generated video data is then encoded and sent back to the device.
[0300] Sentiment Analysis and Feedback
[0301] When a user watches a video, an emotion engine (e.g., OpenCV, Emotion API) runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone. The emotion engine recognizes the user's emotions and sends that data to the server. The server then automatically generates feedback based on the emotion data.
[0302] Video Regeneration
[0303] The feedback obtained by the emotion engine and correction requests entered by the user are sent back to the server. The server analyzes this feedback, adjusts the generative model, and regenerates the video. By repeating this process, the user can refine the video until they are satisfied.
[0304] Storage and provision of final footage
[0305] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[0306] Specific examples
[0307] For example, suppose a user (scriptwriter) inputs a script such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user may provide feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it will automatically generate feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[0308] Prompt Sentence Examples
[0309] "Enter a script: The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[0310] "Feedback: The rain in Scene 3 needs to be stronger."
[0311] "Final video generation completed. Download link: https: / / example.com / download / final_video.mp4"
[0312] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[0313] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0314] Step 1:
[0315] The user uses the device to input a movie script. The script is entered in text format through a dedicated application or a web interface. After input, the device converts the script into JSON format. At this time, the script data is structured by scene and line.
[0316] Input: The script text that the user types into the terminal.
[0317] Data processing: Convert text to JSON format
[0318] Output: Structured JSON data
[0319] Step 2:
[0320] The converted script data is sent to the server using an HTTP POST request. The server is built using a server-side framework such as Flask or Django. The server parses the received JSON data and extracts the script content.
[0321] Input: JSON data sent from the terminal
[0322] Data processing: Parsing JSON data
[0323] Output: Parsed script data
[0324] Step 3:
[0325] The server inputs the analyzed script data into a generative model (e.g., GPT-4) to generate the first movie footage. The generative model uses deep learning techniques to generate video data from the text data. The generated footage is then encoded and converted into a viewable format.
[0326] Input: Parsed script data
[0327] Data processing: Video generation and encoding using generative models
[0328] Output: Viewable video file
[0329] Step 4:
[0330] The generated video data is then sent to the device using an HTTP request. The user watches this video. While watching, the emotion engine runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone.
[0331] Input: Viewable video file
[0332] Data processing: Video file distribution and sentiment analysis
[0333] Output: User emotion data
[0334] Step 5:
[0335] The emotion engine transmits emotional data obtained from the user's facial expressions and voice to the server, which analyzes this data and automatically generates feedback as needed. Feedback is expressed as specific correction requests or improvement suggestions.
[0336] Input: User emotion data
[0337] Data processing: analyzing emotional data and generating feedback
[0338] Output: Feedback data
[0339] Step 6:
[0340] The generated feedback data is then sent back to the server, which analyzes the feedback, adjusts the generative model based on it, and regenerates the video. The generation process is then optimized to reflect the feedback.
[0341] Input: Feedback data
[0342] Data processing: Reflecting feedback and regenerating images
[0343] Output: Regenerated video file
[0344] Step 7:
[0345] The regenerated video file is sent back to the device, and this process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal video is generated.
[0346] Input: Regenerated video file
[0347] Data processing: Distribution and reconfirmation of video files
[0348] Output: The final video file that I was satisfied with
[0349] Step 8:
[0350] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[0351] Input: The final video file you are satisfied with
[0352] Data processing: Saving video files and creating links
[0353] Output: Download link
[0354] Through the above steps, it becomes possible to efficiently generate high-quality video that faithfully reflects the user's intentions and emotions.
[0355] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0356] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0357] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0358] [Second embodiment]
[0359] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0360] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0361] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0362] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0363] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0364] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0365] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0366] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0367] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0368] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0369] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0370] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0371] The present invention relates to a system that allows a user to input a movie script and automatically generates movie images based on the script using a generative model. An embodiment of this system will be described below.
[0372] 1. Entering the script
[0373] The user uses the device to input the screenplay in text format, and the device is provided with a dedicated application or a web interface, allowing the user to write detailed descriptions and dialogue for each scene.
[0374] 2. Submit your script
[0375] The terminal sends the input script to the server, where it is converted into a standard data format such as JSON and sent to the server using an HTTP POST request.
[0376] 3. Initial video generation
[0377] The server inputs the received script data into the generative model to generate the first footage of the movie. The generative model automatically creates images and video sequences for each scene based on the text data of the script.
[0378] 4. First video review and feedback
[0379] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. This feedback is entered in text format, detailing specific corrections and improvements.
[0380] 5. Submitting Feedback
[0381] The device sends the user's feedback to the server, which is also converted to JSON format and sent to the server using an HTTP POST request.
[0382] 6. Brush up the video
[0383] The server then applies the feedback data to the generative model and regenerates the new video. In this process, the parameters of the generative model are adjusted and details are modified according to the user's requests.
[0384] 7. Storage and provision of final footage
[0385] Once the user is satisfied with the final video, it is stored on the server and provided in a format that the user can download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link.
[0386] Specific examples
[0387] For example, suppose a user (scriptwriter) inputs a scene such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. If the user provides feedback after viewing the initial footage, such as "I would like the rain to be a little stronger," the server reflects that feedback in the generative model and generates the footage again. If the user reviews the regenerated footage again and is satisfied, it is saved as the final footage and a download link is provided.
[0388] In this way, a system can be realized that enables a user (screenwriter) to efficiently generate movie images that faithfully reflect his or her intentions.
[0389] The processing flow will be explained below.
[0390] Step 1:
[0391] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[0392] Step 2:
[0393] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[0394] Step 3:
[0395] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[0396] Step 4:
[0397] The server encodes the generated initial video data and sends it to the device, where it is provided in the form of a URL that can be accessed by the user.
[0398] Step 5:
[0399] The user accesses the provided URL using a device and watches the initial video. While checking the video, the user can input any corrections or improvements they would like to see as feedback.
[0400] Step 6:
[0401] The device converts the user's feedback into JSON format and sends it to the server again using an HTTP POST request.
[0402] Step 7:
[0403] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback.
[0404] Step 8:
[0405] The server then encodes the regenerated video and sends it back to the device, and this process can be repeated as many times as necessary.
[0406] Step 9:
[0407] The user checks the regenerated image and repeats steps 5 to 8 until satisfied.
[0408] Step 10:
[0409] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[0410] Example 1
[0411] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0412] Conventional methods require a great deal of time and effort to produce movie footage. Furthermore, the process of creating footage that accurately reflects the intent of the script is difficult to correct or adjust. This makes it difficult to efficiently create footage while maintaining quality during movie production.
[0413] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0414] In this invention, the server includes means for a user to input a movie script in text format using a computer terminal, means for converting the input script into a standard data format and transmitting it to the server via a network, means for the server to generate an initial video of the movie based on the script data using a generative model, means for transmitting the generated initial video to the user terminal and for the user to check the video and input feedback, means for transmitting the feedback input by the user to the server and reflecting the feedback in the generative model to regenerate the video, and means for saving the final generated video and providing it in a format that the user can download. This makes it possible to efficiently visualize a movie script and flexibly modify and adjust it according to user requests.
[0415] A "user" is someone who uses the system to input a movie script, review the generated footage, and provide feedback.
[0416] A "computer terminal" is an electronic device used by a user to input a movie script and feedback.
[0417] "Text format" is a data format that handles character information and allows detailed descriptions of each scene and dialogue.
[0418] A "standard data format" is a format defined to maintain data compatibility, and the JSON format is an example of this.
[0419] A "network" is a communications infrastructure for transmitting and receiving data between computer terminals and servers.
[0420] A "server" is a central computer system that processes and stores data and provides services to users over a network.
[0421] A "generative model" is a machine learning model for automatically generating movie footage based on input script data.
[0422] The "first footage" is the first footage of a movie generated by the generative model based on the script data.
[0423] "Feedback" refers to complaints and requests for corrections that users input after viewing the initial footage, and describes in detail specific corrections and improvements.
[0424] "Regeneration" is the process by which a generative model creates new images based on user feedback.
[0425] "Storage" refers to the act of storing the final generated video on a server.
[0426] "Providing" refers to the act of making the saved final video accessible in the form of a link or the like so that the user can download it.
[0427] This invention relates to a system that automatically generates movie footage based on a movie script input by a user using a generative AI model. Specific embodiments of this system are described below.
[0428] Hardware and Software Configuration
[0429] Hardware
[0430] The system includes the following hardware components:
[0431] Terminal: A device where the user can enter the script for the film, review the resulting footage, and provide feedback. This can typically be a computer, tablet, or smartphone.
[0432] Server: A central computer system for processing and storing data, with powerful processors and large amounts of storage.
[0433] software
[0434] Typical software used in each step of the program is as follows:
[0435] Dedicated application or web interface: An interface for users to enter scripts and provide feedback.
[0436] Standard Data Format Conversion Library: A Python "json" library for converting script data and feedback data into JSON format.
[0437] HTTP communication library: The Python "requests" library used to send data from the device to the server.
[0438] Generative AI model: A model used to generate video based on script data. Examples include OpenAI's GPT-4 and Facebook's Fairseq.
[0439] Program processing flow
[0440] 1. Entering the script
[0441] The user uses a terminal to input a movie script in text format. Detailed descriptions and dialogue for each scene can be added via a dedicated application or a web interface. For example, the user might input a scene in which the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining.
[0442] 2. Submit your script
[0443] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, specifically using the Python "requests" library.
[0444] 3. Initial video generation
[0445] The server then inputs the received script data into a generative AI model, such as OpenAI's GPT-4 or Facebook's Fairseq, to generate the first footage. This automatically creates images and video sequences for each scene based on the script's text data.
[0446] 4. First video review and feedback
[0447] The generated initial video is sent from the server to the device. The user reviews the video and enters feedback about any complaints or requests for corrections. This feedback is also converted to JSON format and sent to the server using an HTTP POST request.
[0448] 5. Submitting Feedback
[0449] The device converts the user's feedback into JSON format and sends it to the server, again using an HTTP POST request.
[0450] 6. Brush up the video
[0451] The server then applies the feedback data to the generative AI model and regenerates a new image. The parameters of the generative model are adjusted, and the image is modified according to the user's request, for example, to make the rain stronger.
[0452] 7. Storage and provision of final footage
[0453] Once the user is satisfied with the final video, the server stores it and provides it in a format that the user can download. The user can obtain the video file via a download link.
[0454] By applying this invention, users can efficiently generate movie images that faithfully reflect their intentions. Furthermore, by incorporating user feedback into the images, it becomes possible to improve the quality of movie production.
[0455] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0456] Step 1:
[0457] A user uses a computer terminal to input a movie script in text format. The user then uses a dedicated application or a web interface to write detailed descriptions and dialogue for each scene. The input data is a text-format script. For example, the user might input "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." Subsequent processing begins based on this input data.
[0458] Step 2:
[0459] The terminal converts the input script into a standard data format (JSON format). The terminal uses Python's "json" library to convert text data into JSON. This conversion generates a script in a structured data format. This JSON data is sent to the server using an HTTP POST request. Specifically, the terminal sends the data to the specified URL on the server via the network.
[0460] Step 3:
[0461] The server parses the received JSON-formatted script data and inputs it into a generative AI model. The server then uses a generative AI model (such as OpenAI's GPT-4 or Facebook's Fairseq) to generate the first footage of the movie based on the script data. In this process, images and video sequences of each scene are automatically created from the text data. The output is a first footage file. The server then compiles this into a single video file.
[0462] Step 4:
[0463] The generated initial video is sent from the server to the device. The device displays this video to the user. The user reviews the video and inputs feedback about any complaints or requests for corrections. For example, a user might input feedback such as "I'd like the rain to be a little stronger." This input data is also in text format and is later converted to JSON format.
[0464] Step 5:
[0465] The device converts the feedback entered by the user into JSON format. The device again uses the Python "json" library to structure the feedback data. This feedback data is then sent to the server using an HTTP POST request. Specifically, the device again sends the feedback data over the network to the specified URL on the server.
[0466] Step 6:
[0467] The server applies the feedback data to the generative model and regenerates a new video. The parameters of the generative model are adjusted, and the video is modified according to the user's requests. This regeneration process produces a video that is an improvement over the initial video. The output is a modified video file.
[0468] Step 7:
[0469] The server saves the final video. The server stores the video file in a secure storage and generates a download link that the user can access. The download link is provided to the user, who clicks on the link to obtain the video file. Specifically, the server generates a link based on the path of the saved video file and sends it to the user's device.
[0470] Through these steps, the system can efficiently visualize a movie script and flexibly make corrections and adjustments according to the user's requests.
[0471] (Application example 1)
[0472] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0473] Conventional systems that automatically generate movie footage based on a movie script input in text format often fail to fully reflect the user's intentions. Another problem is the time and effort required to review the generated footage and provide feedback. Furthermore, they do not provide a realistic viewing environment to enhance the viewing experience.
[0474] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0475] In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating the images, means for saving and providing the final movie images, means for viewing the images using a terminal, and means for providing a realistic viewing experience using a head-mounted display. This enables users to efficiently generate images that reflect their intentions while enjoying a realistic viewing experience.
[0476] A "movie script" is a textual document that details the scenes, characters, dialogue, and action of a movie.
[0477] A "generative model" is a machine learning model for automatically generating images and video sequences based on input text data.
[0478] "Video" refers to visual data in which objects or scenes move over time, and in the case of film, includes scripted sequences.
[0479] "Feedback" refers to opinions and requests for corrections that a user provides regarding a generated video.
[0480] A "terminal" is an electronic device that a user operates, and includes a smartphone, tablet, computer, etc.
[0481] A "head-mounted display" is a display device that is worn by the user on the head, and is used to provide VR environments, etc.
[0482] An "immersive viewing experience" is a visual and auditory experience that immerses the viewer as if they are inside the movie.
[0483] This invention is a system that inputs a movie script in text format and automatically generates movie images based on it. An embodiment for realizing this system is described below.
[0484] First, the user uses the device to input the script for the movie. The device is provided with a dedicated application or web interface, and the user can write detailed descriptions and dialogue for each scene. For example, if the user wants to input a scene in which "the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the user can input it in the following text format:
[0485] "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[0486] The input script is converted into a standard data format (e.g., JSON) and sent to the server using an HTTP POST request. The server then inputs the received script data into a generative model to generate the first footage of the movie. This generative model automatically generates images and video sequences based on the text data entered by the user. For example, a machine learning library such as TensorFlow or PyTorch in Python can be used to build the generative model.
[0487] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. The feedback is again converted into a standard data format and sent to the server using an HTTP POST request. If the user wants to provide feedback such as "I'd like the rain to be a little stronger," they can enter it in the following text format:
[0488] "I want the rain to be a little stronger."
[0489] The server then applies the feedback data to the generative model and regenerates a new image. During this process, the parameters of the generative model are adjusted and details are modified according to the user's requests. The regenerated image is then sent back to the device for the user to review.
[0490] Once the user is satisfied with the final video, it is stored on the server and made available for download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link. It is also possible to provide an immersive viewing experience using a smartphone or head-mounted display. For example, head-mounted displays such as Oculus Rift and HTC Vive can be used.
[0491] This system allows users to efficiently generate video that reflects their own intentions while enjoying a realistic viewing experience.
[0492] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0493] Step 1:
[0494] The user uses the device to input a movie script. The input is in text format, including detailed descriptions and dialogue for each scene. The device receives this input and converts it into a standard data format (such as JSON).
[0495] Step 2:
[0496] The device sends the converted script data to the server using an HTTP POST request. Specifically, a JSON object containing the script data is sent, and the server receives it.
[0497] Step 3:
[0498] The server inputs the received script data into a generative model to generate the first footage of the movie. This generation process uses machine learning libraries (such as TensorFlow and PyTorch) to generate scene images and video sequences from the input text data. The generated footage is stored on the server.
[0499] Step 4:
[0500] The server sends the generated initial video as an HTTP response to the terminal, which receives the video data and plays it on the device to provide it to the user.
[0501] Step 5:
[0502] The user can then use the device to view the generated video and enter feedback. This feedback is in the form of text, including suggestions for improvement or correction, such as "I'd like the rain to be a little stronger." The device then converts this feedback back into JSON format.
[0503] Step 6:
[0504] The terminal sends the converted feedback data to the server using an HTTP POST request, and the server retrieves the received feedback data.
[0505] Step 7:
[0506] The server then applies the feedback data to the generative model and regenerates a new video. At this time, the parameters of the generative model are adjusted based on the user's feedback, and details are corrected. The regenerated video is then stored back on the server.
[0507] Step 8:
[0508] The server sends the regenerated video as an HTTP response to the device, which receives the video data and plays it on the device for presentation to the user. This process is repeated until the user is finally satisfied.
[0509] Step 9:
[0510] When the final video is generated to the user's satisfaction, the server saves the final video and generates a downloadable link, which the device provides to the user, who clicks on the link to obtain the video file.
[0511] Step 10:
[0512] Users can view the final video using a device or a head-mounted display, creating an immersive viewing experience. For example, they can enjoy movies in a 3D environment using devices such as Oculus Rift or HTC Vive.
[0513] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0514] This invention relates to a system that automatically generates movie images based on a movie script input by a user using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described below.
[0515] 1. Entering the script
[0516] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[0517] 2. Submit your script
[0518] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[0519] 3. Initial video generation
[0520] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[0521] 4. First footage review and emotion analysis
[0522] The server encodes the generated initial video data and sends it to the device. When the user watches the video, the emotion engine analyzes the user's facial expressions and voice to recognize the user's emotions.
[0523] 5. Automatic feedback generation
[0524] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "The rain intensity in scene 3 needs to be increased" is automatically generated.
[0525] 6. Submitting Feedback
[0526] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again using an HTTP POST request.
[0527] 7. Brush up the video
[0528] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[0529] 8. Checking the regenerated video and making final adjustments
[0530] The regenerated image is sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[0531] 9. Storage and provision of final footage
[0532] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[0533] Specific examples
[0534] For example, suppose a user (scriptwriter) inputs a scene such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user inputs feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it automatically generates feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[0535] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[0536] The processing flow will be explained below.
[0537] Step 1:
[0538] The user uses a device to input a movie script into a dedicated application or web interface. Specifically, the user writes detailed descriptions and dialogue for each scene in text format and saves it as script data.
[0539] Step 2:
[0540] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, at which point the script data is transferred to the server via the network.
[0541] Step 3:
[0542] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data. This process utilizes natural language processing and image generation technology.
[0543] Step 4:
[0544] The server encodes the generated initial video data and sends it to the device, which is provided with a link to the video file, and the user clicks on the link to watch the video.
[0545] Step 5:
[0546] Users access the provided link using their device and watch the initial video. While reviewing the video, they can enter feedback on corrections and improvements they would like to see. Feedback is written in a specific format, such as "Make the rain stronger in scene 3."
[0547] Step 6:
[0548] When users provide feedback, the emotion engine analyzes their facial expressions and voice to recognize their emotions. The emotion engine collects emotion data in real time while the user is watching the video.
[0549] Step 7:
[0550] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "You need to increase the intensity of the rain in scene 3" is automatically generated.
[0551] Step 8:
[0552] The device converts the user feedback and feedback automatically generated by the emotion engine into JSON format and sends it back to the server using an HTTP POST request.
[0553] Step 9:
[0554] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[0555] Step 10:
[0556] The regenerated image is then sent to the device. The user checks the regenerated image and repeats steps 5 to 9 until they are satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[0557] Step 11:
[0558] The final video that the user is satisfied with is saved by the server and a downloadable link is generated. The server then provides this link to the user's device, allowing the user to retrieve the completed video file. This process effectively generates an optimized movie video that takes the user's emotions into account.
[0559] Example 2
[0560] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0561] The traditional film production process, from scriptwriting to video generation and editing, is complex and time-consuming, often requiring large resources and advanced technical skills. It is also difficult to reflect user emotions in real time, making it impossible to optimize video content based on viewer emotions. Furthermore, there are limited ways to efficiently incorporate feedback, making the process of improving video quality cumbersome. There was a need for a system that could solve these problems and efficiently generate high-quality video.
[0562] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input a movie script using a terminal, a means for converting the input script into JSON format and sending it to the server, a means for the server to convert the script into movie footage using a generative model, a means for the user to check the generated footage and analyze the user's emotions using an emotion engine, a means for automatically generating feedback based on the emotion analysis, a means for converting user or automatically generated feedback into JSON format and sending it to the server, a means for reflecting the input feedback in the generative model and regenerating the footage, and a means for saving and providing the final movie footage. This enables optimization of the footage based on the user's emotions and an efficient feedback process.
[0563] A "user" is a user of the system who inputs a movie script, reviews the generated footage, and provides feedback.
[0564] A "terminal" is an information processing device that a user uses to input a movie script and feedback, and specifically includes a PC, smartphone, tablet, etc.
[0565] A "screenplay" is a document that describes in text form the details and dialogue of each scene of a movie.
[0566] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that represents data as key-value pairs.
[0567] An "HTTP POST request" is a type of HyperText Transfer Protocol used to send data to a server.
[0568] A "server" is a computer system that receives and processes scripts and feedback submitted by users.
[0569] A "generative model" is a machine learning model for automatically generating movie footage from input script data, and includes, for example, generative AI models (e.g., DALL-E and GPT-4).
[0570] "Movie footage" is visual content that is automatically generated by a generative model based on a script.
[0571] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice and recognizes their emotions.
[0572] "Feedback" refers to evaluations of videos and suggestions for improvement provided by the user or the emotion engine.
[0573] "Encoding" is the process of compressing and converting video data into a particular format.
[0574] The "final footage" is the completed film footage that the user has provided feedback on repeatedly and is finally satisfied with.
[0575] "Storage" refers to storing the final video generated in a data storage system.
[0576] A "download link" is a web address provided to users to obtain the final video.
[0577] This invention relates to a system that automatically generates images using a generative AI model based on a movie script entered by the user, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[0578] 1. Entering the script
[0579] The user uses a device to input a movie script into a dedicated application or web interface. The user writes detailed descriptions of each scene and lines in text format. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining."
[0580] 2. Submit your script
[0581] The terminal converts the input script into JSON format using a script parsing module that converts text input into structured data, and then sends the converted JSON data to the server using an HTTP POST request.
[0582] 3. Initial video generation
[0583] The server analyzes the received JSON data and inputs it into a generative AI model (e.g., DALL-E or GPT-4). The generative AI model generates movie images based on the text data of the script. The generated image data is first temporarily saved.
[0584] 4. First footage review and emotion analysis
[0585] The server encodes the generated video and sends it to the device. When the user watches the video, the device's camera and microphone record the user's facial expressions and voice, and the emotion engine analyzes the data. The emotion engine recognizes the user's emotions in real time and determines emotions such as joy, sadness, and surprise.
[0586] 5. Automatic feedback generation
[0587] The server automatically generates feedback based on the analysis results of the emotion engine. For example, if the user looks dissatisfied while watching a video, the server generates specific feedback such as, "The rain intensity in scene 3 needs to be increased."
[0588] 6. Submitting Feedback
[0589] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server as an HTTP POST request.
[0590] 7. Brush up the video
[0591] The server re-inputs the received feedback into the generative model to regenerate the video. The generative model then modifies the video based on the feedback and outputs the refined video data.
[0592] 8. Checking the regenerated video and making final adjustments
[0593] The regenerated image is then sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine continues to monitor the user's emotions during this process and generates further feedback as needed.
[0594] 9. Storage and provision of final footage
[0595] Once the user is satisfied with the final video, it is stored on the server and a downloadable link is generated. The user can obtain this link via their device and download the final video file.
[0596] For example, if a user inputs a prompt such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the server uses the generative AI model to sequentially generate each scene. If the user reviews the footage and provides feedback such as "I would like the rain to be stronger in Scene 3," or if the emotion engine detects dissatisfaction from the user's facial expression, the server will automatically generate feedback such as "The atmosphere in Scene 3 needs to be more dramatic." In this way, the user can generate movie footage that faithfully reflects their intentions in an efficient and emotion-optimized manner.
[0597] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0598] Step 1:
[0599] A user uses a device to input a movie script into a dedicated application or web interface. The input is text data including details and dialogue for each scene of the movie. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The input is a text script sent to the device. The output is the input script stored in the device's memory.
[0600] Step 2:
[0601] The terminal converts the input script into JSON format. Specifically, the script analysis module analyzes the text data and generates structured JSON data. This JSON data is then sent to the server using an HTTP POST request. The text script is used as input, and the data converted to JSON format is generated as output and sent to the server.
[0602] Step 3:
[0603] The server analyzes the received JSON data and inputs it into a generative AI model. The generative model (e.g., DALL-E or GPT-4) initially generates movie footage based on the text data of the script. The JSON data received by the server is used as input, and the initially generated video data is generated as output and temporarily stored on the server.
[0604] Step 4:
[0605] The server encodes the generated video data and sends it to the device via HTTP. Video encoding technologies such as H.264 are used for encoding. When the user watches the video on the device, the device's camera and microphone record the user's facial expressions and voice in real time, and this data is analyzed by the emotion engine. The video data sent from the server is received as input, and the user's emotion data is generated as output.
[0606] Step 5:
[0607] The emotion engine automatically generates feedback based on the analyzed emotional data. For example, specific feedback is created based on the emotions inferred from the user's facial expressions and voice, such as "The rain intensity in scene 3 needs to be reduced." The user's emotional data is used as input, and the text feedback data is generated as output.
[0608] Step 6:
[0609] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again via an HTTP POST request. The user feedback text is used as input, and the data converted into JSON format is generated as output and sent to the server.
[0610] Step 7:
[0611] The server re-inputs the received feedback data into the generative model and regenerates the video. The generative model outputs a modified video based on the previous video and the new feedback data. The feedback data and the initial video data are used as input, and the modified video data is generated as output and temporarily stored on the server.
[0612] Step 8:
[0613] The regenerated video is sent back to the device for the user to review. The emotion engine continues to analyze the user's reactions and generates new feedback as needed. The modified video data is used as input, and the user's emotion data and new feedback are generated as output.
[0614] Step 9:
[0615] The video data that the user is finally satisfied with is stored on the server. The server generates and provides a download link for the final video to the user. The final video data is used as input, and a download link URL is generated as output and sent to the device.
[0616] (Application example 2)
[0617] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0618] Not only does filmmaking require a significant amount of time and effort, but it is also difficult to create footage that reflects the user's intentions and emotions. In particular, generating a satisfactory video based on a user-provided script requires a lot of trial and error, which hinders the creative process. This challenge stems from the lack of a means to quickly and efficiently generate high-quality footage and achieve results that satisfy the user.
[0619] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating images, means for saving and providing the final movie images, means for analyzing the user's facial expressions and voice to acquire emotional data, and means for automatically generating feedback based on the acquired emotional data. This makes it possible to quickly and efficiently generate high-quality images that reflect the user's intentions and emotions.
[0620] A "movie script" is a text document that describes the content of a movie, the dialogue of the characters, and details of the scenes.
[0621] A "generative model" is an artificial intelligence (AI) system that automatically generates images based on input text data.
[0622] "Feedback" refers to the user's evaluation of the generated video and requests for corrections.
[0623] "Regeneration" is the process of regenerating a previously generated image based on feedback and emotional data.
[0624] "Emotional data" is emotional information obtained by analyzing the user's facial expressions and voice.
[0625] "Automatic generation" refers to the process by which the system automatically generates feedback.
[0626] "Terminal" means a user's device used to input scripts and manage feedback.
[0627] "Interface" refers to the operating screen and input form that allows users to input scripts and feedback via a terminal.
[0628] "Storage" refers to the act of storing the final generated video as data.
[0629] This invention relates to a system that allows a user to input a movie script and automatically generates movie images based on that script using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[0630] Hardware and Software
[0631] First, a user inputs a movie script using a device such as a smartphone. The device has a dedicated application or web interface through which the user can write each scene and line in detail.
[0632] The terminal converts the input script into JSON format and sends it using an HTTP POST request to a server, which is built using a server-side framework such as Flask or Django and has a generative model for analyzing the received script data.
[0633] Image Generation
[0634] The server inputs the received script data into a generative model (e.g., GPT-4) to automatically generate the first movie footage. This video generation process uses the model's deep learning technology. The generated video data is then encoded and sent back to the device.
[0635] Sentiment Analysis and Feedback
[0636] When a user watches a video, an emotion engine (e.g., OpenCV, Emotion API) runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone. The emotion engine recognizes the user's emotions and sends that data to the server. The server then automatically generates feedback based on the emotion data.
[0637] Video Regeneration
[0638] The feedback obtained by the emotion engine and correction requests entered by the user are sent back to the server. The server analyzes this feedback, adjusts the generative model, and regenerates the video. By repeating this process, the user can refine the video until they are satisfied.
[0639] Storage and provision of final footage
[0640] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[0641] Specific examples
[0642] For example, suppose a user (scriptwriter) inputs a script such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user may provide feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it will automatically generate feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[0643] Prompt Sentence Examples
[0644] "Enter a script: The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[0645] "Feedback: The rain in Scene 3 needs to be stronger."
[0646] "Final video generation completed. Download link: https: / / example.com / download / final_video.mp4"
[0647] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[0648] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0649] Step 1:
[0650] The user uses the device to input a movie script. The script is entered in text format through a dedicated application or a web interface. After input, the device converts the script into JSON format. At this time, the script data is structured by scene and line.
[0651] Input: The script text that the user types into the terminal.
[0652] Data processing: Convert text to JSON format
[0653] Output: Structured JSON data
[0654] Step 2:
[0655] The converted script data is sent to the server using an HTTP POST request. The server is built using a server-side framework such as Flask or Django. The server parses the received JSON data and extracts the script content.
[0656] Input: JSON data sent from the terminal
[0657] Data processing: Parsing JSON data
[0658] Output: Parsed script data
[0659] Step 3:
[0660] The server inputs the analyzed script data into a generative model (e.g., GPT-4) to generate the first movie footage. The generative model uses deep learning techniques to generate video data from the text data. The generated footage is then encoded and converted into a viewable format.
[0661] Input: Parsed script data
[0662] Data processing: Video generation and encoding using generative models
[0663] Output: Viewable video file
[0664] Step 4:
[0665] The generated video data is then sent to the device using an HTTP request. The user watches this video. While watching, the emotion engine runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone.
[0666] Input: Viewable video file
[0667] Data processing: Video file distribution and sentiment analysis
[0668] Output: User emotion data
[0669] Step 5:
[0670] The emotion engine transmits emotional data obtained from the user's facial expressions and voice to the server, which analyzes this data and automatically generates feedback as needed. Feedback is expressed as specific correction requests or improvement suggestions.
[0671] Input: User emotion data
[0672] Data processing: analyzing emotional data and generating feedback
[0673] Output: Feedback data
[0674] Step 6:
[0675] The generated feedback data is then sent back to the server, which analyzes the feedback, adjusts the generative model based on it, and regenerates the video. The generation process is then optimized to reflect the feedback.
[0676] Input: Feedback data
[0677] Data processing: Reflecting feedback and regenerating images
[0678] Output: Regenerated video file
[0679] Step 7:
[0680] The regenerated video file is sent back to the device, and this process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal video is generated.
[0681] Input: Regenerated video file
[0682] Data processing: Distribution and reconfirmation of video files
[0683] Output: The final video file that I was satisfied with
[0684] Step 8:
[0685] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[0686] Input: The final video file you are satisfied with
[0687] Data processing: Saving video files and creating links
[0688] Output: Download link
[0689] Through the above steps, it becomes possible to efficiently generate high-quality video that faithfully reflects the user's intentions and emotions.
[0690] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0691] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0692] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0693] [Third embodiment]
[0694] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0695] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0696] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0697] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0698] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0699] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0700] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0701] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0702] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0703] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0704] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0705] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0706] The present invention relates to a system that allows a user to input a movie script and automatically generates movie images based on the script using a generative model. An embodiment of this system will be described below.
[0707] 1. Entering the script
[0708] The user uses the device to input the screenplay in text format, and the device is provided with a dedicated application or a web interface, allowing the user to write detailed descriptions and dialogue for each scene.
[0709] 2. Submit your script
[0710] The terminal sends the input script to the server, where it is converted into a standard data format such as JSON and sent to the server using an HTTP POST request.
[0711] 3. Initial video generation
[0712] The server inputs the received script data into the generative model to generate the first footage of the movie. The generative model automatically creates images and video sequences for each scene based on the text data of the script.
[0713] 4. First video review and feedback
[0714] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. This feedback is entered in text format, detailing specific corrections and improvements.
[0715] 5. Submitting Feedback
[0716] The device sends the user's feedback to the server, which is also converted to JSON format and sent to the server using an HTTP POST request.
[0717] 6. Brush up the video
[0718] The server then applies the feedback data to the generative model and regenerates the new video. In this process, the parameters of the generative model are adjusted and details are modified according to the user's requests.
[0719] 7. Storage and provision of final footage
[0720] Once the user is satisfied with the final video, it is stored on the server and provided in a format that the user can download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link.
[0721] Specific examples
[0722] For example, suppose a user (scriptwriter) inputs a scene such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. If the user provides feedback after viewing the initial footage, such as "I would like the rain to be a little stronger," the server reflects that feedback in the generative model and generates the footage again. If the user reviews the regenerated footage again and is satisfied, it is saved as the final footage and a download link is provided.
[0723] In this way, a system can be realized that enables a user (screenwriter) to efficiently generate movie images that faithfully reflect his or her intentions.
[0724] The processing flow will be explained below.
[0725] Step 1:
[0726] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[0727] Step 2:
[0728] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[0729] Step 3:
[0730] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[0731] Step 4:
[0732] The server encodes the generated initial video data and sends it to the device, where it is provided in the form of a URL that can be accessed by the user.
[0733] Step 5:
[0734] The user accesses the provided URL using a device and watches the initial video. While checking the video, the user can input any corrections or improvements they would like to see as feedback.
[0735] Step 6:
[0736] The device converts the user's feedback into JSON format and sends it to the server again using an HTTP POST request.
[0737] Step 7:
[0738] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback.
[0739] Step 8:
[0740] The server then encodes the regenerated video and sends it back to the device, and this process can be repeated as many times as necessary.
[0741] Step 9:
[0742] The user checks the regenerated image and repeats steps 5 to 8 until satisfied.
[0743] Step 10:
[0744] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[0745] Example 1
[0746] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0747] Conventional methods require a great deal of time and effort to produce movie footage. Furthermore, the process of creating footage that accurately reflects the intent of the script is difficult to correct or adjust. This makes it difficult to efficiently create footage while maintaining quality during movie production.
[0748] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0749] In this invention, the server includes means for a user to input a movie script in text format using a computer terminal, means for converting the input script into a standard data format and transmitting it to the server via a network, means for the server to generate an initial video of the movie based on the script data using a generative model, means for transmitting the generated initial video to the user terminal and for the user to check the video and input feedback, means for transmitting the feedback input by the user to the server and reflecting the feedback in the generative model to regenerate the video, and means for saving the final generated video and providing it in a format that the user can download. This makes it possible to efficiently visualize a movie script and flexibly modify and adjust it according to user requests.
[0750] A "user" is someone who uses the system to input a movie script, review the generated footage, and provide feedback.
[0751] A "computer terminal" is an electronic device used by a user to input a movie script and feedback.
[0752] "Text format" is a data format that handles character information and allows detailed descriptions of each scene and dialogue.
[0753] A "standard data format" is a format defined to maintain data compatibility, and the JSON format is an example of this.
[0754] A "network" is a communications infrastructure for transmitting and receiving data between computer terminals and servers.
[0755] A "server" is a central computer system that processes and stores data and provides services to users over a network.
[0756] A "generative model" is a machine learning model for automatically generating movie footage based on input script data.
[0757] The "first footage" is the first footage of a movie generated by the generative model based on the script data.
[0758] "Feedback" refers to complaints and requests for corrections that users input after viewing the initial footage, and describes in detail specific corrections and improvements.
[0759] "Regeneration" is the process by which a generative model creates new images based on user feedback.
[0760] "Storage" refers to the act of storing the final generated video on a server.
[0761] "Providing" refers to the act of making the saved final video accessible in the form of a link or the like so that the user can download it.
[0762] This invention relates to a system that automatically generates movie footage based on a movie script input by a user using a generative AI model. Specific embodiments of this system are described below.
[0763] Hardware and Software Configuration
[0764] Hardware
[0765] The system includes the following hardware components:
[0766] Terminal: A device where the user can enter the script for the film, review the resulting footage, and provide feedback. This can typically be a computer, tablet, or smartphone.
[0767] Server: A central computer system for processing and storing data, with powerful processors and large amounts of storage.
[0768] software
[0769] Typical software used in each step of the program is as follows:
[0770] Dedicated application or web interface: An interface for users to enter scripts and provide feedback.
[0771] Standard Data Format Conversion Library: A Python "json" library for converting script data and feedback data into JSON format.
[0772] HTTP communication library: The Python "requests" library used to send data from the device to the server.
[0773] Generative AI model: A model used to generate video based on script data. Examples include OpenAI's GPT-4 and Facebook's Fairseq.
[0774] Program processing flow
[0775] 1. Entering the script
[0776] The user uses a terminal to input a movie script in text format. Detailed descriptions and dialogue for each scene can be added via a dedicated application or a web interface. For example, the user might input a scene in which the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining.
[0777] 2. Submit your script
[0778] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, specifically using the Python "requests" library.
[0779] 3. Initial video generation
[0780] The server then inputs the received script data into a generative AI model, such as OpenAI's GPT-4 or Facebook's Fairseq, to generate the first footage. This automatically creates images and video sequences for each scene based on the script's text data.
[0781] 4. First video review and feedback
[0782] The generated initial video is sent from the server to the device. The user reviews the video and enters feedback about any complaints or requests for corrections. This feedback is also converted to JSON format and sent to the server using an HTTP POST request.
[0783] 5. Submitting Feedback
[0784] The device converts the user's feedback into JSON format and sends it to the server, again using an HTTP POST request.
[0785] 6. Brush up the video
[0786] The server then applies the feedback data to the generative AI model and regenerates a new image. The parameters of the generative model are adjusted, and the image is modified according to the user's request, for example, to make the rain stronger.
[0787] 7. Storage and provision of final footage
[0788] Once the user is satisfied with the final video, the server stores it and provides it in a format that the user can download. The user can obtain the video file via a download link.
[0789] By applying this invention, users can efficiently generate movie images that faithfully reflect their intentions. Furthermore, by incorporating user feedback into the images, it becomes possible to improve the quality of movie production.
[0790] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0791] Step 1:
[0792] A user uses a computer terminal to input a movie script in text format. The user then uses a dedicated application or a web interface to write detailed descriptions and dialogue for each scene. The input data is a text-format script. For example, the user might input "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." Subsequent processing begins based on this input data.
[0793] Step 2:
[0794] The terminal converts the input script into a standard data format (JSON format). The terminal uses Python's "json" library to convert text data into JSON. This conversion generates a script in a structured data format. This JSON data is sent to the server using an HTTP POST request. Specifically, the terminal sends the data to the specified URL on the server via the network.
[0795] Step 3:
[0796] The server parses the received JSON-formatted script data and inputs it into a generative AI model. The server then uses a generative AI model (such as OpenAI's GPT-4 or Facebook's Fairseq) to generate the first footage of the movie based on the script data. In this process, images and video sequences of each scene are automatically created from the text data. The output is a first footage file. The server then compiles this into a single video file.
[0797] Step 4:
[0798] The generated initial video is sent from the server to the device. The device displays this video to the user. The user reviews the video and inputs feedback about any complaints or requests for corrections. For example, a user might input feedback such as "I'd like the rain to be a little stronger." This input data is also in text format and is later converted to JSON format.
[0799] Step 5:
[0800] The device converts the feedback entered by the user into JSON format. The device again uses the Python "json" library to structure the feedback data. This feedback data is then sent to the server using an HTTP POST request. Specifically, the device again sends the feedback data over the network to the specified URL on the server.
[0801] Step 6:
[0802] The server applies the feedback data to the generative model and regenerates a new video. The parameters of the generative model are adjusted, and the video is modified according to the user's requests. This regeneration process produces a video that is an improvement over the initial video. The output is a modified video file.
[0803] Step 7:
[0804] The server saves the final video. The server stores the video file in a secure storage and generates a download link that the user can access. The download link is provided to the user, who clicks on the link to obtain the video file. Specifically, the server generates a link based on the path of the saved video file and sends it to the user's device.
[0805] Through these steps, the system can efficiently visualize a movie script and flexibly make corrections and adjustments according to the user's requests.
[0806] (Application example 1)
[0807] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0808] Conventional systems that automatically generate movie footage based on a movie script input in text format often fail to fully reflect the user's intentions. Another problem is the time and effort required to review the generated footage and provide feedback. Furthermore, they do not provide a realistic viewing environment to enhance the viewing experience.
[0809] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0810] In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating the images, means for saving and providing the final movie images, means for viewing the images using a terminal, and means for providing a realistic viewing experience using a head-mounted display. This enables users to efficiently generate images that reflect their intentions while enjoying a realistic viewing experience.
[0811] A "movie script" is a textual document that details the scenes, characters, dialogue, and action of a movie.
[0812] A "generative model" is a machine learning model for automatically generating images and video sequences based on input text data.
[0813] "Video" refers to visual data in which objects or scenes move over time, and in the case of film, includes scripted sequences.
[0814] "Feedback" refers to opinions and requests for corrections that a user provides regarding a generated video.
[0815] A "terminal" is an electronic device that a user operates, and includes a smartphone, tablet, computer, etc.
[0816] A "head-mounted display" is a display device that is worn by the user on the head, and is used to provide VR environments, etc.
[0817] An "immersive viewing experience" is a visual and auditory experience that immerses the viewer as if they are inside the movie.
[0818] This invention is a system that inputs a movie script in text format and automatically generates movie images based on it. An embodiment for realizing this system is described below.
[0819] First, the user uses the device to input the script for the movie. The device is provided with a dedicated application or web interface, and the user can write detailed descriptions and dialogue for each scene. For example, if the user wants to input a scene in which "the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the user can input it in the following text format:
[0820] "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[0821] The input script is converted into a standard data format (e.g., JSON) and sent to the server using an HTTP POST request. The server then inputs the received script data into a generative model to generate the first footage of the movie. This generative model automatically generates images and video sequences based on the text data entered by the user. For example, a machine learning library such as TensorFlow or PyTorch in Python can be used to build the generative model.
[0822] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. The feedback is again converted into a standard data format and sent to the server using an HTTP POST request. If the user wants to provide feedback such as "I'd like the rain to be a little stronger," they can enter it in the following text format:
[0823] "I want the rain to be a little stronger."
[0824] The server then applies the feedback data to the generative model and regenerates a new image. During this process, the parameters of the generative model are adjusted and details are modified according to the user's requests. The regenerated image is then sent back to the device for the user to review.
[0825] Once the user is satisfied with the final video, it is stored on the server and made available for download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link. It is also possible to provide an immersive viewing experience using a smartphone or head-mounted display. For example, head-mounted displays such as Oculus Rift and HTC Vive can be used.
[0826] This system allows users to efficiently generate video that reflects their own intentions while enjoying a realistic viewing experience.
[0827] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0828] Step 1:
[0829] The user uses the device to input a movie script. The input is in text format, including detailed descriptions and dialogue for each scene. The device receives this input and converts it into a standard data format (such as JSON).
[0830] Step 2:
[0831] The device sends the converted script data to the server using an HTTP POST request. Specifically, a JSON object containing the script data is sent, and the server receives it.
[0832] Step 3:
[0833] The server inputs the received script data into a generative model to generate the first footage of the movie. This generation process uses machine learning libraries (such as TensorFlow and PyTorch) to generate scene images and video sequences from the input text data. The generated footage is stored on the server.
[0834] Step 4:
[0835] The server sends the generated initial video as an HTTP response to the terminal, which receives the video data and plays it on the device to provide it to the user.
[0836] Step 5:
[0837] The user can then use the device to view the generated video and enter feedback. This feedback is in the form of text, including suggestions for improvement or correction, such as "I'd like the rain to be a little stronger." The device then converts this feedback back into JSON format.
[0838] Step 6:
[0839] The terminal sends the converted feedback data to the server using an HTTP POST request, and the server retrieves the received feedback data.
[0840] Step 7:
[0841] The server then applies the feedback data to the generative model and regenerates a new video. At this time, the parameters of the generative model are adjusted based on the user's feedback, and details are corrected. The regenerated video is then stored back on the server.
[0842] Step 8:
[0843] The server sends the regenerated video as an HTTP response to the device, which receives the video data and plays it on the device for presentation to the user. This process is repeated until the user is finally satisfied.
[0844] Step 9:
[0845] When the final video is generated to the user's satisfaction, the server saves the final video and generates a downloadable link, which the device provides to the user, who clicks on the link to obtain the video file.
[0846] Step 10:
[0847] Users can view the final video using a device or a head-mounted display, creating an immersive viewing experience. For example, they can enjoy movies in a 3D environment using devices such as Oculus Rift or HTC Vive.
[0848] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0849] This invention relates to a system that automatically generates movie images based on a movie script input by a user using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described below.
[0850] 1. Entering the script
[0851] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[0852] 2. Submit your script
[0853] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[0854] 3. Initial video generation
[0855] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[0856] 4. First footage review and emotion analysis
[0857] The server encodes the generated initial video data and sends it to the device. When the user watches the video, the emotion engine analyzes the user's facial expressions and voice to recognize the user's emotions.
[0858] 5. Automatic feedback generation
[0859] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "The rain intensity in scene 3 needs to be increased" is automatically generated.
[0860] 6. Submitting Feedback
[0861] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again using an HTTP POST request.
[0862] 7. Brush up the video
[0863] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[0864] 8. Checking the regenerated video and making final adjustments
[0865] The regenerated image is sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[0866] 9. Storage and provision of final footage
[0867] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[0868] Specific examples
[0869] For example, suppose a user (scriptwriter) inputs a scene such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user inputs feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it automatically generates feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[0870] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[0871] The processing flow will be explained below.
[0872] Step 1:
[0873] The user uses a device to input a movie script into a dedicated application or web interface. Specifically, the user writes detailed descriptions and dialogue for each scene in text format and saves it as script data.
[0874] Step 2:
[0875] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, at which point the script data is transferred to the server via the network.
[0876] Step 3:
[0877] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data. This process utilizes natural language processing and image generation technology.
[0878] Step 4:
[0879] The server encodes the generated initial video data and sends it to the device, which is provided with a link to the video file, and the user clicks on the link to watch the video.
[0880] Step 5:
[0881] Users access the provided link using their device and watch the initial video. While reviewing the video, they can enter feedback on corrections and improvements they would like to see. Feedback is written in a specific format, such as "Make the rain stronger in scene 3."
[0882] Step 6:
[0883] When users provide feedback, the emotion engine analyzes their facial expressions and voice to recognize their emotions. The emotion engine collects emotion data in real time while the user is watching the video.
[0884] Step 7:
[0885] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "You need to increase the intensity of the rain in scene 3" is automatically generated.
[0886] Step 8:
[0887] The device converts the user feedback and feedback automatically generated by the emotion engine into JSON format and sends it back to the server using an HTTP POST request.
[0888] Step 9:
[0889] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[0890] Step 10:
[0891] The regenerated image is then sent to the device. The user checks the regenerated image and repeats steps 5 to 9 until they are satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[0892] Step 11:
[0893] The final video that the user is satisfied with is saved by the server and a downloadable link is generated. The server then provides this link to the user's device, allowing the user to retrieve the completed video file. This process effectively generates an optimized movie video that takes the user's emotions into account.
[0894] Example 2
[0895] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0896] The traditional film production process, from scriptwriting to video generation and editing, is complex and time-consuming, often requiring large resources and advanced technical skills. It is also difficult to reflect user emotions in real time, making it impossible to optimize video content based on viewer emotions. Furthermore, there are limited ways to efficiently incorporate feedback, making the process of improving video quality cumbersome. There was a need for a system that could solve these problems and efficiently generate high-quality video.
[0897] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input a movie script using a terminal, a means for converting the input script into JSON format and sending it to the server, a means for the server to convert the script into movie footage using a generative model, a means for the user to check the generated footage and analyze the user's emotions using an emotion engine, a means for automatically generating feedback based on the emotion analysis, a means for converting user or automatically generated feedback into JSON format and sending it to the server, a means for reflecting the input feedback in the generative model and regenerating the footage, and a means for saving and providing the final movie footage. This enables optimization of the footage based on the user's emotions and an efficient feedback process.
[0898] A "user" is a user of the system who inputs a movie script, reviews the generated footage, and provides feedback.
[0899] A "terminal" is an information processing device that a user uses to input a movie script and feedback, and specifically includes a PC, smartphone, tablet, etc.
[0900] A "screenplay" is a document that describes in text form the details and dialogue of each scene of a movie.
[0901] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that represents data as key-value pairs.
[0902] An "HTTP POST request" is a type of HyperText Transfer Protocol used to send data to a server.
[0903] A "server" is a computer system that receives and processes scripts and feedback submitted by users.
[0904] A "generative model" is a machine learning model for automatically generating movie footage from input script data, and includes, for example, generative AI models (e.g., DALL-E and GPT-4).
[0905] "Movie footage" is visual content that is automatically generated by a generative model based on a script.
[0906] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice and recognizes their emotions.
[0907] "Feedback" refers to evaluations of videos and suggestions for improvement provided by the user or the emotion engine.
[0908] "Encoding" is the process of compressing and converting video data into a particular format.
[0909] The "final footage" is the completed film footage that the user has provided feedback on repeatedly and is finally satisfied with.
[0910] "Storage" refers to storing the final video generated in a data storage system.
[0911] A "download link" is a web address provided to users to obtain the final video.
[0912] This invention relates to a system that automatically generates images using a generative AI model based on a movie script entered by the user, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[0913] 1. Entering the script
[0914] The user uses a device to input a movie script into a dedicated application or web interface. The user writes detailed descriptions of each scene and lines in text format. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining."
[0915] 2. Submit your script
[0916] The terminal converts the input script into JSON format using a script parsing module that converts text input into structured data, and then sends the converted JSON data to the server using an HTTP POST request.
[0917] 3. Initial video generation
[0918] The server analyzes the received JSON data and inputs it into a generative AI model (e.g., DALL-E or GPT-4). The generative AI model generates movie images based on the text data of the script. The generated image data is first temporarily saved.
[0919] 4. First footage review and emotion analysis
[0920] The server encodes the generated video and sends it to the device. When the user watches the video, the device's camera and microphone record the user's facial expressions and voice, and the emotion engine analyzes the data. The emotion engine recognizes the user's emotions in real time and determines emotions such as joy, sadness, and surprise.
[0921] 5. Automatic feedback generation
[0922] The server automatically generates feedback based on the analysis results of the emotion engine. For example, if the user looks dissatisfied while watching a video, the server generates specific feedback such as, "The rain intensity in scene 3 needs to be increased."
[0923] 6. Submitting Feedback
[0924] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server as an HTTP POST request.
[0925] 7. Brush up the video
[0926] The server re-inputs the received feedback into the generative model to regenerate the video. The generative model then modifies the video based on the feedback and outputs the refined video data.
[0927] 8. Checking the regenerated video and making final adjustments
[0928] The regenerated image is then sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine continues to monitor the user's emotions during this process and generates further feedback as needed.
[0929] 9. Storage and provision of final footage
[0930] Once the user is satisfied with the final video, it is stored on the server and a downloadable link is generated. The user can obtain this link via their device and download the final video file.
[0931] For example, if a user inputs a prompt such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the server uses the generative AI model to sequentially generate each scene. If the user reviews the footage and provides feedback such as "I would like the rain to be stronger in Scene 3," or if the emotion engine detects dissatisfaction from the user's facial expression, the server will automatically generate feedback such as "The atmosphere in Scene 3 needs to be more dramatic." In this way, the user can generate movie footage that faithfully reflects their intentions in an efficient and emotion-optimized manner.
[0932] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0933] Step 1:
[0934] A user uses a device to input a movie script into a dedicated application or web interface. The input is text data including details and dialogue for each scene of the movie. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The input is a text script sent to the device. The output is the input script stored in the device's memory.
[0935] Step 2:
[0936] The terminal converts the input script into JSON format. Specifically, the script analysis module analyzes the text data and generates structured JSON data. This JSON data is then sent to the server using an HTTP POST request. The text script is used as input, and the data converted to JSON format is generated as output and sent to the server.
[0937] Step 3:
[0938] The server analyzes the received JSON data and inputs it into a generative AI model. The generative model (e.g., DALL-E or GPT-4) initially generates movie footage based on the text data of the script. The JSON data received by the server is used as input, and the initially generated video data is generated as output and temporarily stored on the server.
[0939] Step 4:
[0940] The server encodes the generated video data and sends it to the device via HTTP. Video encoding technologies such as H.264 are used for encoding. When the user watches the video on the device, the device's camera and microphone record the user's facial expressions and voice in real time, and this data is analyzed by the emotion engine. The video data sent from the server is received as input, and the user's emotion data is generated as output.
[0941] Step 5:
[0942] The emotion engine automatically generates feedback based on the analyzed emotional data. For example, specific feedback is created based on the emotions inferred from the user's facial expressions and voice, such as "The rain intensity in scene 3 needs to be reduced." The user's emotional data is used as input, and the text feedback data is generated as output.
[0943] Step 6:
[0944] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again via an HTTP POST request. The user feedback text is used as input, and the data converted into JSON format is generated as output and sent to the server.
[0945] Step 7:
[0946] The server re-inputs the received feedback data into the generative model and regenerates the video. The generative model outputs a modified video based on the previous video and the new feedback data. The feedback data and the initial video data are used as input, and the modified video data is generated as output and temporarily stored on the server.
[0947] Step 8:
[0948] The regenerated video is sent back to the device for the user to review. The emotion engine continues to analyze the user's reactions and generates new feedback as needed. The modified video data is used as input, and the user's emotion data and new feedback are generated as output.
[0949] Step 9:
[0950] The video data that the user is finally satisfied with is stored on the server. The server generates and provides a download link for the final video to the user. The final video data is used as input, and a download link URL is generated as output and sent to the device.
[0951] (Application example 2)
[0952] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0953] Not only does filmmaking require a significant amount of time and effort, but it is also difficult to create footage that reflects the user's intentions and emotions. In particular, generating a satisfactory video based on a user-provided script requires a lot of trial and error, which hinders the creative process. This challenge stems from the lack of a means to quickly and efficiently generate high-quality footage and achieve results that satisfy the user.
[0954] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating images, means for saving and providing the final movie images, means for analyzing the user's facial expressions and voice to acquire emotional data, and means for automatically generating feedback based on the acquired emotional data. This makes it possible to quickly and efficiently generate high-quality images that reflect the user's intentions and emotions.
[0955] A "movie script" is a text document that describes the content of a movie, the dialogue of the characters, and details of the scenes.
[0956] A "generative model" is an artificial intelligence (AI) system that automatically generates images based on input text data.
[0957] "Feedback" refers to the user's evaluation of the generated video and requests for corrections.
[0958] "Regeneration" is the process of regenerating a previously generated image based on feedback and emotional data.
[0959] "Emotional data" is emotional information obtained by analyzing the user's facial expressions and voice.
[0960] "Automatic generation" refers to the process by which the system automatically generates feedback.
[0961] "Terminal" means a user's device used to input scripts and manage feedback.
[0962] "Interface" refers to the operating screen and input form that allows users to input scripts and feedback via a terminal.
[0963] "Storage" refers to the act of storing the final generated video as data.
[0964] This invention relates to a system that allows a user to input a movie script and automatically generates movie images based on that script using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[0965] Hardware and Software
[0966] First, a user inputs a movie script using a device such as a smartphone. The device has a dedicated application or web interface through which the user can write each scene and line in detail.
[0967] The terminal converts the input script into JSON format and sends it using an HTTP POST request to a server, which is built using a server-side framework such as Flask or Django and has a generative model for analyzing the received script data.
[0968] Image Generation
[0969] The server inputs the received script data into a generative model (e.g., GPT-4) to automatically generate the first movie footage. This video generation process uses the model's deep learning technology. The generated video data is then encoded and sent back to the device.
[0970] Sentiment Analysis and Feedback
[0971] When a user watches a video, an emotion engine (e.g., OpenCV, Emotion API) runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone. The emotion engine recognizes the user's emotions and sends that data to the server. The server then automatically generates feedback based on the emotion data.
[0972] Video Regeneration
[0973] The feedback obtained by the emotion engine and correction requests entered by the user are sent back to the server. The server analyzes this feedback, adjusts the generative model, and regenerates the video. By repeating this process, the user can refine the video until they are satisfied.
[0974] Storage and provision of final footage
[0975] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[0976] Specific examples
[0977] For example, suppose a user (scriptwriter) inputs a script such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user may provide feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it will automatically generate feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[0978] Prompt Sentence Examples
[0979] "Enter a script: The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[0980] "Feedback: The rain in Scene 3 needs to be stronger."
[0981] "Final video generation completed. Download link: https: / / example.com / download / final_video.mp4"
[0982] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[0983] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0984] Step 1:
[0985] The user uses the device to input a movie script. The script is entered in text format through a dedicated application or a web interface. After input, the device converts the script into JSON format. At this time, the script data is structured by scene and line.
[0986] Input: The script text that the user types into the terminal.
[0987] Data processing: Convert text to JSON format
[0988] Output: Structured JSON data
[0989] Step 2:
[0990] The converted script data is sent to the server using an HTTP POST request. The server is built using a server-side framework such as Flask or Django. The server parses the received JSON data and extracts the script content.
[0991] Input: JSON data sent from the terminal
[0992] Data processing: Parsing JSON data
[0993] Output: Parsed script data
[0994] Step 3:
[0995] The server inputs the analyzed script data into a generative model (e.g., GPT-4) to generate the first movie footage. The generative model uses deep learning techniques to generate video data from the text data. The generated footage is then encoded and converted into a viewable format.
[0996] Input: Parsed script data
[0997] Data processing: Video generation and encoding using generative models
[0998] Output: Viewable video file
[0999] Step 4:
[1000] The generated video data is then sent to the device using an HTTP request. The user watches this video. While watching, the emotion engine runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone.
[1001] Input: Viewable video file
[1002] Data processing: Video file distribution and sentiment analysis
[1003] Output: User emotion data
[1004] Step 5:
[1005] The emotion engine transmits emotional data obtained from the user's facial expressions and voice to the server, which analyzes this data and automatically generates feedback as needed. Feedback is expressed as specific correction requests or improvement suggestions.
[1006] Input: User emotion data
[1007] Data processing: analyzing emotional data and generating feedback
[1008] Output: Feedback data
[1009] Step 6:
[1010] The generated feedback data is then sent back to the server, which analyzes the feedback, adjusts the generative model based on it, and regenerates the video. The generation process is then optimized to reflect the feedback.
[1011] Input: Feedback data
[1012] Data processing: Reflecting feedback and regenerating images
[1013] Output: Regenerated video file
[1014] Step 7:
[1015] The regenerated video file is sent back to the device, and this process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal video is generated.
[1016] Input: Regenerated video file
[1017] Data processing: Distribution and reconfirmation of video files
[1018] Output: The final video file that I was satisfied with
[1019] Step 8:
[1020] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[1021] Input: The final video file you are satisfied with
[1022] Data processing: Saving video files and creating links
[1023] Output: Download link
[1024] Through the above steps, it becomes possible to efficiently generate high-quality video that faithfully reflects the user's intentions and emotions.
[1025] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1026] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1027] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1028] [Fourth embodiment]
[1029] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1030] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1031] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1032] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1033] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1034] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1035] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1036] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1037] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1038] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1040] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1042] The present invention relates to a system that allows a user to input a movie script and automatically generates movie images based on the script using a generative model. An embodiment of this system will be described below.
[1043] 1. Entering the script
[1044] The user uses the device to input the screenplay in text format, and the device is provided with a dedicated application or a web interface, allowing the user to write detailed descriptions and dialogue for each scene.
[1045] 2. Submit your script
[1046] The terminal sends the input script to the server, where it is converted into a standard data format such as JSON and sent to the server using an HTTP POST request.
[1047] 3. Initial video generation
[1048] The server inputs the received script data into the generative model to generate the first footage of the movie. The generative model automatically creates images and video sequences for each scene based on the text data of the script.
[1049] 4. First video review and feedback
[1050] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. This feedback is entered in text format, detailing specific corrections and improvements.
[1051] 5. Submitting Feedback
[1052] The device sends the user's feedback to the server, which is also converted to JSON format and sent to the server using an HTTP POST request.
[1053] 6. Brush up the video
[1054] The server then applies the feedback data to the generative model and regenerates the new video. In this process, the parameters of the generative model are adjusted and details are modified according to the user's requests.
[1055] 7. Storage and provision of final footage
[1056] Once the user is satisfied with the final video, it is stored on the server and provided in a format that the user can download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link.
[1057] Specific examples
[1058] For example, suppose a user (scriptwriter) inputs a scene such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. If the user provides feedback after viewing the initial footage, such as "I would like the rain to be a little stronger," the server reflects that feedback in the generative model and generates the footage again. If the user reviews the regenerated footage again and is satisfied, it is saved as the final footage and a download link is provided.
[1059] In this way, a system can be realized that enables a user (screenwriter) to efficiently generate movie images that faithfully reflect his or her intentions.
[1060] The processing flow will be explained below.
[1061] Step 1:
[1062] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[1063] Step 2:
[1064] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[1065] Step 3:
[1066] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[1067] Step 4:
[1068] The server encodes the generated initial video data and sends it to the device, where it is provided in the form of a URL that can be accessed by the user.
[1069] Step 5:
[1070] The user accesses the provided URL using a device and watches the initial video. While checking the video, the user can input any corrections or improvements they would like to see as feedback.
[1071] Step 6:
[1072] The device converts the user's feedback into JSON format and sends it to the server again using an HTTP POST request.
[1073] Step 7:
[1074] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback.
[1075] Step 8:
[1076] The server then encodes the regenerated video and sends it back to the device, and this process can be repeated as many times as necessary.
[1077] Step 9:
[1078] The user checks the regenerated image and repeats steps 5 to 8 until satisfied.
[1079] Step 10:
[1080] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[1081] Example 1
[1082] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1083] Conventional methods require a great deal of time and effort to produce movie footage. Furthermore, the process of creating footage that accurately reflects the intent of the script is difficult to correct or adjust. This makes it difficult to efficiently create footage while maintaining quality during movie production.
[1084] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1085] In this invention, the server includes means for a user to input a movie script in text format using a computer terminal, means for converting the input script into a standard data format and transmitting it to the server via a network, means for the server to generate an initial video of the movie based on the script data using a generative model, means for transmitting the generated initial video to the user terminal and for the user to check the video and input feedback, means for transmitting the feedback input by the user to the server and reflecting the feedback in the generative model to regenerate the video, and means for saving the final generated video and providing it in a format that the user can download. This makes it possible to efficiently visualize a movie script and flexibly modify and adjust it according to user requests.
[1086] A "user" is someone who uses the system to input a movie script, review the generated footage, and provide feedback.
[1087] A "computer terminal" is an electronic device used by a user to input a movie script and feedback.
[1088] "Text format" is a data format that handles character information and allows detailed descriptions of each scene and dialogue.
[1089] A "standard data format" is a format defined to maintain data compatibility, and the JSON format is an example of this.
[1090] A "network" is a communications infrastructure for transmitting and receiving data between computer terminals and servers.
[1091] A "server" is a central computer system that processes and stores data and provides services to users over a network.
[1092] A "generative model" is a machine learning model for automatically generating movie footage based on input script data.
[1093] The "first footage" is the first footage of a movie generated by the generative model based on the script data.
[1094] "Feedback" refers to complaints and requests for corrections that users input after viewing the initial footage, and describes in detail specific corrections and improvements.
[1095] "Regeneration" is the process by which a generative model creates new images based on user feedback.
[1096] "Storage" refers to the act of storing the final generated video on a server.
[1097] "Providing" refers to the act of making the saved final video accessible in the form of a link or the like so that the user can download it.
[1098] This invention relates to a system that automatically generates movie footage based on a movie script input by a user using a generative AI model. Specific embodiments of this system are described below.
[1099] Hardware and Software Configuration
[1100] Hardware
[1101] The system includes the following hardware components:
[1102] Terminal: A device where the user can enter the script for the film, review the resulting footage, and provide feedback. This can typically be a computer, tablet, or smartphone.
[1103] Server: A central computer system for processing and storing data, with powerful processors and large amounts of storage.
[1104] software
[1105] Typical software used in each step of the program is as follows:
[1106] Dedicated application or web interface: An interface for users to enter scripts and provide feedback.
[1107] Standard Data Format Conversion Library: A Python "json" library for converting script data and feedback data into JSON format.
[1108] HTTP communication library: The Python "requests" library used to send data from the device to the server.
[1109] Generative AI model: A model used to generate video based on script data. Examples include OpenAI's GPT-4 and Facebook's Fairseq.
[1110] Program processing flow
[1111] 1. Entering the script
[1112] The user uses a terminal to input a movie script in text format. Detailed descriptions and dialogue for each scene can be added via a dedicated application or a web interface. For example, the user might input a scene in which the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining.
[1113] 2. Submit your script
[1114] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, specifically using the Python "requests" library.
[1115] 3. Initial video generation
[1116] The server then inputs the received script data into a generative AI model, such as OpenAI's GPT-4 or Facebook's Fairseq, to generate the first footage. This automatically creates images and video sequences for each scene based on the script's text data.
[1117] 4. First video review and feedback
[1118] The generated initial video is sent from the server to the device. The user reviews the video and enters feedback about any complaints or requests for corrections. This feedback is also converted to JSON format and sent to the server using an HTTP POST request.
[1119] 5. Submitting Feedback
[1120] The device converts the user's feedback into JSON format and sends it to the server, again using an HTTP POST request.
[1121] 6. Brush up the video
[1122] The server then applies the feedback data to the generative AI model and regenerates a new image. The parameters of the generative model are adjusted, and the image is modified according to the user's request, for example, to make the rain stronger.
[1123] 7. Storage and provision of final footage
[1124] Once the user is satisfied with the final video, the server stores it and provides it in a format that the user can download. The user can obtain the video file via a download link.
[1125] By applying this invention, users can efficiently generate movie images that faithfully reflect their intentions. Furthermore, by incorporating user feedback into the images, it becomes possible to improve the quality of movie production.
[1126] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1127] Step 1:
[1128] A user uses a computer terminal to input a movie script in text format. The user then uses a dedicated application or a web interface to write detailed descriptions and dialogue for each scene. The input data is a text-format script. For example, the user might input "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." Subsequent processing begins based on this input data.
[1129] Step 2:
[1130] The terminal converts the input script into a standard data format (JSON format). The terminal uses Python's "json" library to convert text data into JSON. This conversion generates a script in a structured data format. This JSON data is sent to the server using an HTTP POST request. Specifically, the terminal sends the data to the specified URL on the server via the network.
[1131] Step 3:
[1132] The server parses the received JSON-formatted script data and inputs it into a generative AI model. The server then uses a generative AI model (such as OpenAI's GPT-4 or Facebook's Fairseq) to generate the first footage of the movie based on the script data. In this process, images and video sequences of each scene are automatically created from the text data. The output is a first footage file. The server then compiles this into a single video file.
[1133] Step 4:
[1134] The generated initial video is sent from the server to the device. The device displays this video to the user. The user reviews the video and inputs feedback about any complaints or requests for corrections. For example, a user might input feedback such as "I'd like the rain to be a little stronger." This input data is also in text format and is later converted to JSON format.
[1135] Step 5:
[1136] The device converts the feedback entered by the user into JSON format. The device again uses the Python "json" library to structure the feedback data. This feedback data is then sent to the server using an HTTP POST request. Specifically, the device again sends the feedback data over the network to the specified URL on the server.
[1137] Step 6:
[1138] The server applies the feedback data to the generative model and regenerates a new video. The parameters of the generative model are adjusted, and the video is modified according to the user's requests. This regeneration process produces a video that is an improvement over the initial video. The output is a modified video file.
[1139] Step 7:
[1140] The server saves the final video. The server stores the video file in a secure storage and generates a download link that the user can access. The download link is provided to the user, who clicks on the link to obtain the video file. Specifically, the server generates a link based on the path of the saved video file and sends it to the user's device.
[1141] Through these steps, the system can efficiently visualize a movie script and flexibly make corrections and adjustments according to the user's requests.
[1142] (Application example 1)
[1143] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1144] Conventional systems that automatically generate movie footage based on a movie script input in text format often fail to fully reflect the user's intentions. Another problem is the time and effort required to review the generated footage and provide feedback. Furthermore, they do not provide a realistic viewing environment to enhance the viewing experience.
[1145] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1146] In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating the images, means for saving and providing the final movie images, means for viewing the images using a terminal, and means for providing a realistic viewing experience using a head-mounted display. This enables users to efficiently generate images that reflect their intentions while enjoying a realistic viewing experience.
[1147] A "movie script" is a textual document that details the scenes, characters, dialogue, and action of a movie.
[1148] A "generative model" is a machine learning model for automatically generating images and video sequences based on input text data.
[1149] "Video" refers to visual data in which objects or scenes move over time, and in the case of film, includes scripted sequences.
[1150] "Feedback" refers to opinions and requests for corrections that a user provides regarding a generated video.
[1151] A "terminal" is an electronic device that a user operates, and includes a smartphone, tablet, computer, etc.
[1152] A "head-mounted display" is a display device that is worn by the user on the head, and is used to provide VR environments, etc.
[1153] An "immersive viewing experience" is a visual and auditory experience that immerses the viewer as if they are inside the movie.
[1154] This invention is a system that inputs a movie script in text format and automatically generates movie images based on it. An embodiment for realizing this system is described below.
[1155] First, the user uses the device to input the script for the movie. The device is provided with a dedicated application or web interface, and the user can write detailed descriptions and dialogue for each scene. For example, if the user wants to input a scene in which "the protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the user can input it in the following text format:
[1156] "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[1157] The input script is converted into a standard data format (e.g., JSON) and sent to the server using an HTTP POST request. The server then inputs the received script data into a generative model to generate the first footage of the movie. This generative model automatically generates images and video sequences based on the text data entered by the user. For example, a machine learning library such as TensorFlow or PyTorch in Python can be used to build the generative model.
[1158] The generated initial video is sent from the server to the device, where the user can review it. The user watches the video and enters feedback if there are any complaints or requests for corrections. The feedback is again converted into a standard data format and sent to the server using an HTTP POST request. If the user wants to provide feedback such as "I'd like the rain to be a little stronger," they can enter it in the following text format:
[1159] "I want the rain to be a little stronger."
[1160] The server then applies the feedback data to the generative model and regenerates a new image. During this process, the parameters of the generative model are adjusted and details are modified according to the user's requests. The regenerated image is then sent back to the device for the user to review.
[1161] Once the user is satisfied with the final video, it is stored on the server and made available for download. The completed video is provided as a download link, and the user can obtain the video file by clicking the link. It is also possible to provide an immersive viewing experience using a smartphone or head-mounted display. For example, head-mounted displays such as Oculus Rift and HTC Vive can be used.
[1162] This system allows users to efficiently generate video that reflects their own intentions while enjoying a realistic viewing experience.
[1163] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1164] Step 1:
[1165] The user uses the device to input a movie script. The input is in text format, including detailed descriptions and dialogue for each scene. The device receives this input and converts it into a standard data format (such as JSON).
[1166] Step 2:
[1167] The device sends the converted script data to the server using an HTTP POST request. Specifically, a JSON object containing the script data is sent, and the server receives it.
[1168] Step 3:
[1169] The server inputs the received script data into a generative model to generate the first footage of the movie. This generation process uses machine learning libraries (such as TensorFlow and PyTorch) to generate scene images and video sequences from the input text data. The generated footage is stored on the server.
[1170] Step 4:
[1171] The server sends the generated initial video as an HTTP response to the terminal, which receives the video data and plays it on the device to provide it to the user.
[1172] Step 5:
[1173] The user can then use the device to view the generated video and enter feedback. This feedback is in the form of text, including suggestions for improvement or correction, such as "I'd like the rain to be a little stronger." The device then converts this feedback back into JSON format.
[1174] Step 6:
[1175] The terminal sends the converted feedback data to the server using an HTTP POST request, and the server retrieves the received feedback data.
[1176] Step 7:
[1177] The server then applies the feedback data to the generative model and regenerates a new video. At this time, the parameters of the generative model are adjusted based on the user's feedback, and details are corrected. The regenerated video is then stored back on the server.
[1178] Step 8:
[1179] The server sends the regenerated video as an HTTP response to the device, which receives the video data and plays it on the device for presentation to the user. This process is repeated until the user is finally satisfied.
[1180] Step 9:
[1181] When the final video is generated to the user's satisfaction, the server saves the final video and generates a downloadable link, which the device provides to the user, who clicks on the link to obtain the video file.
[1182] Step 10:
[1183] Users can view the final video using a device or a head-mounted display, creating an immersive viewing experience. For example, they can enjoy movies in a 3D environment using devices such as Oculus Rift or HTC Vive.
[1184] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1185] This invention relates to a system that automatically generates movie images based on a movie script input by a user using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions. An embodiment of this system will be described below.
[1186] 1. Entering the script
[1187] The user uses a device to input a movie script into a dedicated application or web interface, writing detailed descriptions and dialogue for each scene in text format.
[1188] 2. Submit your script
[1189] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request.
[1190] 3. Initial video generation
[1191] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data.
[1192] 4. First footage review and emotion analysis
[1193] The server encodes the generated initial video data and sends it to the device. When the user watches the video, the emotion engine analyzes the user's facial expressions and voice to recognize the user's emotions.
[1194] 5. Automatic feedback generation
[1195] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "The rain intensity in scene 3 needs to be increased" is automatically generated.
[1196] 6. Submitting Feedback
[1197] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again using an HTTP POST request.
[1198] 7. Brush up the video
[1199] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[1200] 8. Checking the regenerated video and making final adjustments
[1201] The regenerated image is sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[1202] 9. Storage and provision of final footage
[1203] Once the user is satisfied with the final video, the server stores it and generates a downloadable link, which the server provides to the device, allowing the user to obtain the completed video file.
[1204] Specific examples
[1205] For example, suppose a user (scriptwriter) inputs a scene such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user inputs feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it automatically generates feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[1206] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[1207] The processing flow will be explained below.
[1208] Step 1:
[1209] The user uses a device to input a movie script into a dedicated application or web interface. Specifically, the user writes detailed descriptions and dialogue for each scene in text format and saves it as script data.
[1210] Step 2:
[1211] The terminal converts the input script into JSON format and sends it to the server using an HTTP POST request, at which point the script data is transferred to the server via the network.
[1212] Step 3:
[1213] The server analyzes the received script data and inputs it into a generative model, which then automatically generates the first movie footage based on the script text data. This process utilizes natural language processing and image generation technology.
[1214] Step 4:
[1215] The server encodes the generated initial video data and sends it to the device, which is provided with a link to the video file, and the user clicks on the link to watch the video.
[1216] Step 5:
[1217] Users access the provided link using their device and watch the initial video. While reviewing the video, they can enter feedback on corrections and improvements they would like to see. Feedback is written in a specific format, such as "Make the rain stronger in scene 3."
[1218] Step 6:
[1219] When users provide feedback, the emotion engine analyzes their facial expressions and voice to recognize their emotions. The emotion engine collects emotion data in real time while the user is watching the video.
[1220] Step 7:
[1221] Based on the emotional data analyzed by the emotion engine, feedback is automatically generated as needed. For example, if the user looks dissatisfied, specific feedback such as "You need to increase the intensity of the rain in scene 3" is automatically generated.
[1222] Step 8:
[1223] The device converts the user feedback and feedback automatically generated by the emotion engine into JSON format and sends it back to the server using an HTTP POST request.
[1224] Step 9:
[1225] The server analyzes the received feedback data and applies it to the generative model, which then regenerates the video based on the feedback. In this process, the generative process is adjusted based on the user's emotional data.
[1226] Step 10:
[1227] The regenerated image is then sent to the device. The user checks the regenerated image and repeats steps 5 to 9 until they are satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal image is generated.
[1228] Step 11:
[1229] The final video that the user is satisfied with is saved by the server and a downloadable link is generated. The server then provides this link to the user's device, allowing the user to retrieve the completed video file. This process effectively generates an optimized movie video that takes the user's emotions into account.
[1230] Example 2
[1231] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1232] The traditional film production process, from scriptwriting to video generation and editing, is complex and time-consuming, often requiring large resources and advanced technical skills. It is also difficult to reflect user emotions in real time, making it impossible to optimize video content based on viewer emotions. Furthermore, there are limited ways to efficiently incorporate feedback, making the process of improving video quality cumbersome. There was a need for a system that could solve these problems and efficiently generate high-quality video.
[1233] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes a means for a user to input a movie script using a terminal, a means for converting the input script into JSON format and sending it to the server, a means for the server to convert the script into movie footage using a generative model, a means for the user to check the generated footage and analyze the user's emotions using an emotion engine, a means for automatically generating feedback based on the emotion analysis, a means for converting user or automatically generated feedback into JSON format and sending it to the server, a means for reflecting the input feedback in the generative model and regenerating the footage, and a means for saving and providing the final movie footage. This enables optimization of the footage based on the user's emotions and an efficient feedback process.
[1234] A "user" is a user of the system who inputs a movie script, reviews the generated footage, and provides feedback.
[1235] A "terminal" is an information processing device that a user uses to input a movie script and feedback, and specifically includes a PC, smartphone, tablet, etc.
[1236] A "screenplay" is a document that describes in text form the details and dialogue of each scene of a movie.
[1237] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that represents data as key-value pairs.
[1238] An "HTTP POST request" is a type of HyperText Transfer Protocol used to send data to a server.
[1239] A "server" is a computer system that receives and processes scripts and feedback submitted by users.
[1240] A "generative model" is a machine learning model for automatically generating movie footage from input script data, and includes, for example, generative AI models (e.g., DALL-E and GPT-4).
[1241] "Movie footage" is visual content that is automatically generated by a generative model based on a script.
[1242] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice and recognizes their emotions.
[1243] "Feedback" refers to evaluations of videos and suggestions for improvement provided by the user or the emotion engine.
[1244] "Encoding" is the process of compressing and converting video data into a particular format.
[1245] The "final footage" is the completed film footage that the user has provided feedback on repeatedly and is finally satisfied with.
[1246] "Storage" refers to storing the final video generated in a data storage system.
[1247] A "download link" is a web address provided to users to obtain the final video.
[1248] This invention relates to a system that automatically generates images using a generative AI model based on a movie script entered by the user, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[1249] 1. Entering the script
[1250] The user uses a device to input a movie script into a dedicated application or web interface. The user writes detailed descriptions of each scene and lines in text format. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining."
[1251] 2. Submit your script
[1252] The terminal converts the input script into JSON format using a script parsing module that converts text input into structured data, and then sends the converted JSON data to the server using an HTTP POST request.
[1253] 3. Initial video generation
[1254] The server analyzes the received JSON data and inputs it into a generative AI model (e.g., DALL-E or GPT-4). The generative AI model generates movie images based on the text data of the script. The generated image data is first temporarily saved.
[1255] 4. First footage review and emotion analysis
[1256] The server encodes the generated video and sends it to the device. When the user watches the video, the device's camera and microphone record the user's facial expressions and voice, and the emotion engine analyzes the data. The emotion engine recognizes the user's emotions in real time and determines emotions such as joy, sadness, and surprise.
[1257] 5. Automatic feedback generation
[1258] The server automatically generates feedback based on the analysis results of the emotion engine. For example, if the user looks dissatisfied while watching a video, the server generates specific feedback such as, "The rain intensity in scene 3 needs to be increased."
[1259] 6. Submitting Feedback
[1260] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server as an HTTP POST request.
[1261] 7. Brush up the video
[1262] The server re-inputs the received feedback into the generative model to regenerate the video. The generative model then modifies the video based on the feedback and outputs the refined video data.
[1263] 8. Checking the regenerated video and making final adjustments
[1264] The regenerated image is then sent back to the device, and the feedback and regeneration process is repeated until the user is satisfied. The emotion engine continues to monitor the user's emotions during this process and generates further feedback as needed.
[1265] 9. Storage and provision of final footage
[1266] Once the user is satisfied with the final video, it is stored on the server and a downloadable link is generated. The user can obtain this link via their device and download the final video file.
[1267] For example, if a user inputs a prompt such as "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining," the server uses the generative AI model to sequentially generate each scene. If the user reviews the footage and provides feedback such as "I would like the rain to be stronger in Scene 3," or if the emotion engine detects dissatisfaction from the user's facial expression, the server will automatically generate feedback such as "The atmosphere in Scene 3 needs to be more dramatic." In this way, the user can generate movie footage that faithfully reflects their intentions in an efficient and emotion-optimized manner.
[1268] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1269] Step 1:
[1270] A user uses a device to input a movie script into a dedicated application or web interface. The input is text data including details and dialogue for each scene of the movie. For example, a scene might be input as follows: "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The input is a text script sent to the device. The output is the input script stored in the device's memory.
[1271] Step 2:
[1272] The terminal converts the input script into JSON format. Specifically, the script analysis module analyzes the text data and generates structured JSON data. This JSON data is then sent to the server using an HTTP POST request. The text script is used as input, and the data converted to JSON format is generated as output and sent to the server.
[1273] Step 3:
[1274] The server analyzes the received JSON data and inputs it into a generative AI model. The generative model (e.g., DALL-E or GPT-4) initially generates movie footage based on the text data of the script. The JSON data received by the server is used as input, and the initially generated video data is generated as output and temporarily stored on the server.
[1275] Step 4:
[1276] The server encodes the generated video data and sends it to the device via HTTP. Video encoding technologies such as H.264 are used for encoding. When the user watches the video on the device, the device's camera and microphone record the user's facial expressions and voice in real time, and this data is analyzed by the emotion engine. The video data sent from the server is received as input, and the user's emotion data is generated as output.
[1277] Step 5:
[1278] The emotion engine automatically generates feedback based on the analyzed emotional data. For example, specific feedback is created based on the emotions inferred from the user's facial expressions and voice, such as "The rain intensity in scene 3 needs to be reduced." The user's emotional data is used as input, and the text feedback data is generated as output.
[1279] Step 6:
[1280] The device converts the user feedback and automatically generated feedback into JSON format and sends it to the server again via an HTTP POST request. The user feedback text is used as input, and the data converted into JSON format is generated as output and sent to the server.
[1281] Step 7:
[1282] The server re-inputs the received feedback data into the generative model and regenerates the video. The generative model outputs a modified video based on the previous video and the new feedback data. The feedback data and the initial video data are used as input, and the modified video data is generated as output and temporarily stored on the server.
[1283] Step 8:
[1284] The regenerated video is sent back to the device for the user to review. The emotion engine continues to analyze the user's reactions and generates new feedback as needed. The modified video data is used as input, and the user's emotion data and new feedback are generated as output.
[1285] Step 9:
[1286] The video data that the user is finally satisfied with is stored on the server. The server generates and provides a download link for the final video to the user. The final video data is used as input, and a download link URL is generated as output and sent to the device.
[1287] (Application example 2)
[1288] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1289] Not only does filmmaking require a significant amount of time and effort, but it is also difficult to create footage that reflects the user's intentions and emotions. In particular, generating a satisfactory video based on a user-provided script requires a lot of trial and error, which hinders the creative process. This challenge stems from the lack of a means to quickly and efficiently generate high-quality footage and achieve results that satisfy the user.
[1290] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for inputting a movie script, means for converting the input script into movie images using a generative model, means for checking the generated images and inputting feedback, means for reflecting the input feedback in the generative model and regenerating images, means for saving and providing the final movie images, means for analyzing the user's facial expressions and voice to acquire emotional data, and means for automatically generating feedback based on the acquired emotional data. This makes it possible to quickly and efficiently generate high-quality images that reflect the user's intentions and emotions.
[1291] A "movie script" is a text document that describes the content of a movie, the dialogue of the characters, and details of the scenes.
[1292] A "generative model" is an artificial intelligence (AI) system that automatically generates images based on input text data.
[1293] "Feedback" refers to the user's evaluation of the generated video and requests for corrections.
[1294] "Regeneration" is the process of regenerating a previously generated image based on feedback and emotional data.
[1295] "Emotional data" is emotional information obtained by analyzing the user's facial expressions and voice.
[1296] "Automatic generation" refers to the process by which the system automatically generates feedback.
[1297] "Terminal" means a user's device used to input scripts and manage feedback.
[1298] "Interface" refers to the operating screen and input form that allows users to input scripts and feedback via a terminal.
[1299] "Storage" refers to the act of storing the final generated video as data.
[1300] This invention relates to a system that allows a user to input a movie script and automatically generates movie images based on that script using a generative model, and further optimizes the image generation process by combining it with an emotion engine that recognizes the user's emotions.
[1301] Hardware and Software
[1302] First, a user inputs a movie script using a device such as a smartphone. The device has a dedicated application or web interface through which the user can write each scene and line in detail.
[1303] The terminal converts the input script into JSON format and sends it using an HTTP POST request to a server, which is built using a server-side framework such as Flask or Django and has a generative model for analyzing the received script data.
[1304] Image Generation
[1305] The server inputs the received script data into a generative model (e.g., GPT-4) to automatically generate the first movie footage. This video generation process uses the model's deep learning technology. The generated video data is then encoded and sent back to the device.
[1306] Sentiment Analysis and Feedback
[1307] When a user watches a video, an emotion engine (e.g., OpenCV, Emotion API) runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone. The emotion engine recognizes the user's emotions and sends that data to the server. The server then automatically generates feedback based on the emotion data.
[1308] Video Regeneration
[1309] The feedback obtained by the emotion engine and correction requests entered by the user are sent back to the server. The server analyzes this feedback, adjusts the generative model, and regenerates the video. By repeating this process, the user can refine the video until they are satisfied.
[1310] Storage and provision of final footage
[1311] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[1312] Specific examples
[1313] For example, suppose a user (scriptwriter) inputs a script such as, "The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it is raining." The server sends this scenario to the generative model, which generates each scene sequentially. After viewing the first footage, the user may provide feedback such as, "I would like the rain to be stronger in scene 3." If the emotion engine detects dissatisfaction from the user's facial expression, it will automatically generate feedback such as, "The atmosphere in scene 3 needs to be more dramatic."
[1314] Prompt Sentence Examples
[1315] "Enter a script: The protagonist wakes up in the morning, opens the curtains, and looks out the window to see that it's raining."
[1316] "Feedback: The rain in Scene 3 needs to be stronger."
[1317] "Final video generation completed. Download link: https: / / example.com / download / final_video.mp4"
[1318] In this way, the user (screenwriter) can realize a system that can generate movie images that faithfully reflect his or her intentions in a more efficient and emotionally optimized manner by using the emotion engine.
[1319] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1320] Step 1:
[1321] The user uses the device to input a movie script. The script is entered in text format through a dedicated application or a web interface. After input, the device converts the script into JSON format. At this time, the script data is structured by scene and line.
[1322] Input: The script text that the user types into the terminal.
[1323] Data processing: Convert text to JSON format
[1324] Output: Structured JSON data
[1325] Step 2:
[1326] The converted script data is sent to the server using an HTTP POST request. The server is built using a server-side framework such as Flask or Django. The server parses the received JSON data and extracts the script content.
[1327] Input: JSON data sent from the terminal
[1328] Data processing: Parsing JSON data
[1329] Output: Parsed script data
[1330] Step 3:
[1331] The server inputs the analyzed script data into a generative model (e.g., GPT-4) to generate the first movie footage. The generative model uses deep learning techniques to generate video data from the text data. The generated footage is then encoded and converted into a viewable format.
[1332] Input: Parsed script data
[1333] Data processing: Video generation and encoding using generative models
[1334] Output: Viewable video file
[1335] Step 4:
[1336] The generated video data is then sent to the device using an HTTP request. The user watches this video. While watching, the emotion engine runs, analyzing the user's facial expressions and voice in real time using the device's built-in camera and microphone.
[1337] Input: Viewable video file
[1338] Data processing: Video file distribution and sentiment analysis
[1339] Output: User emotion data
[1340] Step 5:
[1341] The emotion engine transmits emotional data obtained from the user's facial expressions and voice to the server, which analyzes this data and automatically generates feedback as needed. Feedback is expressed as specific correction requests or improvement suggestions.
[1342] Input: User emotion data
[1343] Data processing: analyzing emotional data and generating feedback
[1344] Output: Feedback data
[1345] Step 6:
[1346] The generated feedback data is then sent back to the server, which analyzes the feedback, adjusts the generative model based on it, and regenerates the video. The generation process is then optimized to reflect the feedback.
[1347] Input: Feedback data
[1348] Data processing: Reflecting feedback and regenerating images
[1349] Output: Regenerated video file
[1350] Step 7:
[1351] The regenerated video file is sent back to the device, and this process is repeated until the user is satisfied. The emotion engine constantly monitors the user's emotions and makes adjustments until the optimal video is generated.
[1352] Input: Regenerated video file
[1353] Data processing: Distribution and reconfirmation of video files
[1354] Output: The final video file that I was satisfied with
[1355] Step 8:
[1356] Once a satisfactory final video is generated, the server stores it and generates a downloadable link that can be provided to the device, allowing the user to retrieve the completed video file.
[1357] Input: The final video file you are satisfied with
[1358] Data processing: Saving video files and creating links
[1359] Output: Download link
[1360] Through the above steps, it becomes possible to efficiently generate high-quality video that faithfully reflects the user's intentions and emotions.
[1361] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1362] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1363] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1364] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1365] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1366] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1367] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1368] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1369] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1370] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1371] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1372] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1373] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1374] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1375] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1376] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1377] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1378] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1379] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1380] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1381] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1382] The following is further disclosed regarding the above embodiment.
[1383] (Claim 1)
[1384] a means for inputting a motion picture script;
[1385] A means for converting an input script into movie footage using a generative model;
[1386] A means for reviewing the generated video and providing feedback;
[1387] A means for reflecting the input feedback in the generative model and regenerating the image;
[1388] a means of preserving and providing the final footage of the film;
[1389] A system including:
[1390] (Claim 2)
[1391] 10. The system of claim 1, further comprising: means for adjusting the generative process as the generative model reproduces the video based on the feedback.
[1392] (Claim 3)
[1393] 10. The system of claim 1, further comprising means for providing an interface through which a user can input a script and feedback via a terminal.
[1394] "Example 1"
[1395] (Claim 1)
[1396] a means for a user to input a movie script in text form using a computer terminal;
[1397] means for converting the input script into a standard data format and transmitting it to a server via a network;
[1398] A means for the server to generate a first-run video of a movie based on script data using the generative model;
[1399] a means for transmitting the generated initial video to a user terminal, and for the user to check the video and input feedback;
[1400] a means for transmitting feedback input by a user to a server and regenerating the video by reflecting the feedback in a generative model;
[1401] A means for storing the final generated video and providing it in a format that can be downloaded by users;
[1402] A system including:
[1403] (Claim 2)
[1404] 10. The system of claim 1, further comprising: means for adjusting the generative process as the generative model reproduces the video based on the feedback.
[1405] (Claim 3)
[1406] 10. The system of claim 1, further comprising means for providing an interface through which a user can input a script and feedback via a computer terminal.
[1407] "Application Example 1"
[1408] (Claim 1)
[1409] a means for inputting a motion picture script;
[1410] A means for converting an input script into movie footage using a generative model;
[1411] A means for reviewing the generated video and providing feedback;
[1412] A means for reflecting the input feedback in the generative model and regenerating the image;
[1413] a means of preserving and providing the final footage of the film;
[1414] a means for viewing the video using the device;
[1415] A means of providing an immersive viewing experience using a head-mounted display;
[1416] A system including:
[1417] (Claim 2)
[1418] 10. The system of claim 1, further comprising: means for adjusting the generative process as the generative model reproduces the video based on the feedback.
[1419] (Claim 3)
[1420] 10. The system of claim 1, further comprising means for providing an interface through which a user can input a script and feedback via a terminal.
[1421] "Example 2: Combining Emotion Engines"
[1422] (Claim 1)
[1423] a means for a user to input a movie script using a terminal;
[1424] A means to convert the input script into JSON format and send it to the server,
[1425] a means for the server to use the generative model to convert the screenplay into cinematic footage;
[1426] A means for allowing a user to check the generated video and analyze the user's emotions using an emotion engine;
[1427] a means for automatically generating feedback based on sentiment analysis;
[1428] A means of converting user or automatically generated feedback into JSON format and sending it to the server;
[1429] A means for reflecting the input feedback in the generative model and regenerating the image;
[1430] a means of preserving and providing the final footage of the film;
[1431] A system including:
[1432] (Claim 2)
[1433] 10. The system of claim 1, further comprising means for adjusting the generative process based on the sentiment analysis data when the generative model regenerates the video in response to the feedback.
[1434] (Claim 3)
[1435] 10. The system of claim 1, further comprising means for providing an interface through which a user can input a movie script and feedback via a terminal.
[1436] "Application example 2 when combining emotion engines"
[1437] (Claim 1)
[1438] a means for inputting a motion picture script;
[1439] A means for converting an input script into movie footage using a generative model;
[1440] A means for reviewing the generated video and providing feedback;
[1441] A means for reflecting the input feedback in the generative model and regenerating the image;
[1442] a means of preserving and providing the final footage of the film;
[1443] A means for acquiring emotion data by analyzing the user's facial expressions and voice;
[1444] A means for automatically generating feedback based on the acquired emotional data;
[1445] A system including:
[1446] (Claim 2)
[1447] 10. The system of claim 1, further comprising: means for adjusting the generative process as the generative model reproduces the video based on the feedback.
[1448] (Claim 3)
[1449] 10. The system of claim 1, further comprising means for providing an interface through which a user can input a script and feedback via a terminal. [Explanation of symbols]
[1450] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for inputting a motion picture script; A means for converting an input script into movie footage using a generative model; A means for reviewing the generated video and providing feedback; A means for reflecting the input feedback in the generative model and regenerating the image; a means of preserving and providing the final footage of the film; A system including:
2. The system of claim 1 , further comprising: means for adjusting the generative process as the generative model reproduces the image based on the feedback.
3. The system of claim 1 , further comprising means for providing an interface through which a user can input a script and feedback via a terminal.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A