Video generation system, video generation method, and video generation program

The image generation system addresses the burden of preparing two-dimensional drawings by directly generating videos from user queries using AI models, improving usability and efficiency.

JP2025162283AActive Publication Date: 2025-10-27SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024065475
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-15
Publication Date
2025-10-27
Estimated Expiration
2044-04-15

AI Technical Summary

Technical Problem

Conventional image generation technologies require users to prepare two-dimensional line drawings, which is burdensome and limits usability, especially when users cannot provide such drawings.

Method used

An image generation system that acquires user queries to generate video data, utilizing an acquisition unit, scenario generation unit, and video data acquisition unit to create 3D data and video based on user inputs, employing AI models like LLMs for scenario and code generation.

Benefits of technology

Reduces user burden by generating videos directly from queries, enhancing usability and enabling video data acquisition without the need for complex image preparation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025162283000001_ABST
    Figure 2025162283000001_ABST
Patent Text Reader

Abstract

To acquire video data in accordance with an input query from a user.SOLUTION: A video generation system according to the present disclosure comprises: an acquisition unit for acquiring, from a user, an input query relating to video generation; a scenario generation unit for generating scenario data relating to the video generation on the basis of the input query; a code generation unit for generating a code to configure 3D data on the basis of the scenario data; and a video data acquisition unit for acquiring video data on the basis of the code.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an image generation system, an image generation method, and an image generation program. [Background technology]

[0002] There are technologies available for automatically generating images (also called "videos"). For example, there is a technology available for estimating the three-dimensional posture of a virtual person from a two-dimensional line drawing and generating a video (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2003-058906 Summary of the Invention [Problem to be solved by the invention]

[0004] However, there is room for improvement in the conventional technology. For example, in the conventional technology, two-dimensional line drawings, i.e., images, are required to generate a video. Preparing images such as two-dimensional line drawings places a heavy burden on the user, and it is difficult to generate a video when the user is unable to prepare an image. Therefore, there is a need to provide a video generation service that places less burden on the user and has high usability, and there is a need, for example, to obtain video data in response to a query input from a user.

[0005] Therefore, the present disclosure proposes an image generation system, an image generation method, and an image generation program that are capable of acquiring video data in response to a query input by a user. [Means for solving the problem]

[0006] In order to solve the above problems, one embodiment of a video generation system according to the present disclosure includes an acquisition unit that acquires an input query related to video generation from a user, a scenario generation unit that generates scenario data related to video generation based on the input query, a code generation unit that generates code for constructing 3D data based on the scenario data, and a video data acquisition unit that acquires video data based on the code. [Brief explanation of the drawings]

[0007] [Figure 1] FIG. 1 is a diagram illustrating an example of an image generation system according to the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of an image generation system according to the present disclosure. [Figure 3] FIG. 10 is a diagram showing an example of the flow of video generation processing according to the present disclosure. [Figure 4] FIG. 10 is a diagram showing another example of the flow of the video generation process of the present disclosure. [Figure 5] FIG. 10 is a diagram showing an example of the flow of an evaluation process according to the present disclosure. [Figure 6] FIG. 10 is a diagram illustrating an example of a process for generating scenario generation information. [Figure 7] FIG. 10 is a diagram illustrating an example of a scenario data generation process. [Figure 8] FIG. 10 is a diagram illustrating an example of a process for generating information for code generation. [Figure 9] FIG. 10 is a diagram illustrating an example of image quality improvement processing. [Figure 10] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 11] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 12] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 13] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 14] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 15] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 16] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 17] FIG. 10 is a diagram illustrating an example of a process for generating information for sound generation. [Figure 18] FIG. 1 is a diagram illustrating an example of a USD file. [Figure 19] 10 is a flowchart showing a processing procedure executed by the video production system. [Figure 20] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 21] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 22] FIG. 10 is a diagram illustrating an example of a user interface. [Figure 23] FIG. 10 illustrates an example of an editing process. [Figure 24] FIG. 10 illustrates an example of an editing process. [Figure 25] FIG. 10 illustrates an example of an editing process. [Figure 26] FIG. 10 illustrates an example of an editing process. [Figure 27] FIG. 10 illustrates an example of an editing process. [Figure 28] FIG. 10 illustrates an example of an editing process. [Figure 29] FIG. 10 illustrates an example of an editing process. [Figure 30] FIG. 10 is a diagram showing an example of a confirmation operation during editing. [Figure 31] 10 is a flowchart showing a processing procedure executed by the video production system. [Figure 32] FIG. 10 is a diagram illustrating an example of an evaluation process according to a range selection. [Figure 33] FIG. 10 is a conceptual diagram showing an example of processing according to the depth of field. [Figure 34] FIG. 10 is a diagram showing an example of an image during a confirmation operation. [Figure 35] FIG. 10 is a diagram showing an example of highlighting a changed portion. [Figure 36] FIG. 10 is a diagram showing an example of presentation of the relationship between cuts. [Figure 37] FIG. 10 is a diagram illustrating an example of object selection. [Figure 38] FIG. 10 is a diagram illustrating an example of object selection. [Figure 39] FIG. 10 is a diagram illustrating an example of processing using reference data. [Figure 40] 10 is a flowchart showing a flow of processing in response to a user operation. [Figure 41] FIG. 2 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of the information processing device. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the image generation system, image generation method, and image generation program according to the present application are not limited to these embodiments. Furthermore, in the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The present disclosure will be described in the following order: 1. Embodiment 1-1. Overview of the configuration of the video generation system of the present disclosure 1-2. Processing by the image generation system of the present disclosure 1-3.User Interface 1-4. Processing example 1-4-1. Sound generation example 1-4-2. Example of text logo generation 1-4-3. Re-learning example 1-4-4.USD update example 1-4-5. Evaluation example 1-4-6. Example of multiple cut selection 1-4-7.Examples of using time information 1-4-8. Response examples according to the degree of verbalization 1-4-9. Examples of using qualitative values 1-4-10. Example of rendering process during input 1-4-11. Example of checking work when editing 1-4-12. Processing examples according to range selection 1-4-13.Example of processing according to depth of field 1-4-14. Example of playback processing during confirmation work 1-4-15. Highlighting example 1-4-16. Example of showing the relationship between cuts 1-4-17.Example of object selection 1-4-18.Examples of using 3D models 1-4-19.Examples of using reference data 1-4-20.Examples of the benefits of having 3D data 1-5. Example of processing flow from the user's perspective 1-6.About AI models 2. Other embodiments 2-1.Other configuration examples 2-2.Other 3. Effects of this disclosure 4. Hardware Configuration

[0010] <1. Embodiment> <1-1. Overview of the configuration of the video generation system of the present disclosure> Fig. 1 is a diagram illustrating an example of an image generation system according to the present disclosure. The image generation system 1 includes an image generation module 100, an information acquisition module 200, a sensor unit 300, and a client UI display unit 400. Although Fig. 1 illustrates only one of each component, the image generation system 1 may include multiple image generation modules 100, multiple information acquisition modules 200, multiple sensor units 300, and multiple client UI display units 400.

[0011] First, we will explain the configuration of the image generation module 100 that performs image generation processing. The image generation module 100 has an input text analysis unit 110, a sensor analysis unit 120, a prompt etc. generation unit 130, an image generation unit 140, a sound generation unit 150, a text / logo generation unit 160, a composite editing unit 170, an evaluation unit 180, a client UI module 190, etc.

[0012] The input text analysis unit 110 analyzes input text. For example, the input text analysis unit 110 analyzes text input from the information acquisition module 200. The sensor analysis unit 120 analyzes input sensor information. For example, the sensor analysis unit 120 analyzes sensor information acquired from the information acquisition module 200.

[0013] The prompt etc. generation unit 130 generates various information required for generating video (moving image) including prompts etc. to be input to an AI (Artificial Intelligence) model (simply referred to as a "model"), which is a machine learning model described below. For example, the prompt etc. generation unit 130 generates a prompt using a user input and a pre-saved prompt (template, etc.). Note that a prompt is merely one example of information to be input to an AI model (model input information). The model input information to be input to an AI model is not limited to a prompt, and any format of model input information can be employed. Therefore, the "prompt etc. generation unit" may be read as a "model input information etc. generation unit." In FIG. 1, the prompt etc. generation unit 130 includes a scenario-oriented generation unit 131, a video-oriented generation unit 132, a sound-oriented generation unit 133, and a text / logo-oriented generation unit 134.

[0014] The scenario-oriented generation unit 131 generates various information related to the generation of a scenario. The scenario-oriented generation unit 131 generates input information to be input to a model that outputs a scenario. For example, the scenario-oriented generation unit 131 is a scenario generation unit that generates scenario data related to video generation based on an input query. For example, the scenario-oriented generation unit 131 is a first output unit that outputs scenario generation information used by the scenario generation unit to generate scenario data based on an input query.

[0015] The video generation unit 132 generates various information related to the generation of videos. The video generation unit 132 generates input information to be input to a model that outputs code for configuring 3D (three-dimensional) data. For example, the video generation unit 132 is a code generation unit that generates code for configuring 3D data based on scenario data. For example, the video generation unit 132 is a second output unit that outputs code generation information used by the code generation unit to generate code for configuring 3D data based on scenario data.

[0016] The sound generation unit 133 generates various information related to the generation of sound information (audio information). The sound generation unit 133 generates input information to be input to a model that outputs sound. The text / logo generation unit 134 generates various information related to the generation of text and logos. The text / logo generation unit 134 generates input information to be input to a model that outputs at least one of text and logos.

[0017] The video generation unit 140 executes processing related to video generation. The video generation unit 140 is a video acquisition unit that acquires video data based on code. The video generation unit 140 generates video using various information generated by the prompt etc. generation unit 130. For example, the video generation unit 140 is a video generation unit that generates video data based on code. Note that the video generation unit 140 may acquire video data in any manner. For example, the video generation unit 140 may transmit data used to generate the video data to an external service providing device (such as a vendor) that provides a video data generation service, and acquire the video data by receiving the video data generated by the service providing device from the service providing device. In FIG. 1, the video generation unit 140 includes a USD generation unit 141, a rendering unit 142, and a video refinement unit 143.

[0018] The USD generation unit 141 generates various information related to a USD (Universal Scene Description). For example, the USD generation unit 141 generates USD-Python or the like using an AI model such as a Large Language Model (hereinafter also referred to as "LLM"), using a prompt obtained by video prompt generation in the video generation unit 132.

[0019] The rendering unit 142 executes various processes related to rendering. The rendering unit 142 executes a process of rendering the USD generated by the USD generation unit 141.

[0020] The image refinement unit 143 executes various processes for refining the image. The rendering unit 142 improves the quality of the generated image through image refinement processing. For example, the image refinement unit 143 is an image quality improvement unit that executes image quality improvement processing to improve the image quality of video data.

[0021] The sound generation unit 150 executes a process of generating sounds. The sound generation unit 150 uses the prompts obtained by the sound-oriented prompt generation in the sound-oriented generation unit 133 to generate sound information such as background music (BGM), sound effects (SE), narration, and dialogue, based on an AI model such as a Contrastive Learning Model.

[0022] The text / logo generation unit 160 executes a process for generating at least one of text and a logo. The text / logo generation unit 160 generates at least one of text and a logo using the information generated by the text / logo generation unit 134.

[0023] With the above-described configuration, the image generation module 100 generates a prompt for scenario generation by combining a pre-stored prompt with a user's input. The image generation module 100 generates a scenario by inputting the generated prompt into an AI model such as an LLM. The image generation module 100 also generates prompts for generating images, sounds, and text / logos from the generated scenario. The image generation module 100 performs image generation, sound generation, and text / logo generation using prompts, scenarios, etc. for generating images, sounds, and text / logos.

[0024] The composite editing unit 170 executes processes related to editing. For example, the composite editing unit 170 executes a process of combining (combining) the generated video, sound, and text / logo into one video.

[0025] The evaluation unit 180 executes an evaluation process for evaluating various targets. The evaluation unit 180 evaluates the information generated by the above-described configuration. For example, the evaluation unit 180 generates information indicating an evaluation of at least one of the scenario data and the video data.

[0026] The client UI module 190 executes processing related to output on a UI (User Interface) on the client side. For example, the client UI module 190 generates various information related to output on the UI on the client side. In this case, the client UI module 190 executes processing to generate a UI to be displayed on the user side. The client UI module 190 generates various information to be displayed on the client UI display unit 400.

[0027] Furthermore, the information acquisition module 200 acquires various types of information. The information acquisition module 200 has an input text acquisition unit 210, a sensor acquisition unit 220, etc. The input text acquisition unit 210 acquires text information input via a keyboard 320 or a microphone 330. For example, the input text acquisition unit 210 acquires text information input by a user via the keyboard 320 or the microphone 330. For example, the input text acquisition unit 210 is an acquisition unit that acquires an input query related to video generation from a user.

[0028] The sensor acquisition unit 220 acquires information (also referred to as "sensor information") detected by a sensor such as a camera 340 or a motion capture device. The information acquisition module 200 provides (transmits) the acquired various pieces of information to the image generation module 100. Note that the information acquisition module 200 may be integrated with the image generation module 100.

[0029] The sensor unit 300 has various sensors. The sensor unit 300 senses user input. The sensor unit 300 accepts user operations. For example, the sensor unit 300 is a reception unit that accepts video editing-related operations from the user. For example, the sensor unit 300 has a mouse 310, a keyboard 320, a microphone 330, a camera 340, an IMU 350 which is an inertial measurement unit, and the like. In this way, the sensor unit 300 includes, in addition to the mouse 310 and the keyboard 320, a user terminal (such as a smartphone) equipped with the microphone 330, the camera 340, and the IMU 350, and sensors such as motion capture, and senses user input.

[0030] The client UI display unit 400 displays various information to be presented to the client (user). The client UI display unit 400 displays the UI generated by the client UI module 190 on a display (display device) of the client. For example, the client UI display unit 400 is a display control unit that displays a storyboard based on scenario data. The storyboard is configured to display video data for each cut of the video.

[0031] The video generation system 1 may have a hardware configuration as shown in Fig. 2. Fig. 2 is a diagram illustrating an example of a hardware configuration of the video generation system of the present disclosure. In Fig. 2, the video generation system 1 has, as its hardware configuration, a cloud-side computer 10, a client-side computer 20, and a camera / sensor 30 including various sensors such as a camera. The video generation system 1 may also include an information providing device (computer) that provides the computer 10 with information resources 40 such as learning data and an AI model 50.

[0032] 2 is merely an example, and any hardware configuration can be adopted for the image generation system 1 as long as it can execute the desired processing. For example, the computer 10 and the computer 20 may be integrated. Furthermore, the information resource 40 and the AI ​​model 50 may be stored inside the computer 10.

[0033] The computer 10 includes a CPU (Central Processing Unit) 11, a GPU (Graphics Processing Unit) 12, a communication device 13, and a memory / storage 14. For example, the computer 10 corresponds to the image generation module 100 and the information acquisition module 200 in FIG. 1. The computer 10 may be a service providing device (server device) that provides an image generation service. The CPU 11 and the GPU 12 are so-called processors, and perform calculations (arithmetic processing) related to various processes such as image generation.

[0034] The communication device 13 is a communication device having a communication function for transmitting and receiving information to and from the computer 20, an information providing device, etc., and may be, for example, a communication circuit, a NIC (Network Interface Card), etc. The communication device 13 communicates with other devices such as the computer 20 and the information providing device via a predetermined network (such as the Internet). For example, the communication device 13 is connected to the predetermined network by wire or wirelessly, and transmits and receives information to and from other devices such as the computer 20 and the information providing device.

[0035] The memory / storage 14 is a storage device that stores various types of information. The memory / storage 14 is, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 14 stores various types of information used for processing by processors such as the CPU 11 and the GPU 12. The memory / storage 14 may also store information resources 40, AI models 50, etc.

[0036] The computer 20 includes a CPU 21, a GPU 22, a communication device 23, a memory / storage 24, and an IO interface 25. For example, the computer 20 corresponds to the client UI display unit 400 in FIG. 1 . The computer 20 may be a terminal device (such as a personal computer (PC) or a mobile device such as a smartphone) used by a user who uses the video generation service. The CPU 21 and the GPU 22 are so-called processors, and perform calculations (arithmetic processing) related to various processes such as video display. Note that the above is merely an example, and the computer 20 may have any configuration as long as it is capable of performing the desired processing. For example, the computer 20 may perform calculations (arithmetic processing) related to various processes such as video display using circuits such as an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA). Furthermore, the computer 20 may be configured so that a program is directly embedded in the processor circuitry instead of storing the program in a memory (such as the memory / storage 24). In this case, the processor realizes its functions by reading and executing the program embedded in the circuitry. In addition, each processor in this embodiment is not limited to being configured as a single circuit, but may be configured as a single processor by combining multiple independent circuits to realize its functions. Also, like the computer 20, the computer 10 can adopt any configuration as long as it can perform the desired processing.

[0037] The communication device 23 is a communication device having a communication function for transmitting and receiving information to and from the computer 10, the sensor 30, etc., and may be, for example, a communication circuit, a NIC, etc. The communication device 23 communicates with other devices such as the computer 10, the sensor 30, etc. via a predetermined network (such as the Internet). For example, the communication device 23 is connected to the predetermined network by wire or wirelessly, and transmits and receives information to and from other devices such as the computer 10, the sensor 30, etc.

[0038] The memory / storage 24 is a storage device that stores various types of information. The memory / storage 24 is, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 24 stores various types of information that are used for processing by processors such as the CPU 21 and the GPU 22.

[0039] The IO interface 25 is an input / output interface device. The computer 20 receives input from the sensor 30 via the IO interface 25. For example, the computer 20 receives input from an input device such as a keyboard or a mouse via the IO interface 25. The computer 20 also outputs information from a display (display device) and a speaker (audio output device) via the IO interface 25. For example, the computer 20 plays video on the display and speaker via the IO interface 25.

[0040] Various sensors 30 such as cameras sense user input. Various sensors 30 such as cameras accept user operations. For example, the sensor 30 corresponds to the sensor unit 300 in FIG. 1. Furthermore, the information resource 40 includes various information such as training data. For example, the information resource 40 includes training data used for training various AI models such as LLM. The AI ​​model 50 includes information on AI models used in processing related to video generation such as LLM. For example, the AI ​​model 50 includes information on various AI models such as models M1 to M3, which will be described later. As described above, the video generation system 1 may have a configuration other than that shown in FIG. 2.

[0041] <1-2. Processing by the image generation system of the present disclosure> Next, the processing performed by the video generation system will be described. First, an example of the flow of the video generation processing shown in FIG. 3 will be described. FIG. 3 is a diagram showing an example of the flow of the video generation processing of the present disclosure. Note that the processing described below with the video generation system 1 as the processing subject may be performed by any device capable of executing that processing, depending on the device configuration included in the video generation system 1.

[0042] User input information UIN1, denoted as "User's Input" in FIG. 3, corresponds to information input by the user for image generation (also referred to as "input query"). Note that the input query is not limited to text (character information) and any information can be used. The input query may be any information including at least one of text, image, audio, and 3D data.

[0043] The video generation system 1 uses user input information UIN1 to generate scenario generation information (also referred to as "first input information") to be used as input for model M1, denoted as "LLM" in FIG. 3, which will be described later. For example, model M1 is a first model that outputs scenario data in response to the input of the first input information. Any AI model, such as an LLM (large-scale language model), can be used for model M1 as long as it is capable of producing the desired output in response to the input. AI models such as model M1 will be described later.

[0044] The video generation system 1 generates a scenario FD1 by inputting first input information to a model M1 and causing the model M1 to output a scenario FD1, which is scenario data. The video generation system 1 then generates USD generation necessary information SD1, which is code generation information (also referred to as "second input information") to be used as input for a model M3, using the scenario FD1, the output of a model M2 that uses the scenario FD1 as input, and user input information UIN2, etc. For example, the model M3 is a second model that outputs code in response to the input of the second input information. Any AI model, such as a large-scale language model (LLM), can be used for the model M3 as long as it is capable of producing a desired output in response to the input.

[0045] 3 shows only one piece of information SD1 required for USD generation, but there may be multiple pieces of information SD1 required for USD generation depending on, for example, the number of USD files to be generated. For example, there may be multiple pieces of information SD1 required for USD generation depending on the number of USD files to be generated corresponding to the data structure shown in FIG.

[0046] For example, model M2 may be a model that outputs a template or the like corresponding to scenario data in response to input of the scenario data. Any AI model can be adopted for model M2 as long as it is capable of producing the desired output in response to the input. For example, user input information UIN2 may be information for specifying constraints for image generation. Note that image generation system 1 may generate second input information using scenario FD1 and template input information, but this point will be described later.

[0047] The video generation system 1 inputs information SD1 required for USD generation into the model M3 and generates the Python code OD1 by having the model M3 output the Python code OD1, denoted as "python" in Figure 3. For example, the Python code OD1 is (program) code that generates USD format data (also called a "USD file") when executed. Note that Python is merely an example, and any code format can be used as long as it can generate the desired 3DCG data. Also, USD is merely an example, and any format, such as FBX (Film Box), can be used as long as the data is for 3DCG. The video generation system 1 executes the Python code OD1 to generate a USD file OD2, denoted as "USD" in Figure 3.

[0048] The video generation system 1 generates video data MV1, which is denoted as "PreMovie" in Fig. 3, by executing a rendering process PS1, which is denoted as "Renderer" in Fig. 3. For example, the video data MV1 is data (also referred to as "first video data") before executing a refinement process PS2, which will be described later.

[0049] The video generation system 1 generates video data MV2, denoted as "RefinedMovie" in FIG. 3, by executing a refinement process PS2, denoted as "Refiner" in FIG. 3. For example, the refinement process PS2 is a picture quality improvement process that improves the picture quality of the video data. The video data MV2 is data (also referred to as "second video data") obtained after the refinement process PS2 updates the first video data, that is, the video data MV1.

[0050] The video production system 1 executes composite editing PS3 using the video data MV2, user input information UIN3, etc., to generate video data MV3, which is indicated as "FinalMovie" in Fig. 3. For example, the composite editing PS3 executes a process of updating (editing) the video data MV2 in accordance with a user editing instruction indicated by the user input information UIN3, thereby generating video data MV3 in which the video data MV2 has been updated.

[0051] Note that the flow of the video generation process shown in FIG. 3 is merely an example, and the video generation system 1 can employ any processing mode as long as it can generate video data from a user's input query. For example, while FIG. 3 illustrates an example in which the model M1 outputs code (Python code), the model M1 may also be a model that outputs 3DCG data such as a USD file. Furthermore, the video generation system 1 is not limited to the process illustrated in FIG. 3 and may perform various modes of video generation processing. An example of this point will be described with reference to FIG. 4. FIG. 4 is a diagram illustrating another example of the flow of the video generation process of the present disclosure. FIG. 4 differs from FIG. 3 in that sound generation required information SD2 and text logo required information SD3 are generated and used to perform video generation processing. Note that explanations of points similar to those described in FIG. 3 will be omitted where appropriate.

[0052] In FIG. 4, the video generation system 1 uses a scenario FD1, the output of a model M2 that uses the scenario FD1 as input, and user input information UIN2 to generate information required for sound generation SD2, which is information for generating sound to be used as input for a model M4, denoted as "AI" in FIG. 3. For example, the model M4 is a model that outputs various sound data in response to input of the information required for sound generation SD2 and video data MV2. Any AI model can be used for the model M4 as long as it can produce the desired output in response to the input. Note that the model M4 may also be a model that uses only the information required for sound generation SD2 as input.

[0053] The video production system 1 generates sound data corresponding to the video by inputting sound generation necessary information SD2 to the model M4 and causing the model M4 to output sound data AD1 for background music (BGM), sound data AD2 for sound effects, sound data AD3 for narration, etc. The video production system 1 may also generate the sound data AD1, AD2, and AD3 using user input information UIN4. For example, if the model M4 outputs the sound data AD1, AD2, and AD3 as one piece of sound data, the video production system 1 may extract the sound data AD1, AD2, and AD3 from the one piece of sound data output by the model M4 based on the specification in the user input information UIN4, and generate the sound data AD1, AD2, and AD3.

[0054] In Figure 4, video generation system 1 uses scenario FD1, the output of model M2 that uses scenario FD1 as input, and user input information UIN2 to generate information required for text logo generation SD3, which is information for generating a text logo to be used as input for model M5. For example, model M5 is a model that outputs at least one of text and a logo in response to the input of information required for text logo generation SD3. Any AI model can be used for model M5 as long as it can produce the desired output in response to the input.

[0055] The video generation system 1 inputs the information SD3 required for text logo generation into the model M5, and generates text logo data corresponding to the video by having the model M5 output text logo data DI1 for Text, text logo data DI2 for Logo, etc.

[0056] Video generation system 1 generates video data MV3 by executing composite editing PS3 using video data MV2, sound data AD1, AD2, AD3, text logo data DI1, DI2, user input information UIN3, etc. For example, composite editing PS3 executes a process of combining (combining) video data MV2, sound data AD1, AD2, AD3, text logo data DI1, DI2, etc. into one image to generate video data MV3 as a single image.

[0057] 3 and 4 illustrate an example of processing in an initial state where there is no scenario, information required for USD generation, USD, PreMovie, RefinedMovie, or the like. As described above, the video generation system 1 generates prompts for generating a scenario based on a user's input query, provides the prompts to a natural language model to generate a scenario, generates a scenario, generates prompts for outputting code that constitutes a video from text information described in the scenario, provides the prompts to the natural language model to generate code that constitutes the video, and generates a video. In this way, when creating a video, the video generation system 1 generates an effective video storyboard and video by inputting what the user wants to create and their purpose, even if they have no knowledge of 3DCG or video production. Furthermore, the video generation system 1 creates a storyboard, making subsequent editing easier. These points will be described in detail later.

[0058] Furthermore, the image generation system 1 may perform various processes related to image generation. For example, the image generation system 1 may perform evaluation processing on the generated information. In this regard, an example of the flow of the evaluation processing will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the flow of the evaluation processing of the present disclosure.

[0059] In FIG. 5, video generation system 1 receives at least one of scenario FD1, information required for USD generation SD1, information required for sound generation SD2, and information required for text logo generation SD3, and causes model M10 to output evaluation text information EV1 indicating an evaluation of the input information, thereby evaluating the generated information. For example, model M10 outputs an evaluation of input information in response to input of the information. For example, model M10 outputs evaluation text indicating an evaluation of the input scenario FD1 in response to input of scenario FD1. Note that model M10 may be a model that accepts input of scenario FD1, information required for USD generation SD1, information required for sound generation SD2, and information required for text logo generation SD3 separately, or a model that accepts input of a combination of these pieces of information. Furthermore, model M10 may be a model that accepts input of information indicating a video (e.g., captions) in response to input of the video, and outputs an evaluation of the video corresponding to the input information.

[0060] From here, the flow of the above-mentioned processing will be described with specific examples of each process executed by the image generation system 1. Note that explanations of points similar to those described above will be omitted as appropriate.

[0061] For example, the video generation system 1 generates scenario generation information (first input information) as shown in Fig. 6. Fig. 6 is a diagram showing an example of a generation process of scenario generation information. In Fig. 6, the video generation system 1 acquires user input information IDT1 and IDT2 entered by the user into content CT1 as user input information. Content CT1 is content for accepting user input information for each of the questions "What kind of video do you want to create?" and "Style."

[0062] For example, the client UI display unit 400 displays the content CT1, and the sensor unit 300 receives the user input information IDT1 and IDT2 as user input information. For example, the user input information IDT1 and IDT2 correspond to the user input information UIN1 in FIGS.

[0063] In FIG. 6, the client UI display unit 400 displays a question, "What kind of video do you want to make?". In response to the question, "What kind of video do you want to make?", the sensor unit 300 receives user input information IDT1, "a 15-second sneaker commercial video." The client UI display unit 400 also displays a question, "style." In response to the question, "style," the sensor unit 300 receives user input information IDT2, "cinematic."

[0064] The video production system 1 may accept user input information in any manner, or may accept a user selection from multiple options. For example, the video production system 1 may convert information entered by the user via a keyboard or microphone into text information and accept it as user input information. The video production system 1 may also accept, in addition to free text, settings such as the number of seconds for the entire video, style, and camerawork, as well as other files such as images and videos as user input information.

[0065] The video production system 1 generates a prompt PT1, which is scenario generation information (first input information), using user input information IDT1, IDT2, and a template TP1, which is template input information. For example, the template TP1 may be preset or may be selected from a plurality of template candidates. For example, the video production system 1 may select a template corresponding to the user's input information from the plurality of template candidates. For example, the video production system 1 may select a template TP1 related to a movie-style advertisement from the plurality of template candidates based on the content indicated by the user input information IDT1, IDT2.

[0066] For example, the video production system 1 generates prompt PT1 by reflecting user input information IDT1 and IDT2 in template TP1. In FIG. 6, the video production system 1 generates prompt PT1 by adding "cinematic" indicated by input information IDT2 to the style item of the constraint condition and adding "15-second sneaker commercial video" indicated by input information IDT1 to the input sentence. In this way, the video production system 1 generates a prompt for generating a scenario based on the user's input information. Note that the user's input information may be entered on a single screen or by answering several questions; examples of these points will be described later.

[0067] The video production system 1 also generates scenario data as shown in Fig. 7. Fig. 7 is a diagram showing an example of a scenario data generation process. In Fig. 7, the video production system 1 generates scenario data SN1 using a prompt PT1. For example, the scenario data SN1 corresponds to the scenario FD1 in Figs. 3 and 4. The scenario data SN1 includes information such as the number of seconds for each scene, an explanation of the cut, etc., for the opening scene, the scene where sneakers are put on, etc.

[0068] For example, the video generation system 1 generates scenario data SN1 by inputting a prompt PT1 into a model M1, such as an LLM, and having the model M1 output scenario data SN1. In this way, the video generation system 1 generates a scenario by inputting the generated prompt into an AI (such as an LLM). In addition to the information shown in FIG. 7 (also referred to as "scenario information"), the scenario data SN1 also includes information such as the environment, characters, motion, camerawork, lighting, and color. For example, to create distinctive features in the scenario generation, the video generation system 1 can generate a wide variety of scenario variations by inputting the user's past experiential learning data or the learning data of a specific director or person into the model M1 or the like using RAG (Retrieval-Augmented Generation) or fine-tuning.

[0069] Furthermore, the video production system 1 generates code generation information (second input information) as shown in FIG. 8. FIG. 8 is a diagram showing an example of a process for generating code generation information. The video production system 1 generates a prompt PT2, which is code generation information (second input information), using scenario data SN1 and a template TP2, which is template input information. For example, the template TP2 may be preset or may be selected from a plurality of template candidates. For example, the video production system 1 may select a template corresponding to a scenario from the plurality of template candidates. For example, the video production system 1 may select a template TP2 related to a commercial from the plurality of template candidates based on the content indicated by the scenario data SN1.

[0070] For example, the video production system 1 generates prompt PT2 by reflecting scenario data SN1 in template TP2. In Figure 8, the video production system 1 generates prompt PT2 by adding information indicated by scenario data SN1 to the input sentence. In this way, the video production system 1 generates a prompt for video based on a scenario generated by AI.

[0071] For example, FIG. 8 shows an example of generating a prompt for converting a scenario related to a person into USD-Python. Conversion to USD-Python is merely one example of a conversion format, and the conversion is not limited to USD-Python and may be any conversion format. For example, the conversion format may be Python for Blender, USD, or other formats. Furthermore, the video generation system 1 may generate prompts individually for each subject, such as a person, environment, or camerawork, or may generate prompts collectively. Furthermore, the generated prompt may include paths to assets and motions to be used, or may include source code or an API (Application Programming Interface) to be passed (input) to an AI generation algorithm for assets / motions.

[0072] The video generation system 1 then generates a USD-Python file by inputting the generated prompt into an AI (such as an LLM). Note that the file format is not limited to Python format, and the file may be generated in another format such as USD. The video generation system 1 then converts the file into a format that can be rendered, such as a USD file, and performs rendering to generate a PreMovie (a video file such as an mp4).

[0073] The video production system 1 may perform composite editing using a PreMovie, but may also perform refinement processing, which is an example of image quality improvement processing, on the PreMovie, as shown in Fig. 9. Fig. 9 is a diagram showing an example of image quality improvement processing. In Fig. 9, the video production system 1 generates a second video OT1 from the first video IN1 by refinement processing using a model M11, which is a Diffusion model that takes a first video IN1, which is a PreMovie, as input and outputs a second video OT1, which is a RefinedMovie.

[0074] The AI ​​model (such as model M11) used in the refinement process is not limited to the diffusion model; any AI model such as the latent diffusion model (LDM) or the latent consistency model (LCM) can be used. Furthermore, techniques such as AnimateDiff (time direction stabilization) and ControlNet (line art control) may also be used in the refinement process. Through this refinement process, the video generation system 1 can improve the quality of the video while maintaining the consistency of the characters, backgrounds, props, etc.

[0075] Furthermore, the refiner process may use prompts in addition to videos. For example, the model M11 may input prompt IN2 in addition to the first video IN1. For example, when a target such as a woman in her 30s is specified by prompt IN2, the model M11 outputs a second video OT1 in which the part of the first video IN1 that refers to the woman in her 30s has been improved. This allows the video generation system 1 to generate a second video OT1 in which the image quality, etc. of the target specified by prompt IN2 in the first video IN1 has been improved.

[0076] Through the processing of the video generation system 1 described above, it appears to the user that a scenario (storyboard) and animation for each scene are being generated after input, with the processing in between being confined within the system. These processes may involve generating the scenario, all scenes, and refining processes all at once from the user's input text, or the user may input preferences during the process. For example, the video generation system 1 may generate several scenarios with outlines only, then allow the user to select one, and then execute detailed scenario and animation generation processing based on the selected outline scenario. Furthermore, the video generation system 1 may generate several patterns of characters to be generated in the video before video generation after scenario generation, and after the user selects one, execute animation rendering and refinement processing.

[0077] The components of a scenario may include video of the cut, a representative image (such as the first frame of the video), a description of the cut, characters (visuals, setting, etc.), the motion of each character, lighting, camera work, background environmental information, transitions between cuts, dialogue, narration, etc. Some of these are presented to the user, while others are kept for processing purposes without being presented to the user. The scenario is arranged in chronological order by cut.

[0078] Currently, various video generation services are available, including Pika, Runway Gen-2, Lumiere, and Stable Video Diffusion. These generate video using a diffusion model that moves vectors in the spatial and temporal directions from images. These generate video using only 2D images. Video Generation System 1, on the other hand, stores 3D information internally. For example, existing video generation services make it possible to modify only a specified (X,Y) area within a video, but this poses an issue where changing only the color of clothing also changes the motion. On the other hand, Video Generation System 1 stores 3D information internally, making it possible to modify only targeted areas, such as only the motion, lighting, or the color of a person's clothing.

[0079] Furthermore, existing video generation services only generate videos for each cut, and users must ensure the consistency of each cut themselves, but video generation system 1 can consistently carry out everything from the scenario (storyboard) to video generation and editing, making it possible to generate videos with consistency in terms of actors, backgrounds, color grading, etc.

[0080] <1-3. User Interface> The following describes the user interface (UI) for users who use the image generation system 1. Note that explanations of points similar to those described above will be omitted where appropriate.

[0081] As shown in FIG. 10, the video generation system 1 provides the user with content CT11. FIG. 10 is a diagram showing an example of a user interface. The content CT11 is a display screen (content) for receiving user input information. For example, the client UI display unit 400 displays the content CT11. The user inputs text instructing what kind of video to create (the "Prompt" field in FIG. 10) and a style selection (the "Style" field in FIG. 10) as user input information via the content CT11. In this way, the user inputs the text and style selection as user input information. For example, the user inputs text information in the "Prompt" field by referring to example sentences included in the content CT11. For example, the user selects a style to use from multiple style candidates displayed by pressing (clicking, etc.) the downward-facing triangle in the "Style" field.

[0082] After completing the input of the user's input information, the user selects the button labeled "Ask AI Director" in Fig. 10 to instruct the video generation system 1 to generate a video in accordance with the user's input information. As a result, the video generation system 1 executes the process of generating a video in accordance with the user's input information.

[0083] As shown in FIG. 14, the video generation system 1 provides the user with content CT15 related to the generated video. FIG. 14 is a diagram showing an example of a user interface. The content CT15 is a storyboard screen (content) for receiving user operations (instructions) on the generated video. As shown in FIG. 14, the content CT15 is a storyboard screen that displays video data for each cut of the generated video. For example, the client UI display unit 400 displays the content CT15. In this way, the video generation system 1 provides a UI that outputs a storyboard and video in accordance with information input by the user. The user sets the video, content, narration, dialogue, camerawork, background music, lighting, color, etc. for each cut on the storyboard screen.

[0084] The video generation system 1 may accept user input information while asking the user a question. For example, when the button labeled "Ask AI Director" in FIG. 10 is selected, the video generation system 1 accepts user input information through a conversation (dialogue) with the user, as shown in FIGS. 11 to 13. FIGS. 11 to 13 are diagrams showing an example of a user interface. The content CT12 in FIG. 11 is a display screen (content) that presents samples generated in response to the user's input information entered in FIG. 10 and asks the user whether there is any that is close to the image they have in mind. For example, the client UI display unit 400 displays the content CT12.

[0085] The content CT13 in Fig. 12 is a display screen (content) that requests (questions) for detailed targets, etc. in response to the user's response (input information) that none of the samples presented in the content CT12 in Fig. 11 match the image. For example, the client UI display unit 400 displays the content CT13.

[0086] The content CT14 in Fig. 13 is a display screen (content) that presents samples regenerated in response to a user's answer (input information) that specifically specifies a target, etc., and asks the user whether any of them are close to the image they have in mind. For example, the client UI display unit 400 displays the content CT14. Fig. 13 shows a case in which the user places the mouse cursor over the leftmost sample video among the four samples and performs a specifying operation such as clicking, thereby specifying that the leftmost sample video among the four samples is close to the image they have in mind. As a result, the video generation system 1 executes a process of generating a video in accordance with the user's input information that specifies the leftmost sample video among the four samples.

[0087] In this case, the video generation system 1 provides the user with content CT15 related to the generated video, as shown in Fig. 14. For example, the client UI display unit 400 displays the content CT15. In this way, when generating a storyboard and video from the initial input content, if the video generation system 1 does not have enough information, it may collect the necessary information through conversation (dialogue) with the user and work out the details.

[0088] 15 and 16, the video generation system 1 may prompt the user to input a short sentence about the video they want to create, generate several video stories based on the short sentence, and allow the user to select the one they like best. Figures 15 and 16 are diagrams showing an example of a user interface.

[0089] As shown in FIG. 15, the video generation system 1 provides the user with content CT21. The content CT21 is a display screen (content) for receiving input information from the user. For example, the client UI display unit 400 displays the content CT21. The user inputs a sentence (short sentence) indicating what kind of video they want to create as the user's input information via the content CT21. For example, the user inputs text information into an input field in the content CT21.

[0090] After completing the input of the user's input information, the user selects the button labeled "Start" on the right end of the input field in the content CT21 to instruct the video generation system 1 to generate a video story according to the user's input information. This causes the video generation system 1 to execute the process of generating a video story according to the user's input information.

[0091] As shown in Fig. 16, the video generation system 1 provides the user with content CT22 relating to the story of the generated video. The content CT22 in Fig. 16 is a display screen (content) that presents a sample of the story of the video generated in response to the user's input information entered in Fig. 15. For example, the client UI display unit 400 displays the content CT22. For example, if the user likes one of the video story samples, the user selects the button labeled "Continue" on the right end of the display area for that sample, thereby instructing the video generation system 1 to generate a video corresponding to the selected sample.

[0092] If the user does not like any of the video story samples, the video generation system 1 generates another pattern again. For example, if the user does not like any of the video story samples, the user can select the button labeled "Continue" on the right edge of the short sentence display area to instruct the video generation system 1 to generate another pattern of video story sample again. This causes the video generation system 1 to generate another pattern of video story sample.

[0093] <1-4. Processing example> From here, in addition to the specific examples described above, specific examples of each process executed by the video generation system 1 will be described. Note that explanations of points similar to those described above will be omitted as appropriate. Below, specific examples of the generation process of sound, etc., and evaluation process in the processing of the video generation system 1 described above will be explained. Note that explanations of points similar to those described above will be omitted as appropriate.

[0094] <1-4-1. Sound generation example> For example, the video production system 1 generates sound generation information (corresponding to the sound generation required information SD2 in FIGS. 3 and 4) as shown in FIG. 17. FIG. 17 is a diagram showing an example of a process for generating sound generation information. The video production system 1 generates a prompt PT3, which is sound generation information, using scenario data SN1 and a template TP3, which is template input information. For example, the template TP3 may be preset or may be selected from a plurality of template candidates. For example, the video production system 1 may select a template corresponding to a scenario from the plurality of template candidates. For example, the video production system 1 may select a template TP3 related to a commercial from the plurality of template candidates based on the content indicated by the scenario data SN1.

[0095] For example, the video production system 1 generates prompt PT3 by reflecting scenario data SN1 in template TP3. In FIG. 17, the video production system 1 generates prompt PT3 by adding information indicated by scenario data SN1 to an input sentence. For example, the video production system 1 generates a prompt for extracting several key phrases from the scenario in order to generate a sound (such as background music) that better suits the scenario. The video production system 1 may input the generated prompt into an AI (such as an LLM) to obtain keywords and text information necessary for sound generation.

[0096] The video production system 1 generates sound data using the generated prompt PT3. For example, the video production system 1 generates sounds (sound data) such as background music, sound effects, narration, and dialogue based on the generated video and text information written on the storyboard (scenario data SN1, etc.). Furthermore, if necessary words (text information) such as dialogue or narration have already been extracted from the scenario, the video production system 1 may save the words (text information) as sound data for the dialogue, narration, etc., without generating a prompt.

[0097] For example, the video generation system 1 generates sound data from the information required for sound generation (sound generation information), video, and audio data input by the user (sound source, user's voice, humming, etc.). For background music and sound effects, the video generation system 1 may generate sound from text or video using a transformer (model) such as text-to-music generation. The video generation system 1 may also search for contrastively learned sound sources, such as text-to-music estimation, from natural text.

[0098] Furthermore, for dialogue and narration, the video generation system 1 may generate audio using text-to-speech (such as a diffusion model or flow matching) based on the words themselves and text information about the characters obtained from the scenario. Furthermore, when connecting the generated video and audio, the video generation system 1 may incorporate meta information into the video or audio file to indicate the start and end times, volume, etc. of the audio.

[0099] <1-4-2. Example of text logo generation> For example, video production system 1 generates information for generating a text logo (corresponding to information SD3 required for text logo generation in Figures 3 and 4). Based on the generated storyboard (scenario data SN1, etc.), video production system 1 generates text and logo information (captions, titles, logos, descriptions, etc.) to be displayed on the video.

[0100] For example, the generated scenario may clearly state the text sentences to be displayed, but if they are not, the video generation system 1 generates a prompt to generate text information (text logo data, etc.) and sends it to an AI (such as an LLM) to generate the text information. Also, if the user inputs the text to be displayed, the video generation system 1 may use the information input by the user as the text information (text logo data, etc.).

[0101] The font, size, and position of the text display may be determined by any method. For example, the video generation system 1 may determine the font, size, and position of the text display using any AI such as a diffusion model, a variational auto-encoder (VAE), a generative adversarial network (GAN), DALL E, StyleGAN, StyleGAN2, Pix2Pix, TransGAN, or LLM. The font, size, and position of the text display may also be manually set by the user.

[0102] For logos and images, the user may input images or videos in JPEG format, MP4 format, etc. Furthermore, for logos and images, the video generation system 1 may generate a prompt for image generation and send it to any AI, such as a Diffusion model, VAE, GAN, DALL E, StyleGAN, StyleGAN2, Pix2Pix, TransGAN, or LLM, to generate logo information.

[0103] In addition, when linking text or logo information with a video, the video generation system 1 may incorporate meta information into the video or the text or logo itself to clearly indicate (meta information) the start and end times, position, and size of the text or logo.

[0104] <1-4-3. Re-learning example> Furthermore, when generating a scenario or video, the video generation system 1 may generate a scenario or video that can only be produced using specific, replaceable learning data. For example, it is possible to re-train previously produced videos and images as learning data. Re-training using RAG or fine tuning can change the scenario or video to be generated. In other words, the video generation system 1 can re-train the data (history) of personal user's past productions, or generate a scenario or video using a model re-trained using the works of a specific film director as learning data.

[0105] The training data can also be retrained on an individual's PC or on a server. The training data is used to train a number of models, such as LLM and Diffusion models. In cases such as when an overall scenario is to be generated using Director A, but the color of the video is to be generated using a different Director B to generate color grading, the video generation system 1 may retrain specific parts using different training data.

[0106] <1-4-4.USD update example> As described above, the video generation system 1 modifies the 3DCG assets and rendering method that form the basis of the existing video, based on the user's input information and the text information of the generated storyboard.The video generation system 1 then outputs prompts for outputting code that will compose a new video, provides the prompts to a natural language model, outputs the code that will compose the video, and generates the video.This allows the video generation system 1 to modify the video in accordance with the user's input, even if the user has no knowledge of 3DCG or video production.

[0107] In the video generation system 1, after generating a video (video), scenario information, video information, etc. are saved as text or video, so the user can modify this information with input information (text, sensor, etc.).

[0108] For example, in the video production system 1, USD files are saved separately for each asset and motion, as shown in Fig. 18. Fig. 18 is a diagram showing an example of a USD file. For example, in the data structure shown in Fig. 18, a USD in a higher layer may include a path (file path) to a USD in a lower layer.

[0109] For example, the overall USD includes paths to environment asset USD, person asset UDS, camera USD, etc. The environment asset USD also includes paths to building asset USD and prop asset USD. The building asset USD also includes mesh information for the building itself, etc. The person asset USD also includes paths to person mesh information and motion USD. In this way, 3DCG data (3D data) such as a USD file may include multiple data sets. Note that the configuration (data structure) of the USD file shown in FIG. 18 is merely an example, and any configuration can be adopted, and the entire USD (USD file) may be configured as a single block (one data set).

[0110] When modifications are made, the image generation system 1 analyzes the user's modification information using AI (such as LLM) according to the processing flow shown in Fig. 19, and the processing changes depending on whether the USD is replaced or part of the USD is modified. After the modification, rendering and refinement processes are executed. Fig. 19 is a flowchart showing the processing procedures executed by the image generation system. As a specific example, Fig. 19 is a flowchart showing the processing procedures related to rewriting a USD file.

[0111] First, the image generation system 1 receives correction information input from the user (step S101). For example, the sensor unit 300 receives input information instructing the user to make corrections. The image generation system 1 analyzes the input using AI (step S102). For example, the image generation module 100 analyzes the content of the input information instructing the user to make corrections using various models, etc.

[0112] The video generation system 1 recognizes the format of the existing USD file (step S103). For example, the video generation module 100 recognizes the format of the USD file before modification. The video generation system 1 determines whether to modify a portion of the USD (step S104). For example, the video generation module 100 determines whether to modify a portion of the USD based on the content of input information instructing the user to modify and the format of the existing USD file.

[0113] When part of the USD is to be modified (step S104: Yes), the video generation system 1 generates a prompt for generating a modified USD-Python (step S105). For example, when part of the USD is to be modified, the video generation module 100 generates a prompt for generating a modified USD-Python using a template or the like for generating a modified USD-Python.

[0114] The video generation system 1 generates USD-Python using AI (such as LLM) (step S106). For example, the video generation module 100 generates USD-Python by inputting a prompt into a model for generating USD-Python.

[0115] The video generation system 1 replaces the USD file to be modified (step S107). For example, the video generation module 100 replaces the USD file to be modified by reflecting the generated USD-Python in the USD file to be modified. In this way, the video generation system 1 updates at least one of the multiple data sets of the USD file. For example, the video generation system 1 executes a process of updating some of the multiple data sets of the USD file.

[0116] The image generation system 1 executes rendering processing using the updated USD file (step S108). For example, the image generation module 100 executes rendering processing using the rewritten, i.e., corrected, USD file.

[0117] On the other hand, if the image generation system 1 does not modify part of the USD (step S104: No), it generates a prompt for generating a creation USD-Python (step S109). For example, if the image generation module 100 does not modify part of the USD, that is, if a new USD is to be created (generated), it generates a prompt for generating a creation USD-Python using a template or the like for generating a creation USD-Python.

[0118] The video generation system 1 generates USD-Python and USD using AI (such as LLM) (step S110). For example, the video generation module 100 generates USD-Python and USD by inputting a prompt into a model for generating USD-Python.

[0119] The video generation system 1 replaces the USD file to be modified (step S111). For example, the video generation module 100 replaces the generated USD with the USD file to be modified. In this way, the video generation system 1 executes processing to update the USD file. For example, the video generation system 1 executes processing to update all of the multiple data sets in the USD file. Then, the video generation system 1 executes the processing of step S108. For example, the video generation module 100 executes rendering processing using the replaced, i.e., modified, USD file.

[0120] Here, a user interface (UI) related to the above-mentioned corrections will be described. As shown in FIG. 20, the video generation system 1 provides the user with content CT31. FIG. 20 is a diagram showing an example of the user interface. The content CT31 is a display screen (content) for presenting information about the generated video, such as a storyboard, and for accepting correction instructions from the user. For example, the client UI display unit 400 displays the content CT31.

[0121] The user inputs information instructing corrections to the generated video via the content CT 31. For example, when the user clicks on a specific part of the video and inputs corrections in text, the video generation system 1 updates the USD and updates (changes) the video to reflect the corrections.

[0122] 20 shows a case where the user selects the top cut (thumbnail) image and instructs a modification to "child skipping and going to mother." In this case, the video generation system 1 updates the USD of the part of the generated video corresponding to the top cut (thumbnail) image based on the modification instruction to "child skipping and going to mother," and generates a video that reflects the modification content.

[0123] The above UI is merely an example, and the video generation system 1 may accept a user's correction instruction in various ways. For example, as shown in FIG. 21, the video generation system 1 may provide the user with content CT32 and accept the user's correction instruction. FIG. 21 is a diagram showing an example of a user interface. The content CT32 is a display screen (content) for accepting the user's correction instruction. For example, the client UI display unit 400 displays the content CT32.

[0124] Furthermore, the video generation system 1 may provide the user with content CT33 and accept a correction instruction from the user, as shown in Fig. 22. Fig. 22 is a diagram showing an example of a user interface. The content CT33 is a display screen (content) for accepting a correction instruction from the user. For example, the client UI display unit 400 displays the content CT33.

[0125] The user inputs information instructing corrections to the generated video via content CT32 or content CT33. As a result, the video generation system 1 accepts the correction instructions from the user, updates the USD based on the correction instructions, and generates a video that reflects the corrections. For example, the user can select a person or object that appears in the video and edit the motion settings for the selected object using text. The user can also change the asset of the selected object. The user can also edit the camera work using text. In addition to the above, the user can also edit the background and lighting settings of the video using text.

[0126] Although the above description has been given as an example of a moving image, the video generation system 1 may also perform correction processing on sound information (BGM, SE, narration, dialogue, etc.), text logo information, etc. based on user correction instructions.

[0127] <1-4-5. Evaluation example> Furthermore, the video generation system 1 executes an evaluation process. For example, the video generation system 1 generates an evaluation text for scenario data using an AI model such as model M10. For example, the video generation system 1 may regenerate scenario data based on an instruction for editing (correction, etc.) given by a user based on the evaluation text generated for the scenario data. The video generation system 1 presents the generated evaluation text. This allows the user to create a video while viewing the evaluation.

[0128] For example, the video generation system 1 presents an evaluation text about scenario data to a user and receives an instruction to edit the scenario data from the user who has confirmed the presented evaluation text. The video generation system 1 generates scenario data based on the editing instruction received from the user. For example, the video generation system 1 changes (updates) the content of the scenario data based on the editing instruction received from the user. The video generation system 1 generates code based on scenario data generated based on the evaluation text. The video generation system 1 generates video data using the code generated based on the evaluation text.

[0129] The video generation system 1 may automatically generate (update) at least one of the scenario data or the code based on the evaluation text. For example, the video generation system 1 may change the content of the scenario data to correspond to the content indicated by the evaluation text for the scenario data. For example, if the evaluation text for the scenario data indicates that the orientation of a certain character is not good, the video generation system 1 generates scenario data in which the orientation of the character is changed. The above is merely an example, and the video generation system 1 may generate at least one of the scenario data or the code by appropriately using the evaluation text.

[0130] After generating a video (movie), the video generation system 1 stores scenario information, video information, sound information, text / logo information, and the like in text, video files, sound files, image files, and the like. Therefore, the video generation system 1 can perform evaluation processing using this information. Based on user input information, the text information of the generated storyboard, and the evaluation text, the video generation system 1 modifies the 3DCG assets and rendering method that form the source of the existing video, outputs prompts for outputting code that constitutes a new video, provides the prompts to a natural language model, outputs the code that constitutes the video, and generates a video. For example, the video generation system 1 may output prompts for generating evaluation text (corresponding to evaluation text information EV1 in FIG. 5) based on user input information and the text information of the generated storyboard, and provide the prompts to a natural language model to generate the evaluation text. For example, the video generation system 1 may generate evaluation prompts based on scenario information and input them to an AI (such as an LLM) to generate evaluation text indicating the evaluation of the scenario.

[0131] Furthermore, to evaluate the composition of a video, the video generation system 1 acquires a caption for the video by inputting one frame of the video into an AI (such as a Contrastive Captioner Model or an Image Captioning Model). The video generation system 1 then inputs the acquired caption and scenario text into an AI (such as an LLM) to generate an evaluation text indicating an evaluation of the composition of the video. For example, the video generation system 1 may evaluate whether the image is consistent with the caption by comparing the acquired caption and scenario text using an AI (such as an LLM). The video generation system 1 may also accept a specification of the type of evaluator that the user desires to evaluate. For example, the video generation system 1 may perform an evaluation from the perspective of a market strategy, copywriter, video director, or a specific person.

[0132] <1-4-6. Example of selecting multiple cuts> From here, some examples of editing processing will be described. Conventional technology has a problem in that multiple cuts cannot be simultaneously edited on a storyboard. As such, conventional technology has problems with usability, and there is room for improvement in usability. Therefore, the video production system 1 may select multiple cuts and perform editing processing (editing processing) as shown in FIG. 23. FIG. 23 is a diagram showing an example of editing processing. As a specific example, FIG. 23 is a diagram showing an example of editing processing based on the selection of multiple cuts.

[0133] 23, the video production system 1 may provide the user with content CT41 and accept correction instructions according to the user's selection of multiple cuts. The content CT41 is a display screen (content) for accepting the user's correction instructions for a video including multiple cuts (scenes) such as cuts CU1 to CU4. For example, the client UI display unit 400 displays the content CT41.

[0134] The user selects multiple cuts from among the cuts CU1 to CU4 via the content CT41 and inputs information instructing corrections to the selected multiple cuts. For example, the user selects cuts CU1 to CU3 by performing an operation to select the range in which cuts CU1 to CU3 are displayed (such as by surrounding it with a line) among the cuts CU1 to CU4 in the content CT41, or by clicking on each of the cuts CU1 to CU3.

[0135] After selecting cuts CU1 to CU3, the user inputs a prompt (e.g., text information) indicating a correction instruction as user input information, thereby instructing the video production system 1 to correct cuts CU1 to CU3. The video production system 1 performs corrections on cuts CU1 to CU3 in response to the user's correction instruction. This allows the user to select multiple cuts and correct the structure and content of the selected cuts by inputting a prompt. For example, scenario information for each cut includes the content of the cut, the characters, the shooting time of the cut, etc., and the video production system 1 regenerates the scenario by providing the user's input information and scenario information to an AI (e.g., LLM). The video production system 1 performs any correction processing based on the content of the user's correction instruction. For example, the video production system 1 may update the USD file of the video as needed, or may change only the order of the cuts while leaving the USD file unchanged. Through the above-mentioned processing, the video production system 1 can improve usability.

[0136] <1-4-7. Examples of using time information> The conventional technology has an issue in that after a specific cut is corrected, other scenes affected by the correction are not corrected (changed). As such, the conventional technology has usability issues and there is room for improvement in usability. Therefore, the video production system 1 may perform editing processing using time information, as shown in FIG. 24. FIG. 24 is a diagram showing an example of editing processing. As a specific example, FIG. 24 is a diagram showing an example of editing processing based on correction content determined according to time information. For example, the video production system 1 includes AI-generated date and time information (also called "time information" or "date information") for each cut, and determines the cutting of the video and the correction content based on that information.

[0137] In FIG. 24, the video production system 1 manages each cut by associating it with date and time information (time information) corresponding to that cut. For example, time information TI11 indicating 10:31 on December 21, 2023 is associated with cut CU11. Furthermore, time information TI12 indicating 10:41 on December 21, 2023 is associated with cut CU12. In this way, cuts CU11 and CU12 are close in time (nearby). In this case, if cut CU11 is modified, the video production system 1 also reflects the modification in cut CU12. Time information (date information) is associated with each cut in the video data.

[0138] For example, if the clothing of person X in cut CU11 is modified, the video production system 1 also modifies the clothing of person X in cut CU12 in the same way as the clothing of person X in cut CU11. For example, the video production system 1 compares the time information between each cut, and if there is a cut (also called an "influenced cut") whose time difference with the modified cut (also called a "cut to be modified") is within a predetermined range, it determines that the modification based on the modification content of the cut to be modified should also be reflected in the influenced cut.

[0139] Then, the video production system 1 executes corrections to the influencing cuts based on the correction content of the cut to be corrected. Note that the corrections (changes) to the influencing cuts based on the influence between cuts are not limited to human assets, and the video production system 1 may also perform the corrections (changes) to environmental assets (weather, lighting, props such as candles that change over time, etc.).

[0140] For example, cut CU21 is associated with time information TI21 indicating 10:31 on December 21, 2023. Also, cut CU22 is associated with time information TI22 indicating 18:41 on December 21, 2023. In this way, cuts CU21 and CU22 are distant (separate) cuts in terms of time. In this case, if cut CU21 is modified, the video production system 1 does not reflect the modification in cut CU22.

[0141] For example, if the clothing of person X in cut CU21 is modified, the video production system 1 does not modify the clothing of person X in cut CU22 in accordance with the modification of the clothing of person X in cut CU21. For example, the video production system 1 compares the time information between each cut, and if there is no cut (affecting cut) whose time difference with the modified cut (cut to be modified) is within a predetermined range, it determines not to reflect the modification based on the modification content of the cut to be modified in other cuts.

[0142] For example, when generating a scenario, the video generation system 1 generates date and time information for each cut using AI (such as LLM) and saves it as meta information for each cut. This date and time information is used to consider the relationship between previous and next cuts and to understand the season, etc. The generated fictitious date and time information is included to maintain the content of each cut and the relationship between each cut. For example, if the shot is taken in a morning when it is snowing, the video generation system 1 may set the time to 6:30 AM on January 24th. Also, if the shot has a strong relationship with the previous and next cuts, the video generation system 1 may set the time to a nearby time on the same day, such as 6:30 AM on January 24th and 7:00 AM on January 24th.

[0143] For example, the video generation system 1 changes the clothes the user is wearing, the sun's lighting settings, the haze of the air, etc. depending on the season and time. Also, if the date and time information of a target cut and a previous cut is close, changing the target cut will also affect the previous and next cuts. For example, if the same person appears in another cut and the time is close to the target cut, changing the clothes in the target cut will also change the clothes in the other cut. Through the above-mentioned processing, the video generation system 1 can improve usability.

[0144] <1-4-8. Response examples according to the degree of verbalization> In the conventional technology, when the user cannot specifically verbalize the content of the correction, it is difficult to make corrections in accordance with the user's intention. As such, the conventional technology has problems with usability, and there is room for improvement in usability. Therefore, the video generation system 1 may use the process described below to enable corrections in accordance with the user's intention, even when the user cannot specifically verbalize the content of the correction.

[0145] For example, the video production system 1 may change the response method depending on the input sentence. As shown in Fig. 25, the video production system 1 may vary the response depending on the level of abstraction and verbalization of the user's instruction. Fig. 25 is a diagram showing an example of editing processing. As a specific example, Fig. 25 is a diagram showing an example of editing processing depending on the level of verbalization.

[0146] Response example AP1 in FIG. 25 shows a case where a user issues a purpose-based correction instruction such as, "Please make the video so that the viewer feels like they are in their daily lives." In this case, image generation system 1 generates an image with a specific instructional sentence attached. For example, in response to a purpose-based correction instruction such as, "Please make the video so that the viewer feels like they are in their daily lives," image generation system 1 adds a sentence such as, "Lower your hands at a natural angle and move them in accordance with the direction of your body," and presents the video corrected based on the sentence to the user by displaying it on a display or the like.

[0147] Response example AP2 in FIG. 25 illustrates a case where the user instructs to correct an abstract sentence, such as "Keep both arms in a neutral position." In this case, the image generation system 1 presents multiple specific sentences to the user and allows the user to select one. For example, the image generation system 1 presents multiple sentences to the user, such as "Put your hands down and move them in a direction that points up and down," "Scratch your head with your hands and then put your hands down," and "Put your hands in your pockets," by displaying them on a display or the like. The image generation system 1 then presents to the user an image corrected based on the sentence selected by the user from the multiple sentences by displaying it on a display or the like.

[0148] Response example AP3 in Figure 25 shows a case where the user gives a specific instruction to correct a sentence, such as "Please look at the car in front of you on the right in two seconds." In this case, the image generation system 1 presents the image to the user by displaying on a display or the like an image that has been changed (corrected) according to the sentence entered by the user. For example, the difference between a purpose-based sentence, an abstract sentence, and a specific sentence may be determined by an AI (such as an LLM), and the image generation system 1 may change the prompt it generates accordingly and pass the processing to the AI ​​(such as an LLM).

[0149] For example, the video generation system 1 may use a sensor to allow the user to specify motion or camera movement. As shown in Fig. 26, the video generation system 1 may superimpose (overlap) motion information obtained from an arbitrary sensor such as a web camera on the video and accept user corrections. Fig. 26 is a diagram showing an example of editing processing. As a specific example, Fig. 26 is a diagram showing an example of editing processing by superimposing display on video.

[0150] For example, the video generation system 1 displays motion information MT obtained from any sensor, such as a mobile motion capture or a webcam, superimposed on the video MV11 and accepts user correction instructions. In this way, the video generation system 1 displays motion information, including facial expressions, superimposed on the current video, visualizing the difference between the current video and the motion information, and accepts user correction instructions based on this. For example, if a four-second cut is played, the four-second cut will always continue to be played, and the user can change the motion information as many times as they like. For example, if the user finds motion information they like, the video generation system 1 may ultimately generate natural motion information and change the motion of the video.

[0151] The user's motion is tracked using mobile motion capture, a webcam, etc. The user may also use the mouse to select in advance from the UI which character's motion they wish to modify. The size, position, and orientation of the portion of the user's motion to be superimposed are determined based on the size and orientation of the head and body in the video. The user may also adjust the size, position, and orientation in detail using the mouse and keyboard. Motion recording can be performed by repeating the cuts shown in the image above, or by starting recording a few seconds after pressing the start button and adjusting one's position. After filming, the user may press the motion generation button to apply the input motion to the specified character.

[0152] Furthermore, since the captured motion itself may be unnatural, the video generation system 1 may use motion-to-motion AI to estimate or generate motion from the captured motion, convert it into a more natural motion, and then adapt it to the motion of the characters. The video generation system 1 may also estimate or generate motion using input motion and natural language. For example, after inputting motion, the video generation system 1 may submit the motion to AI (such as an LLM or text&motion-to-motion generation model) for processing along with natural language input by the user, such as "Make the movement lively like this."

[0153] Furthermore, the user may change the camerawork by moving their hand as if it were a camera. For example, the image generation system 1 acquires the user's hand movements using a sensor such as a web camera. As shown in FIG. 27, the image generation system 1 presents (superimposed display) the changed camerawork as a frame superimposed on the image. FIG. 27 is a diagram showing an example of editing processing. As a specific example, FIG. 27 is a diagram showing an example of superimposed display of camerawork on the image.

[0154] For example, the video production system 1 displays a frame CW1 indicating the camerawork before the change and a frame CW2 indicating the camerawork after the change, superimposed on the video MV12. Then, when the user instructs the video production system 1 to use the changed camerawork, the video production system 1 executes rendering processing based on the changed camerawork. In this way, when the user approves the changed camerawork, the video is rendered using the changed camerawork.

[0155] The user may also specify the camerawork using an IMU or ImageSLAM of their mobile device (such as a smartphone). For example, the image generation system 1 may present the main character by placing it in real space (such as on a real desk) using AR (Augmented Reality). The image generation system 1 displays the changed camerawork by superimposing a frame on the image, similar to the case shown in FIG. 27.

[0156] As described above, in the video generation system 1, the user may use hand or smartphone movements to determine motion and camera movement. For example, when determining camerawork using hand movements, the user may imagine a person designated with the other hand and determine the relative position of the camera based on the distance between the left hand (camera) and the right hand (person). Furthermore, the video generation system 1 displays the camerawork changed by the user's hand or camera input as a frame superimposed on the video, but the user may then fine-tune the frame and its movement using a mouse. Furthermore, because the input camerawork is manually entered, it may be unnatural, so the video generation system 1 may correct the camerawork to make it more natural using a motion-to-motion generation model, a text&motion-to-motion generation model, or the like. Through the above-described processing, the video generation system 1 can improve usability.

[0157] <1-4-9. Examples of using qualitative values> In conventional technologies, there are cases where consideration is not given to how the system determines the current state and the change history. For example, there are cases where a user wants to make adjustments by comparing the current state with the current state, such as "stand a little further back" or "make it a little brighter," and there is a problem with how the system understands (grasp) the current state. As such, conventional technologies have usability issues, and there is room for improvement in usability. Therefore, the video generation system 1 may be able to appropriately determine the current state and the change history by performing the process described below.

[0158] For example, the video production system 1 may store each point such as motion speed and brightness as a quantitative numerical value and modify the video by comparing the values. The video production system 1 may also perform editing processing using qualitative values, as shown in Fig. 28. Fig. 28 is a diagram showing an example of editing processing. As a specific example, Fig. 28 is a diagram showing an example of editing processing based on modification content determined according to the qualitative values.

[0159] 28, value information VL21 indicates a quantitative value associated with video MV21. For example, value information VL21 includes qualitative values ​​such as the brightness of the entire town, the walking speed of people, the speed at which people move their faces, the position of people, the speed of cars, and the position of cars, and indicates values ​​associated with cuts in video MV21. In this way, video production system 1 stores all values ​​that can be changed as quantitative values ​​and changes the video by comparing them with those values.

[0160] 28, the user gives a correction instruction for video MV21, saying, "Move your face a little more slowly," and video generation system 1 generates video MV22 by slowing down the speed at which the face moves in video MV21. Based on the correction instruction, saying, "Move your face a little more slowly," video generation system 1 generates video MV22 with value information VL22 in which the value of the speed at which the person's face moves, in the value information VL21 of video MV21, is reduced from "21" to "10."

[0161] 28, value information VL22 indicates quantitative values ​​associated with the corrected video MV22. For example, value information VL22 includes qualitative values ​​such as the brightness of the entire town, the walking speed of people, the speed at which people move their faces, the positions of people, the speed of cars, and the positions of cars, and indicates values ​​associated with cuts and the like of video MV22.

[0162] In this way, the video generation system 1 may perform video editing processing using qualitative values. For example, the video generation system 1 may store quantitative values ​​for each point (item), such as movement speed and brightness, and compare these values ​​to modify the video. As described above, classifications such as the brightness of the entire city and people's walking speed, for which qualitative values ​​are stored, are pre-set items. For example, the video generation system 1 may set and modify each item using AI (such as LLM) from natural language, or may directly modify the set value. Through the above-described processing, the video generation system 1 can improve usability.

[0163] An example of a method for acquiring each value is shown below, but the acquisition method is not limited to the above and other acquisition methods may also be used. For example, for lighting, the video generation system 1 acquires the light setting values ​​(position, rotation, intensity, color, etc.) within the 3DCG. For color grading, the video generation system 1 acquires the white balance, color temperature, color cast correction, saturation, exposure, contrast, highlights, shadows, white level, black level, color, LUT settings, etc. set in composite editing. For walking speed, the video generation system 1 acquires the position movement speed of the hip bone of the target 3D model. For face movement speed, the video generation system 1 acquires the rotation speed of the head bone. For position, the video generation system 1 acquires the position of the 3D model.

[0164] <1-4-10. Example of rendering process during input> The conventional technology has an issue in that it is difficult for the user to wait for the rendering time (low usability). As such, the conventional technology has usability issues and there is room for improvement in usability. Therefore, the video generation system 1 may improve the usability of rendering by performing the following process.

[0165] For example, the video production system 1 may start rendering processing in the middle of text input. The video production system 1 may perform rendering processing in the middle of user input, as shown in Fig. 29. Fig. 29 is a diagram showing an example of editing processing. As a specific example, Fig. 29 is a diagram showing an example of editing processing based on rendering processing in the middle of input.

[0166] In FIG. 29, the video generation system 1 executes rendering processing and the like in response to the user's input of text TX31, "scratching his head with his hand," and displays video MV31. For example, the "|" at the end of text TX31 indicates that the user is in the middle of inputting a correction instruction. Once the video generation system 1 is able to understand the text as a sentence, it starts processing in the background. For example, the video generation system 1 starts rendering processing and the like in the background once the user has input up to "scratching his head with his hand." In this way, the video generation system 1 may start rendering processing while the user is still inputting text.

[0167] In FIG. 29, the user changes the sentence from text TX31 to text TX32, which reads "Put your hands down." For example, the "|" at the end of text TX32 indicates that the user is in the middle of inputting a correction instruction. The video generation system 1 executes processing in response to the change from text TX31 to text TX32. For example, when the sentence is changed, the video generation system 1 stops the processing being performed at that time and starts the processing again. For example, the video generation system 1 ends the processing that was being performed based on text TX31, and executes rendering processing and the like based on text TX32 to display video MV32.

[0168] In Fig. 29, the user changes the text TX32 to text TX33, "Put your hands down, in a natural way," completes the sentence, and instructs execution of the process by pressing a start button or the like. The video generation system 1 ends the background process that has begun, and displays video MV33 for which rendering process and the like have been performed based on the text TX33. Through the above-described process, the video generation system 1 can improve usability.

[0169] <1-4-11. Example of checking work when editing> In the prior art, there are problems such as low usability in the confirmation work during editing, and there is room for improvement. Therefore, the video generation system 1 may perform the confirmation work during editing as shown in Fig. 30. Fig. 30 is a diagram showing an example of the editing process. As a specific example, Fig. 30 is a diagram showing an example of the confirmation work during editing.

[0170] A confirmation task CP1 in FIG. 30 shows an example of a confirmation task when it is necessary to check changes in the time axis direction, such as motion or camera work. For example, in confirmation task CP1, the user repeatedly selects a video that is close to ideal. Also, a confirmation task CP2 in FIG. 30 shows an example of a confirmation task other than when it is necessary to check changes in the time axis direction. For example, in confirmation task CP2, the user repeatedly selects an image that is close to ideal.

[0171] For example, the video generation system 1 may render images of all cuts and then render a video starting with the one the user wants to edit. For example, the video generation system 1 may present the generated video to the user and allow the user to decide whether to edit the input information. For example, as shown in confirmation tasks CP1, CP2, etc., when presenting the user with the results of generating several patterns of video with variations based on the user's input information, there is an issue of waiting time for rendering multiple videos. As such, the conventional technology has usability issues and there is room for improvement in usability.

[0172] Therefore, in order to reduce the waiting time, for tasks other than the confirmation task that requires viewing changes along the time axis (corresponding to confirmation task CP2), the video generation system 1 renders only one or several frames from the video and presents candidates as images to the user. This allows the video generation system 1 to reduce the rendering waiting time.

[0173] Furthermore, when the user edits motion or camerawork, the video generation system 1 presents video candidates. For example, methods for selecting a few frames from a video to use for rendering include simply rendering only the first and last two frames of the video, rendering a few frames with large animation changes in the USD file, or having AI select frames to display as highlights from all frames. For example, when the video generation system 1 renders multiple images, the user can check the generated results in the form of a flip book by hovering the mouse cursor over them.

[0174] Furthermore, the video generation system 1 can reduce waiting times during video preview by sequentially starting video generation for each pattern after completing the generation of images for candidate selection. The order in which videos are generated for each pattern can be generated in a random order, or the user can select the rendering order by pressing a button on the UI, or the images can be rendered in order of the length of time the mouse cursor is hovered over them. The above-described process allows the video generation system 1 to improve usability.

[0175] An example of a processing flow related to the above-mentioned checking work will now be described with reference to Fig. 31. Fig. 31 is a flowchart showing the processing procedure executed by the video production system. As a specific example, Fig. 31 is a flowchart showing the processing procedure related to editing processing.

[0176] 31, the video generation system 1 generates several patterns of USD files based on the user's settings (step S201). If image writing of USD files for all patterns has been completed (step S202: Yes), the video generation system 1 ends the process of writing images to files (e.g., steps S202 to S204).

[0177] If image writing of USD files for all patterns has not been completed (step S202: No), the video generation system 1 writes out images of USD files for which writing has not been completed (step S203). The video generation system 1 displays the written images on the UI (step S204). After that, the video generation system 1 starts the processing from step S205 onwards, and repeats the processing of steps S202 to S204 until image writing of files is completed.

[0178] If the video generation system 1 has finished writing the moving images of all patterns of USD files (step S205: Yes), it ends the process of writing the moving images of the files (for example, steps S205 to S207).

[0179] If the video writing of all patterns of USD files has not been completed (step S205: No), the video generation system 1 writes out the video of the USD files that have not been written out (step S206). The video generation system 1 displays the written out video on the UI (step S207). Thereafter, the video generation system 1 repeats the processes of steps S205 to S207 until the video writing of the files is completed.

[0180] <1-4-12. Example of processing according to range selection> In the conventional technology, when a person (also called an "amateur") who has no experience (knowledge) in video (video) generation tries to create a video, it is difficult for them to judge what is good and what is bad, and they often make incorrect judgments. As such, the conventional technology has issues with usability, and there is room for improvement in usability. Therefore, the video generation system 1 may evaluate the generated video, as shown in FIG. 32. FIG. 32 is a diagram showing an example of evaluation processing according to range selection.

[0181] In Fig. 32, the video generation system 1 performs evaluation processing in accordance with the range selection. Video MV40 in Fig. 32 shows a moving image generated in accordance with user input. Video MV41 in Fig. 32 shows a state in which the user specifies a part of a person as the range they want evaluated, and text TX41 indicating the evaluation by the video generation system 1 of that range is superimposed. In Fig. 32, the video generation system 1 makes an evaluation (suggests a correction) for the part of the person specified by the user, indicated by the text TX41, saying, "Showing this person's facial expression will make it easier for the user to understand their emotions."

[0182] Video MV42 in Fig. 32 shows a state in which text TX42 indicating an evaluation by video creation system 1 is further superimposed. In Fig. 32, video creation system 1 performs an evaluation on the part of the person specified by the user, as indicated by text TX42, saying, "From the perspective of production, it would be better to capture the face from the front."

[0183] For example, if the user thinks the evaluation indicated by text TX42 is good, the video generation system 1 receives a correction instruction from the user based on that evaluation. Then, the video generation system 1 generates and displays a plurality of candidate videos MV43, MV44, and MV45 corresponding to the evaluation indicated by text TX42. In this way, the video generation system 1 presents the plurality of candidate videos MV43, MV44, and MV45 corresponding to the evaluation indicated by text TX42 to the user.

[0184] For example, a user specifies a range with a mouse, and the video generation system 1 evaluates the specified range and starts a dialogue (discussion) with the user. In response to the range selected with the mouse, the video generation system 1 starts an AI evaluation of that range. For example, the video generation system 1 performs an evaluation (presents correction suggestions) once every N seconds until the user instructs correction. When the user clicks with the mouse when they think the result is good, the video generation system 1 presents multiple correction suggestions based on the dialogue (discussion) up to that point. Through the above-mentioned processing, the video generation system 1 can improve usability.

[0185] <1-4-13. Processing examples according to depth of field> In conventional technology, when the environmental asset has a large number of polygons or a high texture resolution, making the asset heavy, it is difficult to suppress an increase in the time required for rendering. As such, conventional technology has usability issues and there is room for improvement in usability. Therefore, the video generation system 1 may perform processing according to the depth of field, as shown in FIG. 33. FIG. 33 is a conceptual diagram showing an example of processing according to the depth of field.

[0186] In FIG. 33, the video generation system 1 uses a high number of polygons and high-resolution texture within a predetermined range in front and behind the subject, and a low number of polygons and low-resolution texture within a predetermined range in front and behind the subject, depending on the depth of field. In FIG. 33, objects with a high number of polygons and high-resolution texture (circles close to the subject) are shown with thick hatching, and objects with a low number of polygons and low-resolution texture (circles farther from the subject) are shown with light hatching. In this way, the video generation system 1 changes the number of polygons and texture resolution depending on the depth of field. The video generation system 1 calculates the depth of field using the following equations (1) to (3).

[0187]

number

[0188]

number

[0189]

number

[0190] Equation (1) is a function for calculating the front depth of field (mm). The image generation system 1 calculates the front depth of field using equation (1). For example, in FIG. 33, the front depth of field corresponds to the front side of the subject (the side closer to the camera). Furthermore, equation (2) is a function for calculating the rear depth of field (mm). The image generation system 1 calculates the rear depth of field using equation (2). For example, in FIG. 33, the rear depth of field corresponds to the rear side of the subject (the side farther from the camera). Equation (3) is a function for calculating the depth of field. The image generation system 1 calculates the depth of field by adding the front depth of field and the rear depth of field using equation (3).

[0191] For example, the video generation system 1 replaces parts closer to the camera than the front depth of field and parts farther from the camera than the rear depth of field with parts with a lower number of polygons and texture resolution before rendering. For example, when creating blur in terms of depth of field, the video generation system 1 reduces (lowers) the number of polygons and texture resolution. In this way, the video generation system 1 determines the number of polygons and texture resolution based on criteria such as focal length, F-number, and subject distance. For example, the video generation system 1 may make the above determination when generating a USD file. Through the above-mentioned processing, the video generation system 1 can improve usability.

[0192] <1-4-14. Example of playback processing during confirmation work> In the conventional technology, when the entire generated video is played back for review, it is difficult to suppress the increase in the time required for the review work. As such, the conventional technology has usability issues, and there is room for improvement in usability. Therefore, the video generation system 1 may perform playback suitable for the review work. The video generation system 1 may change the playback mode in response to a user operation. For example, the video generation system 1 may play back the video in a mode similar to a flip book in response to a user operation.

[0193] The video production system 1 may advance the playback of the video by a predetermined number of seconds (for example, 0.5 seconds) with each click or mouse wheel rotation by the user. For example, the video production system 1 may advance frame by frame within a cut according to the movement or position of the mouse or mouse wheel. As shown in FIG. 34, the video production system 1 may advance the playback of the video by the number of seconds corresponding to the mouse position.

[0194] In FIG. 34, the video generation system 1 may provide the user with content CT41 and accept an instruction from the user as to how much to advance the video (number of seconds, number of frames, etc.). The content CT41 is a display screen (content) for accepting the user's instruction as to how much to advance the video. The content CT41 arranges information for specifying the number of seconds to advance the video, superimposed on the video. In FIG. 34, a range from 0 to 3.5 seconds can be specified, with the number of seconds to advance the video increasing from left to right. For example, the client UI display unit 400 displays the content CT41.

[0195] The user inputs information specifying the number of seconds to advance the video when playing the video via the content CT41. In FIG. 34, the user operates the mouse to position the mouse cursor MS in the area marked 1.0, thereby specifying 1.0 seconds as the number of seconds to advance the video. In this case, the video production system 1 advances playback of the video at 1.0 second intervals in accordance with the user's specification of the number of seconds to advance the video. Note that the user may also operate the mouse to position the mouse cursor MS in the area marked 1.0 and click to specify 1.0 seconds as the number of seconds to advance the video. Through the above-described processing, the video production system 1 can improve usability.

[0196] <1-4-15. Highlighting example> In the conventional technology, it may be difficult to identify generated parts of a video that have been changed by editing or the like, making it difficult to prevent an increase in the time required for the confirmation work. As such, the conventional technology has usability issues and there is room for improvement in usability. Therefore, the video generation system 1 may perform highlighting as shown in FIG. 35. FIG. 35 is a diagram showing an example of highlighting changed parts. For example, the video generation system 1 highlights the changed parts from the previous generation result. FIG. 35 explains an example in which a woman in the video has been changed.

[0197] Video MV51 in Fig. 35 shows a first highlighting mode. Video MV51 shows a mode in which the changed parts are highlighted by darkening parts other than the changed parts (by lowering the brightness, etc.). Video generation system 1 generates video MV51 and displays video MV51, thereby highlighting the parts that have been changed by editing, etc.

[0198] 35 shows a second highlighting mode. Video MV52 shows a mode in which the changed parts are highlighted (by adding color, etc.). Video generation system 1 generates video MV52 and displays video MV52 to highlight the parts that have been changed by editing, etc. In this way, videos MV51 and MV52 show cases in which parts of people that have been changed are highlighted.

[0199] 35 shows a third highlighting mode. Video MV53 shows a mode in which the changed portion is highlighted by indicating the time when the change occurred on the seek bar. Video generation system 1 generates video MV53 including a seek bar on which colored points HL531 and HL532 are placed at positions corresponding to the time when the change occurred, and by displaying video MV53, the portion that has been changed by editing or the like is highlighted.

[0200] Video MV54 in Fig. 35 shows a fourth highlighting mode. Video MV54 shows a mode in which the changed portion is highlighted by indicating the time when the change occurred on the seek bar. Video generation system 1 generates video MV54 including a seek bar with a colored bar HL54 positioned in a range corresponding to the time period when the change occurred, and by displaying video MV54, the portion that has been changed by editing or the like is highlighted.

[0201] As described above, when a user configures settings related to video generation and regenerates a video, the generated video must be checked. In a UI that presents multiple generation results, the user has to check each video one by one, which places a heavy burden on the user. Therefore, to reduce the burden of checking the video, the video generation system 1 presents the user with the differences from the previous video generation result. For example, the video generation system 1 presents the user with the differences by highlighting only the parts of the video that have changed or by displaying the time of the change in the video on a seek bar. This allows the user to check only the parts that have changed, and the video generation system 1 can reduce the burden of the checking work. Through the above-described processing, the video generation system 1 can improve usability.

[0202] For example, methods for extracting changed portions of a video include detecting differences from a generated USD file and detecting differences by comparing a rendered video with a previously rendered video frame by frame. For example, in the method of detecting differences from a generated USD file, taking advantage of the characteristic of generating USD files each time video rendering is performed, the video generation system 1 compares the contents of USD files generated before and after a user updates video generation settings to detect changed objects and the time at which the changes occurred. Also, in the method of detecting differences by comparing a rendered video with a previously rendered video frame by frame, the video generation system 1 compares the videos generated before and after a user updates video generation settings to detect changed pixels and the time at which the changes occurred.

[0203] <1-4-16. Example of showing the relationship between cuts> In the conventional technology, when each cut is corrected, there is a problem that the relationship between the previous and next cuts may become unclear. As such, the conventional technology has a problem with usability, and there is room for improvement in usability. Therefore, the video production system 1 may present the relationship between cuts as shown in Fig. 36. Fig. 36 is a diagram showing an example of presenting the relationship between cuts.

[0204] FIG. 36 shows a case where cut CU61 is a modified cut (also called a "target cut"). A previous relationship bar TR60 indicating that a cut precedes cut CU61, which is the target cut, is superimposed on cut CU60. For example, the previous relationship bar TR60 is a triangle with its base on the right and extending to the left. The previous relationship bar TR60 is displayed in such a way that the further away in time it is from the target cut, the longer it extends to the left.

[0205] Furthermore, a subsequent relationship bar TR62 indicating that a cut CU62 is a cut that comes after the target cut CU61 is superimposed on the cut CU62. For example, the subsequent relationship bar TR62 is a triangle with its base on the left side and extending to the right. A subsequent relationship bar TR63 indicating that a cut comes after the target cut is superimposed on the cut CU63 that comes after the cut CU62 that comes after the target cut CU61. For example, the subsequent relationship bar TR63 is a triangle with its base on the left side and extending to the right.

[0206] The subsequent relationship bars TR62 and TR63 are displayed in a manner that the further they are from the target cut in time, the further they extend to the left. In Figure 36, cut CU63 is later than cut CU62, so the subsequent relationship bar TR63 is displayed in a manner that extends further to the right than the subsequent relationship bar TR62. Note that the triangular display manner is merely one example of a display manner, and any display manner can be adopted as long as it is possible to present the time relationship and its amount.

[0207] The video production system 1 plays a video including cuts CU60 to CU63 in response to a user's operation. For example, when displaying cut CU60 in response to a user's operation, the video production system 1 superimposes a previous relationship bar TR60. For example, when displaying cut CU62 in response to a user's operation, the video production system 1 superimposes a next relationship bar TR62. For example, when displaying cut CU63 in response to a user's operation, the video production system 1 superimposes a next relationship bar TR63. In this way, when playing cuts before and after a target cut, the video production system 1 presents the amount of distance from the target cut. In this way, when playing cuts before and after the target cut, the video production system 1 superimposes information indicating the amount and direction of distance on the screen according to the number of seconds from the target cut, allowing the user to recognize whether the cut is before or after the target cut and how far the cut is from the target cut. Through the above-described processing, the video production system 1 can improve usability.

[0208] <1-4-17. Example of object selection> The conventional technology has a problem in that selecting an object during editing is difficult for the user (low usability). As such, the conventional technology has a problem with usability, and there is room for improvement in usability. Therefore, the video generation system 1 may select an object as shown in Fig. 37. Fig. 37 is a diagram showing an example of object selection.

[0209] In the video MV71 of Fig. 37, the user operates the mouse to select the right-hand person of the three people by positioning the mouse cursor MS71 in a range that includes the right-hand person. In this case, the video generation system 1 accepts an operation to select the right-hand person where the mouse cursor MS71 is positioned as the target object. For example, the video generation system 1 identifies the segmentation (range) corresponding to the right-hand person where the mouse cursor MS71 is positioned as the range selected by the user.

[0210] 37, the user operates the mouse to select the central person by positioning the mouse cursor MS72 in a range that includes the central person among the three people. In this case, the image generation system 1 accepts an operation to select the central person where the mouse cursor MS72 is positioned as the target object. For example, the image generation system 1 identifies the segmentation (range) corresponding to the central person where the mouse cursor MS72 is positioned as the range selected by the user.

[0211] 37, the user operates the mouse to select the left person of the three people by positioning the mouse cursor MS73 in a range that includes the left person. In this case, the image generation system 1 accepts an operation to select the left person where the mouse cursor MS73 is positioned as the target object. For example, the image generation system 1 identifies the segmentation (range) corresponding to the left person where the mouse cursor MS73 is positioned as the range selected by the user.

[0212] In this way, the video generation system 1 recognizes the selection of the target object within the range recognized by segmentation. This allows the user to select the object they want to specify with just a click. For example, when a user wants to change the motion or asset of an object displayed in a video, they need to select the object to be changed. If the user can click the object in the video directly, rather than using a UI that requires the user to select the object to be changed from a list of objects displayed in the video, the user can select the object to be changed more intuitively. The click range of an object in a video can be achieved by segmenting the pixel area of ​​each object for each frame of the video.

[0213] The click range of an object may be calculated in the following manner: For example, because object and camera information for each frame of a video is stored in USD, it is possible to map which object a pixel seen from the camera in that frame points to, and therefore the video generation system 1 can reverse-calculate from this to map the coordinates where the user clicked on the video to the object to be edited.

[0214] One method for determining the click range of an object is to associate a person selected by segmenting the object pixel by pixel calculated backward from the camera for each frame of the video with the target USD. For example, the video itself can include information about what is located in this area (x, y coordinates, etc.). For example, the video generation system 1 may select an object as shown in Figure 38.

[0215] Frame FR in FIG. 38 indicates one frame of a video rendered from a camera. USD data OB in FIG. 38 indicates a 3D object and a camera on the USD. Because information about the 3D object and camera is included in the USD, the video generation system 1 can map which location on the 3D object a pixel viewed from a camera (such as camera CM in FIG. 38) points to. This allows the video generation system 1 to map the coordinates where a user clicks on a video to an object. Through the above-described processing, the video generation system 1 can improve usability.

[0216] <1-4-18. 3D model usage example> The above-described processing is merely an example, and the image generation system 1 may execute various processing other than the above-described processing. In this regard, several examples will be described below.

[0217] For example, the video generation system 1 may perform a process of importing characters or props (such as products). A method of importing characters or props (such as products) in the video generation system 1 may involve, for example, registering a 3D model or a three-dimensional drawing of a character or prop to be included in a video and using that model in the video. For example, the video generation system 1 may register a 3D model input by a user and use that model in the video. For example, the video generation system 1 may be capable of incorporating photos of a product taken from various angles. In this case, the video generation system 1 may use technology such as Neural Radiance Fields (NeRF) to generate a 3D model of the product from photos of the product taken from various angles and use that model in the video.

[0218] For example, the video generation system 1 accepts a user operation specifying an object to be changed in video data, and generates code in which the 3D data of the object specified by the user operation has been changed. For example, if a character in a video (also referred to as "character A") is specified by the user as the object to be changed, and the user selects another character (also referred to as "character B") from a registered 3D model as the changed character, the video generation system 1 generates code in which the 3D data of character A in the video has been changed to the 3D data of character B. This allows the video generation system 1 to generate video data in which character A in the video has been changed to character B indicated by the registered 3D model.

[0219] The above-described process is merely an example, and the image generation system 1 may generate a code in which the 3D data of an object indicated by a user operation has been modified by any process. For example, the image generation system 1 may generate a modified code by modifying the 3D data itself of an object indicated by a user operation in the video data. For example, the image generation system 1 may generate a modified code by modifying the external shape (height, etc.) of the 3D data of an object indicated by a user operation in the video data. Furthermore, the image generation system 1 may not need to partially perform refinement processing on registered people and props after rendering. Furthermore, the image generation system 1 may provide a marketplace within the image generation service for selling characters, props, etc.

[0220] <1-4-19. Examples of using reference data> Furthermore, the video production system 1 may refer to the video and story of a previously created (video) project. When a sequel to a previously created project is desired, the video production system 1 may receive input of the project as reference data, as shown in Fig. 39. Fig. 39 is a diagram showing an example of processing using reference data.

[0221] For example, the video production system 1 acquires user input information regarding reference to other projects that the user inputs to the content CT51. The content CT51 is content for accepting user input information regarding an item for specifying with a check mark whether or not to refer to other projects, an item for specifying with a check mark the projects to refer to, an item for specifying with a check mark which information of the reference projects to refer to, etc.

[0222] For example, the client UI display unit 400 displays the content CT51, and the sensor unit 300 accepts user input information. In Fig. 39, the user checks "Use other projects as reference" to select other projects as reference. The user also checks "Product X commercial video" to select the project of a commercial video for Product X as reference.

[0223] The user also checks "Characters" and "Visual Style," choosing to use the characters and visual style of the commercial video project for Product X as a reference. The user also does not check "Story," "Content Style," or "BGM," choosing not to use the story, content style, or background music of the commercial video project for Product X as a reference.

[0224] The video production system 1 generates a prompt PT51, which is scenario generation information (first input information), using user input information accepted by the content CT51. Although not illustrated in Fig. 39, the video production system 1 may also generate the prompt PT51 using template input information such as the template TP1 shown in Fig. 6.

[0225] For example, the video production system 1 generates prompt PT51 by reflecting the characters and visual style of a commercial video project for product X. In FIG. 39, the video production system 1 generates prompt PT51 including a constraint specifying that the character "Mike" from the commercial video for product X and the visual style "cinematic" from the commercial video for product X be used. In this way, the video production system 1 generates a prompt for generating a scenario based on reference data from past projects in accordance with the user's selection. This allows the video production system 1 to generate video data for character A in the video based on the past project specified by the user.

[0226] <1-4-20. Examples of advantages of having 3D data> Because the image generation system 1 has 3D data internally, it has the following functions and advantages. The image generation system 1 has the following functions and advantages with regard to the images (videos) it generates. For example, the image generation system 1 can create more natural images through physical simulation. For example, the image generation system 1 can place an object on a platform, have it bounce, roll, or reproduce the natural swaying of cloth.

[0227] For example, the image generation system 1 can realistically reproduce the reflection of light depending on the lighting and object material. When a hood or mirror is highly reflective, people or objects in front of it will be brighter due to the reflected light. When a cloth is less reflective, people or objects in front of it will not receive much reflected light.

[0228] For example, the video generation system 1 can fix lighting and objects, reducing the possibility of time-series disruptions. For example, the video generation system 1 can reproduce video without disruptions even if the lighting position changes during the video. For example, the video generation system 1 can output video in real time by setting a simple light source. In this case, the video generation system 1 does not need to perform refinement processing.

[0229] For example, by inputting product data and characters as 3D data, the video generation system 1 can faithfully reproduce the product itself within the video. For example, because the video generation system 1 knows the depth, global position, normal, etc., it can reduce the possibility of failure during refiner processing. For example, because the video generation system 1 can fix the 3D position of sound sources such as speakers, it is easy to create interactions with sound. For example, the video generation system 1 can reduce rendering time by performing lighting processing such as Differential Rendering.

[0230] Furthermore, the video production system 1 has the following functions and advantages when it comes to editing (correcting) videos (moving images). For example, the video production system 1 can maintain consistency of the surrounding environment and lighting even when the camera position, angle, and camerawork are significantly corrected. For example, the video production system 1 can specify the position and angle three-dimensionally within a video.

[0231] For example, even after a certain amount of video has been generated, the video generation system 1 can change only specific elements, such as the characters or props, while maintaining the video for the rest of the video. For example, the video generation system 1 can change a character from a realistic human-like figure to a two-thirds-tall character. For example, the video generation system 1 can change an installed signboard from a blackboard type to a plastic board. For example, the video generation system 1 can change only a part of the logo on a product package.

[0232] For example, by viewing a captured scene in 3D, the image generation system 1 can correct images of areas that are not visible in the video frames. For example, the image generation system 1 can place a light source in an area that is not visible. For example, the image generation system 1 can place a person or object in an area that is not visible and display only their shadow in the video frame. For example, the image generation system 1 can place a person or object in an area that is not visible and specify that the person in the image should look into the eyes of the person not in the image.

[0233] In addition to the above, the image generation system 1 has the following functions and advantages: For example, the image generation system 1 can easily convert content into content for 3D devices such as AR, VR (Virtual Reality), SRD (Spatial Reality Display), and 3D displays.

[0234] <1-5. Example of processing flow from the user's perspective> Next, the procedure of information processing by the image generation system 1 in response to a user operation will be described as an example of a processing flow seen from the user's perspective, with reference to Fig. 40. Fig. 40 is a flowchart showing the flow of processing in response to a user operation.

[0235] 40, in the video production system 1, the user presses a project creation button (step S1). Then, in the video production system 1, the user inputs information required for video generation (step S2). For example, the user inputs information including the purpose of creating the video, the message to be conveyed through the video, the target users, the functional features of the product / service to be conveyed, the length of the video, the aspect ratio, etc.

[0236] Then, the user presses the storyboard, video, sound, and text logo creation buttons in the video generation system 1 (step S3). For example, the video generation system 1 may generate everything (including video data, for example) at once in response to user operations, or may present a storyboard and allow the user to make some modifications in response to instructions from the user before generating the video, sound, and text logo.

[0237] Then, the image production system 1 modifies the storyboard, video, sound, and text logo in response to user operations (step S4). For example, the image production system 1 may accept modifications to each of the items in the order of the user's preference.

[0238] Then, the video production system 1 performs export in response to a user operation (step S5). For example, the video production system 1 may perform export in a video file format such as mp4, avi, or mov in response to a user operation. Furthermore, for example, the video production system 1 may perform export in a file format of any video editing software such as Premiere Pro, After Effects, or DaVinci Resolve.

[0239] <1-6. About AI models> Note that the AI ​​models used in the above-described processes are not limited to the examples given in each section, and any internal structure can be adopted as long as it is possible to output desired information in response to input. Any combination of input, output, and internal structure of the AI ​​model can be adopted as long as it is possible to output desired information.

[0240] The input of the AI ​​model may be text, images, audio, 3D data, etc., or a combination thereof. The output of the AI ​​model may be text, images, audio, 3D data, etc. Note that the above-mentioned inputs and outputs are merely examples, and the above-mentioned AI model may have any inputs and outputs.

[0241] Furthermore, the internal structure of the AI ​​model can be any structure depending on the combination of input and output. In other words, the internal structure of the AI ​​model can be any structure as long as it can produce a desired output for the input.

[0242] For example, the AI ​​model may have a structure related to Transformer. For example, the AI ​​model may have a structure related to Transformer and perform processing taking into account context, such as context within data, such as text or time-series data. For example, the AI ​​model may have a self-attention mechanism. For example, the AI ​​model may have any attention mechanism, such as Single-Head Attention or Multi-Head Attention. Note that the AI ​​model does not necessarily have to have an attention mechanism.

[0243] The AI ​​model may have a mechanism for extracting features from an input. For example, the AI ​​model may have an encoder. The AI ​​model may have a mechanism for generating information based on the extracted features. For example, the AI ​​model may have a decoder.

[0244] The AI ​​model may have a structure related to a convolutional neural network (CNN). For example, when processing an image, the AI ​​model may have a structure related to a CNN. For example, the AI ​​model may have at least one of a convolution layer, a pooling layer, a fully connected layer, etc.

[0245] The above-described internal structure is merely an example, and the AI ​​model may have any internal structure. For example, the AI ​​model may have a skip connection. Furthermore, the AI ​​model may have a structure related to a diffusion model.

[0246] Furthermore, the above-described AI model may be generated (trained) by any learning process. The AI ​​model may be a machine learning model trained using any machine learning method. For example, the AI ​​model may be a model generated based on a so-called Foundation Model by fine-tuning the Foundation Model to apply it to a specific task (e.g., scenario data generation, code generation, etc.). For example, an AI model such as the above-described LLM may be a model generated by fine-tuning the Foundation Model to apply it to a specific task.

[0247] The base model here is a model that has been trained to be applicable to various tasks, for example, to be able to perform a wide variety of tasks. For example, the base model is a neural network that has been pre-trained with a large amount of unlabeled data set. Note that the base model may have any structure, such as a Transformer-based architecture. For example, the base model is generated by self-supervised learning using data without correct answer labels. As described above, the base model is fine-tuned so that it can be adapted to a wide range of downstream tasks.

[0248] For example, when applied to a scenario data generation task, the base model is fine-tuned so that it can be adapted to the scenario data generation task, and an AI model (model M1, etc.) adapted to the scenario data generation task is generated. Also, when applied to a code generation task, for example, the base model is fine-tuned so that it can be adapted to the code generation task, and an AI model (model M3, etc.) adapted to the code generation task is generated. Also, when applied to a sound data generation task, for example, the base model is fine-tuned so that it can be adapted to the sound data generation task, and an AI model (model M4, etc.) adapted to the sound data generation task is generated. Also, when applied to a text logo generation task, for example, the base model is fine-tuned so that it can be adapted to the text logo generation task, and an AI model (model M5, etc.) adapted to the text logo generation task is generated.

[0249] For example, an AI model (such as model M1) applied to a scenario data generation task is trained using training data including a combination of input information corresponding to the AI ​​model and scenario data (also referred to as "correct answer information") that is the correct output when the input information is input. The training data, such as the input information and correct answer information, may be data created by a person or data automatically generated by a computer that generates the training data. For example, the scenario data that is the correct answer information may be data created by a person. Below, model M1 will be briefly described as an example. For example, model M1 is trained to output correct answer information corresponding to each piece of input information when that input information is input. For example, model M1 is trained by adjusting (correcting) parameters (connection coefficients) using a method such as backpropagation (error backpropagation) so as to reduce the error between the output of model M1 when certain input information is input and the correct answer information corresponding to that input information. In addition, other AI models such as an AI model applied to a code generation task (e.g., model M3), an AI model applied to a sound data generation task (e.g., model M4), an AI model applied to a text logo generation task (e.g., model M5), an AI model applied to an evaluation task (e.g., model M10), and an AI model applied to an image improvement processing task (e.g., model M11) may also be trained using a similar learning process.

[0250] The above-described learning process is merely an example, and the AI ​​model described above may be trained by any learning process depending on the input, output, and internal structure of the AI ​​model. For example, the AI ​​model may be trained by an unsupervised learning method such as a generative adversarial network (GAN). The AI ​​model may also be trained in a distributed manner without aggregating data, such as federated learning. In this case, each video generation service device (e.g., server) may generate local models collected by that service, and a server (aggregation server) that aggregates information (e.g., parameters) of the local models generated by each video generation service device (e.g., server) may generate a global model using the information on the local models. In this case, the video generation system 1 may receive the global model generated by the aggregation server from the aggregation server and use the received global model as an AI model for processing.

[0251] In this way, the above-mentioned AI models may be generated (learned) by any computer. That is, the learning process to generate the AI ​​models may be performed by any device (computer, etc.) in the video production system 1, or may be performed by a device outside the video production system 1. For example, when a device outside the video production system 1 generates at least one of the above-mentioned AI models, the video production system 1 acquires the AI ​​model from the device outside the video production system 1 and performs processing using the acquired AI model.

[0252] <2. Other embodiments> The processing according to each of the above-described embodiments may be implemented in various different forms (modifications) other than the above-described embodiments and modifications.

[0253] <2-1. Other configuration examples> The above-described configuration of the image generation system 1 is merely an example, and any desired division of functions in the image generation system 1 can be adopted. In other words, the above-described configuration is merely an example, and the image generation system 1 may have any desired division of functions and any desired configuration as long as it can provide the above-described services related to image generation. For example, the image generation system 1 may be configured by a single device (such as a computer) that performs the above-described processing. In this case, one device of the image generation system 1 may have the functions of the image generation module 100, the information acquisition module 200, the sensor unit 300, and the client UI display unit 400. For example, the image generation service provided by the image generation system 1 may be provided to a user as a program such as a tool (AI Assist Creation Tool) that runs on a terminal device (such as the computer 20) used by the user.

[0254] <2-2.Other> Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. Furthermore, the information, including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings, can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0255] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0256] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0257] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0258] 3. Effects of the present disclosure As described above, the video generation system according to the present disclosure (video generation system 1 in the embodiment) includes an acquisition unit (input text acquisition unit 210 in the embodiment), a scenario generation unit (scenario-oriented generation unit 131 in the embodiment), a code generation unit (video-oriented generation unit 132 in the embodiment), and a video data acquisition unit (video generation unit 140 in the embodiment). The acquisition unit acquires an input query related to video generation from a user. The scenario generation unit generates scenario data related to video generation based on the input query. The code generation unit generates code for constructing 3D data based on the scenario data. The video data acquisition unit acquires video data based on the code.

[0259] In this way, the video generation system of the present disclosure generates code for constructing 3D data based on scenario data generated based on an input query from a user, and obtains video data based on the code, thereby being able to obtain video data in response to an input query from a user.

[0260] The video generation system also includes an image quality improvement unit (in this embodiment, an image refinement unit 143). The image quality improvement unit executes image quality improvement processing to improve the image quality of the video data. In this way, the video generation system can obtain high-quality video by improving the image quality of the video data.

[0261] The image quality improvement unit improves the image quality of the video data by performing an image quality improvement process based on the text prompt. In this way, the video generation system can obtain high-quality video by improving the image quality of the video data based on the text prompt.

[0262] The image quality improvement unit also performs the image quality improvement process based on a text prompt that specifies the target of the image quality improvement process in the video data. In this way, the video production system can obtain high-quality video by improving the image quality of the video data for the specified target.

[0263] The image quality improvement unit executes the image quality improvement process based on a text prompt specifying the target determined to require improvement. In this way, the video production system can obtain high-quality video by improving the image quality of the video data for the target determined to require improvement.

[0264] The video generation system also includes a display control unit (in this embodiment, a client UI display unit 400). The display control unit displays a storyboard based on the scenario data. In this way, the video generation system can present information in a manner that is highly convenient for the user by displaying a storyboard based on the scenario data.

[0265] The storyboard is configured to display video data for each cut of the video. In this way, the video production system is able to present information in a manner that is highly convenient for the user by configuring the storyboard to display video data for each cut of the video.

[0266] The video production system also includes a sound production unit (sound production unit 150 in this embodiment). The sound production unit generates sound data corresponding to the video data based on the scenario data and the video data. In this way, the video production system can obtain a video including sound by generating sound data corresponding to the video data.

[0267] The video production system also includes a text production unit (in this embodiment, a text / logo production unit 160). The text production unit produces text data indicating text to be displayed on a video produced by video data, based on the scenario data. In this way, the video production system can obtain a video containing text by producing text data corresponding to the video data.

[0268] The video production system also includes a logo production unit (text / logo production unit 160 in this embodiment). Based on the scenario data, the logo production unit generates logo data indicating a logo to be displayed on a video displayed by the video data. In this way, the video production system can obtain a video including a logo by generating logo data corresponding to the video data.

[0269] Furthermore, the input query includes at least one of text, image, audio, and 3D data. In this way, the video generation system can acquire video data in response to the input query from the user by the input query including at least one of text, image, audio, and 3D data.

[0270] The video generation system also includes a first output unit (a scenario-oriented generation unit 131 in this embodiment). The first output unit outputs scenario generation information used by the scenario generation unit to generate scenario data based on the input query. The scenario generation unit generates scenario data based on the scenario generation information. In this way, the video generation system can acquire video data in response to the input query from the user by generating scenario data based on the scenario generation information generated based on the input query.

[0271] Furthermore, the first output unit generates, as scenario generation information, a first prompt for generating scenario data based on the input query. The scenario generation unit generates the scenario data based on the first prompt. In this way, the video generation system can acquire video data in response to the input query from the user by generating scenario data based on the first prompt.

[0272] Furthermore, the first output unit generates, based on the input query, first input information as scenario generation information to be used as input to a first model for generating scenario data. The scenario generation unit inputs the first input information generated using the input query to the first model and causes the first model to output scenario data, thereby generating scenario data. In this way, the video generation system can acquire video data in response to an input query from a user by generating scenario data using the first model.

[0273] The video generation system also includes a second output unit (video generation unit 132 in this embodiment). The second output unit outputs code generation information used by the code generation unit to generate code for configuring 3D data based on scenario data. The code generation unit generates code based on the code generation information. In this way, the video generation system can acquire video data in response to an input query from a user by generating code based on the code generation information generated based on scenario data.

[0274] Furthermore, the second output unit generates, as code generation information, a second prompt for outputting code for configuring 3D data based on the scenario data. The scenario generation unit generates the code based on the second prompt. In this way, the video generation system can acquire video data in response to an input query from a user by generating the code based on the second prompt.

[0275] Furthermore, the second output unit generates, based on the input query, second input information as code generation information to be used as input to a second model for generating code. The scenario generation unit inputs the second input information generated using the scenario data into the second model and causes the second model to output code, thereby generating code. In this way, the video generation system can acquire video data in response to an input query from a user by generating code using the second model.

[0276] The video production system also includes a reception unit (sensor unit 300 in this embodiment). The reception unit receives video editing operations from a user. The code generation unit generates code based on the editing operations. In this way, the video production system generates code in response to the video editing operations from the user, thereby enabling the video production system to appropriately acquire video corresponding to the user's editing.

[0277] The receiving unit receives an operation specifying a motion or camera movement using a sensor. The code generating unit generates a code corresponding to the motion or camera movement indicated by the operation. In this way, the video generation system generates a code in response to a user's operation specifying a motion or camera movement, thereby enabling the video generation system to appropriately acquire video corresponding to the user's editing.

[0278] The receiving unit also receives an operation to select multiple cuts from the video data. The code generating unit generates code in which portions corresponding to the multiple cuts indicated by the operation have been changed. In this way, the video generation system generates code in response to the user's selection of multiple cuts, thereby enabling appropriate acquisition of video corresponding to the user's editing.

[0279] Furthermore, date information is associated with each cut of the video data. The code generation unit determines the content of the edit indicated by the operation based on the date information of each cut of the video data. In this way, the video generation system can appropriately acquire video corresponding to the user's editing by determining the content of the edit based on the date information of each cut of the video data.

[0280] The receiving unit also receives an operation to specify an object to be changed from among the video data. The code generating unit generates a code in which the 3D data of the object specified by the operation has been changed. In this way, the video generation system can appropriately acquire a video corresponding to the user's editing by generating a code in which the 3D data specified by the user as the object to be changed has been changed.

[0281] The video production system also includes an evaluation unit (evaluation unit 180 in this embodiment). The evaluation unit generates information indicating an evaluation of at least one of the scenario data and the video data. In this way, the video production system can evaluate the generated information by generating information indicating an evaluation of at least one of the scenario data and the video data.

[0282] Furthermore, the code generator generates the code based on the evaluation. In this way, the video generation system generates the code based on the evaluation, thereby enabling appropriate acquisition of information in accordance with the evaluation.

[0283] Furthermore, the scenario generation unit generates scenario data based on the evaluation. In this way, the image generation system generates scenario data based on the evaluation, thereby enabling appropriate acquisition of information in accordance with the evaluation.

[0284] Furthermore, the scenario generation unit generates a code based on the scenario data generated based on the evaluation. In this way, the video generation system generates a code based on the scenario data generated based on the evaluation, thereby enabling appropriate acquisition of information according to the evaluation.

[0285] <4. Hardware Configuration> An information processing device (information appliance) having the image generation module 100, information acquisition module 200, client UI display unit 400, etc. according to each of the above-described embodiments is realized by, for example, a computer 1000 configured as shown in FIG. 41. FIG. 41 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the information processing device. The image generation module 100 according to the embodiment will be described below as an example. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0286] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0287] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0288] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an image generation program according to the present disclosure, which is an example of program data 1450.

[0289] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0290] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disk), magneto-optical recording media such as an MO (Magneto-Optical disk), tape media, magnetic recording media, and semiconductor memories.

[0291] For example, when the computer 1000 functions as the image generation module 100 according to the embodiment, the CPU 1100 of the computer 1000 executes an image generation program loaded onto the RAM 1200, thereby realizing the functions of the control unit 1301, etc. The image generation program according to the present disclosure and data in the storage unit 1302 are stored in the HDD 1400. The CPU 1100 reads and executes program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0292] The present technology can also be configured as follows. (1) an acquisition unit that acquires an input query related to video generation from a user; a scenario generation unit that generates scenario data related to video generation based on the input query; a code generation unit that generates a code for configuring 3D data based on the scenario data; a video data acquisition unit that acquires video data based on the code; A video generation system comprising: (2) an image quality improvement unit that executes an image quality improvement process to improve the image quality of the video data; The image generation system according to (1) further comprises: (3) The image quality improvement unit The image quality of the video data is improved by performing the image quality improvement process based on a text prompt. (2) The image generation system according to (1). (4) The image quality improvement unit The image quality improvement process is performed based on a text prompt that specifies a target of the image quality improvement process among the video data. The image generation system according to (2) or (3). (5) The image quality improvement unit and performing the image quality improvement process based on a text prompt specifying the target determined to require improvement. (4) A video generation system according to (4). (6) a display control unit that displays a storyboard based on the scenario data; The image generation system according to any one of (1) to (5), further comprising: (7) The storyboard is configured to display the video data for each cut of the video. (6) An image generation system according to (6). (8) a sound generation unit that generates sound data corresponding to the video data based on the scenario data and the video data; The image generation system according to any one of (1) to (7), further comprising: (9) a text generation unit that generates text data indicating text to be displayed on a video displayed by the moving image data based on the scenario data; The image generation system according to any one of (1) to (8), further comprising: (10) a logo generation unit that generates logo data indicating a logo to be displayed on a video displayed based on the moving image data, based on the scenario data; The image generation system according to any one of (1) to (9), further comprising: (11) The input query includes at least one of text, image, audio, and 3D data. The image generation system according to any one of (1) to (10). (12) a first output unit that outputs scenario generation information used by the scenario generation unit to generate the scenario data based on the input query; Further provided with The scenario generation unit The scenario data is generated based on the scenario generation information. The image generation system according to any one of (1) to (11). (13) The first output unit generating a first prompt for generating the scenario data based on the input query as the scenario generation information; The scenario generation unit generating the scenario data based on the first prompt; (12) The image generation system according to (12). (14) The first output unit generating, based on the input query, first input information to be used as an input of a first model for generating the scenario data, as the scenario generation information; The scenario generation unit The first input information generated using the input query is input to the first model, and the first model is caused to output the scenario data, thereby generating the scenario data. The image generation system according to (12) or (13). (15) a second output unit that outputs code generation information used by the code generation unit to generate code for configuring the 3D data based on the scenario data; Further provided with The code generation unit Generate the code based on the code generation information The image generation system according to any one of (1) to (14). (16) The second output unit generating, as the code generation information, a second prompt for outputting a code for configuring the 3D data based on the scenario data; The scenario generation unit Generate the code based on the second prompt (15) An image generation system according to (15). (17) The second output unit generating, based on the input query, second input information to be used as an input of a second model for generating the code, as the code generation information; The scenario generation unit The second input information generated using the scenario data is input to the second model, and the second model is caused to output the code, thereby generating the code. The image generation system according to (15) or (16). (18) a reception unit that receives operations related to video editing from a user; Further provided with The code generation unit The code is generated by editing based on the operation. The image generation system according to any one of (1) to (17). (19) The reception unit Accepting the operation specifying motion or camera movement using a sensor; The code generation unit Generate the code corresponding to the motion or camera movement indicated by the operation. (18) The image generation system according to (18). (20) The reception unit accepting the operation of selecting a plurality of cuts from the video data; The code generation unit Generate the code in which the portions corresponding to the plurality of cuts indicated by the operation have been changed The image generation system according to (18) or (19). (twenty one) Date information is associated with each cut of the video data, The code generation unit The content of the editing indicated by the operation is determined based on date information of each cut of the video data. The image generation system according to any one of (18) to (20). (twenty two) The reception unit accepting the operation to designate an object to be changed from among the video data; The code generation unit Generate the code in which the 3D data of the object indicated by the operation has been changed. The image generation system according to any one of (18) to (21). (twenty three) the 3D data comprises a plurality of data sets; The code generation unit generating the code corresponding to at least one of the plurality of data sets by editing based on the operation; The image generation system according to any one of (18) to (22). (twenty four) The code generation unit Depending on the editing content indicated by the operation, one of a process of updating a part of the plurality of data sets and a process of updating the plurality of data sets in its entirety is executed. (23) The image generation system according to (23). (twenty five) an evaluation unit that generates information indicating an evaluation of at least one of the scenario data and the video data; The image generation system according to any one of (1) to (24), further comprising: (26) The code generation unit generating the code based on the evaluation (25) An image generation system according to (25). (27) The scenario generation unit generating the scenario data based on the evaluation; The image generation system according to (25) or (26). (28) The code generation unit generating the code based on the scenario data generated based on the evaluation; (27) An image generation system according to (27). (29) Obtaining an input query for video generation from a user; generating scenario data related to video generation based on the input query; generating a code for constructing 3D data based on the scenario data; acquiring video data based on the code; A video generation method including: (30) Obtaining an input query for video generation from a user; generating scenario data related to video generation based on the input query; generating a code for constructing 3D data based on the scenario data; acquiring video data based on the code; An image generation program that causes a computer to execute the above. [Explanation of symbols]

[0293] 1. Video Generation System 100 Image Generation Module 110 Input text analysis unit 120 Sensor Analysis Unit 130 Prompt Generation Unit 131 Scenario Generation Unit 132 Video Generation Unit 133 Sound Generation Unit 134 Text / Logo Generator 140 Image Generation Unit 141 USD generation section 142 Rendering Department 143 Video Refinement Department 150 Sound Generation Unit 160 Text / Logo Generator 170 Composite Editorial Department 180 Evaluation Department 190 Client UI Module 200 Information Acquisition Module 210 Input text acquisition unit 220 Sensor Acquisition Unit 300 Sensor unit 400 Client UI display section

Claims

1. an acquisition unit that acquires an input query related to video generation from a user; a scenario generation unit that generates scenario data related to video generation based on the input query; a code generation unit that generates a code for configuring 3D data based on the scenario data; a video data acquisition unit that acquires video data based on the code; A video generation system comprising:

2. an image quality improvement unit that executes an image quality improvement process to improve the image quality of the video data; The video production system of claim 1 further comprising:

3. The image quality improvement unit The image quality of the video data is improved by performing the image quality improvement process based on a text prompt. The video production system of claim 2 .

4. The image quality improvement unit The image quality improvement process is performed based on a text prompt that specifies a target of the image quality improvement process among the video data. The video production system of claim 2 .

5. The image quality improvement unit and performing the image quality improvement process based on a text prompt specifying the target determined to require improvement. The video production system of claim 4 .

6. a display control unit that displays a storyboard based on the scenario data; The video production system of claim 1 further comprising:

7. The storyboard is configured to display the video data for each cut of the video. The video production system of claim 6 .

8. a sound generation unit that generates sound data corresponding to the video data based on the scenario data and the video data; The video production system of claim 1 further comprising:

9. a text generation unit that generates text data indicating text to be displayed on a video displayed by the moving image data based on the scenario data; The video production system of claim 1 further comprising:

10. a logo generation unit that generates logo data indicating a logo to be displayed on a video displayed based on the moving image data, based on the scenario data; The video production system of claim 1 further comprising:

11. The input query includes at least one of text, image, audio, and 3D data. The video production system of claim 1 .

12. a first output unit that outputs scenario generation information used by the scenario generation unit to generate the scenario data based on the input query; Further provided with The scenario generation unit The scenario data is generated based on the scenario generation information. The video production system of claim 1 .

13. The first output unit generating a first prompt for generating the scenario data based on the input query as the scenario generation information; The scenario generation unit generating the scenario data based on the first prompt; The video production system of claim 12.

14. The first output unit generating, based on the input query, first input information to be used as an input of a first model for generating the scenario data, as the scenario generation information; The scenario generation unit The first input information generated using the input query is input to the first model, and the first model is caused to output the scenario data, thereby generating the scenario data. The video production system of claim 12.

15. a second output unit that outputs code generation information used by the code generation unit to generate a code for configuring the 3D data based on the scenario data; Further provided with The code generation unit Generate the code based on the code generation information The video production system of claim 1 .

16. The second output section generating, as the code generation information, a second prompt for outputting a code for configuring the 3D data based on the scenario data; The scenario generation unit generating the code based on the second prompt The video production system of claim 15.

17. The second output section generating, as the code generation information, second input information to be used as an input of a second model for generating the code based on the input query; The scenario generation unit The second input information generated using the scenario data is input to the second model, and the second model is caused to output the code, thereby generating the code. The video production system of claim 15.

18. a reception unit that receives operations related to video editing from a user; Further provided with The code generation unit The code is generated by editing based on the operation. The video production system of claim 1 .

19. The reception unit Accepting the operation specifying motion or camera movement using a sensor; The code generation unit Generate the code corresponding to the motion or camera movement indicated by the operation.

20. The video production system of claim 18.

20. The reception unit accepting the operation of selecting a plurality of cuts from the video data; The code generation unit Generate the code in which the portions corresponding to the plurality of cuts indicated by the operation have been changed 20. The video production system of claim 18.

21. Date information is associated with each cut of the video data, The code generation unit The content of the editing indicated by the operation is determined based on date information of each cut of the video data.

20. The video production system of claim 18.

22. The reception unit accepting the operation to designate an object to be changed from among the video data; The code generation unit Generate the code in which the 3D data of the object indicated by the operation has been changed.

20. The video production system of claim 18.

23. the 3D data includes a plurality of data sets; The code generation unit generating the code corresponding to at least one of the plurality of data sets by editing based on the operation; 20. The video production system of claim 18.

24. The code generation unit Depending on the editing content indicated by the operation, one of a process of updating a part of the plurality of data sets and a process of updating the plurality of data sets in its entirety is executed.

24. The video production system of claim 23.

25. an evaluation unit that generates information indicating an evaluation of at least one of the scenario data and the video data; The video production system of claim 1 further comprising:

26. The code generation unit generating the code based on the evaluation 26. The image generation system of claim 25.

27. The scenario generation unit generating the scenario data based on the evaluation; 26. The image generation system of claim 25.

28. The code generation unit generating the code based on the scenario data generated based on the evaluation; 28. The image generation system of claim 27.

29. Obtaining an input query for video generation from a user; generating scenario data related to video generation based on the input query; generating a code for constructing 3D data based on the scenario data; acquiring video data based on the code; A video generation method including:

30. Obtaining an input query for video generation from a user; generating scenario data related to video generation based on the input query; generating a code for constructing 3D data based on the scenario data; acquiring video data based on the code; An image generation program that causes a computer to execute the above.

Citation Information

Patent Citations

  • User-customized meta content providing system based on artificial neural network and method therefor

    KR102508765B1

  • Automated story generation

    US20110249953A1

  • Information processing device, information processing method, and program

    WO2023002659A1

  • Figure animation creation system

    JP2003058906A