Video generation system, video generation method, and video generation program

The image generation system addresses the burden of preparing two-dimensional drawings by generating 3D scene data in response to user queries, facilitating efficient video creation with enhanced usability.

WO2025220361A1PCT designated stage Publication Date: 2025-10-23SONY GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/008918
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-15
Filing Date
2025-03-11
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Conventional image generation systems require users to prepare two-dimensional line drawings, which is burdensome and limits their ability to generate videos, especially when they cannot provide such drawings.

Method used

An image generation system that includes an acquisition unit for user queries, a scenario generation unit, a scene configuration data acquisition unit, and a scene generation unit to generate 3D scene data in response to user inputs, reducing the burden on users and enhancing usability.

Benefits of technology

Enables the generation of video data with reduced user input requirements, allowing for high usability and efficient video creation without the need for complex image preparation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025008918_23102025_PF_FP_ABST
    Figure JP2025008918_23102025_PF_FP_ABST
Patent Text Reader

Abstract

A video generation system according to the present disclosure comprises: an acquisition unit that acquires, from a user, an input query relating to video generation; a scenario generation unit that generates scenario data relating to video generation, on the basis of the input query; a scene configuration data acquisition unit that acquires, on the basis of the scenario data, a plurality of types of scene configuration data constituting 3D scene data; and a scene generation unit that generates 3D scene data relating to video generation on the basis of the plurality of types of scene configuration data.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation system, image generation method, and image generation program

[0001] The present disclosure relates to an image generation system, an image generation method, and an image generation program.

[0002] There have been proposed techniques for automatically generating images (also called "videos"), such as a technique for estimating the three-dimensional posture of a virtual person from a two-dimensional line drawing and generating a video (see, for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2003-058906

[0004] However, there is room for improvement in the conventional technology. For example, in the conventional technology, two-dimensional line drawings, i.e., images, are required to generate a video. Preparing images such as two-dimensional line drawings places a heavy burden on the user, making it difficult to generate a video when the user is unable to prepare an image. Therefore, there is a need for a video generation service that places less burden on the user and has high usability, and there is a need for, for example, a service that generates data related to a video in response to a query input from a user.

[0005] Therefore, the present disclosure proposes an image generation system, an image generation method, and an image generation program that can generate data related to moving images in response to a query input from a user.

[0006] In order to solve the above problems, one embodiment of a video generation system according to the present disclosure includes an acquisition unit that acquires an input query related to video generation from a user, a scenario generation unit that generates scenario data related to video generation based on the input query, a scene configuration data acquisition unit that acquires multiple types of scene configuration data that constitute 3D scene data based on the scenario data, and a scene generation unit that generates 3D scene data related to video generation based on the multiple types of scene configuration data.

[0007] 1 is a diagram illustrating an example of an image generation system according to the present disclosure. FIG. 1 is a diagram illustrating an example of a hardware configuration of an image generation system according to the present disclosure. FIG. 2 is a diagram illustrating an example of a flow of image generation processing according to the present disclosure. FIG. 3 is a diagram illustrating another example of the flow of image generation processing according to the present disclosure. FIG. 4 is a diagram illustrating an example of a flow of evaluation processing according to the present disclosure. FIG. 5 is a diagram illustrating an example of a generation process of information for scenario generation. FIG. 6 is a diagram illustrating an example of a generation process of scenario data. FIG. 7 is a diagram illustrating an example of a generation process of information for code generation. FIG. 8 is a diagram illustrating an example of an image quality improvement process. FIG. 9 is a diagram illustrating an example of a user interface. FIG. 10 is a diagram illustrating an example of a user interface. FIG. 11 is a diagram illustrating an example of a user interface. FIG. 12 is a diagram illustrating an example of a user interface. FIG. 13 is a diagram illustrating an example of a user interface. FIG. 14 is a diagram illustrating an example of a user interface. FIG. 15 is a diagram illustrating an example of a user interface. FIG. 16 is a diagram illustrating an example of a user interface. FIG. 17 is a diagram illustrating an example of a user interface. 1 is a conceptual diagram showing an example of processing according to depth of field. FIG. 1 is a diagram showing an example of video during a confirmation operation. FIG. 2 is a diagram showing an example of highlighting a changed portion. FIG. 3 is a diagram showing an example of presenting a relationship between cuts. FIG. 4 is a diagram showing an example of object selection. FIG. 5 is a diagram showing an example of object selection. FIG. 6 is a diagram showing an example of processing using reference data. FIG. 7 is a flowchart showing a flow of processing according to a user operation. FIG. 8 is a hardware configuration diagram showing an example of a computer that realizes the functions of an information processing device. FIG. 9 is a diagram showing another example of an image generation system of the present disclosure. FIG. 10 is a diagram showing an example of the flow of image generation processing of the present disclosure. FIG. 11 is a diagram showing an example of multiple types of scene configuration data of the present disclosure. FIG. 12 is a diagram showing an example of the flow of editing processing of the present disclosure.

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the image generation system, image generation method, and image generation program according to the present application are not limited to these embodiments. Furthermore, in the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The present disclosure will be described in the following order of items: 1. Embodiment 1-1. Overview of the configuration of the video generation system of the present disclosure 1-2. Processing by the video generation system of the present disclosure 1-3. User interface 1-4. Processing examples 1-4-1. Sound generation example 1-4-2. Text logo generation example 1-4-3. Re-learning example 1-4-4. USD update example 1-4-5. Evaluation example 1-4-6. Multiple cut selection example 1-4-7. Use of time information example 1-4-8. Response example according to the degree of verbalization 1-4-9. Use of qualitative values ​​example 1-4-10. Rendering processing example during input 1-4-11. Checking work example during editing 1-4-12. Processing example according to range selection 1-4-13. Processing example according to depth of field 1-4-14. 1-4-14. Example of playback processing during confirmation work 1-4-15. Example of highlighting 1-4-16. Example of presenting relationships between cuts 1-4-17. Example of object selection 1-4-18. Example of using 3D models 1-4-19. Example of using reference data 1-4-20. Example of advantages of having 3D data 1-5. Example of processing flow from the user's perspective 1-6. Other configurations and processing examples of video generation system 1-7. Regarding AI models 2. Other embodiments 2-1. Other configuration examples 2-2. Other 3. Effects of the present disclosure 4. Hardware configuration

[0010] <1. Embodiment> <1-1. Overview of the Configuration of the Image Generation System of the Present Disclosure> Fig. 1 is a diagram illustrating an example of an image generation system of the present disclosure. The image generation system 1 includes an image generation module 100, an information acquisition module 200, a sensor unit 300, and a client UI display unit 400. Note that while Fig. 1 illustrates only one of each component, the image generation system 1 may include multiple image generation modules 100, multiple information acquisition modules 200, multiple sensor units 300, and multiple client UI display units 400.

[0011] First, the configuration of the image generation module 100 that performs image generation processing will be described. The image generation module 100 includes an input text analysis unit 110, a sensor analysis unit 120, a prompt generation unit 130, an image generation unit 140, a sound generation unit 150, a text / logo generation unit 160, a composite editing unit 170, an evaluation unit 180, and a client UI module 190.

[0012] The input text analysis unit 110 analyzes input text. For example, the input text analysis unit 110 analyzes text input from the information acquisition module 200. The sensor analysis unit 120 analyzes input sensor information. For example, the sensor analysis unit 120 analyzes sensor information acquired from the information acquisition module 200.

[0013] The prompt etc. generation unit 130 generates various information required to generate video (movie) including prompts etc. to be input into an AI (Artificial Intelligence) model (also simply referred to as a "model"), which is a machine learning model described below. For example, the prompt etc. generation unit 130 generates a prompt using a user input and a pre-saved prompt (template, etc.). Note that a prompt is merely one example of information to be input into an AI model (model input information). The model input information to be input into an AI model is not limited to a prompt, and any form of model input information can be used. Therefore, the "prompt etc. generation unit" may be read as a "model input information etc. generation unit." In FIG. 1 , the prompt etc. generation unit 130 includes a scenario-oriented generation unit 131, a video-oriented generation unit 132, a sound-oriented generation unit 133, and a text / logo-oriented generation unit 134.

[0014] The scenario generation unit 131 generates various information related to the generation of a scenario. The scenario generation unit 131 generates input information to be input to a model that outputs a scenario. For example, the scenario generation unit 131 is a scenario generation unit that generates scenario data related to video generation based on an input query. For example, the scenario generation unit 131 is a first output unit that outputs scenario generation information used by the scenario generation unit to generate scenario data based on an input query.

[0015] The video generation unit 132 generates various information related to the generation of videos. The video generation unit 132 generates input information to be input to a model that outputs code for constructing 3D (three-dimensional) data. For example, the video generation unit 132 is a code generation unit that generates code for constructing 3D data based on scenario data. For example, the video generation unit 132 is a second output unit that outputs code generation information used by the code generation unit to generate code for constructing 3D data based on scenario data.

[0016] The sound generation unit 133 generates various information related to the generation of sound information (audio information). The sound generation unit 133 generates input information to be input to a model that outputs sound. The text / logo generation unit 134 generates various information related to the generation of text and logos. The text / logo generation unit 134 generates input information to be input to a model that outputs at least one of text and logos.

[0017] The video generation unit 140 executes processing related to video generation. The video generation unit 140 is a video acquisition unit that acquires video data based on a code. The video generation unit 140 generates video using various information generated by the prompt generation unit 130. For example, the video generation unit 140 is a video generation unit that generates video data based on a code. Note that the video generation unit 140 may acquire video data in any manner. For example, the video generation unit 140 may acquire video data by transmitting data used to generate the video data to an external service providing device (such as a vendor) that provides a video data generation service, and receiving the video data generated by the service providing device from the service providing device. In FIG. 1 , the video generation unit 140 includes a USD generation unit 141, a rendering unit 142, and a video refinement unit 143.

[0018] The USD generation unit 141 generates various information related to a Universal Scene Description (USD). For example, the USD generation unit 141 generates USD-Python or the like using an AI model such as a Large Language Model (hereinafter also referred to as "LLM"), using a prompt obtained by video prompt generation in the video generation unit 132.

[0019] The rendering unit 142 executes various processes related to rendering, such as rendering the USD generated by the USD generation unit 141.

[0020] The image refinement unit 143 executes various processes for refining the image. The rendering unit 142 improves the quality of the generated image through image refinement processing. For example, the image refinement unit 143 is an image quality improvement unit that executes image quality improvement processing to improve the image quality of video data.

[0021] The sound generation unit 150 executes a process of generating sounds. The sound generation unit 150 generates sound information such as background music (BGM), sound effects (SE), narration, and dialogue using an AI model such as a contrastive learning model, using prompts obtained by the sound-oriented prompt generation in the sound-oriented generation unit 133.

[0022] The text / logo generator 160 executes a process for generating at least one of text and a logo. The text / logo generator 160 generates at least one of text and a logo using the information generated by the text / logo generator 134.

[0023] The image generation module 100, with the above-described configuration, generates a prompt for generating a scenario by combining a pre-stored prompt with a user's input. The image generation module 100 generates a scenario by inputting the generated prompt into an AI model such as an LLM. The image generation module 100 also generates prompts for generating images, sounds, and text / logos from the generated scenario. The image generation module 100 performs image generation, sound generation, and text / logo generation using prompts, scenarios, etc. for generating images, sounds, and text / logos.

[0024] The composite editing unit 170 executes processes related to editing, such as combining (combining) the generated video, sound, and text / logo into one video.

[0025] The evaluation unit 180 executes an evaluation process for evaluating various targets. The evaluation unit 180 evaluates the information generated by the above-described configuration. For example, the evaluation unit 180 generates information indicating an evaluation of at least one of the scenario data and the video data.

[0026] The client UI module 190 executes processing related to output on a UI (User Interface) on the client side. For example, the client UI module 190 generates various information related to output on the UI on the client side. In this case, the client UI module 190 executes processing to generate a UI to be displayed on the user side. The client UI module 190 generates various information to be displayed on the client UI display unit 400.

[0027] Furthermore, the information acquisition module 200 acquires various types of information. The information acquisition module 200 includes an input text acquisition unit 210, a sensor acquisition unit 220, etc. The input text acquisition unit 210 acquires text information input via a keyboard 320 or a microphone 330. For example, the input text acquisition unit 210 acquires text information input by a user via the keyboard 320 or the microphone 330. For example, the input text acquisition unit 210 is an acquisition unit that acquires an input query related to video generation from a user.

[0028] The sensor acquisition unit 220 acquires information (also referred to as "sensor information") detected by a sensor such as a camera 340 or a motion capture device. The information acquisition module 200 provides (transmits) the acquired various pieces of information to the image generation module 100. Note that the information acquisition module 200 may be integrated with the image generation module 100.

[0029] The sensor unit 300 has various sensors. The sensor unit 300 senses user input. The sensor unit 300 accepts user operations. For example, the sensor unit 300 is a reception unit that accepts video editing operations from the user. For example, the sensor unit 300 has a mouse 310, a keyboard 320, a microphone 330, a camera 340, an IMU 350 which is an inertial measurement unit, and the like. In this way, the sensor unit 300 includes, in addition to the mouse 310 and the keyboard 320, a user terminal (such as a smartphone) equipped with the microphone 330, the camera 340, and the IMU 350, and sensors such as motion capture, and senses user input.

[0030] The client UI display unit 400 displays various information to be presented to the client (user). The client UI display unit 400 displays the UI generated by the client UI module 190 on a display (display device) of the client. For example, the client UI display unit 400 is a display control unit that displays a storyboard based on scenario data. The storyboard is configured to display video data for each cut of the video.

[0031] The video generation system 1 may have a hardware configuration as shown in Fig. 2. Fig. 2 is a diagram illustrating an example of a hardware configuration of the video generation system of the present disclosure. In Fig. 2, the video generation system 1 has, as its hardware configuration, a cloud-side computer 10, a client-side computer 20, a camera / sensor 30 including various sensors such as a camera, and the like. The video generation system 1 may also include an information providing device (computer) that provides information resources 40 such as learning data and an AI model 50 to the computer 10.

[0032] 2 is merely an example, and any hardware configuration can be adopted for the video generation system 1 as long as it can execute the desired processing. For example, the computer 10 and the computer 20 may be integrated. Furthermore, the information resource 40 and the AI ​​model 50 may be stored inside the computer 10.

[0033] The computer 10 includes a CPU (Central Processing Unit) 11, a GPU (Graphics Processing Unit) 12, a communication device 13, and a memory / storage 14. For example, the computer 10 corresponds to the image generation module 100 and the information acquisition module 200 in FIG. 1 . The computer 10 may be a service providing device (server device) that provides an image generation service. The CPU 11 and the GPU 12 are so-called processors, and execute calculations (arithmetic operations) related to various processes such as image generation.

[0034] The communication device 13 is a communication device having a communication function for transmitting and receiving information to and from the computer 20, an information providing device, etc., and may be, for example, a communication circuit, a NIC (Network Interface Card), etc. The communication device 13 communicates with other devices such as the computer 20 and the information providing device via a predetermined network (such as the Internet). For example, the communication device 13 is connected to the predetermined network via a wired or wireless connection, and transmits and receives information to and from other devices such as the computer 20 and the information providing device.

[0035] The memory / storage 14 is a storage device that stores various types of information. The memory / storage 14 is, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 14 stores various types of information used for processing by processors such as the CPU 11 and the GPU 12. The memory / storage 14 may also store information resources 40, AI models 50, etc.

[0036] The computer 20 includes a CPU 21, a GPU 22, a communication device 23, a memory / storage 24, and an IO interface 25. For example, the computer 20 corresponds to the client UI display unit 400 in FIG. 1 . The computer 20 may also be a terminal device (such as a personal computer (PC) or a mobile device such as a smartphone) used by a user who uses the video generation service. The CPU 21 and the GPU 22 are so-called processors, and perform calculations (arithmetic processing) related to various processes such as video display. Note that the above is merely an example, and the computer 20 can have any configuration as long as it can perform the desired processing. For example, the computer 20 may perform calculations (arithmetic processing) related to various processes such as video display using circuits such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). Furthermore, the computer 20 may be configured so that programs are directly embedded in the processor circuitry instead of storing the programs in memory (such as the memory / storage 24). In this case, the processor realizes its functions by reading and executing the programs embedded in the circuitry. In addition, each processor in this embodiment is not limited to being configured as a single circuit, but may be configured as a single processor by combining multiple independent circuits to realize its functions. Furthermore, like the computer 20, the computer 10 can also adopt any configuration as long as it is capable of performing the desired processing.

[0037] The communication device 23 is a communication device having a communication function for transmitting and receiving information to and from the computer 10, the sensor 30, etc., and may be, for example, a communication circuit, a NIC, etc. The communication device 23 communicates with other devices such as the computer 10 and the sensor 30 via a predetermined network (such as the Internet). For example, the communication device 23 is connected to the predetermined network by wire or wirelessly, and transmits and receives information to and from other devices such as the computer 10 and the sensor 30.

[0038] The memory / storage 24 is a storage device that stores various types of information. The memory / storage 24 is, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 24 stores various types of information that are used for processing by processors such as the CPU 21 and the GPU 22.

[0039] The IO interface 25 is an input / output interface device. The computer 20 receives input from the sensor 30 via the IO interface 25. For example, the computer 20 receives input from an input device such as a keyboard or a mouse via the IO interface 25. The computer 20 also outputs information from a display (display device) and a speaker (audio output device) via the IO interface 25. For example, the computer 20 plays video on the display and speaker via the IO interface 25.

[0040] Various sensors 30, such as cameras, sense user input. Various sensors 30, such as cameras, accept user operations. For example, the sensor 30 corresponds to the sensor unit 300 in FIG. 1. Furthermore, the information resource 40 includes various information such as training data. For example, the information resource 40 includes training data used to train various AI models, such as LLM. The AI ​​model 50 includes information on AI models used in processing related to video generation, such as LLM. For example, the AI ​​model 50 includes information on various AI models, such as models M1 to M3, which will be described later. As described above, the video generation system 1 may have a configuration other than that shown in FIG. 2.

[0041] <1-2. Processing by the video generation system of the present disclosure> Processing by the video generation system will now be described. First, an example of the flow of the video generation process shown in FIG. 3 will be described. FIG. 3 is a diagram showing an example of the flow of the video generation process of the present disclosure. Note that the processes described below with the video generation system 1 as the processing subject may be performed by any device capable of executing the processes, depending on the device configuration included in the video generation system 1.

[0042] User input information UIN1, denoted as "User's Input" in FIG. 3, corresponds to information input by a user for generating an image (also referred to as an "input query"). Note that the input query is not limited to text (character information) and any information can be used. The input query may be any information including at least one of text, image, audio, and 3D data.

[0043] The video generation system 1 uses user input information UIN1 to generate scenario generation information (also referred to as "first input information") to be used as input for model M1, denoted as "LLM" in FIG. 3, which will be described later. For example, model M1 is a first model that outputs scenario data in response to the input of the first input information. Any AI model, such as an LLM (large-scale language model), can be used for model M1 as long as it is capable of producing the desired output in response to the input. AI models such as model M1 will be described later.

[0044] The video generation system 1 generates a scenario FD1 by inputting first input information to a model M1 and causing the model M1 to output a scenario FD1, which is scenario data. The video generation system 1 then generates USD generation necessary information SD1, which is code generation information (also referred to as "second input information") to be used as input for a model M3, using the scenario FD1, the output of a model M2 that uses the scenario FD1 as input, and user input information UIN2, etc. For example, the model M3 is a second model that outputs code in response to the input of the second input information. Any AI model, such as an LLM (large-scale language model), can be used for the model M3, as long as it is capable of producing a desired output in response to the input.

[0045] 3 illustrates only one piece of information SD1 required for generating USD, but there may be a plurality of pieces of information SD1 required for generating USD depending on the number of USD files to be generated. For example, there may be a plurality of pieces of information SD1 required for generating USD depending on the number of USD files to be generated corresponding to the data structure shown in FIG.

[0046] For example, model M2 may be a model that outputs a template or the like corresponding to scenario data in response to input of the scenario data. Any AI model can be adopted as model M2 as long as it is capable of producing the desired output in response to the input. For example, user input information UIN2 may be information for specifying constraints for image generation. Note that image generation system 1 may generate second input information using scenario FD1 and template input information, but this point will be described later.

[0047] The video production system 1 generates the Python code OD1 by inputting the USD generation necessary information SD1 into the model M3 and causing the model M3 to output the Python code OD1, denoted as "python" in FIG. 3 . For example, the Python code OD1 is (program) code that generates USD format data (also referred to as a "USD file") when executed. Note that Python is merely an example, and any code format, not limited to Python, can be used as long as it can generate the desired 3DCG data. Also, USD is merely an example, and any format, such as FBX (Film Box), can be used for 3DCG data. The video production system 1 executes the Python code OD1 to generate the USD file OD2, denoted as "USD" in FIG. 3 .

[0048] The video generation system 1 generates video data MV1, which is denoted as "PreMovie" in Fig. 3, by executing a rendering process PS1, which is denoted as "Renderer" in Fig. 3. For example, the video data MV1 is data (also referred to as "first video data") before executing a refinement process PS2, which will be described later.

[0049] The video generation system 1 generates video data MV2, denoted as "RefinedMovie" in Fig. 3, by executing a refinement process PS2, denoted as "Refiner" in Fig. 3. For example, the refinement process PS2 is a picture quality improvement process that improves the picture quality of the video data. The video data MV2 is data (also referred to as "second video data") obtained after the refinement process PS2 updates the first video data, i.e., video data MV1.

[0050] The video production system 1 executes composite editing PS3 using the video data MV2, user input information UIN3, etc., to generate video data MV3, which is denoted as "FinalMovie" in Fig. 3. For example, the composite editing PS3 executes a process of updating (editing) the video data MV2 in accordance with a user editing instruction indicated by the user input information UIN3, thereby generating video data MV3 in which the video data MV2 has been updated.

[0051] The flow of the video generation process shown in FIG. 3 is merely an example, and the video generation system 1 can employ any processing mode as long as it can generate video data from a user's input query. For example, while FIG. 3 illustrates an example in which the model M1 outputs code (Python code), the model M1 may also output 3DCG data such as a USD file. Furthermore, the video generation system 1 may perform various modes of video generation processing, not limited to the processing illustrated in FIG. 3 . An example of this point will be described using FIG. 4 . FIG. 4 is a diagram illustrating another example of the flow of the video generation process of the present disclosure. FIG. 4 differs from FIG. 3 in that sound generation required information SD2 and text logo required information SD3 are generated and used to perform video generation processing. Note that explanations of points similar to those described in FIG. 3 will be omitted where appropriate.

[0052] In FIG. 4, the video generation system 1 uses a scenario FD1, the output of a model M2 that uses the scenario FD1 as input, and user input information UIN2 to generate sound generation information SD2, which is information for generating sound to be used as input for a model M4, denoted as "AI" in FIG. 3. For example, the model M4 is a model that outputs various sound data in response to input of the sound generation information SD2 and video data MV2. Any AI model can be used for the model M4 as long as it is capable of producing the desired output in response to the input. Note that the model M4 may also be a model that receives only the sound generation information SD2 as input.

[0053] The video production system 1 generates sound data corresponding to a video by inputting sound generation necessary information SD2 to the model M4 and causing the model M4 to output sound data AD1 for background music (BGM), sound effects AD2, and narration sound data AD3. The video production system 1 may also generate the sound data AD1, AD2, and AD3 using user input information UIN4. For example, if the model M4 outputs the sound data AD1, AD2, and AD3 as a single piece of sound data, the video production system 1 may extract the sound data AD1, AD2, and AD3 from the single piece of sound data output by the model M4 based on the specifications in the user input information UIN4, and generate the sound data AD1, AD2, and AD3.

[0054] 4 , video generation system 1 uses scenario FD1, the output of model M2 that uses scenario FD1 as input, and user input information UIN2 to generate text logo generation information SD3, which is used as input for model M5. For example, model M5 is a model that outputs at least one of text and a logo in response to the input of text logo generation information SD3. Any AI model can be used for model M5 as long as it can produce the desired output in response to the input.

[0055] Video generation system 1 inputs information SD3 required for text logo generation into model M5 and causes model M5 to output text logo data DI1 for Text, text logo data DI2 for Logo, etc., thereby generating text logo data corresponding to the video.

[0056] Video production system 1 generates video data MV3 by executing composite editing PS3 using video data MV2, sound data AD1, AD2, AD3, text logo data DI1, DI2, user input information UIN3, etc. For example, composite editing PS3 combines (combines) video data MV2, sound data AD1, AD2, AD3, text logo data DI1, DI2, etc. into one image to generate video data MV3 as a single image.

[0057] 3 and 4 illustrate an example of processing in an initial state where there is no scenario, information required for USD generation, USD, PreMovie, RefinedMovie, or the like. As described above, the video production system 1 generates prompts for generating a scenario based on a user's input query, provides the prompts to a natural language model to generate a scenario, generates a prompt for outputting code constituting a video from text information described in the scenario, provides the prompts to the natural language model to generate code constituting the video, and generates a video. In this way, when creating a video, the video production system 1 generates an effective video storyboard and video by inputting what the user wants to create and their purpose, even without knowledge of 3DCG or video production. Furthermore, creating a storyboard in the video production system 1 makes subsequent editing easier. These points will be described in detail later.

[0058] Furthermore, the image generation system 1 may perform various processes related to image generation. For example, the image generation system 1 may perform evaluation processing on the generated information. In this regard, an example of the flow of the evaluation processing will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the flow of the evaluation processing of the present disclosure.

[0059] 5 , the video generation system 1 receives input of at least one of a scenario FD1, information required for USD generation SD1, information required for sound generation SD2, and information required for text logo generation SD3, and causes the model M10 to output evaluation text information EV1 indicating an evaluation of the input information, thereby evaluating the generated information. For example, the model M10 outputs an evaluation of input information in response to input of the information. For example, the model M10 outputs evaluation text indicating an evaluation of the input scenario FD1 in response to input of the scenario FD1. Note that the model M10 may be a model that accepts input of the scenario FD1, information required for USD generation SD1, information required for sound generation SD2, and information required for text logo generation SD3 separately, or a model that accepts input of a combination of these pieces of information. Furthermore, the model M10 may be a model that accepts input of information indicating a video (e.g., captions) in response to input of the video, and outputs an evaluation of the video corresponding to the input information.

[0060] From here, a specific example of each process in the above-mentioned process flow executed by the image generation system 1 will be described. Note that explanations of points similar to those described above will be omitted as appropriate.

[0061] For example, the video production system 1 generates scenario generation information (first input information) as shown in Fig. 6. Fig. 6 is a diagram showing an example of a process for generating scenario generation information. In Fig. 6, the video production system 1 acquires user input information IDT1 and IDT2 entered by the user into content CT1 as user input information. Content CT1 is content for receiving user input information for each of the questions "What kind of video do you want to create?" and "Style."

[0062] For example, the client UI display unit 400 displays the content CT1, and the sensor unit 300 receives the user input information IDT1 and IDT2 as user input information. For example, the user input information IDT1 and IDT2 correspond to the user input information UIN1 in FIGS.

[0063] 6, the client UI display unit 400 displays a question, "What kind of video do you want to make?". In response to the question, "What kind of video do you want to make?", the sensor unit 300 accepts user input information IDT1, "a 15-second sneaker commercial video." The client UI display unit 400 also displays a question, "style." In response to the question, "style," the sensor unit 300 accepts user input information IDT2, "cinematic."

[0064] The video production system 1 may accept user input information in any manner, or may accept a user selection from multiple options. For example, the video production system 1 may convert information entered by the user via a keyboard or microphone into text information and accept it as user input information. Furthermore, the video production system 1 may accept, in addition to free text, settings such as the number of seconds for the entire video, style, and camerawork, as well as other files such as images and videos, as user input information.

[0065] The video production system 1 generates a prompt PT1, which is scenario generation information (first input information), using the user input information IDT1 and IDT2 and a template TP1, which is template input information. For example, the template TP1 may be preset or may be selected from a plurality of template candidates. For example, the video production system 1 may select a template corresponding to the user's input information from the plurality of template candidates. For example, the video production system 1 may select a template TP1 related to a movie-style advertisement from the plurality of template candidates based on the content indicated by the user input information IDT1 and IDT2.

[0066] For example, the video production system 1 generates a prompt PT1 by reflecting user input information IDT1 and IDT2 in a template TP1. In Fig. 6, the video production system 1 generates the prompt PT1 by adding "cinematic" indicated by the input information IDT2 to the style item of the constraints and adding "a 15-second sneaker commercial video" indicated by the input information IDT1 to the input sentence. In this way, the video production system 1 generates a prompt for generating a scenario based on the information input by the user. Note that the user's input information may be input on a single screen or by answering several questions; examples of these points will be described later.

[0067] The video production system 1 also generates scenario data as shown in Fig. 7. Fig. 7 is a diagram showing an example of a scenario data generation process. In Fig. 7, the video production system 1 generates scenario data SN1 using a prompt PT1. For example, the scenario data SN1 corresponds to the scenario FD1 in Figs. 3 and 4. The scenario data SN1 includes information such as the number of seconds for each scene, an explanation of the cut, etc., for each scene, such as the opening scene and the scene where sneakers are put on.

[0068] For example, the video production system 1 generates scenario data SN1 by inputting a prompt PT1 into a model M1, such as an LLM, and outputting scenario data SN1 from the model M1. In this manner, the video production system 1 generates a scenario by inputting the generated prompt into an AI (such as an LLM). In addition to the information shown in FIG. 7 (also referred to as "scenario information"), the scenario data SN1 also includes information such as the environment, characters, motion, camerawork, lighting, and color. For example, to create unique scenarios, the video production system 1 can generate a variety of scenario variations by inputting a user's past experiential learning data or learning data from a specific director or person into the model M1 using Retrieval-Augmented Generation (RAG) or fine-tuning.

[0069] Furthermore, the video production system 1 generates code generation information (second input information) as shown in FIG. 8 . FIG. 8 is a diagram showing an example of a process for generating code generation information. The video production system 1 generates a prompt PT2, which is code generation information (second input information), using scenario data SN1 and a template TP2, which is template input information. For example, the template TP2 may be preset or may be selected from multiple template candidates. For example, the video production system 1 may select a template corresponding to a scenario from the multiple template candidates. For example, the video production system 1 may select a template TP2 related to a commercial from the multiple template candidates based on the content indicated by the scenario data SN1.

[0070] For example, the video production system 1 generates a prompt PT2 by reflecting scenario data SN1 in a template TP2. In Fig. 8, the video production system 1 generates the prompt PT2 by adding information indicated by the scenario data SN1 to an input sentence. In this way, the video production system 1 generates a prompt for a video based on a scenario generated by AI.

[0071] For example, FIG. 8 shows an example of prompt generation for converting a scenario related to a person into USD-Python. Conversion to USD-Python is merely one example of a conversion format, and the conversion is not limited to USD-Python and may be any conversion format. For example, the conversion format may be Python for Blender, USD, or other formats. Furthermore, the video production system 1 may generate prompts individually for each subject, such as a person, environment, or camerawork, or may generate prompts collectively. The generated prompt may include paths to assets and motions to be used, or may include source code or an API (Application Programming Interface) to be passed (input) to the AI ​​generation algorithm for assets and motions.

[0072] The video production system 1 then generates a USD-Python file by inputting the generated prompt into an AI (such as an LLM). Note that the file format is not limited to Python format, and the file may be generated in another format such as USD. The video production system 1 then converts the file into a format that can be rendered, such as a USD file, and performs rendering to generate a PreMovie (a video file such as MP4).

[0073] The video production system 1 may perform composite editing using a PreMovie, but may also perform refinement processing, which is an example of image quality improvement processing, on the PreMovie, as shown in Fig. 9. Fig. 9 is a diagram showing an example of image quality improvement processing. In Fig. 9, the video production system 1 generates the second video OT1 from the first video IN1 by refinement processing using a model M11, which is a diffusion model that receives a first video IN1, which is a PreMovie, as input and outputs a second video OT1, which is a RefinedMovie.

[0074] The AI ​​model (such as model M11) used in the refiner process is not limited to the diffusion model; any AI model such as the latent diffusion model (LDM) or the latent consistency model (LCM) can be used. Furthermore, the refiner process may use techniques such as AnimateDiff (time direction stabilization) or ControlNet (line art control). Through this refiner process, the video production system 1 can improve the quality of the video while maintaining the consistency of the characters, backgrounds, props, and the like.

[0075] Furthermore, the refiner process may use prompts in addition to videos. For example, the model M11 may input prompt IN2 in addition to the first video IN1. For example, when a target such as a woman in her 30s is specified by prompt IN2, the model M11 outputs a second video OT1 in which the portion of the first video IN1 that represents the woman in her 30s has been improved. This allows the video generation system 1 to generate a second video OT1 in which the image quality, etc., of the target specified by prompt IN2 in the first video IN1 has been improved.

[0076] Through the processing of the video production system 1 described above, it appears to the user that a scenario (storyboard) and animation for each cut are being generated after input, with the processing in between being confined within the system. These processes may involve generating the scenario, all cuts, and refinement processing all at once from the user's input text, or the user may input preferences during the processing. For example, the video production system 1 may generate several scenarios with outlines only, then allow the user to select one, and then execute detailed scenario and animation generation processing based on the selected outline scenario. Furthermore, the video production system 1 may generate several patterns of characters to be generated in the video before video generation after scenario generation, and after the user selects one, execute animation rendering and refinement processing.

[0077] The components of a scenario may include video of the cut, a representative image (such as the first frame of the video), a description of the cut, characters (visuals, setting, etc.), the motion of each character, lighting, camera work, background environmental information, transitions between cuts, dialogue, narration, etc. Some of these are presented to the user, while others are kept for processing purposes without being presented to the user. The scenario is arranged in chronological order by cut.

[0078] Currently, various video generation services are available, including Pika, Runway Gen-2, Lumiere, and Stable Video Diffusion. These generate video using a diffusion model that moves vectors in the spatial and temporal directions from images. These generate video using only 2D images. On the other hand, video generation system 1 stores 3D information internally. For example, existing video generation services allow you to modify only a specified (X, Y) area within a video, but there is an issue that changing only the color of clothing also changes the motion. On the other hand, video generation system 1 stores 3D information internally, making it possible to modify only targeted areas, such as only the motion, only the lighting, or only the color of a person's clothing.

[0079] Furthermore, existing video generation services only generate videos for each cut, and users must ensure the consistency of each cut themselves, but video generation system 1 can consistently carry out everything from the scenario (storyboard) to video generation and editing, making it possible to generate videos with consistency in terms of the actors, backgrounds, color grading, etc.

[0080] <1-3. User Interface> Hereinafter, we will describe the user interface (UI) for users who use the image generation system 1. Note that explanations of points similar to those described above will be omitted where appropriate.

[0081] As shown in FIG. 10 , the video generation system 1 provides the user with content CT11. FIG. 10 is a diagram illustrating an example of a user interface. The content CT11 is a display screen (content) for accepting user input information. For example, the client UI display unit 400 displays the content CT11. The user inputs text instructing the type of video to be created (the “Prompt” column in FIG. 10 ) and a style selection (the “Style” column in FIG. 10 ) as user input information via the content CT11. In this manner, the user inputs the text and style selection as user input information. For example, the user inputs text information in the “Prompt” column by referring to example sentences, etc., included in the content CT11. For example, the user selects a style to use from multiple style candidates displayed by, for example, pressing (clicking, etc.) the downward-pointing triangle in the “Style” column.

[0082] After completing the input of the user's input information, the user selects the button labeled "Ask AI Director" in Fig. 10 to instruct the video production system 1 to generate a video in accordance with the user's input information. As a result, the video production system 1 executes the process of generating a video in accordance with the user's input information.

[0083] As shown in FIG. 14 , the video generation system 1 provides the user with content CT15 related to the generated video. FIG. 14 is a diagram showing an example of a user interface. The content CT15 is a storyboard screen (content) for receiving user operations (instructions) on the generated video. As shown in FIG. 14 , the content CT15 is a storyboard screen that displays video data for each cut of the generated video. For example, the client UI display unit 400 displays the content CT15. In this way, the video generation system 1 provides a UI that outputs a storyboard and video in accordance with information input by the user. The user sets the video, content, narration, dialogue, camerawork, background music, lighting, color, and the like for each cut on the storyboard screen.

[0084] The video generation system 1 may accept user input information while asking the user a question. For example, when the button labeled "Ask AI Director" in FIG. 10 is selected, the video generation system 1 accepts user input information through a conversation (dialogue) with the user, as shown in FIGS. 11 to 13. FIGS. 11 to 13 are diagrams showing an example of a user interface. Content CT12 in FIG. 11 is a display screen (content) that presents samples generated in response to the user's input information entered in FIG. 10 and asks the user whether there is one that is close to their image. For example, the client UI display unit 400 displays the content CT12.

[0085] Content CT13 in Fig. 12 is a display screen (content) that requests (questions) for detailed targets, etc. in response to a user's response (input information) that none of the samples presented in content CT12 in Fig. 11 match the image. For example, the client UI display unit 400 displays content CT13.

[0086] Content CT14 in Figure 13 is a display screen (content) that presents samples regenerated in response to a user's response (input information) that specifically specifies a target, etc., and asks the user whether any of them are close to the image they had in mind. For example, the client UI display unit 400 displays content CT14. Figure 13 shows a case in which the user positions the mouse cursor over the leftmost sample video of the four samples and performs a designation operation such as clicking, thereby designating that the leftmost sample video of the four samples is close to the image they had in mind. As a result, the video generation system 1 executes a video generation process in accordance with the user's input information that designates the leftmost sample video of the four samples.

[0087] In this case, the video generation system 1 provides the user with content CT15 related to the generated video, as shown in Fig. 14. For example, the client UI display unit 400 displays the content CT15. In this way, when generating a storyboard and video using the initial input content, if the video generation system 1 does not have enough necessary information, it may collect the necessary information through conversation (dialogue) with the user and work out the details.

[0088] 15 and 16, the video production system 1 may prompt the user to input a short sentence about the video they want to create, generate several video stories based on the short sentence, and allow the user to select the one they like best. Figures 15 and 16 are diagrams showing examples of a user interface.

[0089] As shown in FIG. 15 , the video generation system 1 provides the user with content CT21. The content CT21 is a display screen (content) for receiving information input by the user. For example, the client UI display unit 400 displays the content CT21. The user inputs a sentence (short sentence) indicating what kind of video they want to create as user input information via the content CT21. For example, the user inputs text information into an input field in the content CT21.

[0090] After completing the input of the user's input information, the user selects the button labeled "Start" on the right end of the input field in content CT21 to instruct the video production system 1 to generate a video story in accordance with the user's input information. This causes the video production system 1 to execute the process of generating a video story in accordance with the user's input information.

[0091] As shown in Fig. 16, the video generation system 1 provides the user with content CT22 relating to the story of the generated video. The content CT22 in Fig. 16 is a display screen (content) that presents a sample of the story of the video generated in response to the user's input information entered in Fig. 15. For example, the client UI display unit 400 displays the content CT22. For example, if the user likes one of the video story samples, the user selects the button labeled "Continue" on the right edge of the display area for that sample, thereby instructing the video generation system 1 to generate a video corresponding to the selected sample.

[0092] If the user does not like any of the video story samples, the video production system 1 generates another pattern. For example, if the user does not like any of the video story samples, the user can select the button labeled "Continue" on the right edge of the short sentence display area to instruct the video production system 1 to generate another pattern of video story sample. This causes the video production system 1 to generate another pattern of video story sample.

[0093] <1-4. Processing Examples> From here, in addition to the specific examples described above, specific examples of each process executed by the video production system 1 will be described. Note that explanations of points similar to those described above will be omitted as appropriate. Below, specific examples of the generation process of sound, etc., and evaluation process in the processing of the video production system 1 described above will be described. Note that explanations of points similar to those described above will be omitted as appropriate.

[0094] <1-4-1. Example of Sound Generation> For example, the video production system 1 generates sound generation information (corresponding to the sound generation required information SD2 in FIGS. 3 and 4) as shown in FIG. 17. FIG. 17 is a diagram showing an example of a process for generating sound generation information. The video production system 1 generates a prompt PT3, which is sound generation information, using scenario data SN1 and a template TP3, which is template input information. For example, the template TP3 may be preset or may be selected from multiple template candidates. For example, the video production system 1 may select a template corresponding to a scenario from the multiple template candidates. For example, the video production system 1 may select a template TP3 related to a commercial from the multiple template candidates based on the content indicated by the scenario data SN1.

[0095] For example, the video production system 1 generates a prompt PT3 by reflecting the scenario data SN1 in a template TP3. In Fig. 17, the video production system 1 generates the prompt PT3 by adding information indicated by the scenario data SN1 to an input sentence. For example, the video production system 1 generates a prompt for extracting several key phrases from a scenario in order to generate a sound (such as background music) that better suits the scenario. The video production system 1 may input the generated prompt into an AI (such as an LLM) to obtain keywords and text information necessary for sound generation.

[0096] The video production system 1 generates sound data using the generated prompt PT3. For example, the video production system 1 generates sounds (sound data) such as background music, sound effects, narration, and dialogue based on the generated video and text information described in the storyboard (scenario data SN1, etc.). Furthermore, if necessary words (text information) such as dialogue or narration have already been extracted from the scenario, the video production system 1 may save the words (text information) as sound data for the dialogue, narration, etc., without generating a prompt.

[0097] For example, the video generation system 1 generates sound data from the obtained information required for sound generation (sound generation information), video, and audio data input by the user (sound source, user's voice, humming, etc.). For background music and sound effects, the video generation system 1 may generate sound from text or video using a transformer (model) such as text-to-music generation. The video generation system 1 may also search for contrastively learned sound sources, such as text-to-music estimation, from natural text.

[0098] Furthermore, for dialogue and narration, the video production system 1 may generate audio using text-to-speech (such as a diffusion model or flow matching) based on the words themselves and text information about the characters obtained from the scenario. Furthermore, when connecting the generated video and audio, the video production system 1 may incorporate meta information into the video or audio file to indicate the start and end times, volume, etc. of the audio.

[0099] <1-4-2. Example of Text Logo Generation> For example, video production system 1 generates information for generating a text logo (corresponding to information required for text logo generation SD3 in Figures 3 and 4). Based on the generated storyboard (scenario data SN1, etc.), video production system 1 generates text and logo information (captions, titles, logos, descriptions, etc.) to be displayed on the video.

[0100] For example, the generated scenario may clearly state the text to be displayed, but if it does not, the video generation system 1 generates a prompt to generate text information (text logo data, etc.) and sends it to an AI (such as an LLM) to generate the text information. Also, if the user inputs the text to be displayed, the video generation system 1 may use the information input by the user as the text information (text logo data, etc.).

[0101] The font, size, and position of the text display may be determined by any method. For example, the video generation system 1 may determine the font, size, and position of the text display using any AI such as a diffusion model, a variational auto-encoder (VAE), generative adversarial networks (GAN), DALL E, StyleGAN, StyleGAN2, Pix2Pix, TransGAN, or LLM. The font, size, and position of the text display may also be manually set by the user.

[0102] For logos and images, the user may input images or videos in JPEG format, MP4 format, etc. For logos and images, the video generation system 1 may generate a prompt for image generation and send it to any AI such as a Diffusion model, VAE, GAN, DALL E, StyleGAN, StyleGAN2, Pix2Pix, TransGAN, or LLM to generate logo information.

[0103] In addition, when linking text or logo information with a video, the video generation system 1 may incorporate meta information into the video or the text or logo itself to clearly indicate (meta information) the start and end times, position, and size of the text or logo.

[0104] <1-4-3. Re-learning Example> Furthermore, when generating a scenario or video, the video generation system 1 may generate a scenario or video that can only be produced using specific, replaceable learning data. For example, it is possible to re-learn using previously produced videos and images as learning data. Re-learning using RAG or fine tuning can change the scenario or video to be generated. In other words, the video generation system 1 can re-learn data (history) of personal user's past productions, or can generate videos using a model re-learned using the work of a specific film director as learning data.

[0105] Furthermore, the learning data can be retrained on an individual's PC or on a server. The learning data is used to train a number of models, such as the LLM and Diffusion model. In cases such as when an overall scenario is to be generated using Director A, but the color of the video is to be generated using a different Director B to generate color grading, the video generation system 1 may retrain specific parts using different learning data.

[0106] <1-4-4. USD Update Example> As described above, the video generation system 1 modifies the 3DCG assets and rendering method that are the basis of the existing video, based on the user's input information and the text information of the generated storyboard. The video generation system 1 then outputs a prompt for outputting code that will compose a new video, provides the prompt to a natural language model, outputs the code that will compose the video, and generates the video. This allows the video generation system 1 to modify the video in accordance with the user's input, even if the user has no knowledge of 3DCG or video production.

[0107] In the video generation system 1, after generating a video (video), scenario information, video information, etc. are saved as text or video, so the user can modify this information using input information (text, sensors, etc.).

[0108] For example, in the video production system 1, USD files are stored separately for each asset and motion, as shown in Fig. 18. Fig. 18 is a diagram showing an example of a USD file. For example, in the data structure shown in Fig. 18, a USD in a higher layer may include a path (file path) to a USD in a lower layer.

[0109] For example, the entire USD includes paths to the environment assets USD, person assets UDS, camera USD, etc. The environment assets USD also include paths to the building assets USD and prop assets USD. The building assets USD also include mesh information of the building itself, etc. The person assets USD also include paths to mesh information of a person and motion USD. In this way, data for 3DCG (3D data) such as a USD file may include multiple data sets. Note that the configuration (data structure) of the USD file shown in FIG. 18 is merely an example, and any configuration can be adopted, and the entire USD (USD file) may be configured as a single block (one data set).

[0110] When a correction is made, the video generation system 1 analyzes the user's correction information using AI (such as LLM) according to the processing flow shown in Fig. 19, and the processing changes depending on whether the USD is replaced or part of the USD is corrected. After the correction, rendering processing and refinement processing are executed. Fig. 19 is a flowchart showing the processing procedure executed by the video generation system. As a specific example, Fig. 19 is a flowchart showing the processing procedure related to rewriting a USD file.

[0111] First, the image generation system 1 accepts correction information input from the user (step S101). For example, the sensor unit 300 accepts input information instructing the user to make corrections. The image generation system 1 analyzes the input using AI (step S102). For example, the image generation module 100 analyzes the content of the input information instructing the user to make corrections using various models, etc.

[0112] The video generation system 1 recognizes the format of the existing USD file (step S103). For example, the video generation module 100 recognizes the format of the USD file before modification. The video generation system 1 determines whether to modify a part of the USD (step S104). For example, the video generation module 100 determines whether to modify a part of the USD based on the content of input information instructing modification by the user and the format of the existing USD file.

[0113] When a part of the USD is to be modified (step S104: Yes), the image generation system 1 generates a prompt for generating a modified USD-Python (step S105). For example, when a part of the USD is to be modified, the image generation module 100 generates a prompt for generating the modified USD-Python using a template or the like for generating the modified USD-Python.

[0114] The image generation system 1 generates USD-Python using AI (LLM, etc.) (step S106). For example, the image generation module 100 generates USD-Python by inputting a prompt into a model for generating USD-Python.

[0115] The video production system 1 replaces the USD file to be modified (step S107). For example, the video production module 100 replaces the USD file to be modified by reflecting the generated USD-Python in the USD file to be modified. In this way, the video production system 1 updates at least one of the multiple data sets of the USD file. For example, the video production system 1 executes a process of updating some of the multiple data sets of the USD file.

[0116] The video generation system 1 executes the rendering process using the updated USD file (step S108). For example, the video generation module 100 executes the rendering process using the rewritten, i.e., corrected, USD file.

[0117] On the other hand, if the image generation system 1 does not modify a part of the USD (step S104: No), it generates a prompt for generating the USD-Python for creation (step S109). For example, if the image generation module 100 does not modify a part of the USD, that is, if the USD is newly created (generated), it generates a prompt for generating the USD-Python for creation using a template or the like for generating the USD-Python for creation.

[0118] The video generation system 1 generates USD-Python and USD using AI (LLM, etc.) (step S110). For example, the video generation module 100 generates USD-Python and USD by inputting a prompt into a model for generating USD-Python.

[0119] The video generation system 1 replaces the USD file to be modified (step S111). For example, the video generation module 100 replaces the generated USD with the USD file to be modified. In this way, the video generation system 1 executes a process to update the USD file. For example, the video generation system 1 executes a process to update all of the multiple data sets of the USD file. Then, the video generation system 1 executes the process of step S108. For example, the video generation module 100 executes a rendering process using the replaced, i.e., modified, USD file.

[0120] Here, the user interface (UI) related to the above-mentioned corrections will be described. As shown in FIG. 20, the video generation system 1 provides the user with content CT31. FIG. 20 is a diagram showing an example of the user interface. The content CT31 is a display screen (content) for presenting information about the generated video, such as a storyboard, and for accepting correction instructions from the user. For example, the client UI display unit 400 displays the content CT31.

[0121] The user inputs information instructing corrections to the generated video via the content CT31. For example, when the user clicks on a specific part of the video and inputs corrections in text, the video generation system 1 updates the USD and updates (changes) the video to reflect the corrections.

[0122] 20 shows a case where the user selects the top cut (thumbnail) image and instructs a modification to "the child skips and runs to his mother." In this case, the video production system 1 updates the USD for the portion of the generated video corresponding to the top cut (thumbnail) image based on the modification instruction to "the child skips and runs to his mother," and generates a video that reflects the modification.

[0123] The above UI is merely an example, and the video production system 1 may accept user correction instructions in various ways. For example, as shown in FIG. 21 , the video production system 1 may provide content CT32 to the user and accept the user's correction instructions. FIG. 21 is a diagram showing an example of a user interface. The content CT32 is a display screen (content) for accepting the user's correction instructions. For example, the client UI display unit 400 displays the content CT32.

[0124] Furthermore, the video production system 1 may provide the user with content CT33 and accept a user's correction instruction, as shown in Fig. 22. Fig. 22 is a diagram showing an example of a user interface. The content CT33 is a display screen (content) for accepting the user's correction instruction. For example, the client UI display unit 400 displays the content CT33.

[0125] The user inputs information instructing corrections to the generated video via content CT32 or content CT33. As a result, the video generation system 1 accepts the correction instructions from the user, updates the USD based on the correction instructions, and generates a video that reflects the corrections. For example, the user can select a person or object that appears in the video and edit the motion settings for the selected object using text. The user can also change the asset of the selected object. The user can also edit the camerawork using text. In addition to the above, the user can also edit the background and lighting settings of the video using text.

[0126] Although the above description has been given as an example of a moving image, the video generation system 1 may also perform correction processing on sound information (BGM, SE, narration, dialogue, etc.), text logo information, etc. based on user correction instructions.

[0127] <1-4-5. Evaluation Example> The video production system 1 also executes an evaluation process. For example, the video production system 1 generates an evaluation text for scenario data using an AI model such as model M10. For example, the video production system 1 may regenerate scenario data based on an instruction for editing (correction, etc.) given by a user based on the evaluation text generated for the scenario data. The video production system 1 presents the generated evaluation text. This allows the user to create a video while viewing the evaluation.

[0128] For example, the video generation system 1 presents an evaluation text about scenario data to a user and receives an instruction to edit the scenario data from the user who has confirmed the presented evaluation text. The video generation system 1 generates scenario data based on the editing instruction received from the user. For example, the video generation system 1 changes (updates) the content of the scenario data based on the editing instruction received from the user. The video generation system 1 generates code based on scenario data generated based on the evaluation text. The video generation system 1 generates video data using the code generated based on the evaluation text.

[0129] The video generation system 1 may automatically generate (update) at least one of the scenario data or the code based on the evaluation text. For example, the video generation system 1 may change the content of the scenario data to correspond to the content indicated by the evaluation text for the scenario data. For example, if the evaluation text for the scenario data indicates that the orientation of a certain character is not good, the video generation system 1 generates scenario data in which the orientation of the character is changed. Note that the above is merely an example, and the video generation system 1 may generate at least one of the scenario data or the code by appropriately using the evaluation text.

[0130] In the video generation system 1, after generating a video (video), scenario information, video information, sound information, text / logo information, etc. are saved as text, video files, sound files, image files, etc. Therefore, the video generation system 1 is capable of performing evaluation processing using this information. The video generation system 1 modifies the 3DCG assets and rendering method that form the basis of the existing video based on user input information, text information of the generated storyboard, and evaluation text, outputs prompts for outputting code that constitutes a new video, provides the prompts to a natural language model, outputs the code that constitutes the video, and generates a video. For example, the video generation system 1 may output prompts for generating evaluation text (corresponding to evaluation text information EV1 in FIG. 5 ) based on user input information and text information of the generated storyboard, and provides the prompts to a natural language model to generate the evaluation text. For example, the video generation system 1 may generate evaluation prompts based on scenario information and input them to an AI (such as an LLM) to generate evaluation text indicating the evaluation of the scenario.

[0131] Furthermore, to evaluate the composition of a video, the video generation system 1 acquires a caption for the video by inputting one frame of the video into AI (such as a Contrastive Captioner Model or an Image Captioning Model). Then, the video generation system 1 generates an evaluation text indicating an evaluation of the composition of the video by inputting the acquired caption and a scenario sentence into AI (such as an LLM). For example, the video generation system 1 may evaluate whether the image is consistent with the caption by comparing the acquired caption and the scenario sentence using AI (such as an LLM). The video generation system 1 may also accept a specification of the type of evaluator the user desires to have evaluate the video. For example, the video generation system 1 may perform evaluation from the perspective of a market strategy, a copywriter, a video director, a specific person, or the like.

[0132] <1-4-6. Example of Selection of Multiple Cuts> From here, several examples of editing processing will be described. In the conventional technology, there is a problem that multiple cuts cannot be simultaneously edited on a storyboard. As such, the conventional technology has a problem with usability, and there is room for improvement in usability. Therefore, the video production system 1 may select multiple cuts and perform editing processing (editing processing) as shown in FIG. 23. FIG. 23 is a diagram showing an example of editing processing. As a specific example, FIG. 23 is a diagram showing an example of editing processing based on the selection of multiple cuts.

[0133] 23, the video production system 1 may provide the user with content CT41 and accept correction instructions in response to the user's selection of multiple cuts. The content CT41 is a display screen (content) for accepting the user's correction instructions for a video including multiple cuts (scenes) such as cuts CU1 to CU4. For example, the client UI display unit 400 displays the content CT41.

[0134] The user selects multiple cuts from cuts CU1 to CU4 via content CT41 and inputs information instructing corrections to the selected multiple cuts. For example, the user selects cuts CU1 to CU3 by performing an operation to select the range in which cuts CU1 to CU3 are displayed (such as by surrounding it with a line) among cuts CU1 to CU4 in content CT41, or by clicking on each of cuts CU1 to CU3.

[0135] After selecting cuts CU1 to CU3, the user inputs a prompt (e.g., text information) indicating a correction instruction as user input information, thereby instructing the video production system 1 to perform a correction on cuts CU1 to CU3. The video production system 1 then executes corrections on cuts CU1 to CU3 in response to the user's correction instruction. This allows the user to select multiple cuts and correct the structure and content of the selected cuts by inputting a prompt. For example, scenario information for each cut includes the content of the cut, the characters, and the time of filming of the cut. The video production system 1 regenerates the scenario by providing the user's input information and scenario information to an AI (e.g., an LLM). The video production system 1 executes any correction process based on the content of the user's correction instruction. For example, the video production system 1 may update the USD file of the video as needed, or may simply change the order of the cuts while leaving the USD file unchanged. Through the above-described process, the video production system 1 can improve usability.

[0136] <1-4-7. Example of Use of Time Information> Conventional technology has an issue in that after a specific cut is modified, other scenes affected by the modification are not modified (changed). As such, conventional technology has usability issues and there is room for improvement in usability. Therefore, the video production system 1 may perform editing processing using time information, as shown in FIG. 24. FIG. 24 is a diagram showing an example of editing processing. As a specific example, FIG. 24 is a diagram showing an example of editing processing based on modification content determined according to time information. For example, the video production system 1 includes AI-generated date and time information (also referred to as "time information" or "date information") for each cut, and determines the cutting of the video and the modification content based on that information.

[0137] 24, the video production system 1 manages each cut by associating it with date and time information (time information) corresponding to that cut. For example, time information TI11 indicating 10:31 on December 21, 2023 is associated with cut CU11. Furthermore, time information TI12 indicating 10:41 on December 21, 2023 is associated with cut CU12. In this way, cut CU11 and cut CU12 are cuts that are close in time (nearby). In this case, if cut CU11 is modified, the video production system 1 also reflects the modification in cut CU12. Time information (date information) is associated with each cut of the video data.

[0138] For example, if the clothing of person X in cut CU11 is modified, the video production system 1 also modifies the clothing of person X in cut CU12 in the same way as the clothing of person X in cut CU11. For example, the video production system 1 compares the time information between each cut, and if there is a cut (also called an "influenced cut") whose time difference with the modified cut (also called a "cut to be modified") is within a predetermined range, it determines that the modification based on the modification content of the cut to be modified should also be reflected in the influenced cut.

[0139] Then, the video production system 1 executes corrections to the influencing cuts based on the correction content of the cut to be corrected. Note that the corrections (changes) to the influencing cuts based on the influence between cuts are not limited to human assets, and the video production system 1 may also perform the corrections (changes) to environmental assets (weather, lighting, props such as candles that change over time, etc.).

[0140] For example, cut CU21 is associated with time information TI21 indicating 10:31 on December 21, 2023. Furthermore, cut CU22 is associated with time information TI22 indicating 18:41 on December 21, 2023. In this way, cut CU21 and cut CU22 are cuts that are distant in time. In this case, if cut CU21 is modified, the video production system 1 does not reflect the modification in cut CU22.

[0141] For example, if the clothing of person X in cut CU21 is modified, the video production system 1 does not modify the clothing of person X in cut CU22 in accordance with the modification of the clothing of person X in cut CU21. For example, the video production system 1 compares time information between each cut, and if there is no cut (affecting cut) whose time difference with the modified cut (cut to be modified) is within a predetermined range, it determines that the modification based on the modification content of the cut to be modified will not be reflected in other cuts.

[0142] For example, when generating a scenario, the video production system 1 generates date and time information for each cut using AI (such as LLM) and saves it as meta information for each cut. This date and time information is used to consider the relationship between previous and next cuts and to understand the season, etc. The generated fictitious date and time information is included to maintain the content of each cut and the relationship between each cut. For example, if the shot is taken in a morning when it is snowing, the video production system 1 may set the time to 6:30 AM on January 24th. Furthermore, if the shot has a strong relationship with the previous and next cuts, the video production system 1 may set the time to a nearby time on the same day, such as 6:30 AM on January 24th and 7:00 AM on January 24th.

[0143] For example, the video production system 1 changes the clothes worn by the user, the sun's lighting settings, the haze of the air, etc. depending on the season and time. Also, if the date and time information of a target cut and a previous cut is close, changing the target cut will also affect the previous and next cuts. For example, if the same person appears in another cut and the time is close to the target cut, changing the clothes in the target cut will also change the clothes in the other cut. Through the above-mentioned processing, the video production system 1 can improve usability.

[0144] <1-4-8. Response Examples Depending on the Degree of Verbalization> Conventional techniques have the problem that it is difficult to make corrections in line with the user's intentions when the user cannot specifically verbalize the content of the corrections. As such, conventional techniques have usability issues and there is room for improvement in usability. Therefore, the image generation system 1 may use the process described below to enable corrections in line with the user's intentions even when the user cannot specifically verbalize the content of the corrections.

[0145] For example, the video production system 1 may change the response method depending on the input sentence. As shown in Fig. 25, the video production system 1 may vary the response depending on the level of abstraction and verbalization of the user's instruction. Fig. 25 is a diagram showing an example of editing processing. As a specific example, Fig. 25 is a diagram showing an example of editing processing depending on the level of verbalization.

[0146] 25 shows a case where a user issues a purpose-based correction instruction such as, "Please make the video so that the viewer feels like they are in their daily lives." In this case, the video generation system 1 generates a video accompanied by a specific instruction sentence. For example, in response to a purpose-based correction instruction such as, "Please make the video so that the viewer feels like they are in their daily lives," the video generation system 1 adds a sentence such as, "Lower your hands at a natural angle and move them in accordance with the direction of your body," and presents the video corrected based on the sentence to the user by displaying it on a display or the like.

[0147] 25 illustrates a case where the user instructs to correct an abstract sentence, such as "Keep both arms in a neutral position." In this case, the image generation system 1 presents a plurality of specific sentences to the user and allows the user to select one. For example, the image generation system 1 presents a plurality of sentences, such as "Put your hands down and move them in a direction that points up and down in the air," "Scratch your head with your hands and then lower your hands," and "Put your hands in your pockets," to the user by displaying them on a display or the like. The image generation system 1 then presents a corrected image based on the sentence selected by the user from the plurality of sentences to the user by displaying it on a display or the like.

[0148] Response example AP3 in Figure 25 shows a case where the user issues a specific instruction to correct a sentence, such as "Please look at the car in front of you on the right in two seconds." In this case, the image generation system 1 presents the image to the user by displaying on a display or the like an image that has been changed (corrected) according to the sentence entered by the user. For example, the difference between a purpose-based sentence, an abstract sentence, and a specific sentence may be determined by AI (such as an LLM), and the image generation system 1 may change the prompt it generates accordingly and pass the processing to the AI ​​(such as an LLM).

[0149] For example, the video production system 1 may use a sensor to allow the user to specify motion or camera movement. As shown in Fig. 26 , the video production system 1 may superimpose (overlap) motion information obtained from an arbitrary sensor such as a web camera on the video and accept user corrections. Fig. 26 is a diagram showing an example of editing processing. As a specific example, Fig. 26 is a diagram showing an example of editing processing using superimposed display on the video.

[0150] For example, the video production system 1 displays motion information MT obtained from any sensor, such as a mobile motion capture device or a web camera, superimposed on the video MV11 and accepts user correction instructions. In this way, the video production system 1 displays motion information, including facial expressions, superimposed on the current video, thereby visualizing the difference between the current video and the motion information and accepting user correction instructions based on this. For example, if a four-second cut is played, the four-second cut is always played, and the user can change the motion information as many times as they like. For example, if the user finds motion information they like, the video production system 1 may ultimately generate natural motion information and change the motion of the video.

[0151] The user's motion is tracked using a mobile motion capture device, a web camera, or the like. The user may also select in advance from the UI using a mouse which character's motion to modify. The size, position, and orientation of the overlaid portion of the user's motion display are determined based on the size and orientation of the head and body in the video. The user may also adjust the size, position, and orientation in detail using a mouse and keyboard. Motion recording may be performed by repeatedly repeating the cuts shown in the image above, or by starting recording a few seconds after pressing the start button and adjusting one's position. After filming, the user may press a motion generation button to apply the input motion to the specified character.

[0152] Furthermore, since the captured motion itself may be unnatural, the video generation system 1 may use a motion-to-motion AI to estimate or generate motion from the captured motion, convert it into a more natural motion, and then adapt it to the motion of the characters. The video generation system 1 may also estimate or generate motion using input motion and natural language. For example, after inputting motion, the video generation system 1 may submit the motion to an AI (such as an LLM or a text & motion to motion generation model) for processing along with natural language input by the user, such as "Make the movement lively like this."

[0153] Furthermore, the user may change the camerawork by moving their hand as if it were a camera. For example, the image production system 1 acquires the user's hand movements using a sensor such as a web camera. As shown in Fig. 27 , the image production system 1 presents (superimposed display) the changed camerawork as a frame superimposed on the image. Fig. 27 is a diagram showing an example of editing processing. As a specific example, Fig. 27 is a diagram showing an example of superimposed display of camerawork on the image.

[0154] For example, the video production system 1 displays a frame CW1 indicating the camerawork before the change and a frame CW2 indicating the camerawork after the change, superimposed on the video MV12. Then, when the user instructs the video production system 1 to use the changed camerawork, the video production system 1 executes rendering processing based on the changed camerawork. In this way, when the user approves the changed camerawork, the video is rendered using the changed camerawork.

[0155] The user may also specify camerawork using an IMU or ImageSLAM of their mobile device (such as a smartphone). For example, the image generation system 1 may present the main character by placing it in real space (such as on a real desk) using AR (Augmented Reality). The image generation system 1 displays the changed camerawork by superimposing a frame on the image, similar to the case shown in FIG. 27 .

[0156] As described above, in the video generation system 1, a user may determine motion and camera movement using hand or smartphone movements. For example, when a user determines camerawork using hand movements, the user may imagine a person designated with the other hand and determine the relative position of the camera based on the distance between the left hand (camera) and the right hand (person). Furthermore, the video generation system 1 displays the camerawork changed by the user's hand or camera input as a frame superimposed on the video, but the user may then fine-tune the frame and its movement using a mouse. Furthermore, because the input camerawork is manually entered, it may be unnatural. Therefore, the video generation system 1 may correct the camerawork to make it more natural using a motion-to-motion generation model, a text&motion-to-motion generation model, or the like. Through the above-described processing, the video generation system 1 can improve usability.

[0157] <1-4-9. Example of Use of Qualitative Values> In conventional technologies, there are cases where consideration is not given to how the system determines the current state and change history. For example, there are cases where a user wants to compare the current state with the current state and make adjustments, such as "stand a little further back" or "make it a little brighter," and there are issues with how the system understands (grasps) the current state. As such, conventional technologies have usability issues, and there is room for improvement in usability. Therefore, the video generation system 1 may be able to appropriately determine the current state and change history by performing the processing described below.

[0158] For example, the video production system 1 may store each point, such as motion speed and brightness, as a quantitative numerical value and modify the video by comparing the values. The video production system 1 may also perform editing processing using qualitative values, as shown in Fig. 28. Fig. 28 is a diagram showing an example of editing processing. As a specific example, Fig. 28 is a diagram showing an example of editing processing based on modification content determined according to the qualitative values.

[0159] 28, value information VL21 indicates quantitative values ​​associated with video MV21. For example, value information VL21 includes qualitative values ​​such as the brightness of the entire town, the walking speed of people, the speed at which people move their faces, the positions of people, the speed of cars, and the positions of cars, and indicates values ​​associated with cuts of video MV21. In this way, video production system 1 stores all values ​​that can be changed as quantitative values ​​and changes the video by comparing them with those values.

[0160] 28, the user gives a correction instruction for video MV21, saying, "Move your face a little more slowly," and video production system 1 generates video MV22 by slowing down the speed at which the face moves in video MV21. Based on the correction instruction, saying, "Move your face a little more slowly," video production system 1 generates video MV22 with value information VL22 in which the value of the speed at which the person's face moves, in the value information VL21 of video MV21, is reduced from "21" to "10."

[0161] 28, the value information VL22 indicates quantitative values ​​associated with the corrected video MV22. For example, the value information VL22 includes qualitative values ​​such as the brightness of the entire town, the walking speed of people, the speed at which people move their faces, the positions of people, the speed of cars, and the positions of cars, and indicates values ​​associated with cuts and the like of the video MV22.

[0162] In this way, the video production system 1 may perform video editing processing using qualitative values. For example, the video production system 1 stores quantitative values ​​for each point (item), such as movement speed and brightness, and compares the values ​​to modify the video. As described above, classifications such as the brightness of the entire city and people's walking speed, for which qualitative values ​​are stored, are pre-set items. For example, the video production system 1 may set and modify each item using AI (such as LLM) from natural language, or may directly modify the setting value. Through the above-described processing, the video production system 1 can improve usability.

[0163] An example of a method for acquiring each value is shown below, but the acquisition method is not limited to the above and other acquisition methods may also be used. For example, for lighting, the video production system 1 acquires the setting values ​​(position, rotation, intensity, color, etc.) of the light in the 3DCG. For color grading, the video production system 1 acquires the white balance, color temperature, color cast correction, saturation, exposure, contrast, highlights, shadows, white level, black level, color, LUT settings, etc. set in composite editing. For walking speed, the video production system 1 acquires the position movement speed of the waist bone of the target 3D model. For face movement speed, the video production system 1 acquires the rotation speed of the head bone. For position, the video production system 1 acquires the position of the 3D model.

[0164] <1-4-10. Example of rendering process during input> Conventional techniques have the problem that it is difficult for users to wait for rendering time (low usability). As such, conventional techniques have usability issues and there is room for improvement in usability. Therefore, the video generation system 1 may improve usability related to rendering by performing the process described below.

[0165] For example, the video production system 1 may start rendering processing in the middle of text input. The video production system 1 may perform rendering processing in the middle of user input, as shown in Fig. 29. Fig. 29 is a diagram showing an example of editing processing. As a specific example, Fig. 29 is a diagram showing an example of editing processing based on rendering processing in the middle of input.

[0166] In FIG. 29 , the video generation system 1 executes rendering processing and the like in response to the user's input of text TX31, "scratching his head with his hand," and displays video MV31. For example, the "|" at the end of text TX31 indicates that the user is in the middle of inputting a correction instruction. Once the video generation system 1 is able to understand the text as a sentence, it starts processing in the background. For example, the video generation system 1 starts rendering processing and the like in the background once the user has input up to "scratching his head with his hand." In this way, the video generation system 1 may start rendering processing while the user is still inputting text.

[0167] In Figure 29, the user changes the sentence from text TX31 to text TX32, which reads "Put your hands down." For example, the "|" at the end of text TX32 indicates that the user is in the middle of inputting a correction instruction. The video generation system 1 executes processing in response to the change from text TX31 to text TX32. For example, when the sentence is changed, the video generation system 1 stops the processing being performed at that time and starts the processing again. For example, the video generation system 1 ends the processing that was being performed based on text TX31, and executes rendering processing and the like based on text TX32 to display video MV32.

[0168] 29, the user changes the text TX32 to text TX33, "Put your hands down, in a natural way," completes the sentence, and instructs execution of the process by pressing a start button or the like. The video production system 1 ends the background process that has begun, and displays video MV33, for which rendering process and the like have been performed based on the text TX33. Through the above-described process, the video production system 1 can improve usability.

[0169] <1-4-11. Example of Confirmation Work During Editing> In conventional technology, there are issues such as low usability when it comes to confirmation work during editing, and there is room for improvement. Therefore, the video production system 1 may perform confirmation work during editing as shown in Fig. 30. Fig. 30 is a diagram showing an example of editing processing. As a specific example, Fig. 30 is a diagram showing an example of confirmation work during editing.

[0170] Confirmation task CP1 in Fig. 30 shows an example of a confirmation task when it is necessary to check changes in the time axis direction, such as motion or camera work. For example, in confirmation task CP1, the user repeatedly selects a video that is close to the ideal. Also, confirmation task CP2 in Fig. 30 shows an example of a confirmation task other than when it is necessary to check changes in the time axis direction. For example, in confirmation task CP2, the user repeatedly selects an image that is close to the ideal.

[0171] For example, the video production system 1 may render images of all cuts and then render a video starting with the one the user wants to edit. For example, the video production system 1 may present the generated video to the user and allow the user to decide whether to edit the input information. For example, as shown in confirmation tasks CP1, CP2, etc., when presenting the user with the results of generating several video patterns that incorporate variations in response to the user's input information, there is an issue of waiting time for rendering multiple videos. Thus, the conventional technology has usability issues and there is room for improvement in usability.

[0172] Therefore, in order to reduce this waiting time, for tasks other than the confirmation task that requires viewing changes along the time axis (corresponding to confirmation task CP2), the video production system 1 renders only one or several frames from the video and presents candidates as images to the user, thereby enabling the video production system 1 to reduce the rendering waiting time.

[0173] Furthermore, when the user edits motion or camerawork, the video generation system 1 presents video candidates. For example, methods for selecting several frames from a video to be used for rendering include simply rendering only the first and last two frames of the video, rendering several frames with large animation changes in the USD file, and having AI select frames to display as highlights from all frames. For example, when the video generation system 1 renders multiple images, the user can check the generated results in the form of a flip book by hovering the mouse cursor over them.

[0174] Furthermore, the video generation system 1 can reduce waiting times during video preview by sequentially starting video generation for each pattern after completing the generation of images for candidate selection. The order in which videos are generated for each pattern can be generated in a random order, or the user can select the rendering order by pressing a button on the UI, or the images can be rendered in order of the length of time the mouse cursor is hovered over them. The above-described process allows the video generation system 1 to improve usability.

[0175] An example of a processing flow related to the above-mentioned confirmation work will now be described with reference to Fig. 31. Fig. 31 is a flowchart showing a processing procedure executed by the video production system. As a specific example, Fig. 31 is a flowchart showing a processing procedure related to editing processing.

[0176] 31, the video production system 1 generates several patterns of USD files from the user's settings (step S201). If image writing of all patterns of USD files has been completed (step S202: Yes), the video production system 1 ends the process of writing images of files (e.g., steps S202 to S204).

[0177] If image writing of all patterns of USD files has not been completed (step S202: No), the video production system 1 writes out images of the USD files for which writing has not been completed (step S203). The video production system 1 displays the written images on the UI (step S204). Thereafter, the video production system 1 starts the processing from step S205 onwards, and repeats the processing of steps S202 to S204 until image writing of the files is completed.

[0178] If the video production system 1 has finished writing the moving images of all patterns of USD files (step S205: Yes), it ends the process of writing the moving images of the files (for example, steps S205 to S207).

[0179] If the video writing of all patterns of USD files has not been completed (step S205: No), the video production system 1 writes out the videos of the USD files for which writing has not been completed (step S206). The video production system 1 displays the written out videos on the UI (step S207). Thereafter, the video production system 1 repeats the processes of steps S205 to S207 until the video writing of the files is completed.

[0180] <1-4-12. Example of Processing in Accordance with Range Selection> In conventional technology, when a person (also called an "amateur") with no experience (knowledge) in video (movie) generation attempts to create a video, it is difficult for them to determine what is good and what is bad, and they often make incorrect decisions. As such, conventional technology has issues with usability, and there is room for improvement in usability. Therefore, the video generation system 1 may evaluate the generated video, as shown in FIG. 32. FIG. 32 is a diagram showing an example of evaluation processing in response to range selection.

[0181] In Fig. 32, the video production system 1 performs evaluation processing in response to the range selection. Video MV40 in Fig. 32 shows a moving image generated in response to user input. Video MV41 in Fig. 32 shows a state in which the user specifies a part of a person as the range they want evaluated, and text TX41 indicating the evaluation by the video production system 1 of that range is superimposed. In Fig. 32, the video production system 1 performs an evaluation (suggests a revision) on the part of the person specified by the user, as indicated by the text TX41, saying, "Showing this person's facial expression will make it easier for the user to understand their emotions."

[0182] Video MV42 in Fig. 32 shows a state in which text TX42 indicating an evaluation by video production system 1 is further superimposed. In Fig. 32, video production system 1 performs an evaluation on the part of the person specified by the user, as indicated by text TX42, saying, "From the perspective of production, it would be better to capture the face from directly in front."

[0183] For example, if the user thinks the evaluation indicated by text TX42 is good, video production system 1 receives a correction instruction from the user based on that evaluation. Then, video production system 1 generates and displays multiple candidate videos MV43, MV44, and MV45 corresponding to the evaluation indicated by text TX42. In this way, video production system 1 presents the multiple candidate videos MV43, MV44, and MV45 corresponding to the evaluation indicated by text TX42 to the user.

[0184] For example, a user specifies a range with a mouse, and the image generation system 1 evaluates the specified range and starts a dialogue (discussion) with the user. In response to the range selected with the mouse, the image generation system 1 starts an AI evaluation of that range. For example, the image generation system 1 performs an evaluation (presents a correction proposal) once every N seconds until the user instructs correction. When the user clicks with the mouse at a point they think is good, the image generation system 1 presents multiple correction suggestions based on the dialogue (discussion) up to that point. Through the above-described processing, the image generation system 1 can improve usability.

[0185] <1-4-13. Example of Processing According to Depth of Field> Conventional techniques have the problem that it is difficult to suppress an increase in the time required for rendering when the environmental asset has a large number of polygons or a high texture resolution, resulting in a heavy asset. As such, conventional techniques have usability issues and there is room for improvement in usability. Therefore, the video generation system 1 may perform processing according to the depth of field, as shown in FIG. 33. FIG. 33 is a conceptual diagram showing an example of processing according to the depth of field.

[0186] In Figure 33, the video generation system 1 uses a high number of polygons and high-resolution texture within a predetermined range in front of and behind the subject, and a low number of polygons and low-resolution texture within a predetermined range in front of and behind the subject, depending on the depth of field. In Figure 33, objects with a high number of polygons and high-resolution texture (circles close to the subject) are indicated with dark hatching, and objects with a low number of polygons and low-resolution texture (circles farther from the subject) are indicated with light hatching. In this way, the video generation system 1 changes the number of polygons and texture resolution depending on the depth of field. The video generation system 1 calculates the depth of field using the following equations (1) to (3).

[0187]

[0188]

[0189]

[0190] Equation (1) is a function for calculating the front depth of field (mm). The image generation system 1 calculates the front depth of field using equation (1). For example, in FIG. 33 , the front depth of field corresponds to the front side of the subject (the side closer to the camera). Furthermore, equation (2) is a function for calculating the rear depth of field (mm). The image generation system 1 calculates the rear depth of field using equation (2). For example, in FIG. 33 , the rear depth of field corresponds to the rear side of the subject (the side farther from the camera). Equation (3) is a function for calculating the depth of field. The image generation system 1 calculates the depth of field by adding the front depth of field and the rear depth of field using equation (3).

[0191] For example, the video generation system 1 replaces the polygon count and texture resolution of portions closer to the camera than the forward depth of field and portions farther from the camera than the rear depth of field with those of a lower polygon count and texture resolution before rendering. For example, when creating a blur in terms of depth of field, the video generation system 1 reduces (lowers) the polygon count and texture resolution. In this way, the video generation system 1 determines the polygon count and texture resolution based on criteria such as focal length, F-number, and subject distance. For example, the video generation system 1 may make the above determination when generating a USD file. Through the above-described processing, the video generation system 1 can improve usability.

[0192] <1-4-14. Example of Playback Processing During Confirmation Work> In conventional technology, when all generated videos are played back and confirmed, it is difficult to suppress an increase in the time required for the confirmation work. As such, conventional technology has issues with usability, and there is room for improvement in usability. Therefore, the video generation system 1 may perform playback appropriate for the confirmation work. The video generation system 1 may change the playback mode in response to a user operation. For example, the video generation system 1 may play back the video in a mode similar to a flip book in response to a user operation.

[0193] The video production system 1 may advance playback of the video by a predetermined number of seconds (e.g., 0.5 seconds) with each click or mouse wheel rotation by the user. For example, the video production system 1 may advance frame by frame within a cut according to the movement or position of the mouse or mouse wheel. As shown in Fig. 34, the video production system 1 may advance playback of the video by the number of seconds corresponding to the mouse position.

[0194] In FIG. 34 , the video generation system 1 may provide content CT41 to the user and accept an instruction from the user regarding the degree to which the video should be advanced (number of seconds, number of frames, etc.). The content CT41 is a display screen (content) for accepting the user's instruction regarding the degree to which the video should be advanced. The content CT41 is superimposed on the video and places information for specifying the number of seconds to advance the video. In FIG. 34 , a range of 0 to 3.5 seconds can be specified, with the number of seconds to advance the video increasing from left to right. For example, the client UI display unit 400 displays the content CT41.

[0195] The user inputs information specifying the number of seconds to advance the video when playing the video via content CT41. In FIG. 34 , the user operates the mouse to position the mouse cursor MS in the area marked 1.0, thereby specifying 1.0 seconds as the number of seconds to advance the video. In this case, the video production system 1 advances the video at 1.0 second intervals in response to the user's specification of the number of seconds to advance the video. Note that the user may also operate the mouse to position the mouse cursor MS in the area marked 1.0 and click to specify 1.0 seconds as the number of seconds to advance the video. Through the above-described processing, the video production system 1 can improve usability.

[0196] <1-4-15. Highlighting Example> In conventional technology, it may be difficult to identify generated parts of a video that have been changed by editing or the like, making it difficult to prevent an increase in the time required for the confirmation process. As such, conventional technology has usability issues and there is room for improvement in usability. Therefore, the video generation system 1 may perform highlighting as shown in FIG. 35. FIG. 35 is a diagram showing an example of highlighting changed parts. For example, the video generation system 1 highlights the changed parts from the previous generation result. FIG. 35 explains an example in which a woman in a video has been changed.

[0197] 35 shows a first highlighting mode. Video MV51 shows a mode in which changed parts are highlighted by darkening parts other than the changed parts (for example, by lowering the brightness). Video production system 1 generates video MV51 and displays video MV51 to highlight parts that have been changed by editing or the like.

[0198] 35 shows a second highlighting mode. Video MV52 shows a mode in which the changed parts are highlighted (by adding color, for example) to highlight the changed parts. Video production system 1 generates video MV52 and displays video MV52 to highlight the parts that have been changed by editing, etc. In this way, videos MV51 and MV52 show a case in which the changed parts of people are highlighted.

[0199] 35 shows a third highlighting mode. Video MV53 shows a mode in which the changed portion is highlighted by indicating the time when the change occurred on the seek bar. Video production system 1 generates video MV53 including a seek bar on which colored points HL531 and HL532 are positioned corresponding to the time when the change occurred, and highlights the portion that has been changed by editing or the like by displaying video MV53.

[0200] 35 shows a fourth highlighting mode. Video MV54 shows a mode in which the changed portion is highlighted by indicating the time when the change occurred on the seek bar. Video production system 1 generates video MV54 including a seek bar with a colored bar HL54 positioned in a range corresponding to the time period when the change occurred, and displays video MV54 to highlight the portion that has been changed by editing or the like.

[0201] As described above, when a user configures settings related to video generation and regenerates a video, the generated video must be reviewed. In a UI that presents multiple generation results, the user must review each video one by one, placing a heavy burden on the user. Therefore, to reduce the burden of reviewing the video, the video generation system 1 presents the user with the differences from the previous video generation result. For example, the video generation system 1 presents the user with the differences by highlighting only the changed parts of the video or by displaying the time of the change in the video on a seek bar. This allows the user to review only the changed parts, and the video generation system 1 reduces the burden of the reviewing work. Through the above-described processing, the video generation system 1 can improve usability.

[0202] For example, methods for extracting changed portions of a video include detecting differences from a generated USD file and detecting differences by comparing a rendered video with a previously rendered video frame by frame. For example, in the method of detecting differences from a generated USD file, taking advantage of the characteristic of generating USD files each time and performing video rendering, the video generation system 1 compares the contents of the USD files generated before and after a user updates the video generation settings, and detects changed objects and the time at which the changes occurred. Also, in the method of detecting differences by comparing a rendered video with a previously rendered video frame by frame, the video generation system 1 compares the videos generated before and after a user updates the video generation settings, and detects changed pixels and the time at which the changes occurred.

[0203] <1-4-16. Example of Presentation of Relationships Between Cuts> Conventional techniques have the problem that when each cut is corrected, the relationship between the previous and next cuts may become unclear. As such, conventional techniques have usability issues, and there is room for improvement in usability. Therefore, the video production system 1 may present the relationship between cuts as shown in FIG. 36. FIG. 36 is a diagram showing an example of presenting the relationship between cuts.

[0204] 36 shows a case where cut CU61 is a corrected cut (also referred to as a "target cut"). A previous relationship bar TR60 indicating that the cut is a cut before the target cut is superimposed on cut CU60 before the target cut, cut CU61. For example, the previous relationship bar TR60 is a triangle with its base on the right and extending to the left. The previous relationship bar TR60 is displayed in such a way that the further away in time it is from the target cut, the further it extends to the left.

[0205] Furthermore, a cut CU62 that comes after the cut CU61, which is the target cut, is superimposed with a subsequent relationship bar TR62 indicating that the cut comes after the target cut. For example, the subsequent relationship bar TR62 is a triangle with its base on the left side and extending to the right. A subsequent relationship bar TR63 that comes after the cut CU62 that comes after the cut CU61, which is the target cut, is superimposed with a subsequent relationship bar TR63 indicating that the cut comes after the target cut. For example, the subsequent relationship bar TR63 is a triangle with its base on the left side and extending to the right.

[0206] The subsequent relationship bars TR62, TR63 are displayed in a manner that the further they extend to the left the further they are from the target cut. In Figure 36, cut CU63 is later than cut CU62, so the subsequent relationship bar TR63 is displayed in a manner that extends further to the right than the subsequent relationship bar TR62. Note that the triangular display manner is merely one example of a display manner, and any display manner can be used as long as it is possible to present the time relationship and its amount.

[0207] The video production system 1 plays a moving image including cut CUs 60 to 63 in response to a user's operation. For example, when displaying cut CU 60 in response to a user's operation, the video production system 1 superimposes a previous relationship bar TR60. For example, when displaying cut CU 62 in response to a user's operation, the video production system 1 superimposes a next relationship bar TR62. For example, when displaying cut CU 63 in response to a user's operation, the video production system 1 superimposes a next relationship bar TR63. In this way, when playing cuts before and after a target cut, the video production system 1 presents the amount of distance from the target cut. In this way, when playing cuts before and after the target cut, the video production system 1 superimposes information indicating the amount and direction of distance on the screen according to the number of seconds from the target cut, allowing the user to recognize whether the cut is before or after the target cut and how far the cut is from the target cut. Through the above-described processing, the video production system 1 can improve usability.

[0208] <1-4-17. Example of Object Selection> Conventional techniques have the problem that selecting an object during editing is difficult for the user (low usability). As such, conventional techniques have usability issues and there is room for improvement in usability. Therefore, the video generation system 1 may select an object as shown in FIG. 37. FIG. 37 is a diagram showing an example of object selection.

[0209] 37 , the user operates the mouse to select the right-hand person by positioning the mouse cursor MS71 in a range that includes the right-hand person of the three people. In this case, the video production system 1 accepts an operation to select the right-hand person where the mouse cursor MS71 is positioned as the target object. For example, the video production system 1 identifies the segmentation (range) corresponding to the right-hand person where the mouse cursor MS71 is positioned as the range selected by the user.

[0210] 37 , the user selects the central person by operating the mouse to position the mouse cursor MS72 in a range that includes the central person among the three people. In this case, the video production system 1 accepts an operation to select the central person where the mouse cursor MS72 is positioned as the target object. For example, the video production system 1 identifies the segmentation (range) corresponding to the central person where the mouse cursor MS72 is positioned as the range selected by the user.

[0211] 37 , the user operates the mouse to select the left person by positioning the mouse cursor MS73 in a range that includes the left person among the three people. In this case, the video production system 1 accepts an operation to select the left person where the mouse cursor MS73 is positioned as the target object. For example, the video production system 1 identifies the segmentation (range) corresponding to the left person where the mouse cursor MS73 is positioned as the range selected by the user.

[0212] In this way, the video generation system 1 recognizes the selection of the target object within the range recognized by segmentation. This allows the user to select the object they want to specify with just a click. For example, when a user wants to change the motion or asset of an object displayed in a video, they need to select the object to be changed. If the user can directly click on the object in the video, rather than using a UI that requires the user to select the object to be changed from a list of objects displayed in the video, the user can more intuitively select the object to be changed. The click range of an object in a video can be achieved by segmenting the pixel area of ​​each object for each frame of the video.

[0213] The click range of an object may be determined by the following method: For example, because object and camera information for each frame of a video is stored in the USD, it is possible to map which object a pixel seen from the camera in that frame points to, and the video generation system 1 can then reverse-calculate from that information to map the coordinates where the user clicked on the video to the object to be edited.

[0214] One method for determining the click range of an object is to associate a person selected by segmenting the object pixel by pixel calculated backward from the camera for each frame of the video with the target USD. For example, the video itself may include information about what is located in a particular area (e.g., x, y coordinates). For example, the video generation system 1 may select an object as shown in FIG. 38.

[0215] Frame FR in FIG. 38 indicates one frame of a video rendered from a camera. USD data OB in FIG. 38 indicates a 3D object and a camera on the USD. Because information about the 3D object and the camera is included in the USD, the video production system 1 can map the location on the 3D object that a pixel viewed from a camera (such as camera CM in FIG. 38 ) points to. This allows the video production system 1 to map the coordinates where a user clicks on a video to an object. Through the above-described processing, the video production system 1 can improve usability.

[0216] <1-4-18. Examples of Using 3D Models> The above-described processing is merely an example, and the image generation system 1 may execute various processes other than the above-described processing. In this regard, several examples will be described below.

[0217] For example, the video generation system 1 may perform a process of importing characters or props (such as products). A method of importing characters or props (such as products) in the video generation system 1 may involve, for example, registering a 3D model or a three-dimensional drawing of a character or prop that the user wishes to include in the video and using the model in the video. For example, the video generation system 1 may register a 3D model input by a user and use the model in the video. For example, the video generation system 1 may be capable of incorporating photos of a product taken from various angles. In this case, the video generation system 1 may use technology such as Neural Radiance Fields (NeRF) to generate a 3D model of the product from photos of the product taken from various angles and use the model in the video.

[0218] For example, the video generation system 1 accepts a user operation specifying an object to be changed in video data, and generates code in which the 3D data of the object specified by the user operation has been changed. For example, if a user specifies a character in a video (also referred to as "character A") as the object to be changed, and the user selects another character (also referred to as "character B") from a registered 3D model as the new character, the video generation system 1 generates code in which the 3D data of character A in the video has been changed to the 3D data of character B. This allows the video generation system 1 to generate video data in which character A in the video has been changed to character B represented by the registered 3D model.

[0219] The above-described process is merely an example, and the video production system 1 may generate a code in which the 3D data of an object indicated by a user's operation has been modified by any process. For example, the video production system 1 may generate a modified code by modifying the 3D data itself of an object indicated by a user's operation in the video data. For example, the video production system 1 may generate a modified code by modifying the external shape (height, etc.) of the 3D data of an object indicated by a user's operation in the video data. Furthermore, the video production system 1 may not need to partially perform refinement processing on registered people and props after rendering. Furthermore, the video production system 1 may provide a marketplace within the video production service for selling characters, props, etc.

[0220] <1-4-19. Example of Use of Reference Data> Furthermore, the video production system 1 may refer to videos and stories from previously created (video) projects. As shown in Fig. 39, when a sequel to a previously created project is to be used, the video production system 1 may accept input of that project as reference data. Fig. 39 is a diagram showing an example of processing using reference data.

[0221] For example, the video production system 1 acquires user input information regarding reference to other projects that the user inputs to the content CT51. The content CT51 is content for receiving user input information regarding an item for specifying with a check mark whether or not to refer to other projects, an item for specifying with a check mark the projects to refer to, an item for specifying with a check mark which information of the reference projects to refer to, etc.

[0222] For example, the client UI display unit 400 displays content CT51, and the sensor unit 300 accepts user input information. In Fig. 39, the user checks "Use other projects as reference" to select other projects. The user also checks "Product X Commercial Video" to select the project for a commercial video for Product X as reference.

[0223] The user also checks "Characters" and "Visual Style," choosing to use the characters and visual style of the commercial video project for Product X as a reference. The user also does not check "Story," "Content Style," or "BGM," choosing not to use the story, content style, or BGM of the commercial video project for Product X as a reference.

[0224] The video production system 1 generates a prompt PT51, which is scenario generation information (first input information), using user input information accepted by the content CT51. Note that, although not illustrated in Fig. 39, the video production system 1 may also generate the prompt PT51 using template input information such as the template TP1 shown in Fig. 6.

[0225] For example, the video production system 1 generates prompt PT51 by reflecting the characters and visual style of a commercial video project for product X. In FIG. 39 , the video production system 1 generates prompt PT51 including a constraint specifying that the character "Mike" from the commercial video for product X and the visual style "cinematic" from the commercial video for product X be used. In this way, the video production system 1 generates a prompt for generating a scenario based on reference data from past projects in accordance with the user's selection. This allows the video production system 1 to generate video data for character A in the video based on the past project specified by the user.

[0226] <1-4-20. Examples of Advantages of Having 3D Data> Because the image generation system 1 has 3D data internally, it has the following functions and advantages. The image generation system 1 has the following functions and advantages with regard to the images (videos) it generates. For example, the image generation system 1 can create more natural images through physical simulation. For example, the image generation system 1 can place an object on a platform, have it bounce, roll, or reproduce the natural swaying of cloth.

[0227] For example, the image generation system 1 can realistically reproduce the reflection of light depending on the lighting and object material. When a hood or mirror is highly reflective, people or objects in front of it will be brighter due to the reflected light. When a fabric is less reflective, people or objects in front of it will not receive much reflected light.

[0228] For example, the video generation system 1 can fix lighting and objects, reducing the possibility of time-series disruptions. For example, the video generation system 1 can reproduce video without disruptions even if the lighting position changes during the video. For example, the video generation system 1 can output video in real time by setting a simple light source. In this case, the video generation system 1 does not need to perform refinement processing.

[0229] For example, the video generation system 1 can faithfully reproduce the product itself within the video by inputting product data and characters as 3D data. For example, the video generation system 1 can determine the depth, global position, normal, etc., thereby reducing the possibility of failure during refinement processing. For example, the video generation system 1 can fix the three-dimensional position of sound sources such as speakers, making it easier to create interactions with sound. For example, the video generation system 1 can reduce rendering time by performing lighting processing such as Differential Rendering.

[0230] Furthermore, the video production system 1 has the following functions and advantages when it comes to editing (correcting) videos (moving images). For example, the video production system 1 can maintain consistency of the surrounding environment, lighting, etc., even when the camera position, angle, or camerawork is significantly corrected. For example, the video production system 1 can specify the position and angle three-dimensionally within a video.

[0231] For example, even after a certain amount of video has been generated, the video generation system 1 can change only specific elements, such as the characters or props, while maintaining the video for the rest of the video. For example, the video generation system 1 can change a character from a realistic human-like figure to a two-headed character. For example, the video generation system 1 can change an installed signboard from a blackboard type to a plastic board. For example, the video generation system 1 can change only a part of the logo on a product package.

[0232] For example, by viewing a captured scene in three dimensions, the image generation system 1 can correct images of areas that are not shown in video frames. For example, the image generation system 1 can install a light source in an area that is not shown. For example, the image generation system 1 can install a person or object in an area that is not shown and display only their shadow in the video frame. For example, the image generation system 1 can place a person or object in an area that is not shown and specify that the person in the image should look into the eyes of the person not in the image.

[0233] In addition to the above, the image generation system 1 has the following functions and advantages: For example, the image generation system 1 can easily convert content into content for 3D devices such as AR, VR (Virtual Reality), SRD (Spatial Reality Display), and 3D displays.

[0234] <1-5. Example of processing flow from the user's perspective> Next, as an example of a processing flow from the user's perspective, a procedure for information processing by the image generation system 1 in response to a user operation will be described with reference to Fig. 40. Fig. 40 is a flowchart showing the flow of processing in response to a user operation.

[0235] 40 , in the video production system 1, the user presses a project creation button (step S1). Then, in the video production system 1, the user inputs information required for video generation (step S2). For example, the user inputs information including the purpose of creating the video, the message to be conveyed through the video, the target users, the functional features of the product / service to be conveyed, the length of the video, the aspect ratio, etc.

[0236] Then, in the video production system 1, the user presses the storyboard, video, sound, and text logo creation buttons (step S3). For example, the video production system 1 may generate everything at once (including video data, for example) in response to user operation, or may present a storyboard and allow the user to make some modifications in response to instructions from the user before generating the video, sound, and text logo.

[0237] Then, the video production system 1 modifies the storyboard, video, sound, and text logo in response to user operations (step S4). For example, the video production system 1 may accept modifications to each of the items in the order of the user's preference.

[0238] Then, the video production system 1 performs export in response to a user operation (step S5). For example, the video production system 1 may perform export in a video file format such as mp4, avi, or mov in response to a user operation. Furthermore, for example, the video production system 1 may perform export in a file format of any video editing software such as Premiere Pro, After Effects, or DaVinci Resolve.

[0239] <1-6. Other Configurations and Processing Examples of the Video Creation System> Note that the configuration of the video creation system 1 described above is merely an example, and any configuration and processing can be adopted as long as the desired video (video) can be acquired. For example, in the above example, a case where USD data is used has been described as an example, but the video creation system may perform processing similar to the above processing without using USD data.

[0240] In this regard, the configuration and processing of the video production system 1A will be described below as an example. Note that the configuration and processing of the video production system 1A described below are the same as those of the video production system 1, except that processing is performed without using USD data, and therefore, explanations of the same aspects as those of the video production system 1 will be omitted as appropriate.

[0241] Fig. 42 is a diagram showing another example of an image generation system according to the present disclosure. The image generation system 1A includes an image generation module 100A, an information acquisition module 200, a sensor unit 300, and a client UI display unit 400. Although Fig. 42 shows only one of each component, the image generation system 1A may include multiple image generation modules 100A, multiple information acquisition modules 200, multiple sensor units 300, and multiple client UI display units 400.

[0242] First, the configuration of image generation module 100A that performs image generation processing will be described. Note that descriptions of the same aspects of image generation module 100A as those of image generation module 100 will be omitted where appropriate. Image generation module 100A includes an input text analysis unit 110, a sensor analysis unit 120, a prompt etc. generation unit 130A, an image generation unit 140A, a sound generation unit 150, a text / logo generation unit 160, a composite editing unit 170, an evaluation unit 180, a client UI module 190, etc.

[0243] The prompt etc. generation unit 130A generates various information required to generate images (videos) including prompts etc. to be input to the AI ​​model, similar to the prompt etc. generation unit 130. For example, the prompt etc. generation unit 130A generates a prompt using a user input and a pre-stored prompt (template, etc.). In FIG. 42 , the prompt etc. generation unit 130A has a scenario-oriented generation unit 131A, an image etc. generation unit 135, and a pre-processing unit 136.

[0244] The scenario-oriented generation unit 131A generates various types of information related to the generation of a scenario, similar to the scenario-oriented generation unit 131. The scenario-oriented generation unit 131A generates input information to be input to a model that outputs a scenario. For example, the scenario-oriented generation unit 131A functions as a scenario generation unit that generates scenario data related to video generation based on an input query.

[0245] The video etc. generation unit 135 generates various information related to the generation of video etc. Like the video etc. generation unit 132, the sound generation unit 133, and the text / logo generation unit 134, the video etc. generation unit 135 generates various information related to the generation of 3D (three-dimensional) data, sound information (also simply referred to as "sound"), text, and logos.

[0246] The video etc. generation unit 135 generates a prompt to be input to a large-scale language model (e.g., model M21). The video etc. generation unit 135 generates prompt data (also simply referred to as a "prompt") to be input to the large-scale language model to acquire scene-constituting data. The video etc. generation unit 135 generates a prompt to cause the large-scale language model to output query data (also simply referred to as a "query") for acquiring the scene-constituting data.

[0247] The video etc. generation unit 135 generates prompts to be input to the large-scale language model regarding the generation of video. The video etc. generation unit 135 generates prompts to be input to the large-scale language model regarding the generation of sound information (sound). The video etc. generation unit 135 generates prompts to be input to the large-scale language model regarding the generation of text / logo etc. generation unit 134.

[0248] The pre-processing unit 136 acquires various pieces of information used to generate 3D scene data. The pre-processing unit 136 functions as a scene configuration data acquisition unit that acquires multiple types of scene configuration data that make up the 3D scene data based on scenario data. For example, the multiple types of scene configuration data include background information, cast information including the cast and their positions, lighting information, motion information, camerawork information, sound information such as background music and sound effects, dialogue and narration information, text information, and logo information.

[0249] The preprocessing unit 136 functions as a query data acquisition unit that acquires query data for acquiring scene configuration data based on scenario data. The query data is data output by a large-scale language model (e.g., model M21) based on the scenario data.

[0250] For example, the query data is data output by inputting prompt data, which instructs a large-scale language model to output multiple types of scene-constituting data based on scenario data, into the large-scale language model. For example, the query data is data indicating, in text format, information necessary for acquiring multiple types of scene-constituting data. The query data includes necessary acquisition information indicating, in text format, information corresponding to each of the multiple types of scene-constituting data.

[0251] The prompt data also includes a plurality of pieces of instruction information for instructing the output of each of the plurality of types of scene-constituting data. For example, the prompt data includes a plurality of pieces of instruction information for instructing the selection of each of the plurality of types of scene-constituting data.

[0252] The preprocessing unit 136 acquires scene construction data corresponding to the query data. The preprocessing unit 136 acquires at least one type of scene construction data from a database storing data used as scene construction data. For example, the preprocessing unit 136 searches the database using the query, selects scene construction data corresponding to the query, and acquires the selected scene construction data from the database.

[0253] The pre-processing unit 136 uses the query data to generate at least one type of scene configuration data from among multiple types of scene configuration data. The pre-processing unit 136 functions as a layout generating unit that generates layout information based on the scene configuration data including at least one of background information and cast information.

[0254] The video generation unit 140A executes processing related to video generation. The video generation unit 140A is a video acquisition unit that acquires video data based on 3D scene data. The video generation unit 140A generates video using various information generated by the prompt generation unit 130A. For example, the video generation unit 140A is a video generation unit that generates video data based on 3D scene data. Note that the video generation unit 140A may acquire video data in any manner. For example, the video generation unit 140A may acquire video data by transmitting data used to generate the video data to an external service providing device (such as a vendor) that provides a video data generation service, and then receiving the video data generated by the service providing device from the service providing device. In FIG. 42 , the video generation unit 140A includes a 3D scene generation unit 141A, a rendering unit 142, and a video refinement unit 143.

[0255] The 3D scene generation unit 141A is a scene generation unit that generates various information related to 3D scenes. For example, the 3D scene generation unit 141A generates 3D scene data related to video generation based on multiple types of scene configuration data. The 3D scene generation unit 141A generates 3D scene data based on layout information.

[0256] The 3D scene generation unit 141A generates 3D scene data based on the layout information and scene-constituting data other than the scene-constituting data used in the layout information, among the multiple types of scene-constituting data. Each of the multiple types of scene-constituting data includes at least one of data in a data file format corresponding to the type, text format data, and parameter data. For example, the data file format is a file format for 3D data.

[0257] The rendering unit 142 executes various processes related to rendering, such as rendering the 3D scene data generated by the 3D scene generation unit 141A.

[0258] The sound generation unit 150 executes a process for generating sound. The sound generation unit 150 generates sound information using the sound information and dialogue narration information acquired by the pre-processing unit 136. The sound generation unit 150 uses the sound information and dialogue narration information acquired by the pre-processing unit 136 as sound information.

[0259] The text / logo generator 160 executes a process for generating at least one of text and a logo. The text / logo generator 160 generates logo text information using the logo information acquired by the preprocessor 136. The text / logo generator 160 uses the logo information acquired by the preprocessor 136 as the logo text information.

[0260] With the above-described configuration, the image generation module 100A generates a prompt for scenario generation by combining a pre-stored prompt with a user's input. The image generation module 100A generates a scenario by inputting the generated prompt into an AI model such as an LLM. The image generation module 100A also acquires multiple types of scene configuration data that constitute 3D scene data based on the generated scenario. The image generation module 100A performs image generation, sound generation, and text / logo generation using multiple types of scene configuration data.

[0261] From here, an example of processing by the video generation system 1A will be described. An example of the flow of video generation processing by the video generation system 1A will be described below using FIG. 43. FIG. 43 is a diagram showing an example of the flow of video generation processing of the present disclosure. Note that explanations of points similar to those of the video generation system 1 will be omitted as appropriate. For example, the process of generating a scenario FD1 from user input information UIN1 and a model M1 is similar to the process of the video generation system 1 shown in FIGS. 3 and 4, etc., and therefore explanations thereof will be omitted.

[0262] The video production system 1A generates a prompt PT21, which is prompt data represented as "Prompt" in Fig. 43. In Fig. 43, the video production system 1A generates the prompt PT21 from a scenario FD1. The video production system 1A generates the prompt PT21 from the scenario FD1 using an AI model (hereinafter also referred to as "model M20") that takes the scenario FD1 as input and outputs the prompt PT21. Any AI model, such as an LLM (large scale language model), can be used as the model M20, as long as it is capable of producing the desired output in response to the input.

[0263] The prompt PT21 includes a plurality of pieces of instruction information instructing the output of each of a plurality of types of scene-constituting data. The prompt PT21 includes a plurality of pieces of instruction information instructing the selection of each of a plurality of types of scene-constituting data. For example, the prompt PT21 includes a plurality of pieces of instruction information instructing each of background selection, cast and position selection, motion selection, camerawork selection, lighting selection, sound selection, and text logo selection. Note that while FIG. 43 illustrates one prompt PT21, the prompt PT21 may be divided into a plurality of files.

[0264] For example, prompt PT21 includes a plurality of prompt data including instruction information for background selection, cast and position selection, motion selection, camerawork selection, lighting selection, sound selection, and text logo selection. For example, prompt PT21 includes a plurality of prompts such as a prompt for background selection, a prompt for cast and position selection, a prompt for motion selection, a prompt for camerawork selection, a prompt for lighting selection, a prompt for sound selection, and a prompt for text logo selection. For example, sound selection includes selection of at least one of background music, sound effects, and dialogue narration.

[0265] The video generation system 1A generates necessary information by inputting a prompt PT21 to the model M21 and causing the model M21 to output a query for outputting information (also referred to as "necessary information"). The model M21 is a large-scale language model that outputs a query (necessary information) for acquiring scene configuration data in response to the input of the prompt. Any AI model, such as an LLM (large-scale language model), can be used for the model M21 as long as it can produce the desired output in response to the input.

[0266] In FIG. 43, the video production system 1A inputs a prompt PT21 to a model M21 and causes the model M21 to output the necessary information ND1 to ND9, thereby generating the necessary information ND1 to ND9.

[0267] The required information ND1 is background required information used to acquire background data. The video production system 1A inputs a prompt including instruction information for selecting a background from the prompt PT21 to the model M21, and causes the model M21 to output the required information ND1, thereby acquiring the required information ND1.

[0268] For example, the necessary acquisition information ND1 is necessary background acquisition information including information for acquiring background data (also simply referred to as "background") from a background DB (for example, background designation information such as background #1). For example, the background DB, database DB1 shown in Fig. 43, is a database that stores a background list of background data used as background information for a video.

[0269] The required information ND2 is cast and position required information used to acquire data on the cast and their positions. The video production system 1A inputs a prompt from the prompt PT21, which includes instruction information for selecting a cast member and selecting the position of that cast member, to the model M21, and causes the model M21 to output the required information ND2, thereby acquiring the required information ND2.

[0270] For example, the necessary acquisition information ND2 is necessary cast / position acquisition information that includes information for acquiring cast data (also simply referred to as "cast") from the cast DB (for example, information specifying the cast, such as cast #1). For example, the database DB2 shown in Figure 43, which is a cast DB, is a database that stores a cast list of cast data used as information about the cast in the video.

[0271] The required acquisition information ND3 is lighting required acquisition information used to acquire lighting data. The video production system 1A inputs a prompt including instruction information for selecting lighting from the prompt PT21 to the model M21, and causes the model M21 to output the required acquisition information ND3, thereby acquiring the required acquisition information ND3.

[0272] For example, the required acquisition information ND3 is lighting acquisition required information that includes information (e.g., lighting specification information such as lighting setting #1) for acquiring lighting setting data (also simply referred to as "lighting settings") from a lighting DB. For example, the lighting DB is a database that stores a lighting setting list of lighting setting data used as lighting information for a video. Note that when the video production system 1A generates lighting settings from the required acquisition information ND3, the video production system 1A does not need to include a lighting DB.

[0273] The required information ND4 is motion acquisition required information used to acquire motion data. The video production system 1A inputs a prompt including instruction information for instructing motion selection from the prompt PT21 to the model M21, and causes the model M21 to output the required information ND4, thereby acquiring the required information ND4.

[0274] For example, the required acquisition information ND4 is motion acquisition required information that includes information (for example, information specifying a motion such as motion #1) for acquiring motion data (also simply referred to as "motion") from a motion DB. For example, the motion DB is a database that stores a motion list of motion data used as information on the motion of a video. Note that when the video production system 1A generates motion data from the required acquisition information ND4, the video production system 1A does not need to include a motion DB.

[0275] The required information ND5 is camerawork required information used to acquire camerawork data. The video production system 1A inputs a prompt including instruction information for selecting camerawork from the prompt PT21 to the model M21, and causes the model M21 to output the required information ND5, thereby acquiring the required information ND5.

[0276] For example, the required acquisition information ND5 is camerawork acquisition required information that includes information (e.g., camerawork specification information such as camerawork setting #1) for acquiring camerawork setting data (also simply referred to as "camerawork setting") from a camerawork DB. For example, the camerawork DB is a database that stores a camerawork setting list of camerawork setting data used as information on the camerawork of the video. Note that when the video production system 1A generates camerawork settings from the required acquisition information ND5, the video production system 1A does not need to include a camerawork DB.

[0277] The required acquisition information ND6 is BGM / SE required acquisition information used to acquire data for at least one of BGM and SE. The video production system 1A inputs a prompt including instruction information for selecting at least one of BGM and SE from the prompt PT21 to the model M21, and causes the model M21 to output the required acquisition information ND6, thereby acquiring the required acquisition information ND6.

[0278] For example, the required acquisition information ND6 is BGM / SE required acquisition information that includes information (e.g., BGM / SE designation information such as BGM / SE#1) for acquiring at least one of BGM and SE data (also simply referred to as "BGM / SE") from a BGM / SEDB. For example, the BGM / SEDB is a database that stores a BGM / SE list of BGM / SE data used as BGM / SE information for video. Note that when the video production system 1A generates BGM / SE from the required acquisition information ND6, the video production system 1A does not need to include a BGM / SEDB.

[0279] The required information ND7 is dialogue narration required information used to acquire dialogue narration data. The video production system 1A inputs a prompt including instruction information for selecting dialogue narration from the prompt PT21 to the model M21, and causes the model M21 to output the required information ND7, thereby acquiring the required information ND7.

[0280] For example, the necessary acquisition information ND7 is necessary dialogue narration acquisition information that includes information (e.g., information specifying a dialogue narration such as dialogue narration #1) for acquiring dialogue narration data (also simply referred to as "dialogue narration") from a dialogue narration DB. For example, the dialogue narration DB is a database that stores a dialogue narration list of dialogue narration data used as dialogue narration information for a video. Note that when the video production system 1A generates a dialogue narration from the necessary acquisition information ND7, the video production system 1A does not need to include a dialogue narration DB.

[0281] The required information ND8 is text-acquisition required information used to acquire text data. The video production system 1A inputs a prompt including instruction information for selecting text from the prompt PT21 to the model M21, and causes the model M21 to output the required information ND8, thereby acquiring the required information ND8.

[0282] For example, the required acquisition information ND8 is text acquisition required information that includes information (e.g., text specification information such as text #1) for acquiring text data (also simply referred to as "text") from a text DB. For example, the text DB is a database that stores a text list of text data used as information on the text of a video. Note that when the video production system 1A generates text from the required acquisition information ND8, the video production system 1A does not need to include a text DB.

[0283] The required information ND9 is logo acquisition required information used to acquire logo data. The video production system 1A acquires the required information ND9 by inputting a prompt including instruction information for selecting a logo from the prompt PT21 to the model M21 and causing the model M21 to output the required information ND9.

[0284] For example, the required information ND9 is logo acquisition required information that includes information for acquiring logo data (also simply referred to as "logo") from a logo DB (for example, information specifying a logo such as logo #1). For example, the logo DB is a database that stores a logo list of logo data used as logo information for a video. Note that when the video production system 1A generates a logo from the required information ND9, the video production system 1A does not need to include a logo DB.

[0285] In the above example, the process of generating each of the necessary information ND1 to ND9 using a prompt corresponding to each piece of necessary information ND1 to ND9 has been described as an example, but the video production system 1A may also generate the necessary information ND1 to ND9 collectively. For example, the video production system 1A may input a prompt PT21 that causes the model M21 to output the necessary information ND1 to ND9, and cause the model M21 to output the necessary information ND1 to ND9, thereby generating the necessary information ND1 to ND9 collectively. In this way, the video production system 1A may use one prompt data to generate multiple queries (necessary information) for acquiring each of multiple types of scene-constituting data.

[0286] Then, the video production system 1A acquires multiple types of scene constituting data that make up the 3D scene data using the necessary information ND1 to ND9, etc. In Fig. 43, the video production system 1A acquires multiple types of scene constituting data such as scene constituting data CD1 to CD7, etc. using the necessary information ND1 to ND9, etc.

[0287] The scene-constituting data CD1 is background information that constitutes the 3D scene data. For example, the scene-constituting data CD1 is scene-constituting data that constitutes the background of a 3D scene. For example, the scene-constituting data CD1 is a file format of the 3D data. For example, the scene-constituting data CD1 can be in any format, such as ply or fbx.

[0288] The video production system 1A acquires scene constituting data CD1 through background acquisition processing PS11 using necessary acquisition information ND1. The video production system 1A acquires the background specified by the necessary acquisition information ND1 from a background DB. In the example shown in FIG. 44 , the video production system 1A executes background acquisition processing PS11 to acquire the background of office #1 from a DB (e.g., a background DB) based on necessary acquisition information ND1, which is background acquisition necessary information. As a result, the video production system 1A acquires scene constituting data CD1, which is data on the background of office #1. FIG. 44 is a diagram showing an example of multiple types of scene constituting data of the present disclosure.

[0289] The scene configuration data CD2 is cast and position information that constitutes the 3D scene data. For example, the scene configuration data CD2 is scene configuration data that constitutes the cast of a 3D scene. The scene configuration data CD2 includes data in a 3D data file format as cast data. For example, the scene configuration data CD2 includes data in any format, such as fbx. The scene configuration data CD2 also includes text format data as cast position data.

[0290] The video production system 1A acquires scene configuration data CD2 through a cast and position acquisition process PS12 using the required acquisition information ND2. The video production system 1A acquires the cast specified by the required acquisition information ND2 from a cast DB. The video production system 1A also generates position information indicating the position of the cast specified by the required acquisition information ND2. In the example shown in FIG. 44, the video production system 1A acquires the cast member "Chris" from a DB (e.g., a cast DB) based on the required acquisition information ND2, which is the cast and position acquisition information, and executes a cast and position acquisition process PS12 to generate the position information "in front of the desk." As a result, the video production system 1A acquires configuration data CD2 including data on the cast member "Chris" used as a cast member and her position "in front of the desk."

[0291] The scene configuration data CD3 is lighting information that constitutes 3D scene data. For example, the scene configuration data CD3 is scene configuration data that constitutes the lighting of a 3D scene. The scene configuration data CD3 can be in any format, such as fbx or text.

[0292] The video production system 1A acquires scene configuration data CD3 through a lighting setting acquisition process PS13 using the required acquisition information ND3. The video production system 1A acquires the lighting setting specified by the required acquisition information ND3 from a lighting setting list (e.g., a lighting DB). In the example shown in FIG. 44 , the video production system 1A executes a lighting setting acquisition process PS13 to acquire lighting setting #1 from the lighting setting list based on the required acquisition information ND3, which is the lighting acquisition required information. As a result, the video production system 1A acquires scene configuration data CD3, which is data for lighting setting #1.

[0293] The scene-constituting data CD4 is motion information that constitutes 3D scene data. For example, the scene-constituting data CD4 is scene-constituting data that constitutes the motion of a 3D scene. For example, the scene-constituting data CD4 is a file format of 3D data. For example, the scene-constituting data CD4 can be in any format, such as bvh or fbx.

[0294] The video production system 1A acquires scene constituting data CD4 through a motion generation / acquisition process PS14 using the necessary acquisition information ND4. For example, the video production system 1A acquires the motion specified by the necessary acquisition information ND4 from the motion DB. In the example shown in FIG. 44 , the video production system 1A executes a motion generation / acquisition process PS14 to acquire (generate) a motion using text2motion based on the necessary acquisition information ND4, which is the motion acquisition necessary information. As a result, the video production system 1A acquires scene constituting data CD4, which is data for the motion "pumping one's fist in joy."

[0295] The scene configuration data CD5 is camerawork information that constitutes 3D scene data. For example, the scene configuration data CD5 is scene configuration data that constitutes camerawork in a 3D scene. The scene configuration data CD5 can be in any format, such as text or parameters.

[0296] The video production system 1A acquires scene construction data CD5 through a camerawork setting acquisition process PS15 using the required acquisition information ND5. The video production system 1A acquires the camerawork setting specified by the required acquisition information ND5 from a camerawork setting list (e.g., a camerawork DB). In the example shown in FIG. 44, the video production system 1A executes a camerawork setting acquisition process PS15 to acquire camerawork setting #1 from the camerawork setting list based on the required acquisition information ND5, which is the camerawork required acquisition information. As a result, the video production system 1A acquires scene construction data CD5, which is data for the camerawork "close-up, center."

[0297] The scene-constituting data CD6 is sound data that constitutes the video data. For example, the scene-constituting data CD6 is scene-constituting data that constitutes the sound of the video data. Note that the scene-constituting data CD6 may also be used as scene-constituting data that constitutes 3D scene data. The scene-constituting data CD6 includes various pieces of information related to sound, such as sound information, dialogue narration information, etc.

[0298] The video production system 1A acquires sound information from the scene constituting data CD6 through a BGM / SE generation and acquisition process PS16 using the required acquisition information ND6. For example, the video production system 1A acquires the BGM / SE specified by the required acquisition information ND6 from the BGM / SEDB. In the example shown in FIG. 44 , the video production system 1A executes a BGM / SE generation and acquisition process PS16 that acquires (generates) sound using text2sound based on the required acquisition information ND6, which is the BGM / SE acquisition required information. As a result, the video production system 1A acquires sound information from the scene constituting data CD6, including data for the BGM / SE "Music of Victory Won from Rock Bottom."

[0299] Furthermore, the video production system 1A acquires dialogue narration information from the scene construction data CD6 through a TTS setting acquisition process PS17 using the necessary acquisition information ND7. For example, the video production system 1A acquires dialogue narration specified by the necessary acquisition information ND7 from a dialogue narration DB. In the example shown in FIG. 44 , the video production system 1A inputs the necessary acquisition information ND7, which is dialogue narration acquisition information, into a Text To Speech (TTS) acquisition model and executes a TTS setting acquisition process PS17 to cause the TTS acquisition model to generate and output voice. For example, the TTS acquisition model is an AI model that outputs (generates) voice in response to text input. As a result, the video production system 1A acquires dialogue narration information from the scene construction data CD6, which includes the spoken text "Your victory" and the voice tone "A male voice with a calm atmosphere, expressing joy."

[0300] The scene configuration data CD7 is logo text data that constitutes the video data. For example, the scene configuration data CD7 is scene configuration data that constitutes the logo text of the video data. Note that the scene configuration data CD7 may also be used as scene configuration data that constitutes 3D scene data. The scene configuration data CD7 includes various information related to the logo text data, such as text information and logo information.

[0301] The video production system 1A acquires text information from the scene constituting data CD7 through a text setting and acquisition process PS18 using the necessary acquisition information ND8. For example, the video production system 1A acquires text specified by the necessary acquisition information ND8 from a text DB. In the example shown in FIG. 44 , the video production system 1A executes a text setting and acquisition process PS18 to acquire (generate) text based on the necessary acquisition information ND8, which is the text acquisition necessary information. As a result, the video production system 1A acquires scene constituting data CD7 including data such as the display text "You win," the display position "center," and the font "Gothic." For example, the text information from the scene constituting data CD7 may be included in the necessary acquisition information ND8. In this case, the video production system 1A acquires the necessary acquisition information ND8 as text information from the scene constituting data CD7.

[0302] Furthermore, the video production system 1A acquires logo information from the scene constituting data CD7 through a logo setting and acquisition process PS19 using the necessary acquisition information ND9. For example, the video production system 1A acquires the logo specified by the necessary acquisition information ND9 from a logo DB. In the example shown in FIG. 44 , the video production system 1A executes the logo setting and acquisition process PS19 to acquire (generate) a logo based on the necessary acquisition information ND9, which is the logo acquisition necessary information. As a result, the video production system 1A acquires the scene constituting data CD7, which includes data for the file "image.png" and the display position "center." For example, the logo information from the scene constituting data CD7 may be included in the necessary acquisition information ND9. In this case, the video production system 1A acquires the necessary acquisition information ND9 as the logo information from the scene constituting data CD7.

[0303] Then, the video production system 1A generates 3D scene data using scene constituting data CD1 to CD7, etc. In Fig. 43, the video production system 1A generates 3D scene data TD1, which is expressed as "3DScene" in Fig. 43, using scene constituting data CD1 to CD5 out of the scene constituting data CD1 to CD7.

[0304] In the example of Figure 43, the video production system 1A executes a layout generation process PS20 using scene-constituting data CD1, which is background information, and scene-constituting data CD2, which is cast and position information. The layout generation process PS20 determines a layout for the 3D scene data, such as the positions of the background indicated by the scene-constituting data CD1 and the cast indicated by the scene-constituting data CD2. Then, based on the determined background and cast layout, the layout generation process PS20 generates layout data (also referred to as "layout LO1") for the 3D scene including the arranged background and cast. In this way, the video production system 1A generates layout data LO1 in which the background indicated by the scene-constituting data CD1 and the cast indicated by the scene-constituting data CD2 are arranged.

[0305] Then, the video production system 1A generates 3D scene data TD1 using the layout data LO1 and the scene configuration data CD3 to CD5. The video production system 1A generates 3D scene data TD1 using files (3D data, etc.) such as the layout data LO1 and the scene configuration data CD3 to CD5, text, setting values, etc.

[0306] For example, the video creation system 1A causes a program (such as a 3D scene creation program) that creates 3D scene data to read files (such as 3D data) such as layout data LO1 and scene configuration data CD3 to CD5, as well as text and setting values, etc. In this way, the video creation system 1A generates the 3D scene data TD1 by causing the program into which the files, text, setting values, etc. have been read to output the 3D scene data TD1.

[0307] The above-described process is merely an example, and the video production system 1A may generate 3D scene data by any process as long as it is possible to generate 3D scene data using scene constituting data CD1 to CD7, etc. For example, the video production system 1A may generate 3D scene data TD1 using scene constituting data CD1 to CD5 without executing the layout generation process PS20.

[0308] The video production system 1A executes a rendering process PS1 denoted as "Rendering" in Fig. 43 to generate video data MV1 denoted as "PreMovie" in Fig. 43. In the rendering process PS1 in Fig. 43, the video production system 1A generates video data MV1 by rendering using 3D scene data TD1.

[0309] The video production system 1A generates video data MV10, which is represented as "EditingPreMovie" in Figure 43, by adding the sound of scene composition data CD6, which is sound data, and the logo text of scene composition data CD7, which is logo text data, to video data MV1, which is a pre-movie.

[0310] For example, the video production system 1A presents video data MV10, which is a preview movie for editing, to a user and accepts editing from the user, thereby executing a process of updating (editing) the video data MV10. For example, the video production system 1A presents the video data MV10 to the user by displaying the video data MV10 on the client UI display unit 400. As shown in FIG. 45 , the video production system 1A accepts editing of the 3D scene data TD1, scene constitution data CD6, and scene constitution data CD7 from the user. FIG. 45 is a diagram showing an example of the flow of the editing process of the present disclosure.

[0311] The video production system 1A executes a 3D scene data editing process PS21, denoted as "3DScene editing" in FIG. 45, in response to an editing instruction from a user. For example, the video production system 1A reflects the editing process PS21 in the video data MV10 and presents the updated video data MV10 to the user. The video production system 1A also reflects the editing process PS21 in the video data MV1 (3D scene data TD1). As a result, for example, the video production system 1A updates the 3D scene data TD1 in accordance with the editing process PS21.

[0312] Furthermore, the video production system 1A executes an editing process PS22 of sound data (such as scene constituting data CD6), which is indicated as "sound editing" in FIG. 45, in response to an editing instruction from the user. For example, the video production system 1A reflects the editing process PS22 in moving image data MV10 and presents the updated moving image data MV10 to the user. The video production system 1A also reflects the editing process PS22 in sound data (such as scene constituting data CD6). As a result, for example, the video production system 1A updates the scene constituting data CD6 in response to the editing process PS22.

[0313] Furthermore, the video production system 1A executes an editing process PS23 of logo text data (such as scene constitution data CD7), denoted as "logo text editing" in FIG. 45, in response to an editing instruction from the user. For example, the video production system 1A reflects the editing process PS23 in the video data MV10 and presents the updated video data MV10 to the user. The video production system 1A also reflects the editing process PS23 in the logo text data (such as scene constitution data CD7). As a result, for example, the video production system 1A updates the scene constitution data CD7 in response to the editing process PS23.

[0314] The video production system 1A executes a refinement process PS2 denoted as "Refiner" in Fig. 43 to generate video data MV2 denoted as "RefinedMovie" in Fig. 43. The video production system 1A adds the sound of the scene constituting data CD6, which is sound data, and the logo text of the scene constituting data CD7, which is logo text data, to the video data MV2 to generate video data MV3 denoted as "FinalMovie" in Fig. 43.

[0315] As described above, the video creation system 1A acquires multiple types of scene configuration data based on a scenario generated based on a user's input query, and generates 3D scene data related to video creation based on the multiple types of scene configuration data. This allows the video creation system 1A to generate data related to a video in response to a user's input query. The video creation system 1A also acquires video data based on the 3D scene data. This allows the video creation system 1A to generate video data in response to a user's input query. Note that the flow of the video creation process shown in FIG. 43 is merely an example, and the video creation system 1A can employ any processing mode as long as it can generate video data from a user's input query.

[0316] <1-7. Regarding AI Models> Note that the AI ​​models used in the above-described processes are not limited to the examples described in each section, and any internal structure can be adopted as long as desired information can be output in response to input. Any combination of input, output, and internal structure of the AI ​​model can be adopted as long as desired information can be output.

[0317] The input of the AI ​​model may be text, images, audio, 3D data, etc., or a combination thereof. The output of the AI ​​model may be text, images, audio, 3D data, etc. Note that the above-mentioned inputs and outputs are merely examples, and the above-mentioned AI model may have any inputs and outputs.

[0318] Furthermore, the internal structure of the AI ​​model can be any structure depending on the combination of input and output. In other words, the internal structure of the AI ​​model can be any structure as long as it can produce a desired output for the input.

[0319] For example, the AI ​​model may have a structure related to Transformer. For example, the AI ​​model may have a structure related to Transformer and perform processing taking into account context, such as context within data, such as text or time-series data. For example, the AI ​​model may have a self-attention mechanism. For example, the AI ​​model may have any attention mechanism, such as single-head attention or multi-head attention. Note that the AI ​​model does not necessarily have to have an attention mechanism.

[0320] The AI ​​model may have a mechanism for extracting features from an input. For example, the AI ​​model may have an encoder. The AI ​​model may have a mechanism for generating information based on the extracted features. For example, the AI ​​model may have a decoder.

[0321] The AI ​​model may have a structure related to a convolutional neural network (CNN). For example, when processing an image, the AI ​​model may have a structure related to a CNN. For example, the AI ​​model may have at least one of a convolution layer, a pooling layer, a fully connected layer, etc.

[0322] The above-described internal structure is merely an example, and the AI ​​model may have any internal structure. For example, the AI ​​model may have a skip connection. Furthermore, the AI ​​model may have a structure related to a diffusion model.

[0323] Furthermore, the above-described AI model may be generated (trained) by any learning process. The AI ​​model may be a machine learning model trained using any machine learning method. For example, the AI ​​model may be a model generated based on a so-called Foundation Model by fine-tuning the Foundation Model to apply it to a specific task (e.g., scenario data generation, code generation, etc.). For example, an AI model such as the above-described LLM may be a model generated by fine-tuning the Foundation Model to apply it to a specific task.

[0324] The base model here is a model that has been trained to be applicable to various tasks, for example, to be able to perform a wide variety of tasks. For example, the base model is a neural network that has been pre-trained with a large amount of unlabeled data set. Note that the base model may have any structure, such as a Transformer-based architecture. For example, the base model is generated by self-supervised learning using data without correct answer labels. As described above, the base model is fine-tuned so that it can be adapted to a wide range of downstream tasks.

[0325] For example, when applied to a scenario data generation task, the base model is fine-tuned to be adaptable to the scenario data generation task, and an AI model (model M1, etc.) adapted to the scenario data generation task is generated. Also, when applied to a code generation task, the base model is fine-tuned to be adaptable to the code generation task, and an AI model (model M3, etc.) adapted to the code generation task is generated. Also, when applied to a sound data generation task, the base model is fine-tuned to be adaptable to the sound data generation task, and an AI model (model M4, etc.) adapted to the sound data generation task is generated. Also, when applied to a text logo generation task, the base model is fine-tuned to be adaptable to the text logo generation task, and an AI model (model M5, etc.) adapted to the text logo generation task is generated.

[0326] For example, an AI model (such as model M1) applied to a scenario data generation task is trained using training data including a combination of input information corresponding to the AI ​​model and scenario data (also referred to as "correct answer information") that is the correct output when the input information is input. The training data, such as the input information and correct answer information, may be data created by a person, or may be data automatically generated by a computer that generates the training data. For example, the scenario data that is the correct answer information may be data created by a person. Below, model M1 will be briefly described as an example. For example, model M1 is trained to output correct answer information corresponding to each piece of input information when that input information is input. For example, model M1 is trained by adjusting (correcting) parameters (connection coefficients) using a method such as backpropagation (error backpropagation) so as to reduce the error between the output of model M1 when certain input information is input and the correct answer information corresponding to that input information. In addition, other AI models such as an AI model applied to a code generation task (e.g., model M3), an AI model applied to a sound data generation task (e.g., model M4), an AI model applied to a text logo generation task (e.g., model M5), an AI model applied to an evaluation task (e.g., model M10), and an AI model applied to an image improvement processing task (e.g., model M11) may also be trained using a similar learning process.

[0327] The above-described learning process is merely an example, and the above-described AI model may be trained by any learning process depending on the input, output, and internal structure of the AI ​​model. For example, the AI ​​model may be trained using an unsupervised learning method such as a generative adversarial network (GAN). The AI ​​model may also be trained in a distributed state without aggregating data, such as federated learning. In this case, each video generation service device (e.g., server) may generate local models collected by the service, and a server (aggregation server) that aggregates information (e.g., parameters) of the local models generated by each video generation service device (e.g., server) may generate a global model using the information on the local models. In this case, the video generation system 1 may receive the global model generated by the aggregation server from the aggregation server and use the received global model as an AI model for processing.

[0328] In this way, the above-described AI models may be generated (learned) by any computer. That is, the learning process for generating the AI ​​models may be performed by any device (computer, etc.) in the video production system 1, or may be performed by a device outside the video production system 1. For example, when a device outside the video production system 1 generates at least one of the above-described AI models, the video production system 1 acquires the AI ​​model from the device outside the video production system 1 and performs processing using the acquired AI model.

[0329] 2. Other Embodiments The processing according to each of the above-described embodiments may be implemented in various different forms (modifications) other than the above-described embodiments and modifications.

[0330] 2-1. Other Configuration Examples The configurations of the image generation systems 1 and 1A described above are merely examples, and any desired division of functions can be adopted in the image generation systems 1 and 1A. In other words, the configurations described above are merely examples, and the image generation systems 1 and 1A may have any desired division of functions and configuration as long as they can provide the services related to image generation described above. For example, the image generation system 1 may be configured by a single device (such as a computer) that performs the above-described processing. In this case, one device of the image generation system 1 may have the functions of the image generation module 100, the information acquisition module 200, the sensor unit 300, and the client UI display unit 400. For example, the image generation service provided by the image generation system 1 may be provided to a user as a program such as a tool (AI Assist Creation Tool) that runs on a terminal device (such as the computer 20) used by the user.

[0331] <2-2. Others> Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0332] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0333] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0334] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0335] 3. Effects of the Present Disclosure As described above, the video generation system (video generation system 1 in the embodiment) according to the present disclosure includes an acquisition unit (in the embodiment, the input text acquisition unit 210), a scenario generation unit (in the embodiment, the scenario-oriented generation unit 131), a code generation unit (in the embodiment, the video-oriented generation unit 132), and a video data acquisition unit (in the embodiment, the video generation unit 140). The acquisition unit acquires an input query related to video generation from a user. The scenario generation unit generates scenario data related to video generation based on the input query. The code generation unit generates code for constructing 3D data based on the scenario data. The video data acquisition unit acquires video data based on the code.

[0336] In this way, the video generation system of the present disclosure generates code for constructing 3D data based on scenario data generated based on an input query from a user, and obtains video data based on the code, thereby generating data related to videos in response to an input query from a user.

[0337] The video generation system also includes an image quality improvement unit (in this embodiment, an image refinement unit 143). The image quality improvement unit executes image quality improvement processing to improve the image quality of the video data. In this way, the video generation system can obtain high-quality video by improving the image quality of the video data.

[0338] The image quality improvement unit improves the image quality of the video data by performing an image quality improvement process based on the text prompt. In this way, the video generation system can obtain high-quality video by improving the image quality of the video data based on the text prompt.

[0339] The image quality improvement unit also performs the image quality improvement process based on a text prompt that specifies the target of the image quality improvement process in the video data. In this way, the video production system can obtain high-quality video by improving the image quality of the video data for the specified target.

[0340] The image quality improvement unit executes the image quality improvement process based on a text prompt specifying the target determined to require improvement. In this way, the video production system can obtain high-quality video by improving the image quality of the video data for the target determined to require improvement.

[0341] The video generation system also includes a display control unit (client UI display unit 400 in this embodiment). The display control unit displays a storyboard based on the scenario data. In this way, the video generation system can present information in a manner that is highly convenient for the user by displaying a storyboard based on the scenario data.

[0342] The storyboard is configured to display video data for each cut of the video. In this way, the video production system is able to present information in a manner that is highly convenient for the user by configuring the storyboard to display video data for each cut of the video.

[0343] The video production system also includes a sound production unit (sound production unit 150 in this embodiment). The sound production unit generates sound data corresponding to the video data based on the scenario data and the video data. In this way, the video production system can obtain a video including sound by generating sound data corresponding to the video data.

[0344] The video production system also includes a text production unit (in this embodiment, a text / logo production unit 160). The text production unit generates text data indicating text to be displayed on the video produced by the video data, based on the scenario data. In this way, the video production system can obtain a video containing text by generating text data corresponding to the video data.

[0345] The video production system also includes a logo production unit (text / logo production unit 160 in this embodiment). Based on the scenario data, the logo production unit generates logo data indicating a logo to be displayed on the video displayed by the video data. In this way, the video production system can obtain a video including a logo by generating logo data corresponding to the video data.

[0346] Furthermore, the input query includes at least one of text, image, audio, and 3D data. In this way, the video generation system can generate data related to video in response to the input query from the user, since the input query includes at least one of text, image, audio, and 3D data.

[0347] The video generation system also includes a first output unit (in this embodiment, a scenario-oriented generation unit 131). The first output unit outputs scenario generation information used by the scenario generation unit to generate scenario data based on the input query. The scenario generation unit generates scenario data based on the scenario generation information. In this way, the video generation system can generate data related to a video in response to an input query from a user by generating scenario data based on the scenario generation information generated based on the input query.

[0348] Furthermore, the first output unit generates, as scenario generation information, a first prompt for generating scenario data based on the input query. The scenario generation unit generates the scenario data based on the first prompt. In this way, the video generation system can generate data related to moving images in response to the input query from the user by generating scenario data based on the first prompt.

[0349] Furthermore, the first output unit generates, based on the input query, first input information as scenario generation information to be used as input to a first model for generating scenario data. The scenario generation unit inputs the first input information generated using the input query to the first model and causes the first model to output scenario data, thereby generating scenario data. In this way, the video generation system can generate data related to moving images in response to an input query from a user by generating scenario data using the first model.

[0350] The video generation system also includes a second output unit (video generation unit 132 in this embodiment). The second output unit outputs code generation information used by the code generation unit to generate code for configuring 3D data based on scenario data. The code generation unit generates code based on the code generation information. In this way, the video generation system can generate data related to video in response to a query input from a user by generating code based on the code generation information generated based on the scenario data.

[0351] The second output unit generates, as code generation information, a second prompt for outputting a code for constructing 3D data, based on the scenario data. The scenario generation unit generates the code based on the second prompt. In this way, the video generation system can generate data related to video in response to a query input from a user by generating the code based on the second prompt.

[0352] The second output unit generates, based on the input query, second input information as code generation information to be used as input to a second model for generating code. The scenario generation unit generates code by inputting the second input information generated using the scenario data to the second model and causing the second model to output code. In this way, the video generation system can generate data related to video in response to an input query from a user by generating code using the second model.

[0353] The video production system also includes a reception unit (sensor unit 300 in this embodiment). The reception unit receives video editing operations from a user. The code generation unit generates code based on the editing operations. In this way, the video production system generates code in response to the video editing operations from the user, thereby enabling appropriate acquisition of video corresponding to the user's editing.

[0354] The receiving unit receives an operation specifying a motion or camera movement using a sensor. The code generating unit generates a code corresponding to the motion or camera movement indicated by the operation. In this way, the video generation system generates a code in response to a user's operation specifying a motion or camera movement, thereby enabling the video generation system to appropriately acquire video corresponding to the user's editing.

[0355] The receiving unit also receives an operation to select multiple cuts from the video data. The code generating unit generates code in which portions corresponding to the multiple cuts indicated by the operation have been changed. In this way, the video generation system generates code in response to the user's selection of multiple cuts, thereby enabling appropriate acquisition of video corresponding to the user's editing.

[0356] Furthermore, date information is associated with each cut of the video data. The code generation unit determines the content of the edit indicated by the operation based on the date information of each cut of the video data. In this way, the video generation system can appropriately acquire video corresponding to the user's editing by determining the content of the edit based on the date information of each cut of the video data.

[0357] The receiving unit also receives an operation to specify an object to be changed from among the video data. The code generating unit generates a code in which the 3D data of the object specified by the operation has been changed. In this way, the video generation system can appropriately acquire a video corresponding to the user's editing by generating a code in which the 3D data specified by the user as the object to be changed has been changed.

[0358] The video production system also includes an evaluation unit (evaluation unit 180 in this embodiment). The evaluation unit generates information indicating an evaluation of at least one of the scenario data and the video data. In this way, the video production system can evaluate the generated information by generating information indicating an evaluation of at least one of the scenario data and the video data.

[0359] Furthermore, the code generator generates the code based on the evaluation. In this way, the video generation system generates the code based on the evaluation, thereby enabling appropriate acquisition of information in accordance with the evaluation.

[0360] Furthermore, the scenario generation unit generates scenario data based on the evaluation. In this way, the image generation system generates scenario data based on the evaluation, thereby enabling appropriate acquisition of information in accordance with the evaluation.

[0361] Furthermore, the scenario generation unit generates a code based on the scenario data generated based on the evaluation. In this way, the video generation system generates a code based on the scenario data generated based on the evaluation, thereby enabling appropriate acquisition of information according to the evaluation.

[0362] 4. Hardware Configuration An information processing device (information equipment) having the image generation module 100, information acquisition module 200, client UI display unit 400, etc. according to each of the above-described embodiments is realized by, for example, a computer 1000 configured as shown in FIG. 41 . FIG. 41 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the information processing device. The following description will be given using the image generation module 100 according to the embodiment as an example. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.

[0363] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0364] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .

[0365] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an image generation program according to the present disclosure, which is an example of program data 1450.

[0366] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.

[0367] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0368] For example, when the computer 1000 functions as the image generation module 100 according to the embodiment, the CPU 1100 of the computer 1000 executes an image generation program loaded onto the RAM 1200, thereby realizing the functions of the control unit 1301, etc. The image generation program according to the present disclosure and data in the storage unit 1302 are stored in the HDD 1400. The CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.

[0369] Note that the present technology can also be configured as follows. (1) A video generation system comprising: an acquisition unit that acquires an input query related to video generation from a user; a scenario generation unit that generates scenario data related to video generation based on the input query; a scene construction data acquisition unit that acquires multiple types of scene construction data that constitute 3D scene data based on the scenario data; and a scene generation unit that generates 3D scene data related to video generation based on the multiple types of scene construction data. (2) The video generation system described in (1), further comprising: a query data acquisition unit that acquires query data for acquiring the scene construction data based on the scenario data, wherein the scene construction data acquisition unit acquires scene construction data corresponding to the query data. (3) The video generation system described in (2), in which the query data is data output by a large-scale language model based on the scenario data. (4) The video generation system described in (3), in which the query data is data output by inputting prompt data instructed to output the multiple types of scene construction data based on the scenario data to the large-scale language model. (5) The video production system according to (4), wherein the prompt data includes a plurality of pieces of instruction information instructing to output each of the plurality of types of scene construction data. (6) The video production system according to (5), wherein the prompt data includes the plurality of pieces of instruction information instructing to select each of the plurality of types of scene construction data. (7) The video production system according to any one of (2) to (6), wherein the scene construction data acquisition unit acquires at least one type of scene construction data from a database in which data used as scene construction data is stored. (8) The video production system according to any one of (2) to (7), wherein the scene construction data acquisition unit generates at least one type of scene construction data from the plurality of types of scene construction data using the query data.(9) The video production system according to any one of (2) to (8), wherein the query data is data indicating, in text form, information required for acquiring the multiple types of scene-constituting data. (10) The video production system according to (9), wherein the query data includes necessary acquisition information indicating, in text form, information corresponding to each of the multiple types of scene-constituting data. (11) The video production system according to any one of (1) to (10), further comprising: a video data acquisition unit that acquires video data based on the 3D scene data. (12) The video production system according to any one of (1) to (11), wherein the multiple types of scene-constituting data include at least two of background information, cast information including cast members and their positions, lighting information, motion information, and camerawork information. (13) The video production system according to (12), further comprising: a layout generation unit that generates layout information based on scene-constituting data including at least one of the background information and the cast information, wherein the scene generation unit generates the 3D scene data based on the layout information. (14) The video generation system according to (13), wherein the scene generation unit generates the 3D scene data based on the layout information and scene configuration data other than the scene configuration data used in the layout information, among the multiple types of scene configuration data. (15) The video generation system according to any one of (1) to (14), wherein each of the multiple types of scene configuration data includes at least one of data in a data file format corresponding to the type, text format data, and parameter data. (16) The video generation system according to (15), wherein the data file format is a 3D data file format. (17) The video generation system according to any one of (1) to (16), wherein the input query includes at least one of text, image, audio, and 3D data. (18) The video generation system according to any one of (1) to (17), wherein the scenario generation unit generates the scenario data based on a prompt generated based on the input query.(19) A video generation method comprising: acquiring an input query related to video generation from a user; generating scenario data related to video generation based on the input query; acquiring multiple types of scene configuration data that constitute 3D scene data based on the scenario data; and generating 3D scene data related to video generation based on the multiple types of scene configuration data. (20) A video generation program that causes a computer to execute the following operations: acquiring an input query related to video generation from a user; generating scenario data related to video generation based on the input query; acquiring multiple types of scene configuration data that constitute 3D scene data based on the scenario data; and generating 3D scene data related to video generation based on the multiple types of scene configuration data.

[0370] 1, 1A Video generation system 100, 100A Video generation module 110 Input text analysis unit 120 Sensor analysis unit 130, 130A Prompt etc. generation unit 131, 131A Scenario-oriented generation unit 132 Video-oriented generation unit 133 Sound-oriented generation unit 134 Text / logo-oriented generation unit 135 Video etc. generation unit 136 Preprocessing unit 140, 140A Video generation unit 141 USD generation unit 141A 3D scene generation unit 142 Rendering unit 143 Video refinement unit 150 Sound generation unit 160 Text / logo generation unit 170 Composite editing unit 180 Evaluation unit 190 Client UI module 200 Information acquisition module 210 Input text acquisition unit 220 Sensor acquisition unit 300 Sensor unit 400 Client UI display unit

Claims

1. A video generation system comprising: an acquisition unit that acquires an input query related to video generation from a user; a scenario generation unit that generates scenario data related to video generation based on the input query; a scene configuration data acquisition unit that acquires multiple types of scene configuration data that constitute 3D scene data based on the scenario data; and a scene generation unit that generates 3D scene data related to video generation based on the multiple types of scene configuration data.

2. The video generation system of claim 1, further comprising a query data acquisition unit that acquires query data for acquiring the scene configuration data based on the scenario data, wherein the scene configuration data acquisition unit acquires scene configuration data corresponding to the query data.

3. The video generation system according to claim 2, wherein the query data is data output by a large-scale language model based on the scenario data.

4. The video generation system described in claim 3, wherein the query data is data output by inputting prompt data into the large-scale language model instructing it to output the multiple types of scene configuration data based on the scenario data.

5. The video generation system according to claim 4, wherein the prompt data includes a plurality of pieces of instruction information for instructing the output of each of the plurality of types of scene-constituting data.

6. The video generation system according to claim 5, wherein the prompt data includes the plurality of pieces of instruction information that instruct the selection of each of the plurality of types of scene-constituting data.

7. The video generation system according to claim 2, wherein the scene construction data acquisition unit acquires at least one type of scene construction data from a database in which data used as scene construction data is stored.

8. The video generation system according to claim 2, wherein the scene construction data acquisition unit uses the query data to generate at least one type of scene construction data from among the plurality of types of scene construction data.

9. The video generation system according to claim 2, wherein the query data is data indicating, in text format, information required for acquiring the plurality of types of scene configuration data.

10. The video generation system according to claim 9, wherein the query data includes necessary acquisition information indicating, in text format, information corresponding to each of the plurality of types of scene configuration data.

11. The video generation system according to claim 1, further comprising: a video data acquisition unit that acquires video data based on the 3D scene data.

12. The video production system according to claim 1, wherein the multiple types of scene configuration data include at least two of background information, cast information including cast members and their positions, lighting information, motion information, and camerawork information.

13. The video generation system of claim 12, further comprising a layout generation unit that generates layout information based on scene configuration data including at least one of the background information and the cast information, wherein the scene generation unit generates the 3D scene data based on the layout information.

14. The video generation system according to claim 13, wherein the scene generation unit generates the 3D scene data based on the layout information and on scene configuration data other than the scene configuration data used in the layout information, among the multiple types of scene configuration data.

15. The video generation system according to claim 1, wherein each of the plurality of types of scene constituting data includes at least one of data in a data file format corresponding to the type, data in a text format, and parameter data.

16. The image generation system according to claim 15, wherein the data file format is a file format for 3D data.

17. The video generation system of claim 1, wherein the input query includes at least one of text, image, audio, and 3D data.

18. The video generation system according to claim 1, wherein the scenario generation unit generates the scenario data based on a prompt generated based on the input query.

19. A video generation method comprising: acquiring an input query related to video generation from a user; generating scenario data related to video generation based on the input query; acquiring multiple types of scene configuration data that constitute 3D scene data based on the scenario data; and generating 3D scene data related to video generation based on the multiple types of scene configuration data.

20. A video generation program that causes a computer to execute the following steps: acquiring an input query related to video generation from a user; generating scenario data related to video generation based on the input query; acquiring multiple types of scene configuration data that constitute 3D scene data based on the scenario data; and generating 3D scene data related to video generation based on the multiple types of scene configuration data.

Citation Information

Patent Citations

  • User-customized meta content providing system based on artificial neural network and method therefor

    KR102508765B1