Information processing system, information processing method, and information processing program
The information processing system addresses the challenge of synchronizing video playback with music by using AI models to acquire music data corresponding to video scenes, ensuring precise synchronization.
Patent Information
- Application Number
- PCT/JP2025/015020
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2025-04-17
- Publication Date
- 2025-10-30
AI Technical Summary
Conventional technologies face difficulties in acquiring music data corresponding to a video scene, making it challenging to synchronize video playback with music effectively.
An information processing system that includes a scene data acquisition unit to acquire scene data related to a video scene and a music acquisition unit to acquire music data based on the scene data, utilizing AI models for generating and synchronizing video and music.
Enables the acquisition of music data that accurately corresponds to video scenes, facilitating effective synchronization of video playback with music.
Smart Images

Figure JP2025015020_30102025_PF_FP_ABST
Abstract
Description
Information processing system, information processing method, and information processing program
[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing program.
[0002] There are known techniques for playing video (also called "moving images") in synchronization with the playback of music. For example, there is known a system for generating a scenario to be used for playing video in synchronization with the playback of music. For example, there is provided a technique for generating a scenario in which scenes constituting a video are associated with each segment of music (see, for example, Patent Document 1).
[0003] Japanese Patent Application Laid-Open No. 2015-12322
[0004] However, there is room for improvement in the conventional technology. For example, it is difficult to acquire music data corresponding to a video scene with the conventional technology. Therefore, it is desirable to acquire music data corresponding to a video scene.
[0005] Therefore, the present disclosure proposes an information processing system, an information processing method, and an information processing program that are capable of acquiring music data corresponding to a video scene.
[0006] In order to solve the above problem, an information processing system according to one embodiment of the present disclosure includes a scene data acquisition unit that acquires scene data related to a video scene, and a music acquisition unit that acquires music data based on the scene data.
[0007] 1 is a diagram illustrating an example of an image generation system according to the present disclosure. FIG. 1 is a diagram illustrating an example of a hardware configuration related to the image generation system according to the present disclosure. FIG. 2 is a diagram illustrating an example of the flow of image generation processing according to the present disclosure. FIG. 3 is a diagram illustrating another example of the flow of image generation processing according to the present disclosure. FIG. 4 is a diagram illustrating an example of the flow of evaluation processing according to the present disclosure. FIG. 5 is a diagram illustrating an example of generation processing of scenario generation information. FIG. 6 is a diagram illustrating an example of scenario data generation processing. FIG. 7 is a diagram illustrating an example of code generation information generation processing. FIG. 8 is a diagram illustrating an example of image quality improvement processing. FIG. 9 is a diagram illustrating an example of a user interface. FIG. 10 is a diagram illustrating an example of a user interface. FIG. 11 is a diagram illustrating an example of a user interface. FIG. 12 is a diagram illustrating an example of a user interface. FIG. 13 is a diagram illustrating an example of a user interface. FIG. 14 is a diagram illustrating an example of a user interface. FIG. 15 is a diagram illustrating an example of a sound generation unit. FIG. 16 is a diagram illustrating an example of a sound generation unit. FIG. 17 is a diagram illustrating an example of a music acquisition processing. FIG. 18 is a diagram illustrating an example of a music acquisition information generation processing. FIG. 19 is a diagram illustrating an example of a music acquisition information generation processing. FIG. 20 is a diagram illustrating an example of a music acquisition information generation processing. FIG. 21 is a diagram illustrating an example of scenario data for learning. FIG. 22 is a diagram illustrating an example of music text data for learning. FIG. 23 is a diagram illustrating an example of a music acquisition information generation processing. FIG. 24 is a diagram illustrating an example of a character description. FIG. 1 is a diagram showing an example of a method for calculating the intensity of a video scene. FIG. 2 is a diagram showing an example of a method for comparing the intensity of a video scene. FIG. 3 is a diagram showing an example of a process for generating information for music acquisition. FIG. 4 is a diagram showing an example of a music acquisition process. FIG. 5 is a diagram showing an example of a music acquisition process. FIG. 6 is a diagram showing an example of a user interface. FIG. 7 is a diagram showing an example of a music editing process. FIG. 8 is a diagram showing an example of information for music editing. FIG. 9 is a diagram showing an example of a music editing process. FIG. 10 is a diagram showing an example of a music editing process. FIG. 11 is a diagram showing an example of an analysis result of a music analysis model. FIG. 12 is a diagram showing an example of a music control model. FIG. 13 is a diagram showing an example of a music editing process. A flowchart showing a flow of processing in response to a user operation. FIG. 14 is a hardware configuration diagram showing an example of a computer that realizes the functions of an information processing device.
[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the information processing system, information processing method, and information processing program according to the present application are not limited to these embodiments. In addition, in the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.
[0009] The present disclosure will be described in the following order: 1. Embodiment 1-1. Overview of the configuration of the video generation system of the present disclosure 1-2. Processing by the video generation system of the present disclosure 1-3. User interface 1-4. Processing examples 1-4-1. Sound generation example 1-4-2. Text logo generation example 1-4-3. Re-learning example 1-4-4. USD update example 1-4-5. Music acquisition example 1-5. Processing flow example from the user's perspective 1-6. Regarding the AI model 2. Other embodiments 2-1. Other configuration examples 2-2. Other 3. Effects of the present disclosure 4. Hardware configuration
[0010] <1. Embodiment> <1-1. Overview of the Configuration of the Image Generation System of the Present Disclosure> Fig. 1 is a diagram illustrating an example of an image generation system of the present disclosure. The image generation system 1 includes an image generation module 100, an information acquisition module 200, a sensor unit 300, and a client UI output unit 400. Note that while Fig. 1 illustrates only one of each component, the image generation system 1 may include multiple image generation modules 100, multiple information acquisition modules 200, multiple sensor units 300, and multiple client UI output units 400.
[0011] First, the configuration of the image generation module 100 that performs image generation processing will be described. The image generation module 100 includes an input text analysis unit 110, a sensor analysis unit 120, a prompt generation unit 130, an image generation unit 140, a sound generation unit 150, a text / logo generation unit 160, a composite editing unit 170, an evaluation unit 180, and a client UI module 190.
[0012] The input text analysis unit 110 analyzes input text. For example, the input text analysis unit 110 analyzes text input from the information acquisition module 200. The sensor analysis unit 120 analyzes input sensor information. For example, the sensor analysis unit 120 analyzes sensor information acquired from the information acquisition module 200.
[0013] The prompt etc. generation unit 130 generates various information required to generate video (movie) including prompts etc. to be input into an AI (Artificial Intelligence) model (also simply referred to as a "model"), which is a machine learning model described below. For example, the prompt etc. generation unit 130 generates a prompt using a user input and a pre-saved prompt (template, etc.). Note that a prompt is merely one example of information to be input into an AI model (model input information). The model input information to be input into an AI model is not limited to a prompt, and any form of model input information can be used. Therefore, the "prompt etc. generation unit" may be read as a "model input information etc. generation unit." In FIG. 1 , the prompt etc. generation unit 130 includes a scenario-oriented generation unit 131, a video-oriented generation unit 132, a sound-oriented generation unit 133, and a text / logo-oriented generation unit 134.
[0014] The scenario generation unit 131 generates various information related to the generation of a scenario. The scenario generation unit 131 generates input information to be input to a model that outputs a scenario. For example, the scenario generation unit 131 is a scenario generation unit that generates scenario data related to video generation based on an input query. For example, the scenario generation unit 131 is a first output unit that outputs scenario generation information used by the scenario generation unit to generate scenario data based on an input query.
[0015] The video generation unit 132 generates various information related to the generation of videos. The video generation unit 132 generates input information to be input to a model that outputs code for constructing 3D (three-dimensional) data. For example, the video generation unit 132 is a code generation unit that generates code for constructing 3D data based on scenario data. For example, the video generation unit 132 is a second output unit that outputs code generation information used by the code generation unit to generate code for constructing 3D data based on scenario data.
[0016] The sound generation unit 133 generates various information related to the generation of sound information (audio information). The sound generation unit 133 generates input information to be input to a model that outputs sound. The text / logo generation unit 134 generates various information related to the generation of text and logos. The text / logo generation unit 134 generates input information to be input to a model that outputs at least one of text and logos.
[0017] The video generation unit 140 executes processing related to video generation. The video generation unit 140 is a video acquisition unit that acquires video data based on a code. The video generation unit 140 generates video using various information generated by the prompt generation unit 130. For example, the video generation unit 140 is a video generation unit that generates video data based on a code. Note that the video generation unit 140 may acquire video data in any manner. For example, the video generation unit 140 may acquire video data by transmitting data used to generate the video data to an external service providing device (such as a vendor) that provides a video data generation service, and receiving the video data generated by the service providing device from the service providing device. In FIG. 1 , the video generation unit 140 includes a USD generation unit 141, a rendering unit 142, and a video refinement unit 143.
[0018] The USD generation unit 141 generates various information related to a Universal Scene Description (USD). For example, the USD generation unit 141 generates USD-Python or the like using an AI model such as a Large Language Model (hereinafter also referred to as "LLM"), using a prompt obtained by video prompt generation in the video generation unit 132.
[0019] The rendering unit 142 executes various processes related to rendering, such as rendering the USD generated by the USD generation unit 141.
[0020] The image refinement unit 143 executes various processes for refining the image. The rendering unit 142 improves the quality of the generated image through image refinement processing. For example, the image refinement unit 143 is an image quality improvement unit that executes image quality improvement processing to improve the image quality of video data.
[0021] The sound generation unit 150 executes a process of generating sounds. The sound generation unit 150 generates sound information such as background music (BGM), sound effects (SE), narration, and dialogue using an AI model such as a contrastive learning model, using prompts obtained by the sound-oriented prompt generation in the sound-oriented generation unit 133.
[0022] The text / logo generator 160 executes a process for generating at least one of text and a logo. The text / logo generator 160 generates at least one of text and a logo using the information generated by the text / logo generator 134.
[0023] The image generation module 100, with the above-described configuration, generates a prompt for generating a scenario by combining a pre-stored prompt with a user's input. The image generation module 100 generates a scenario by inputting the generated prompt into an AI model such as an LLM. The image generation module 100 also generates prompts for generating images, sounds, and text / logos from the generated scenario. The image generation module 100 performs image generation, sound generation, and text / logo generation using prompts, scenarios, etc. for generating images, sounds, and text / logos.
[0024] The composite editing unit 170 executes processes related to editing, such as combining (combining) the generated video, sound, and text / logo into one video.
[0025] The evaluation unit 180 executes an evaluation process for evaluating various targets. The evaluation unit 180 evaluates the information generated by the above-described configuration. For example, the evaluation unit 180 generates information indicating an evaluation of at least one of the scenario data and the video data.
[0026] The client UI module 190 executes processing related to output on a client-side user interface (UI). For example, the client UI module 190 generates various information related to output on the client-side UI. In this case, the client UI module 190 executes processing to generate a UI to be displayed on the user side. The client UI module 190 generates various information to be displayed on the client UI output unit 400.
[0027] Furthermore, the information acquisition module 200 acquires various types of information. The information acquisition module 200 includes an input text acquisition unit 210, a sensor acquisition unit 220, etc. The input text acquisition unit 210 acquires text information input via a keyboard 320 or a microphone 330. For example, the input text acquisition unit 210 acquires text information input by a user via the keyboard 320 or the microphone 330. For example, the input text acquisition unit 210 is an acquisition unit that acquires an input query related to video generation from a user.
[0028] The sensor acquisition unit 220 acquires information (also referred to as "sensor information") detected by a sensor such as a camera 340 or a motion capture device. The information acquisition module 200 provides (transmits) the acquired various pieces of information to the image generation module 100. Note that the information acquisition module 200 may be integrated with the image generation module 100.
[0029] The sensor unit 300 has various sensors. The sensor unit 300 senses user input. The sensor unit 300 accepts user operations. For example, the sensor unit 300 is a reception unit that accepts video editing operations from the user. For example, the sensor unit 300 has a mouse 310, a keyboard 320, a microphone 330, a camera 340, an IMU 350 which is an inertial measurement unit, and the like. In this way, the sensor unit 300 includes, in addition to the mouse 310 and the keyboard 320, a user terminal (such as a smartphone) equipped with the microphone 330, the camera 340, and the IMU 350, and sensors such as motion capture, and senses user input.
[0030] The client UI output unit 400 displays various information to be presented to the client (user). The client UI output unit 400 displays the UI generated by the client UI module 190 on a display (display device) of the client. For example, the client UI output unit 400 is a display control unit that displays a storyboard based on scenario data. The storyboard is configured to display video data for each cut of the video.
[0031] The video generation system 1 may have a hardware configuration as shown in Fig. 2. Fig. 2 is a diagram illustrating an example of a hardware configuration of the video generation system of the present disclosure. In Fig. 2, the video generation system 1 has, as its hardware configuration, a cloud-side computer 10, a client-side computer 20, a camera / sensor 30 including various sensors such as a camera, and the like. The video generation system 1 may also include an information providing device (computer) that provides information resources 40 such as learning data and an AI model 50 to the computer 10.
[0032] 2 is merely an example, and any hardware configuration can be adopted for the video generation system 1 as long as it can execute the desired processing. For example, the computer 10 and the computer 20 may be integrated. Furthermore, the information resource 40 and the AI model 50 may be stored inside the computer 10.
[0033] The computer 10 includes a CPU (Central Processing Unit) 11, a GPU (Graphics Processing Unit) 12, a communication device 13, and a memory / storage 14. For example, the computer 10 corresponds to the image generation module 100 and the information acquisition module 200 in FIG. 1 . The computer 10 may be a service providing device (server device) that provides an image generation service. The CPU 11 and the GPU 12 are so-called processors, and execute calculations (arithmetic operations) related to various processes such as image generation.
[0034] The communication device 13 is a communication device having a communication function for transmitting and receiving information to and from the computer 20, an information providing device, etc., and may be, for example, a communication circuit, a NIC (Network Interface Card), etc. The communication device 13 communicates with other devices such as the computer 20 and the information providing device via a predetermined network (such as the Internet). For example, the communication device 13 is connected to the predetermined network via a wired or wireless connection, and transmits and receives information to and from other devices such as the computer 20 and the information providing device.
[0035] The memory / storage 14 is a storage device that stores various types of information. The memory / storage 14 is, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 14 stores various types of information used for processing by processors such as the CPU 11 and the GPU 12. The memory / storage 14 may also store information resources 40, AI models 50, etc.
[0036] The computer 20 includes a CPU 21, a GPU 22, a communication device 23, a memory / storage 24, and an IO interface 25. For example, the computer 20 corresponds to the client UI output unit 400 in FIG. 1 . The computer 20 may also be a terminal device (such as a personal computer (PC) or a mobile device such as a smartphone) used by a user who uses the video generation service. The CPU 21 and the GPU 22 are so-called processors, and perform calculations (arithmetic processing) related to various processes such as video display. Note that the above is merely an example, and the computer 20 can have any configuration as long as it can perform the desired processing. For example, the computer 20 may perform calculations (arithmetic processing) related to various processes such as video display using circuits such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). Furthermore, the computer 20 may be configured so that a program is directly embedded in the processor circuitry instead of storing the program in a memory (such as the memory / storage 24). In this case, the processor realizes its functions by reading and executing the program embedded in the circuitry. In addition, each processor in this embodiment is not limited to being configured as a single circuit, but may be configured as a single processor by combining multiple independent circuits to realize its functions. Furthermore, like the computer 20, the computer 10 can also adopt any configuration as long as it is capable of performing the desired processing.
[0037] The communication device 23 is a communication device having a communication function for transmitting and receiving information to and from the computer 10, the sensor 30, etc., and may be, for example, a communication circuit, a NIC, etc. The communication device 23 communicates with other devices such as the computer 10 and the sensor 30 via a predetermined network (such as the Internet). For example, the communication device 23 is connected to the predetermined network by wire or wirelessly, and transmits and receives information to and from other devices such as the computer 10 and the sensor 30.
[0038] The memory / storage 24 is a storage device that stores various types of information. The memory / storage 24 is, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 24 stores various types of information that are used for processing by processors such as the CPU 21 and the GPU 22.
[0039] The IO interface 25 is an input / output interface device. The computer 20 receives input from the sensor 30 via the IO interface 25. For example, the computer 20 receives input from an input device such as a keyboard or a mouse via the IO interface 25. The computer 20 also outputs information from a display (display device) and a speaker (audio output device) via the IO interface 25. For example, the computer 20 plays video on the display and speaker via the IO interface 25.
[0040] Various sensors 30, such as cameras, sense user input. Various sensors 30, such as cameras, accept user operations. For example, the sensor 30 corresponds to the sensor unit 300 in FIG. 1. Furthermore, the information resource 40 includes various information such as training data. For example, the information resource 40 includes training data used to train various AI models, such as LLM. The AI model 50 includes information on AI models used in processing related to video generation, such as LLM. For example, the AI model 50 includes information on various AI models, such as models M1 to M3, which will be described later. As described above, the video generation system 1 may have a configuration other than that shown in FIG. 2.
[0041] <1-2. Processing by the video generation system of the present disclosure> Processing by the video generation system will now be described. First, an example of the flow of the video generation process shown in FIG. 3 will be described. FIG. 3 is a diagram showing an example of the flow of the video generation process of the present disclosure. Note that the processes described below with the video generation system 1 as the processing subject may be performed by any device capable of executing the processes, depending on the device configuration included in the video generation system 1.
[0042] User input information UIN1, denoted as "User's Input" in FIG. 3, corresponds to information input by a user for generating an image (also referred to as an "input query"). Note that the input query is not limited to text (character information) and any information can be used. The input query may be any information including at least one of text, image, audio, and 3D data.
[0043] The video generation system 1 uses user input information UIN1 to generate scenario generation information (also referred to as "first input information") to be used as input for model M1, denoted as "LLM" in FIG. 3, which will be described later. For example, model M1 is a first model that outputs scenario data in response to the input of the first input information. Any AI model, such as an LLM (large-scale language model), can be used for model M1 as long as it is capable of producing the desired output in response to the input. AI models such as model M1 will be described later.
[0044] The video generation system 1 generates a scenario FD1 by inputting first input information to a model M1 and causing the model M1 to output a scenario FD1, which is scenario data. The video generation system 1 then generates USD generation necessary information SD1, which is code generation information (also referred to as "second input information") to be used as input for a model M3, using the scenario FD1, the output of a model M2 that uses the scenario FD1 as input, and user input information UIN2, etc. For example, the model M3 is a second model that outputs code in response to the input of the second input information. Any AI model, such as an LLM (large-scale language model), can be used for the model M3, as long as it is capable of producing a desired output in response to the input.
[0045] 3 illustrates only one piece of information SD1 required for generating USD, but there may be a plurality of pieces of information SD1 required for generating USD depending on the number of USD files to be generated. For example, there may be a plurality of pieces of information SD1 required for generating USD depending on the number of USD files to be generated corresponding to the data structure shown in FIG.
[0046] For example, model M2 may be a model that outputs a template or the like corresponding to scenario data in response to input of the scenario data. Any AI model can be adopted as model M2 as long as it is capable of producing the desired output in response to the input. For example, user input information UIN2 may be information for specifying constraints for image generation. Note that image generation system 1 may generate second input information using scenario FD1 and template input information, but this point will be described later.
[0047] The video production system 1 generates the Python code OD1 by inputting the USD generation necessary information SD1 into the model M3 and causing the model M3 to output the Python code OD1, denoted as "python" in FIG. 3 . For example, the Python code OD1 is (program) code that generates USD format data (also referred to as a "USD file") when executed. Note that Python is merely an example, and any code format, not limited to Python, can be used as long as it can generate the desired 3DCG data. Also, USD is merely an example, and any format, such as FBX (Film Box), can be used for 3DCG data. The video production system 1 executes the Python code OD1 to generate the USD file OD2, denoted as "USD" in FIG. 3 .
[0048] The video generation system 1 generates video data MV1, which is denoted as "PreMovie" in Fig. 3, by executing a rendering process PS1, which is denoted as "Renderer" in Fig. 3. For example, the video data MV1 is data (also referred to as "first video data") before executing a refinement process PS2, which will be described later.
[0049] The video generation system 1 generates video data MV2, denoted as "RefinedMovie" in Fig. 3, by executing a refinement process PS2, denoted as "Refiner" in Fig. 3. For example, the refinement process PS2 is a picture quality improvement process that improves the picture quality of the video data. The video data MV2 is data (also referred to as "second video data") obtained after the refinement process PS2 updates the first video data, i.e., video data MV1.
[0050] The video production system 1 executes composite editing PS3 using the video data MV2, user input information UIN3, etc., to generate video data MV3, which is denoted as "FinalMovie" in Fig. 3. For example, the composite editing PS3 executes a process of updating (editing) the video data MV2 in accordance with a user editing instruction indicated by the user input information UIN3, thereby generating video data MV3 in which the video data MV2 has been updated.
[0051] The flow of the video generation process shown in FIG. 3 is merely an example, and the video generation system 1 can employ any processing mode as long as it can generate video data from a user's input query. For example, while FIG. 3 illustrates an example in which the model M1 outputs code (Python code), the model M1 may also output 3DCG data such as a USD file. Furthermore, the video generation system 1 may perform various modes of video generation processing, not limited to the processing illustrated in FIG. 3 . An example of this point will be described using FIG. 4 . FIG. 4 is a diagram illustrating another example of the flow of the video generation process of the present disclosure. FIG. 4 differs from FIG. 3 in that sound generation required information SD2 and text logo required information SD3 are generated and used to perform video generation processing. Note that explanations of points similar to those described in FIG. 3 will be omitted where appropriate.
[0052] In FIG. 4, the video generation system 1 uses a scenario FD1, the output of a model M2 that uses the scenario FD1 as input, and user input information UIN2 to generate sound generation information SD2, which is information for generating sound to be used as input for a model M4, denoted as "AI" in FIG. 3. For example, the model M4 is a model that outputs various sound data in response to input of the sound generation information SD2 and video data MV2. Any AI model can be used for the model M4 as long as it is capable of producing the desired output in response to the input. Note that the model M4 may also be a model that receives only the sound generation information SD2 as input.
[0053] The video production system 1 generates sound data corresponding to a video by inputting sound generation necessary information SD2 to the model M4 and causing the model M4 to output sound data AD1 for background music (BGM), sound effects AD2, and narration sound data AD3. The video production system 1 may also generate the sound data AD1, AD2, and AD3 using user input information UIN4. For example, if the model M4 outputs the sound data AD1, AD2, and AD3 as a single piece of sound data, the video production system 1 may extract the sound data AD1, AD2, and AD3 from the single piece of sound data output by the model M4 based on the specifications in the user input information UIN4, and generate the sound data AD1, AD2, and AD3.
[0054] 4 , video generation system 1 uses scenario FD1, the output of model M2 that uses scenario FD1 as input, and user input information UIN2 to generate text logo generation information SD3, which is used as input for model M5. For example, model M5 is a model that outputs at least one of text and a logo in response to the input of text logo generation information SD3. Any AI model can be used for model M5 as long as it can produce the desired output in response to the input.
[0055] Video generation system 1 inputs information SD3 required for text logo generation into model M5 and causes model M5 to output text logo data DI1 for Text, text logo data DI2 for Logo, etc., thereby generating text logo data corresponding to the video.
[0056] Video production system 1 generates video data MV3 by executing composite editing PS3 using video data MV2, sound data AD1, AD2, AD3, text logo data DI1, DI2, user input information UIN3, etc. For example, composite editing PS3 combines (combines) video data MV2, sound data AD1, AD2, AD3, text logo data DI1, DI2, etc. into one image to generate video data MV3 as a single image.
[0057] 3 and 4 illustrate an example of processing in an initial state where there is no scenario, information required for USD generation, USD, PreMovie, RefinedMovie, or the like. As described above, the video production system 1 generates prompts for generating a scenario based on a user's input query, provides the prompts to a natural language model to generate a scenario, generates a prompt for outputting code constituting a video from text information described in the scenario, provides the prompts to the natural language model to generate code constituting the video, and generates a video. In this way, when creating a video, the video production system 1 generates an effective video storyboard and video by inputting what the user wants to create and their purpose, even without knowledge of 3DCG or video production. Furthermore, creating a storyboard in the video production system 1 makes subsequent editing easier. These points will be described in detail later.
[0058] Furthermore, the image generation system 1 may perform various processes related to image generation. For example, the image generation system 1 may perform evaluation processing on the generated information. In this regard, an example of the flow of the evaluation processing will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the flow of the evaluation processing of the present disclosure.
[0059] 5 , the video generation system 1 receives input of at least one of a scenario FD1, information required for USD generation SD1, information required for sound generation SD2, and information required for text logo generation SD3, and causes the model M10 to output evaluation text information EV1 indicating an evaluation of the input information, thereby evaluating the generated information. For example, the model M10 outputs an evaluation of input information in response to input of the information. For example, the model M10 outputs evaluation text indicating an evaluation of the input scenario FD1 in response to input of the scenario FD1. Note that the model M10 may be a model that accepts input of the scenario FD1, information required for USD generation SD1, information required for sound generation SD2, and information required for text logo generation SD3 separately, or a model that accepts input of a combination of these pieces of information. Furthermore, the model M10 may be a model that accepts input of information indicating a video (e.g., captions) in response to input of the video, and outputs an evaluation of the video corresponding to the input information.
[0060] From here, a specific example of each process in the above-mentioned process flow executed by the image generation system 1 will be described. Note that explanations of points similar to those described above will be omitted as appropriate.
[0061] For example, the video production system 1 generates scenario generation information (first input information) as shown in Fig. 6. Fig. 6 is a diagram showing an example of a process for generating scenario generation information. In Fig. 6, the video production system 1 acquires user input information IDT1 and IDT2 entered by the user into content CT1 as user input information. Content CT1 is content for receiving user input information for each of the questions "What kind of video do you want to create?" and "Style."
[0062] For example, the client UI output unit 400 displays the content CT1, and the sensor unit 300 receives the user input information IDT1 and IDT2 as user input information. For example, the user input information IDT1 and IDT2 correspond to the user input information UIN1 in FIGS.
[0063] 6, the client UI output unit 400 displays a question, "What kind of video do you want to make?". In response to the question, "What kind of video do you want to make?", the sensor unit 300 accepts user input information IDT1, "a 15-second sneaker commercial video." The client UI output unit 400 also displays a question, "style." In response to the question, "style," the sensor unit 300 accepts user input information IDT2, "cinematic."
[0064] The video production system 1 may accept user input information in any manner, or may accept a user selection from multiple options. For example, the video production system 1 may convert information entered by the user via a keyboard or microphone into text information and accept it as user input information. Furthermore, the video production system 1 may accept, in addition to free text, settings such as the number of seconds for the entire video, style, and camerawork, as well as other files such as images and videos, as user input information.
[0065] The video production system 1 generates a prompt PT1, which is scenario generation information (first input information), using the user input information IDT1 and IDT2 and a template TP1, which is template input information. For example, the template TP1 may be preset or may be selected from a plurality of template candidates. For example, the video production system 1 may select a template corresponding to the user's input information from the plurality of template candidates. For example, the video production system 1 may select a template TP1 related to a movie-style advertisement from the plurality of template candidates based on the content indicated by the user input information IDT1 and IDT2.
[0066] For example, the video production system 1 generates a prompt PT1 by reflecting user input information IDT1 and IDT2 in a template TP1. In Fig. 6, the video production system 1 generates the prompt PT1 by adding "cinematic" indicated by the input information IDT2 to the style item of the constraints and adding "a 15-second sneaker commercial video" indicated by the input information IDT1 to the input sentence. In this way, the video production system 1 generates a prompt for generating a scenario based on the information input by the user. Note that the user's input information may be input on a single screen or by answering several questions; examples of these points will be described later.
[0067] The video production system 1 also generates scenario data as shown in Fig. 7. Fig. 7 is a diagram showing an example of a scenario data generation process. In Fig. 7, the video production system 1 generates scenario data SN1 using a prompt PT1. For example, the scenario data SN1 corresponds to the scenario FD1 in Figs. 3 and 4. The scenario data SN1 includes information such as the number of seconds for each scene, an explanation of the cut, etc., for each scene, such as the opening scene and the scene where sneakers are put on.
[0068] For example, the video production system 1 generates scenario data SN1 by inputting a prompt PT1 into a model M1, such as an LLM, and outputting scenario data SN1 from the model M1. In this manner, the video production system 1 generates a scenario by inputting the generated prompt into an AI (such as an LLM). In addition to the information shown in FIG. 7 (also referred to as "scenario information"), the scenario data SN1 also includes information such as the environment, characters, motion, camerawork, lighting, and color. For example, to create unique scenarios, the video production system 1 can generate a variety of scenario variations by inputting a user's past experiential learning data or learning data from a specific director or person into the model M1 using Retrieval-Augmented Generation (RAG) or fine-tuning.
[0069] Furthermore, the video production system 1 generates code generation information (second input information) as shown in FIG. 8 . FIG. 8 is a diagram showing an example of a process for generating code generation information. The video production system 1 generates a prompt PT2, which is code generation information (second input information), using scenario data SN1 and a template TP2, which is template input information. For example, the template TP2 may be preset or may be selected from multiple template candidates. For example, the video production system 1 may select a template corresponding to a scenario from the multiple template candidates. For example, the video production system 1 may select a template TP2 related to a commercial from the multiple template candidates based on the content indicated by the scenario data SN1.
[0070] For example, the video production system 1 generates a prompt PT2 by reflecting scenario data SN1 in a template TP2. In Fig. 8, the video production system 1 generates the prompt PT2 by adding information indicated by the scenario data SN1 to an input sentence. In this way, the video production system 1 generates a prompt for a video based on a scenario generated by AI.
[0071] For example, FIG. 8 shows an example of prompt generation for converting a scenario related to a person into USD-Python. Conversion to USD-Python is merely one example of a conversion format, and the conversion is not limited to USD-Python and may be any conversion format. For example, the conversion format may be Python for Blender, USD, or other formats. Furthermore, the video production system 1 may generate prompts individually for each subject, such as a person, environment, or camerawork, or may generate prompts collectively. The generated prompt may include paths to assets and motions to be used, or may include source code or an API (Application Programming Interface) to be passed (input) to the AI generation algorithm for assets and motions.
[0072] The video production system 1 then generates a USD-Python file by inputting the generated prompt into an AI (such as an LLM). Note that the file format is not limited to Python format, and the file may be generated in another format such as USD. The video production system 1 then converts the file into a format that can be rendered, such as a USD file, and performs rendering to generate a PreMovie (a video file such as MP4).
[0073] The video production system 1 may perform composite editing using a PreMovie, but may also perform refinement processing, which is an example of image quality improvement processing, on the PreMovie, as shown in Fig. 9. Fig. 9 is a diagram showing an example of image quality improvement processing. In Fig. 9, the video production system 1 generates the second video OT1 from the first video IN1 by refinement processing using a model M11, which is a diffusion model that receives a first video IN1, which is a PreMovie, as input and outputs a second video OT1, which is a RefinedMovie.
[0074] The AI model (such as model M11) used in the refiner process is not limited to the diffusion model; any AI model such as the latent diffusion model (LDM) or the latent consistency model (LCM) can be used. Furthermore, the refiner process may use techniques such as AnimateDiff (time direction stabilization) or ControlNet (line art control). Through this refiner process, the video production system 1 can improve the quality of the video while maintaining the consistency of the characters, backgrounds, props, and the like.
[0075] Furthermore, the refiner process may use prompts in addition to videos. For example, the model M11 may input prompt IN2 in addition to the first video IN1. For example, when a target such as a woman in her 30s is specified by prompt IN2, the model M11 outputs a second video OT1 in which the portion of the first video IN1 that represents the woman in her 30s has been improved. This allows the video generation system 1 to generate a second video OT1 in which the image quality, etc., of the target specified by prompt IN2 in the first video IN1 has been improved.
[0076] Through the processing of the video production system 1 described above, it appears to the user that a scenario (storyboard) and animation for each cut are being generated after input, with the processing in between being confined within the system. These processes may involve generating the scenario, all cuts, and refinement processing all at once from the user's input text, or the user may input preferences during the processing. For example, the video production system 1 may generate several scenarios with outlines only, then allow the user to select one, and then execute detailed scenario and animation generation processing based on the selected outline scenario. Furthermore, the video production system 1 may generate several patterns of characters to be generated in the video before video generation after scenario generation, and after the user selects one, execute animation rendering and refinement processing.
[0077] The components of a scenario may include video of the cut, a representative image (such as the first frame of the video), a description of the cut, characters (visuals, setting, etc.), the motion of each character, lighting, camera work, background environmental information, transitions between cuts, dialogue, narration, etc. Some of these are presented to the user, while others are kept for processing purposes without being presented to the user. The scenario is arranged in chronological order by cut.
[0078] Currently, various video generation services are available, including Pika, Runway Gen-2, Lumiere, and Stable Video Diffusion. These generate video using a diffusion model that moves vectors in the spatial and temporal directions from images. These generate video using only 2D images. On the other hand, video generation system 1 stores 3D information internally. For example, existing video generation services allow you to modify only a specified (X, Y) area within a video, but there is an issue that changing only the color of clothing also changes the motion. On the other hand, video generation system 1 stores 3D information internally, making it possible to modify only targeted areas, such as only the motion, only the lighting, or only the color of a person's clothing.
[0079] Furthermore, existing video generation services only generate videos for each cut, and users must ensure the consistency of each cut themselves, but video generation system 1 can consistently carry out everything from the scenario (storyboard) to video generation and editing, making it possible to generate videos with consistency in terms of the actors, backgrounds, color grading, etc.
[0080] <1-3. User Interface> Hereinafter, we will describe the user interface (UI) for users who use the image generation system 1. Note that explanations of points similar to those described above will be omitted where appropriate.
[0081] As shown in FIG. 10 , the video generation system 1 provides the user with content CT11. FIG. 10 is a diagram illustrating an example of a user interface. The content CT11 is a display screen (content) for accepting user input information. For example, the client UI output unit 400 displays the content CT11. The user inputs text instructing the type of video to be created (the “Prompt” column in FIG. 10 ) and a style selection (the “Style” column in FIG. 10 ) as user input information via the content CT11. In this manner, the user inputs the text and style selection as user input information. For example, the user inputs text information in the “Prompt” column by referring to example sentences, etc., included in the content CT11. For example, the user selects a style to use from multiple style candidates displayed by, for example, pressing (clicking, etc.) the downward-pointing triangle in the “Style” column.
[0082] After completing the input of the user's input information, the user selects the button labeled "Ask AI Director" in Fig. 10 to instruct the video production system 1 to generate a video in accordance with the user's input information. As a result, the video production system 1 executes the process of generating a video in accordance with the user's input information.
[0083] As shown in FIG. 14 , the video generation system 1 provides the user with content CT15 related to the generated video. FIG. 14 is a diagram showing an example of a user interface. The content CT15 is a storyboard screen (content) for receiving user operations (instructions) on the generated video. As shown in FIG. 14 , the content CT15 is a storyboard screen that displays video data for each cut of the generated video. For example, the client UI output unit 400 displays the content CT15. In this way, the video generation system 1 provides a UI that outputs a storyboard and video in accordance with information input by the user. The user sets the video, content, narration, dialogue, camerawork, background music, lighting, color, and the like for each cut on the storyboard screen.
[0084] The video generation system 1 may accept user input information while asking the user a question. For example, when the button labeled "Ask AI Director" in FIG. 10 is selected, the video generation system 1 accepts user input information through a conversation (dialogue) with the user, as shown in FIGS. 11 to 13. FIGS. 11 to 13 are diagrams showing an example of a user interface. Content CT12 in FIG. 11 is a display screen (content) that presents samples generated in response to the user's input information entered in FIG. 10 and asks the user whether there is one that is close to their image. For example, the client UI output unit 400 displays the content CT12.
[0085] Content CT13 in Fig. 12 is a display screen (content) that requests (questions) for detailed targets, etc. in response to a user's response (input information) that none of the samples presented in content CT12 in Fig. 11 match the image. For example, the client UI output unit 400 displays content CT13.
[0086] Content CT14 in Figure 13 is a display screen (content) that presents samples regenerated in response to a user's response (input information) that specifically specifies a target, etc., and asks the user whether any of them are close to the image they have in mind. For example, the client UI output unit 400 displays content CT14. Figure 13 shows a case in which the user positions the mouse cursor over the leftmost sample video of the four samples and performs a designation operation such as clicking, thereby designating that the leftmost sample video of the four samples is close to the image they have in mind. As a result, the video generation system 1 executes a video generation process in accordance with the user's input information that designates the leftmost sample video of the four samples.
[0087] In this case, the video generation system 1 provides the user with content CT15 related to the generated video, as shown in Fig. 14. For example, the client UI output unit 400 displays the content CT15. In this way, when generating a storyboard and video using the initial input content, if the video generation system 1 does not have enough necessary information, it may collect the necessary information through conversation (dialogue) with the user and work out the details.
[0088] 15 and 16, the video production system 1 may prompt the user to input a short sentence about the video they want to create, generate several video stories based on the short sentence, and allow the user to select the one they like best. Figures 15 and 16 are diagrams showing examples of a user interface.
[0089] As shown in FIG. 15 , the video generation system 1 provides content CT21 to the user. The content CT21 is a display screen (content) for receiving information input by the user. For example, the client UI output unit 400 displays the content CT21. The user inputs a sentence (short sentence) indicating what kind of video they want to create as user input information via the content CT21. For example, the user inputs text information into an input field in the content CT21.
[0090] After completing the input of the user's input information, the user selects the button labeled "Start" on the right end of the input field in content CT21 to instruct the video production system 1 to generate a video story in accordance with the user's input information. This causes the video production system 1 to execute the process of generating a video story in accordance with the user's input information.
[0091] As shown in Fig. 16, the video generation system 1 provides the user with content CT22 relating to the story of the generated video. The content CT22 in Fig. 16 is a display screen (content) that presents a sample of the story of the video generated in response to the user's input information entered in Fig. 15. For example, the client UI output unit 400 displays the content CT22. For example, if the user likes one of the video story samples, the user selects the button labeled "Continue" on the right edge of the display area for that sample, thereby instructing the video generation system 1 to generate a video corresponding to the selected sample.
[0092] If the user does not like any of the video story samples, the video production system 1 generates another pattern. For example, if the user does not like any of the video story samples, the user can select the button labeled "Continue" on the right edge of the short sentence display area to instruct the video production system 1 to generate another pattern of video story sample. This causes the video production system 1 to generate another pattern of video story sample.
[0093] <1-4. Processing Examples> From here, in addition to the specific examples described above, specific examples of each process executed by the video production system 1 will be described. Note that explanations of points similar to those described above will be omitted as appropriate. Below, specific examples of the generation process of sound, etc., and evaluation process in the processing of the video production system 1 described above will be described. Note that explanations of points similar to those described above will be omitted as appropriate.
[0094] <1-4-1. Example of Sound Generation> For example, the video production system 1 generates sound generation information (corresponding to the sound generation required information SD2 in FIGS. 3 and 4) as shown in FIG. 17. FIG. 17 is a diagram showing an example of a process for generating sound generation information. The video production system 1 generates a prompt PT3, which is sound generation information, using scenario data SN1 and a template TP3, which is template input information. For example, the template TP3 may be preset or may be selected from multiple template candidates. For example, the video production system 1 may select a template corresponding to a scenario from the multiple template candidates. For example, the video production system 1 may select a template TP3 related to a commercial from the multiple template candidates based on the content indicated by the scenario data SN1.
[0095] For example, the video production system 1 generates a prompt PT3 by reflecting the scenario data SN1 in a template TP3. In Fig. 17, the video production system 1 generates the prompt PT3 by adding information indicated by the scenario data SN1 to an input sentence. For example, the video production system 1 generates a prompt for extracting several key phrases from a scenario in order to generate a sound (such as background music) that better suits the scenario. The video production system 1 may input the generated prompt into an AI (such as an LLM) to obtain keywords and text information necessary for sound generation.
[0096] The video production system 1 generates sound data using the generated prompt PT3. For example, the video production system 1 generates sounds (sound data) such as background music, sound effects, narration, and dialogue based on the generated video and text information described in the storyboard (scenario data SN1, etc.). Furthermore, if necessary words (text information) such as dialogue or narration have already been extracted from the scenario, the video production system 1 may save the words (text information) as sound data for the dialogue, narration, etc., without generating a prompt.
[0097] For example, the video generation system 1 generates sound data from the obtained information required for sound generation (sound generation information), video, and audio data input by the user (sound source, user's voice, humming, etc.). For background music and sound effects, the video generation system 1 may generate sound from text or video using a transformer (model) such as text-to-music generation. The video generation system 1 may also search for contrastively learned sound sources, such as text-to-music estimation, from natural text.
[0098] Furthermore, for dialogue and narration, the video production system 1 may generate audio using text-to-speech (such as a diffusion model or flow matching) based on the words themselves and text information about the characters obtained from the scenario. Furthermore, when connecting the generated video and audio, the video production system 1 may incorporate meta information into the video or audio file to indicate the start and end times, volume, etc. of the audio.
[0099] <1-4-2. Example of Text Logo Generation> For example, video production system 1 generates information for generating a text logo (corresponding to information required for text logo generation SD3 in Figures 3 and 4). Based on the generated storyboard (scenario data SN1, etc.), video production system 1 generates text and logo information (captions, titles, logos, descriptions, etc.) to be displayed on the video.
[0100] For example, the generated scenario may clearly state the text to be displayed, but if it does not, the video generation system 1 generates a prompt to generate text information (text logo data, etc.) and sends it to an AI (such as an LLM) to generate the text information. Also, if the user inputs the text to be displayed, the video generation system 1 may use the information input by the user as the text information (text logo data, etc.).
[0101] The font, size, and position of the text display may be determined by any method. For example, the video generation system 1 may determine the font, size, and position of the text display using any AI such as a diffusion model, a variational auto-encoder (VAE), generative adversarial networks (GAN), DALL E, StyleGAN, StyleGAN2, Pix2Pix, TransGAN, or LLM. The font, size, and position of the text display may also be manually set by the user.
[0102] For logos and images, the user may input images or videos in JPEG format, MP4 format, etc. For logos and images, the video generation system 1 may generate a prompt for image generation and send it to any AI such as a Diffusion model, VAE, GAN, DALL E, StyleGAN, StyleGAN2, Pix2Pix, TransGAN, or LLM to generate logo information.
[0103] In addition, when linking text or logo information with a video, the video generation system 1 may incorporate meta information into the video or the text or logo itself to clearly indicate (meta information) the start and end times, position, and size of the text or logo.
[0104] <1-4-3. Re-learning Example> Furthermore, when generating a scenario or video, the video generation system 1 may generate a scenario or video that can only be produced using specific, replaceable learning data. For example, it is possible to re-learn using previously produced videos and images as learning data. Re-learning using RAG or fine tuning can change the scenario or video to be generated. In other words, the video generation system 1 can re-learn data (history) of personal user's past productions, or can generate videos using a model re-learned using the work of a specific film director as learning data.
[0105] Furthermore, the learning data can be retrained on an individual's PC or on a server. The learning data is used to train a number of models, such as the LLM and Diffusion model. In cases such as when an overall scenario is to be generated using Director A, but the color of the video is to be generated using a different Director B to generate color grading, the video generation system 1 may retrain specific parts using different learning data.
[0106] <1-4-4. USD Update Example> As described above, the video generation system 1 modifies the 3DCG assets and rendering method that are the basis of the existing video, based on the user's input information and the text information of the generated storyboard. The video generation system 1 then outputs a prompt for outputting code that will compose a new video, provides the prompt to a natural language model, outputs the code that will compose the video, and generates the video. This allows the video generation system 1 to modify the video in accordance with the user's input, even if the user has no knowledge of 3DCG or video production.
[0107] In the video generation system 1, after generating a video (video), scenario information, video information, etc. are saved as text or video, so the user can modify this information using input information (text, sensors, etc.).
[0108] For example, in the video production system 1, USD files are stored separately for each asset and motion, as shown in Fig. 18. Fig. 18 is a diagram showing an example of a USD file. For example, in the data structure shown in Fig. 18, a USD in a higher layer may include a path (file path) to a USD in a lower layer.
[0109] For example, the entire USD includes paths to the environment assets USD, person assets UDS, camera USD, etc. The environment assets USD also include paths to the building assets USD and prop assets USD. The building assets USD also include mesh information of the building itself, etc. The person assets USD also include paths to mesh information of a person and motion USD. In this way, data for 3DCG (3D data) such as a USD file may include multiple data sets. Note that the configuration (data structure) of the USD file shown in FIG. 18 is merely an example, and any configuration can be adopted, and the entire USD (USD file) may be configured as a single block (one data set).
[0110] <1-4-5. Example of Music Acquisition> The video production system 1 acquires music data based on scene data related to a video scene. For example, the video production system 1 acquires music data based on information (e.g., scenario data) used to generate video data as scene data. For example, the video production system 1 acquires music data based on information related to video data as scene data. This allows the video production system 1 to acquire music data corresponding to a video scene. An example of music data acquisition by the video production system 1 will be described below.
[0111] Fig. 19 is a diagram showing an example of a sound generation unit. In Fig. 19, the sound generation unit 133 includes a scene data acquisition unit 133A and a music generation unit 133B. The scene data acquisition unit 133A acquires scene data related to video scenes. The music generation unit 133B generates music text data related to music based on the scene data. For example, the music generation unit 133B generates music text data based on scenario data.
[0112] FIG. 20 is a diagram showing an example of a sound generation unit. In FIG. 20 , the sound generation unit 150 includes a music acquisition unit 151 and a music editing unit 152. The music acquisition unit 151 acquires music data based on scene data. Specifically, the music acquisition unit 151 acquires music data by searching for music data based on the scene data. More specifically, the music acquisition unit 151 searches for music data based on music text data generated based on the scene data. For example, the music acquisition unit 151 acquires music data based on scenario data. For example, the music acquisition unit 151 acquires music data by searching for music data based on music text data generated based on the scenario data. The music editing unit 152 edits the music data and generates edited music data, which is the edited music data.
[0113] Note that scenario data is an example of scene data, and scene data may be any information related to video data. For example, scene data may be information displayed on a storyboard. For example, scene data may be text data indicating a scene description, narration, dialogue, or camerawork that describes a video scene. Note that scene data may also be the video data itself. Furthermore, scene data may be each of multiple video scenes included in the video displayed by the video data. For example, scene data may be each of multiple video scenes corresponding to each of multiple scene descriptions included in the scenario data. Furthermore, scene data may be 3D information related to the environment, characters, motion, camerawork, lighting, color, or 3D. For example, scene data may be text data indicating 3D information related to the environment, characters, motion, camerawork, lighting, color, or 3D.
[0114] The music acquisition unit 151 may acquire music data by generating the music data based on scene data. More specifically, the music acquisition unit 151 generates the music data based on music text data generated based on scene data. For example, the music acquisition unit 151 searches for music data based on music text data generated based on scenario data. For example, the music acquisition unit 151 generates music data by inputting the music text data into a text-music generation model that generates music data from text data and outputting the music data from the text-music generation model. For example, the music acquisition unit 151 generates the music data using a text-music generation model such as MusicLM, MusicGen, or Mubert Render.
[0115] FIG. 21 is a diagram showing an example of music acquisition processing. In FIG. 21, the scene data acquisition unit 133A acquires scenario data SN2 related to video generation as scene data. The scene data acquisition unit 133A also acquires condition information CN1 related to music. For example, the scene data acquisition unit 133A acquires condition information CN1 input by a user. For example, the scene data acquisition unit 133A acquires, as condition information CN1, text data indicating the playback time of the music (e.g., 30 to 60 seconds), genre, instruments, mood, etc.
[0116] 21 , the music-specific generation unit 133B generates music text data MT1 based on scenario data SN2. For example, the music-specific generation unit 133B inputs scenario data SN2 into a model M6 and causes the model M6 to output music text data MT1, thereby generating the music text data MT1. For example, the music-specific generation unit 133B generates a music description that describes the music as the music text data MT1. In FIG. 21 , the model M6 is a large-scale language model (LLM). For example, the music-specific generation unit 133B inputs scenario data SN2 and condition information CN1 into the model M6 and causes the model M6 to output music text data MT1, thereby generating the music text data MT1.
[0117] 21 , the song acquisition unit 151 acquires song data MU1 based on song text data MT1. Specifically, the song acquisition unit 151 acquires song data MU1 by searching for song data MU1 based on the song text data MT1. For example, the song acquisition unit 151 searches for song data MU1 based on song descriptions as the song text data MT1. For example, the song acquisition unit 151 acquires song data MU1 by executing song search processing PS4 based on the song text data MT1.
[0118] 21, the music acquisition unit 151 acquires music data based on scene data and condition information. For example, the music acquisition unit 151 acquires music data based on scenario data SN2 and condition information CN2.
[0119] 21 , the music editing unit 152 edits music data MU1 to generate edited music data MU2 (not shown), which is the edited music data. For example, the music editing unit 152 executes music editing process PS5 based on music data MU1 to generate edited music data MU2.
[0120] 21 , the client UI output unit 400 outputs a video displayed by video data corresponding to the scene data in association with an edited musical piece corresponding to the edited musical piece data MU2. Specifically, the client UI output unit 400 outputs a video and an edited musical piece in association with each other on a storyboard that displays information for each video scene. For example, when a video is played on the storyboard, the client UI output unit 400 outputs the video and the edited musical piece in synchronization with each other. For example, when a video of a specific video scene is played on the storyboard, the client UI output unit 400 outputs the video of the specific video scene in synchronization with the edited musical piece corresponding to the specific video scene.
[0121] FIG. 22 is a diagram showing an example of a process for generating music acquisition information. In FIG. 22, the music-specific generation unit 133B generates a prompt PT4 for generating music text data MT1 based on scenario data SN2. For example, the music-specific generation unit 133B generates the prompt PT4, which is music acquisition information, using scenario data SN2 and a template TP4. In FIG. 22, the music-specific generation unit 133B generates the prompt PT4 by reflecting the scenario data SN2 in the template TP4. For example, the template TP4 may be preset or may be selected from multiple template candidates. For example, the video production system 1 may select a template corresponding to the scenario data from the multiple template candidates.
[0122] 23 is a diagram showing an example of a process for generating information for music acquisition. In FIG. 23, the music-oriented generation unit 133B inputs the prompt PT4 to a model M6 and causes the model M6 to output music text data MT1, thereby generating music text data MT1.
[0123] For example, the scene data acquisition unit 133A acquires, as scene data, scene descriptions that describe video scenes. For example, the scene data acquisition unit 133A acquires each of the multiple scene descriptions included in the scenario data. The music generation unit 133B generates multiple prompts for generating multiple pieces of music text data corresponding to each of the multiple scene descriptions based on each of the multiple scene descriptions. The music acquisition unit 151 acquires multiple pieces of music data corresponding to each of the multiple video scenes based on each of the multiple scene descriptions. Specifically, the music acquisition unit 151 acquires multiple pieces of music data based on each of the multiple pieces of music text data generated based on each of the multiple scene descriptions. For example, the music acquisition unit 151 acquires each of the multiple pieces of music data by searching for each of the multiple pieces of music text data based on each of the multiple pieces of music text data. In this way, the music acquisition unit 151 acquires music data based on the scene descriptions. This allows the video production system 1 to acquire multiple pieces of music data corresponding to each of the multiple video scenes included in the video.
[0124] Fig. 24 shows an example of a process for generating music acquisition information. Fig. 23 illustrates the case where music text data MT1 is generated by inputting prompt PT4 into model M6. Fig. 24 differs from Fig. 23 in that music text data MT2 is generated using model M7 that is fine-tuned based on paired data, which is a set of training scenario data and training music text data. In Fig. 24, model M7 is a large-scale language model (LLM).
[0125] 24 , the generation unit for music 133B inputs scenario data SN3 to a model M7 that has been trained in advance based on paired data, which is a set of training scenario data and training music text data, and has the model M7 output music text data MT2, thereby generating music text data MT2. Model M7 is trained in advance so that it outputs training music text data when training scenario data is input. In this way, the generation unit for music 133B inputs scene data to a model that has been trained in advance based on paired data, which is a set of training scene data and training music text data, and has the model output music text data, thereby generating music text data.
[0126] 24 , the music-oriented generation unit 133B generates a prompt PT5 including tag data indicating music metadata. For example, the music metadata is information indicating the music style, genre, BPM (Beats Per Minute), key, music description, playback time, or strength of the music. This enables the video generation system 1 to generate music text data including information indicating the music style, genre, BPM (Beats Per Minute), key, music description, playback time, or strength of the music. Furthermore, the video generation system 1 can acquire music data based on the information indicating the music style, genre, BPM (Beats Per Minute), key, music description, playback time, or strength of the music.
[0127] FIG. 25 is a diagram showing an example of scenario data for training. Scenario data SN4 shown in FIG. 25 is an example of scenario data for training. FIG. 26 is a diagram showing an example of song text data for training. Song text data MT3 shown in FIG. 26 is an example of song text data for training. For example, the song-oriented generation unit 133B generates a song description by inputting a scene description into model M7 that has been re-trained based on paired data that is a combination of each of a plurality of scene descriptions SN41 to SN44 included in scenario data SN4 and each of a plurality of song descriptions MT31 to MT35 included in song text data MT3, and having model M7 output a song description.
[0128] Note that the paired data is not limited to a pair of one scene description and one piece of music text data. For example, the paired data may be a pair of one scene description and multiple pieces of music text data. Furthermore, the paired data may be a pair of multiple scene descriptions and one piece of music text data. Furthermore, the paired data may be a pair of multiple scene descriptions and multiple pieces of music text data. For example, the music generation unit 133B may input scenario data SN3 to a model M7 that has been retrained based on paired data that is a pair of scenario data SN4 and music text data MT3, and have the model M7 output music text data MT2, thereby generating music text data MT2.
[0129] FIG. 27 illustrates an example of a process for generating music acquisition information. In FIG. 27 , the music-specific generation unit 133B generates tag data TG1 indicating music metadata as music text data. For example, the music-specific generation unit 133B generates tag data TG1 to be commonly assigned to both music text data MT41 corresponding to a first video scene SC1 and music text data MT51 corresponding to a second video scene SC2, which is the scene following the first video scene SC1. The music-specific generation unit 133B also generates new music text data MT42 including the music text data MT41 and the tag data TG1. The music-specific generation unit 133B also generates new music text data MT52 including the music text data MT51 and the tag data TG1. The music acquisition unit 151 acquires music data based on the tag data. For example, the music acquisition unit 151 searches for music data based on the new music text data MT42 and the new music text data MT52, each including the tag data TG1. This allows the video production system 1 to acquire music data having a style or genre that is common to the first video scene SC1 and the second video scene SC2, which is the scene following the first video scene SC1.
[0130] FIG. 28 illustrates an example of a process for generating music acquisition information. In FIG. 28 , the music-specific generation unit 133B generates music text data MT6 common to a first video scene SC1 and a second video scene SC2, which is the scene following the first video scene SC1. The music-specific generation unit 133B also generates tag data TG2 corresponding to the first video scene SC1 and tag data TG3 corresponding to the second video scene SC2. The music-specific generation unit 133B also generates new music text data MT61 including the music text data MT6 and the tag data TG2. The music-specific generation unit 133B also generates new music text data MT62 including the music text data MT6 and the tag data TG3. The music acquisition unit 151 acquires music data based on the tag data. For example, the music acquisition unit 151 searches for music data based on the new music text data MT61 including the tag data TG2. The music acquisition unit 151 also searches for music data based on the new music text data MT62 including the tag data TG3. This allows the video production system 1 to, for example, vary the mood of the music data corresponding to the first video scene SC1 from the mood of the music data corresponding to the second video scene SC2, which is the scene following the first video scene SC1, depending on the video scene.
[0131] Fig. 29 is a diagram showing an example of a character description. In Fig. 29, the scene data acquisition unit 133A acquires, as scene data, a character description ST1 that explains a character CL1 that appears in a video scene. The music acquisition unit 151 acquires music data based on the character description ST1. This allows the video production system 1 to acquire music data corresponding to the character that appears in the video scene.
[0132] The scene data acquisition unit 133A acquires, as scene data, cut information relating to a cut in which a specific character appears among multiple cuts included in a video scene. The music acquisition unit 151 acquires music data based on the cut information. This allows the video production system 1 to acquire music data corresponding to the characters appearing in the video scene.
[0133] FIG. 30 is a diagram illustrating an example of a method for calculating the intensity of a video scene. In FIG. 30, the scene data acquisition unit 133A acquires, as scene data, information regarding the intensity of a video scene calculated based on information regarding the rate of change in the video scene. For example, the information regarding the rate of change in the video scene is information regarding the movement speed of a character appearing in the video scene or the movement speed of a camera in the video scene. In FIG. 30, the scene data acquisition unit 133A acquires 3D information regarding the 3D of a character CL2 appearing in a video scene SC11. For example, the scene data acquisition unit 133A acquires information indicating the positions of body parts (e.g., arms, legs, etc.) of the character CL2 at each time. Furthermore, the scene data acquisition unit 133A calculates the speed of movement of the body parts of the character CL2 based on the information indicating the positions of the body parts (e.g., arms, legs, etc.) of the character CL2 at each time. The scene data acquisition unit 133A also calculates the intensity of the video scene SC11 by integrating the speed v(t) of the movement of the body parts of the character CL2 over a predetermined time. For example, the scene data acquisition unit 133A calculates the intensity of the video scene SC11 based on the calculation formula expressed as Equation (1) below.
[0134]
[0135] The music acquisition unit 151 also acquires music data based on information about the intensity of a video scene. This allows the video production system 1 to estimate the excitement of a video scene based on the intensity of the video scene. The video production system 1 can also acquire music data according to the excitement of the video scene.
[0136] FIG. 31 is a diagram illustrating an example of a method for comparing the intensities of video scenes. In FIG. 31 , the scene data acquisition unit 133A calculates the first intensity of a first video scene to be 0.2. The scene data acquisition unit 133A also calculates the intensity of a second video scene, which is the scene following the first video scene, to be 0.4. The music acquisition unit 151 acquires music data based on a comparison result between the first intensity of the first video scene and the second intensity of the second video scene, which is the scene following the first video scene. For example, the music acquisition unit 151 acquires music data based on the magnitude of the difference between the first intensity of the first video scene and the second intensity of the second video scene. This allows the video production system 1 to estimate the excitement of a video scene based on the comparison result between the intensity of a video scene and the intensity of the next video scene. Furthermore, the video production system 1 can acquire music data according to the excitement of a video scene.
[0137] FIG. 32 illustrates an example of a process for generating music acquisition information. In FIG. 32 , the music-specific generation unit 133B extracts image features FT1 indicating characteristics of video data MV3 from video data MV3 corresponding to scene data SN5, and generates music text data MT4 based on the image features FT1. For example, the music-specific generation unit 133B inputs the video data MV3 into a feature extraction model that extracts image features from images and causes the feature extraction model to output the image features FT1, thereby extracting the image features FT1 from the video data MV3. The music-specific generation unit 133B also generates feature text data, which is text data corresponding to the image features. For example, the music-specific generation unit 133B inputs the image features FT1 into a model M8 trained to output feature text data when an image feature is input, and causes the model M8 to output feature text data FT11 corresponding to the image features FT1, thereby generating feature text data FT11. Furthermore, the music generation unit 133B inputs scenario data SN5 and feature text data FT11 into a model M7 that has been trained in advance based on pair data that is a set of learning scenario data and learning feature text data, and learning music text data, and causes the model M7 to output music text data MT4, thereby generating music text data MT4.
[0138] 32 , the music-oriented generation unit 133B acquires 3D information CU1 related to 3D in video data MV3 corresponding to scene data SN5, and generates music text data MT4 based on the 3D information CU1. For example, the music-oriented generation unit 133B acquires 3D information CU1 stored in the video generation system 1. The music-oriented generation unit 133B also generates 3D text data, which is text data corresponding to the 3D information. For example, the music-oriented generation unit 133B inputs the 3D information CU1 to a model M8 that has been trained to output 3D text data when 3D information is input, and causes the model M8 to output 3D text data CU2 corresponding to the 3D information CU1, thereby generating the 3D text data CU2. In addition, the music generation unit 133B inputs scenario data SN5 and 3D text data CU2 into a model M7 that has been pre-trained based on pair data that is a set of learning scenario data and learning 3D text data, and learning music text data, and causes model M7 to output music text data MT4, thereby generating music text data MT4.
[0139] FIG. 33 illustrates an example of a song acquisition process. In FIG. 33 , the song acquisition unit 151 acquires song data MU2 by inputting song text data MT7 into the song retrieval model M9 and causing the song retrieval model M9 to output song data MU2. For example, the song retrieval model M9 may be a search system pre-trained to output song data corresponding to the song text data when the song text data is input. For example, the song retrieval model M9 may be a search system trained to output a predetermined number of candidate song data items with the highest similarity to the song text data as search results based on the similarity between the song text data and the song data. Note that the song retrieval model M9 may be a known search system that searches for song data corresponding to text from text. For example, the song acquisition unit 151 acquires three candidate song data items MU21 to MU23 as search results for the song data MU2.
[0140] FIG. 34 illustrates an example of a song acquisition process. In FIG. 34 , the song acquisition unit 151 extracts metadata MT8 from song text data MT7. For example, the song acquisition unit 151 extracts tag data included in the song text data MT7 as metadata MT8. The song acquisition unit 151 inputs the metadata MT8 into a song retrieval model M91 and causes the song retrieval model M91 to output song data MU3, thereby acquiring the song data MU3. For example, the song retrieval model M91 may be a search system pre-trained to output song data corresponding to the input metadata. For example, the song retrieval model M91 may be a search system trained to output a predetermined number of candidate song data items with the highest similarity to the metadata as search results based on the similarity between the metadata and the song data. Note that the song retrieval model M91 may be a known search system that searches for song data corresponding to text from text. Note that the song acquisition unit 151 may acquire song data based on both the song text data MT7 and the metadata MT8. For example, the music acquisition unit 151 acquires three candidate music data MU31 to MU33 as search results for the music data MU3.
[0141] Fig. 35 is a diagram showing an example of a user interface. In Fig. 35, the client UI output unit 400 displays content CT3, which is a UI for allowing the user to select whether to place emphasis on metadata (Tag shown in Fig. 35) or song text data (Language shown in Fig. 35) in a search. The client UI output unit 400 displays content CT3 including a slider SL1 that allows the user to select whether to place emphasis on metadata (Tag shown in Fig. 35) or song text data (Language shown in Fig. 35) in a search.
[0142] FIG. 36 is a diagram showing an example of music editing processing. In FIG. 36 , the music editing unit 152 generates edited music data by editing the length of music data to match the length of video data corresponding to scene data. For example, the music editing unit 152 acquires the playback time of video data held in the video production system 1. The music editing unit 152 also generates edited music data by editing the playback time of the music data to match the playback time of the video data based on the playback time of the video data and the playback time of the music data. In FIG. 36 , the music editing unit 152 deletes music data corresponding to chorus 1, thereby editing the playback time of the music data to match the playback time of the video data.
[0143] 36, the music editing unit 152 generates edited music data by performing switching editing for the transition between first music data corresponding to a first video scene and second music data corresponding to a second video scene that is the scene following the first video scene. For example, the music editing unit 152 performs editing to add a sound effect to the music data that indicates the excitement when switching from the first video scene to the second video scene.
[0144] FIG. 37 is a diagram showing an example of music editing information. In FIG. 37 , the music editing unit 152 generates music editing information SN5 for generating edited music data based on scenario data related to video generation, and then generates edited music data based on the music editing information SN5. For example, the music editing unit 152 inputs scenario data into a music editing information generation model that has been trained in advance to output music editing information when scenario data is input, and then causes the music editing information generation model to output music editing information, thereby generating music editing information SN5. For example, the music editing unit 152 generates music editing information in which an instruction statement instructing editing is inserted between scene descriptions included in the scenario data. In FIG. 37 , the music editing unit 152 generates music editing information SN5 in which an instruction statement AT1 instructing a transition to be performed between scene descriptions included in the scenario data is inserted. Furthermore, the music editing unit 152 generates music editing information SN5 in which an instruction statement AT2 instructing a change of music is inserted between scene descriptions included in the scenario data. This allows the video production system 1 to acquire music data that has been edited in accordance with the music editing information.
[0145] FIG. 38 illustrates an example of music editing processing. In FIG. 38 , the music editing unit 152 calculates the scene similarity between first scene data relating to a first video scene and second scene data relating to a second video scene, and generates edited music data by performing transition editing based on the first similarity. For example, the music editing unit 152 calculates the scene similarity between a first scene description corresponding to the first video scene and a second scene description corresponding to the second video scene. The music editing unit 152 also determines whether the scene similarity exceeds a predetermined threshold. If the music editing unit 152 determines that the scene similarity exceeds the predetermined threshold, it performs transition editing. This allows the video production system 1 to, for example, naturally transition from one music data to the next, depending on how the video scene transitions from one video scene to the next.
[0146] FIG. 39 illustrates an example of music editing processing. In FIG. 39 , the music editing unit 152 calculates the scene similarity between third scene data relating to a third video scene SC5 and fourth scene data relating to a fourth video scene SC6, which is different from the third video scene SC5. For example, the music editing unit 152 calculates the scene similarity between scenario data SN51 corresponding to the third video scene SC5 and scenario data SN61 corresponding to the fourth video scene SC6. If the scene similarity exceeds a predetermined threshold, the music editing unit 152 edits the third music data corresponding to the third video scene SC5 and acquires the edited third music data as fourth music data corresponding to the fourth video scene SC6. For example, the music editing unit 152 edits the third music data by changing the instrument corresponding to the third music data (e.g., a violin) to another instrument (e.g., a piano), thereby generating fourth music data. This allows the video production system 1 to control the music so as to maintain consistency in the story of the video, for example, by making it possible to play similar music in the same scene.
[0147] FIG. 40 illustrates an example of music editing processing. In FIG. 40, the music editing unit 152 generates edited music data based on information regarding text data indicating text to be displayed on video displayed by video data corresponding to scene data. In the diagram on the left side of FIG. 40, a caption in a Mincho font is displayed on the video. In this case, the music editing unit 152 may edit the music data to create an elegant and smooth melody that corresponds to the Mincho font. In the diagram on the right side of FIG. 40, a caption in a Gothic font is displayed on the video. In this case, the music editing unit 152 may edit the music data to create a melody with a strong impact that corresponds to the Gothic font. Note that the music editing unit 152 may edit the music data based on the size of the text to be displayed on the video. For example, the music data may be edited based on the text with the largest font size among multiple texts to be displayed on the video. This allows the video production system 1 to obtain music data that corresponds to the text to be displayed on the video.
[0148] 41 is a diagram showing an example of the analysis results of the music analysis model. The music editing unit 152 acquires the analysis results of the music data by inputting the music data into the music analysis model and having the music analysis model output the analysis results of the music data. The music acquisition unit 151 re-acquires the music data based on the analysis results of the music data.
[0149] 42 is a diagram showing an example of a music control model. The music editing unit 152 inputs user input information UI5 and image features FT2 extracted from video data into the sound control model M92, and generates edited music data by outputting the edited music data from the sound control model M92. For example, the sound control model M92 is a model that adjusts the intensity of music data to match the intensity of video scenes included in the video data.
[0150] FIG. 43 is a diagram showing an example of music editing processing. In FIG. 43, the music editing unit 152 acquires 3D information corresponding to video data. The music editing unit 152 generates sound effects based on the 3D information. FIG. 43 shows an object entering the camera's field of view from outside. At this time, the music editing unit 152 generates sound effects based on position information of the object. This allows the video production system 1 to generate sound effects based on information about objects that are not shown on the screen.
[0151] <1-5. Example of processing flow from the user's perspective> Next, as an example of a processing flow from the user's perspective, a procedure for information processing by the image generation system 1 in response to a user operation will be described with reference to Fig. 44. Fig. 44 is a flowchart showing the flow of processing in response to a user operation.
[0152] 44, in the video production system 1, the user presses a project creation button (step S1). Then, in the video production system 1, the user inputs information required for video generation (step S2). For example, the user inputs information including the purpose of creating the video, the message to be conveyed through the video, the target users, the functional features of the product / service to be conveyed, the length of the video, the aspect ratio, etc.
[0153] Then, in the video production system 1, the user presses the storyboard, video, sound, and text logo creation buttons (step S3). For example, the video production system 1 may generate everything at once (including video data, for example) in response to user operation, or may present a storyboard and allow the user to make some modifications in response to instructions from the user before generating the video, sound, and text logo.
[0154] Then, the video production system 1 modifies the storyboard, video, sound, and text logo in response to user operations (step S4). For example, the video production system 1 may accept modifications to each of the items in the order of the user's preference.
[0155] Then, the video production system 1 performs export in response to a user operation (step S5). For example, the video production system 1 may perform export in a video file format such as mp4, avi, or mov in response to a user operation. Furthermore, for example, the video production system 1 may perform export in a file format of any video editing software such as Premiere Pro, After Effects, or DaVinci Resolve.
[0156] <1-6. Regarding AI Models> Note that the AI models used in the above-described processes are not limited to the examples described in each section, and any internal structure can be adopted as long as desired information can be output in response to input. Any combination of input, output, and internal structure of the AI model can be adopted as long as desired information can be output.
[0157] The input of the AI model may be text, images, audio, 3D data, etc., or a combination thereof. The output of the AI model may be text, images, audio, 3D data, etc. Note that the above-mentioned inputs and outputs are merely examples, and the above-mentioned AI model may have any inputs and outputs.
[0158] Furthermore, the internal structure of the AI model can be any structure depending on the combination of input and output. In other words, the internal structure of the AI model can be any structure as long as it can produce a desired output for the input.
[0159] For example, the AI model may have a structure related to Transformer. For example, the AI model may have a structure related to Transformer and perform processing taking into account context, such as context within data, such as text or time-series data. For example, the AI model may have a self-attention mechanism. For example, the AI model may have any attention mechanism, such as single-head attention or multi-head attention. Note that the AI model does not necessarily have to have an attention mechanism.
[0160] The AI model may have a mechanism for extracting features from an input. For example, the AI model may have an encoder. The AI model may have a mechanism for generating information based on the extracted features. For example, the AI model may have a decoder.
[0161] The AI model may have a structure related to a convolutional neural network (CNN). For example, when processing an image, the AI model may have a structure related to a CNN. For example, the AI model may have at least one of a convolution layer, a pooling layer, a fully connected layer, etc.
[0162] The above-described internal structure is merely an example, and the AI model may have any internal structure. For example, the AI model may have a skip connection. Furthermore, the AI model may have a structure related to a diffusion model.
[0163] Furthermore, the above-described AI model may be generated (trained) by any learning process. The AI model may be a machine learning model trained using any machine learning method. For example, the AI model may be a model generated based on a so-called Foundation Model by fine-tuning the Foundation Model to apply it to a specific task (e.g., scenario data generation, code generation, etc.). For example, an AI model such as the above-described LLM may be a model generated by fine-tuning the Foundation Model to apply it to a specific task.
[0164] The base model here is a model that has been trained to be applicable to various tasks, for example, to be able to perform a wide variety of tasks. For example, the base model is a neural network that has been pre-trained with a large amount of unlabeled data set. Note that the base model may have any structure, such as a Transformer-based architecture. For example, the base model is generated by self-supervised learning using data without correct answer labels. As described above, the base model is fine-tuned so that it can be adapted to a wide range of downstream tasks.
[0165] For example, when applied to a scenario data generation task, the base model is fine-tuned to be adaptable to the scenario data generation task, and an AI model (model M1, etc.) adapted to the scenario data generation task is generated. Also, when applied to a code generation task, the base model is fine-tuned to be adaptable to the code generation task, and an AI model (model M3, etc.) adapted to the code generation task is generated. Also, when applied to a sound data generation task, the base model is fine-tuned to be adaptable to the sound data generation task, and an AI model (model M4, etc.) adapted to the sound data generation task is generated. Also, when applied to a text logo generation task, the base model is fine-tuned to be adaptable to the text logo generation task, and an AI model (model M5, etc.) adapted to the text logo generation task is generated.
[0166] For example, an AI model (such as model M1) applied to a scenario data generation task is trained using training data including a combination of input information corresponding to the AI model and scenario data (also referred to as "correct answer information") that is the correct output when the input information is input. The training data, such as the input information and correct answer information, may be data created by a person, or may be data automatically generated by a computer that generates the training data. For example, the scenario data that is the correct answer information may be data created by a person. Below, model M1 will be briefly described as an example. For example, model M1 is trained to output correct answer information corresponding to each piece of input information when that input information is input. For example, model M1 is trained by adjusting (correcting) parameters (connection coefficients) using a method such as backpropagation (error backpropagation) so as to reduce the error between the output of model M1 when certain input information is input and the correct answer information corresponding to that input information. In addition, other AI models such as an AI model applied to a code generation task (e.g., model M3), an AI model applied to a sound data generation task (e.g., model M4), an AI model applied to a text logo generation task (e.g., model M5), an AI model applied to an evaluation task (e.g., model M10), and an AI model applied to an image improvement processing task (e.g., model M11) may also be trained using a similar learning process.
[0167] The above-described learning process is merely an example, and the above-described AI model may be trained by any learning process depending on the input, output, and internal structure of the AI model. For example, the AI model may be trained using an unsupervised learning method such as a generative adversarial network (GAN). The AI model may also be trained in a distributed state without aggregating data, such as federated learning. In this case, each video generation service device (e.g., server) may generate local models collected by the service, and a server (aggregation server) that aggregates information (e.g., parameters) of the local models generated by each video generation service device (e.g., server) may generate a global model using the information on the local models. In this case, the video generation system 1 may receive the global model generated by the aggregation server from the aggregation server and use the received global model as an AI model for processing.
[0168] In this way, the above-described AI models may be generated (learned) by any computer. That is, the learning process for generating the AI models may be performed by any device (computer, etc.) in the video production system 1, or may be performed by a device outside the video production system 1. For example, when a device outside the video production system 1 generates at least one of the above-described AI models, the video production system 1 acquires the AI model from the device outside the video production system 1 and performs processing using the acquired AI model.
[0169] 2. Other Embodiments The processing according to each of the above-described embodiments may be implemented in various different forms (modifications) other than the above-described embodiments and modifications.
[0170] 2-1. Other Configuration Examples The above-described configuration of the image generation system 1 is merely an example, and any desired division of functions in the image generation system 1 can be adopted. In other words, the above-described configuration is merely an example, and the image generation system 1 may have any desired division of functions and any desired configuration as long as it can provide the services related to image generation described above. For example, the image generation system 1 may be configured by a single device (e.g., a computer) that performs the above-described processing. In this case, one device of the image generation system 1 may have the functions of the image generation module 100, the information acquisition module 200, the sensor unit 300, and the client UI output unit 400. For example, the image generation service provided by the image generation system 1 may be provided to a user as a program such as a tool (AI Assist Creation Tool) that runs on a terminal device (e.g., computer 20) used by the user.
[0171] <2-2. Others> Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0172] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0173] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.
[0174] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0175] 3. Effects of the Present Disclosure As described above, the information processing system (video production system 1 in the embodiment) according to the present disclosure includes the scene data acquisition unit 133A and the music acquisition unit 151. The scene data acquisition unit 133A acquires scene data related to video scenes. The music acquisition unit 151 acquires music data based on the scene data. This allows the information processing system to acquire music data corresponding to video scenes.
[0176] Furthermore, the scene data acquisition unit 133A acquires scenario data related to video generation as scene data. The music acquisition unit 151 acquires music data based on the scenario data. This allows the information processing system to acquire music data corresponding to video scenes based on the scenario data.
[0177] The music acquisition unit 151 also acquires music data by searching for music data based on scene data, thereby enabling the information processing system to search for and acquire music data corresponding to video scenes.
[0178] The music acquisition unit 151 also acquires music data by generating it based on scene data, allowing the information processing system to generate and acquire music data according to video scenes.
[0179] The scene data acquisition unit 133A also acquires condition information related to music. The music acquisition unit 151 acquires music data based on the scene data and the condition information. This allows the information processing system to acquire music data according to the condition information.
[0180] The information processing system also includes a music generation unit 133B. The music generation unit 133B generates music text data for a music piece based on scene data. The music acquisition unit 151 acquires music data based on the music text data. This allows the information processing system to acquire music data corresponding to the music text data.
[0181] Furthermore, the music-oriented generation unit 133B generates a music description that describes the music as music text data. The music acquisition unit 151 acquires music data based on the music description. This allows the information processing system to acquire music data corresponding to the music description.
[0182] Furthermore, the music generation unit 133B generates tag data indicating metadata for the music as music text data. The music acquisition unit 151 acquires music data based on the tag data. This allows the information processing system to acquire music data corresponding to the music metadata.
[0183] Furthermore, the music generation unit 133B generates a first prompt for generating music text data based on the scene data, inputs the first prompt to the first model, and causes the first model to output the music text data, thereby generating the music text data. This allows the information processing system to generate music text data according to the scene data.
[0184] Furthermore, the music-oriented generation unit 133B generates music text data by inputting scene data to a second model that has been trained in advance based on paired data that is a set of training scene data and training music text data, and having the second model output music text data, thereby enabling the information processing system to generate music text data corresponding to the scene data.
[0185] Furthermore, the music-oriented generation unit 133B extracts image features that indicate the characteristics of the video data from the video data corresponding to the scene data, and generates music text data based on the image features, thereby enabling the information processing system to generate music text data that corresponds to the characteristics of the video data.
[0186] The music-oriented generation unit 133B also acquires 3D information related to the 3D of the video data corresponding to the scene data, and generates music text data based on the 3D information, thereby enabling the information processing system to generate music text data that matches the 3D characteristics of the video data.
[0187] The information processing system also includes a music editing unit 152. The music editing unit 152 edits music data and generates edited music data, which is the edited music data. This allows the information processing system to obtain music data that better matches the video scene.
[0188] The music editing unit 152 also generates edited music data by editing the length of the music data to match the length of the video data corresponding to the scene data, thereby enabling the information processing system to play music that matches the video.
[0189] The music editing unit 152 also generates edited music data by editing the transition between first music data corresponding to a first video scene and second music data corresponding to a second video scene that follows the first video scene, thereby enabling the information processing system to, for example, make the transition from one music data piece to the next music data piece appear natural.
[0190] The information processing system also includes a client UI output unit 400. The client UI output unit 400 outputs a video displayed by video data corresponding to the scene data and an edited music piece corresponding to the edited music data in association with each other. This allows the information processing system to simultaneously play the video and the music piece corresponding to the video.
[0191] The client UI output unit 400 also outputs the video and the edited music in association with each other on a storyboard that displays information for each video scene, allowing the information processing system to simultaneously play the video and the music corresponding to the video on the storyboard that displays information for each video scene.
[0192] Furthermore, when a video is played on the storyboard, the client UI output unit 400 outputs the video and the edited music in synchronization with each other, thereby enabling the information processing system to simultaneously play the video and the music corresponding to the video when the video is played on the storyboard.
[0193] 4. Hardware Configuration An information processing device (information appliance) having the image generation module 100, information acquisition module 200, client UI output unit 400, etc. according to each of the above-described embodiments is realized by, for example, a computer 1000 configured as shown in FIG. 45 . FIG. 45 is a hardware configuration diagram showing an example of the computer 1000 that realizes the functions of the information processing device. The following description will be given using the image generation module 100 according to the embodiment as an example. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, a HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.
[0194] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0195] The ROM 1300 stores boot programs such as a Basic Input Output System (BIOS) that is executed by the CPU 1100 when the computer 1000 is started, and programs that depend on the hardware of the computer 1000 .
[0196] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an image generation program according to the present disclosure, which is an example of program data 1450.
[0197] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0198] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from an input device such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to an output device such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.
[0199] For example, when the computer 1000 functions as the image generation module 100 according to the embodiment, the CPU 1100 of the computer 1000 executes an image generation program loaded onto the RAM 1200, thereby realizing the functions of the control unit 1301, etc. The image generation program according to the present disclosure and data in the storage unit 1302 are stored in the HDD 1400. The CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, the CPU 1100 may obtain these programs from another device via an external network 1550.
[0200] The present technology can also be configured as follows. (1) An information processing system including: a scene data acquisition unit that acquires scene data related to a video scene; and a music acquisition unit that acquires music data based on the scene data. (2) The information processing system described in (1), in which the scene data acquisition unit acquires scenario data related to video generation as the scene data, and the music acquisition unit acquires the music data based on the scenario data. (3) The information processing system described in (1) or (2), in which the scene data acquisition unit acquires, as the scene data, a scene description that describes the video scene, and the music acquisition unit acquires the music data based on the scene description. (4) The information processing system described in any one of (1) to (3), in which the scene data acquisition unit acquires, as the scene data, a character description that describes a character appearing in the video scene, and the music acquisition unit acquires the music data based on the character description. (5) The information processing system according to any one of (1) to (4), wherein the scene data acquisition unit acquires, as the scene data, information about the intensity of the video scene calculated based on information about the rate of change in the video scene, and the music acquisition unit acquires the music data based on the information about the intensity of the video scene. (6) The information processing system according to (5), wherein the information about the rate of change in the video scene is information about the movement speed of a character appearing in the video scene or the movement speed of a camera in the video scene. (7) The information processing system according to (5) or (6), wherein the music acquisition unit acquires the music data based on a comparison result between a first intensity of a first video scene and a second intensity of a second video scene that is the next scene after the first video scene.(8) The information processing system according to any one of (1) to (7), wherein the scene data acquisition unit acquires, as the scene data, cut information relating to a cut in which a specific character appears from among a plurality of cuts included in the video scene, and the music acquisition unit acquires the music data based on the cut information. (9) The information processing system according to any one of (1) to (8), wherein the music acquisition unit acquires the music data by searching for the music data based on the scene data. (10) The information processing system according to any one of (1) to (8), wherein the music acquisition unit acquires the music data by generating the music data based on the scene data. (11) The information processing system according to any one of (1) to (10), wherein the scene data acquisition unit acquires condition information relating to music, and the music acquisition unit acquires the music data based on the scene data and the condition information. (12) The information processing system according to any one of (1) to (11), further comprising a music-oriented generation unit that generates music text data related to a music piece based on the scene data, wherein the music acquisition unit acquires the music data based on the music text data. (13) The information processing system according to (12), wherein the music-oriented generation unit generates, as the music text data, a music description that describes the music, and the music acquisition unit acquires the music data based on the music description. (14) The information processing system according to (12) or (13), wherein the music-oriented generation unit generates, as the music text data, tag data that indicates metadata of the music piece, and the music acquisition unit acquires the music data based on the tag data. (15) The information processing system according to any one of (12) to (14), wherein the music generation unit generates a first prompt for generating the music text data based on the scene data, inputs the first prompt to a first model, and causes the first model to output the music text data, thereby generating the music text data.(16) The information processing system of any one of (12) to (14), wherein the generation unit for music generates the music text data by inputting the scene data to a second model that has been trained in advance based on paired data that is a set of training scene data and training music text data, and having the second model output the music text data. (17) The information processing system of any one of (12) to (16), wherein the generation unit for music extracts image features that indicate characteristics of video data from video data corresponding to the scene data, and generates the music text data based on the image features. (18) The information processing system of any one of (12) to (17), wherein the generation unit for music obtains 3D information related to 3D of video data corresponding to the scene data, and generates the music text data based on the 3D information. (19) The information processing system of (1), further comprising a music editing unit that edits the music data and generates edited music data that is the edited music data. (20) The information processing system of (19), wherein the music editing unit generates music editing information for generating the edited music data based on scenario data related to video generation, and generates the edited music data based on the music editing information. (21) The information processing system of (19) or (20), wherein the music editing unit generates the edited music data based on information related to text data indicating text to be displayed on video displayed by video data corresponding to the scene data. (22) The information processing system of any one of (19) to (21), wherein the music editing unit generates the edited music data by editing the length of the music data to match the length of the video data corresponding to the scene data. (23) The information processing system of any one of (19) to (22), wherein the music editing unit generates the edited music data by performing switching editing regarding a switch between first music data corresponding to a first video scene and second music data corresponding to a second video scene that is the scene following the first video scene.(24) The information processing system of (23), wherein the music editing unit calculates a first similarity between first scene data related to the first video scene and second scene data related to the second video scene, and generates the edited music data by performing the transition editing based on the first similarity. (25) The information processing system of any one of (19) to (24), wherein the music editing unit calculates a second similarity between third scene data related to a third video scene and fourth scene data related to a fourth video scene different from the third video scene, and if the second similarity exceeds a predetermined threshold, edits the third music data corresponding to the third video scene, and acquires the edited third music data as fourth music data corresponding to the fourth video scene. (26) The information processing system of (19), further comprising a client UI output unit that outputs a video displayed by video data corresponding to the scene data and the edited music corresponding to the edited music data in association with each other. (27) The information processing system described in (26), wherein the client UI output unit outputs the video and the edited music in association with each other on a storyboard displaying information for each video scene. (28) The information processing system described in (27), wherein the client UI output unit outputs the video and the edited music in synchronization with each other when the video is played on the storyboard. (29) An information processing method executed by a computer, comprising: acquiring scene data related to a video scene; and acquiring music data based on the scene data. (30) An information processing program causing a computer to acquire scene data related to a video scene; and acquiring music data based on the scene data.
[0201] 1 Video generation system 100 Video generation module 110 Input text analysis unit 120 Sensor analysis unit 130 Prompt etc. generation unit 131 Scenario generation unit 132 Video generation unit 133 Sound generation unit 133A Scene data acquisition unit 133B Music generation unit 134 Text / logo generation unit 140 Video generation unit 141 USD generation unit 142 Rendering unit 143 Video refinement unit 150 Sound generation unit 151 Music acquisition unit 152 Music editing unit 160 Text / logo generation unit 170 Composite editing unit 180 Evaluation unit 190 Client UI module 200 Information acquisition module 210 Input text acquisition unit 220 Sensor acquisition unit 300 Sensor unit 400 Client UI output unit
Claims
1. An information processing system comprising: a scene data acquisition unit that acquires scene data relating to a video scene; and a music acquisition unit that acquires music data based on the scene data.
2. The information processing system according to claim 1, wherein the scene data acquisition unit acquires scenario data relating to video generation as the scene data, and the music acquisition unit acquires the music data based on the scenario data.
3. The information processing system according to claim 1, wherein the music acquisition unit acquires the music data by searching for the music data based on the scene data.
4. The information processing system according to claim 1, wherein the music acquisition unit acquires the music data by generating the music data based on the scene data.
5. An information processing system as described in claim 1, wherein the scene data acquisition unit acquires condition information related to music, and the music acquisition unit acquires the music data based on the scene data and the condition information.
6. The information processing system according to claim 1, further comprising a music generation unit that generates music text data related to music based on the scene data, and the music acquisition unit acquires the music data based on the music text data.
7. The information processing system of claim 6, wherein the music generation unit generates a music description that describes the music as the music text data, and the music acquisition unit acquires the music data based on the music description.
8. An information processing system as described in claim 6, wherein the music generation unit generates tag data indicating metadata of the music as the music text data, and the music acquisition unit acquires the music data based on the tag data.
9. The information processing system of claim 6, wherein the music generation unit generates a first prompt for generating the music text data based on the scene data, inputs the first prompt into a first model, and causes the first model to output the music text data, thereby generating the music text data.
10. The information processing system of claim 6, wherein the music generation unit generates the music text data by inputting the scene data into a second model that has been trained in advance based on paired data that is a set of training scene data and training music text data, and having the second model output the music text data.
11. The information processing system of claim 6, wherein the music generation unit extracts image features indicating characteristics of the video data from the video data corresponding to the scene data, and generates the music text data based on the image features.
12. The information processing system according to claim 6, wherein the music-oriented generation unit acquires 3D information relating to 3D of the video data corresponding to the scene data, and generates the music text data based on the 3D information.
13. The information processing system according to claim 1, further comprising a music editing unit that edits the music data and generates edited music data, which is the edited music data.
14. An information processing system according to claim 13, wherein the music editing unit generates the edited music data by editing the length of the music data to match the length of the video data corresponding to the scene data.
15. An information processing system as described in claim 13, wherein the music editing unit generates the edited music data by performing switching editing regarding the switching between first music data corresponding to a first video scene and second music data corresponding to a second video scene that is the scene following the first video scene.
16. An information processing system according to claim 13, further comprising a client UI output unit that outputs a video displayed by video data corresponding to said scene data in association with an edited music piece corresponding to said edited music piece data.
17. An information processing system according to claim 16, wherein the client UI output unit outputs the video and the edited music in association with each other on a storyboard displaying information for each video scene.
18. The information processing system according to claim 17, wherein the client UI output unit outputs the video and the edited music in synchronization with each other when the video is played on the storyboard.
19. An information processing method executed by a computer, comprising: acquiring scene data relating to a video scene; and acquiring music data based on the scene data.
20. An information processing program that causes a computer to acquire scene data relating to a video scene, and acquire music data based on the scene data.
Citation Information
Patent Citations
Video generation method and device, computer equipment and storage medium
CN117082304A
Video editing device and control method therefor
JP2013042215A
Method for composing music based on image and apparatus therefor
KR102390951B1
Method and apparatus for generating music
US20200051536A1