Information processing system, information processing method, and information processing program

The information processing system addresses the lack of editing support in video content generation by using AI models to generate support information and example responses, enhancing the editing process and improving video production efficiency.

WO2025220311A1PCT designated stage Publication Date: 2025-10-23SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/004852
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-18
Filing Date
2025-02-14
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Conventional technologies for generating video content lack appropriate support for content editing and production, particularly in scenarios involving virtual person postures from two-dimensional drawings, necessitating improved content creation assistance.

Method used

An information processing system that includes an acquisition unit for background and content information, and a processing unit that generates support information using AI models to assist in content creation, such as video generation, sound production, and text/logo creation, integrating user inputs and pre-stored prompts to enhance editing capabilities.

Benefits of technology

The system effectively supports content creation by providing tailored advice and example responses based on user inputs, enhancing the editing process with AI-driven assistance, thereby improving the quality and efficiency of video production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025004852_23102025_PF_FP_ABST
    Figure JP2025004852_23102025_PF_FP_ABST
Patent Text Reader

Abstract

An information processing system according to the present disclosure comprises: an acquisition unit that acquires background information relating to content production and content information being edited; and a processing unit that outputs assistance information pertaining to the content production on the basis of the background information and the content information being edited.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing system, information processing method, and information processing program

[0001] The present disclosure relates to an information processing system, an information processing method, and an information processing program.

[0002] There are currently available technologies for automatically generating video content (also called "moving images"). For example, there is a technology for estimating the three-dimensional posture of a virtual person from a two-dimensional line drawing and generating a moving image (see, for example, Patent Document 1).

[0003] Japanese Patent Application Laid-Open No. 2003-058906

[0004] However, there is room for improvement in the conventional technology. For example, although the conventional technology can generate videos from storyboards, it does not particularly consider editing content such as videos, and there is room for improvement in content production. Therefore, there is a need for appropriate support for content production, such as content editing.

[0005] Therefore, the present disclosure proposes an information processing system, an information processing method, and an information processing program that can appropriately support content creation.

[0006] In order to solve the above problems, one form of information processing system according to the present disclosure includes an acquisition unit that acquires background information related to content production and content information being edited, and a processing unit that outputs support information related to the content production based on the background information and the content information being edited.

[0007] 1 is a diagram illustrating an example of an information processing system of the present disclosure. FIG. 1 is a diagram illustrating an example of a hardware configuration related to the information processing system of the present disclosure. FIG. 1 is a diagram illustrating an example of a user interface. A flowchart showing the procedure of a first process. FIG. 1 is a diagram illustrating an example of a user interface. FIG. 1 is a diagram illustrating an example of a user interface. FIG. 1 is a diagram illustrating an overview of the first process. FIG. 1 is a diagram illustrating an example of a prompt generation process. FIG. 1 is a diagram illustrating an example of a support information generation process. A flowchart showing the procedure of a second process. FIG. 1 is a diagram illustrating an example of content. FIG. 1 is a diagram illustrating an example of a support information generation process. FIG. 1 is a diagram illustrating an example of a user interface. FIG. 1 is a diagram illustrating an example of a user interface. FIG. 1 is a diagram illustrating an example of an intervention timing. A conceptual diagram illustrating an example of application of Calm Technology. FIG. 2 is a hardware configuration diagram illustrating an example of a computer that realizes the functions of an information processing device.

[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Note that the information processing system, information processing method, and information processing program according to the present application are not limited to these embodiments. In addition, in the following embodiments, the same components are designated by the same reference numerals, and redundant description will be omitted.

[0009] The present disclosure will be described in the following order of items: 1. Embodiment 1-1. Overview of the configuration of the information processing system of the present disclosure 1-2. First example 1-2-1. First processing 1-2-1-1. Procedure of the first processing 1-2-1-2. Overview of the first processing 1-3. Second example 1-3-1. Second processing 1-3-2-1. Procedure of the second processing 1-3-2-2. Overview of the second processing 1-4. Other examples 1-5. Regarding the AI ​​model 2. Other embodiments 2-1. Other configuration examples 2-2. Other 3. Hardware configuration

[0010] <1. Embodiments> <1-1. Overview of Configuration of Information Processing System of the Present Disclosure> In the following embodiments, a case where video data (also simply referred to as "video") is generated as an example of content information (also referred to as "content") generated by the information processing system 1 is shown, but the content generated by the information processing system 1 is not limited to video. For example, the content generated by the information processing system 1 may be various types of content such as animation, music data, image data, etc. Furthermore, for example, the content generated by the information processing system 1 may be document data, PDF data, presentation (slide) data, etc.

[0011] Fig. 1 is a diagram illustrating an example of an information processing system according to the present disclosure. The information processing system 1 includes an image generation module 100, an information acquisition module 200, a sensor unit 300, and a client UI display unit 400. Although Fig. 1 illustrates only one of each component, the information processing system 1 may include multiple image generation modules 100, multiple information acquisition modules 200, multiple sensor units 300, and multiple client UI display units 400.

[0012] First, the configuration of the image generation module 100 that performs image generation processing will be described. The image generation module 100 includes an information analysis unit 110, a sensor analysis unit 120, a prompt generation unit 130, an image generation unit 140, a sound generation unit 150, a text / logo generation unit 160, a composite editing unit 170, an evaluation unit 180, and a client UI module 190.

[0013] The information analysis unit 110 analyzes information. The information analysis unit 110 analyzes input text. For example, the information analysis unit 110 analyzes information input from the information acquisition module 200. The sensor analysis unit 120 analyzes input sensor information. For example, the sensor analysis unit 120 analyzes sensor information acquired from the information acquisition module 200.

[0014] The prompt etc. generation unit 130 generates various types of information such as prompts to be input into an AI (Artificial Intelligence) model (also simply referred to as a "model"), which is a machine learning model described below. For example, the prompt etc. generation unit 130 generates various types of information required to generate a video (video) including prompts etc. to be input into the model. For example, the prompt etc. generation unit 130 generates information (also referred to as "support information") that supports a user who creates content.

[0015] For example, the prompt etc. generation unit 130 generates a prompt using a user input and a pre-saved prompt (template, etc.). Note that a prompt is merely an example of information (model input information) to be input to an AI model, and the model input information to be input to an AI model is not limited to a prompt, and any form of model input information can be adopted, and the "prompt etc. generation unit" may be read as a "model input information etc. generation unit." In FIG. 1 , the prompt etc. generation unit 130 includes a scenario-oriented generation unit 131, a video-oriented generation unit 132, a sound-oriented generation unit 133, a text / logo-oriented generation unit 134, a text answer generation unit 135, an example answer generation unit 136, and an example answer prompt generation unit 137.

[0016] The scenario-oriented generation unit 131 generates various information related to the generation of a scenario. The scenario-oriented generation unit 131 generates input information to be input to a model that outputs a scenario. For example, the scenario-oriented generation unit 131 generates scenario data related to video generation based on an input query. For example, the scenario-oriented generation unit 131 outputs scenario generation information used to generate scenario data based on the input query.

[0017] The video generation unit 132 generates various information related to the generation of videos. The video generation unit 132 generates input information to be input to a model that outputs code for constructing 3D (three-dimensional) data. For example, the video generation unit 132 generates code for constructing 3D data based on scenario data. For example, the video generation unit 132 outputs code generation information used to generate code for constructing 3D data based on the scenario data.

[0018] The sound generation unit 133 generates various information related to the generation of sound information (audio information). The sound generation unit 133 generates input information to be input to a model that outputs sound. The text / logo generation unit 134 generates various information related to the generation of text and logos. The text / logo generation unit 134 generates input information to be input to a model that outputs at least one of text and logos.

[0019] The text answer generation unit 135 is a processing unit that outputs support information for content creation based on background information related to content creation and content information being edited. The text answer generation unit 135 outputs the generated support information. In this way, the information processing system 1 can appropriately support content creation by outputting support information based on the background information and content information being edited.

[0020] The background information regarding content production includes information regarding the purpose of content creation, scenario information regarding the entire content and the selected cut, the content data itself in the process of being created, content information regarding the selected cut and the content before and after it, 3DCG models appearing in the selected cut and their position information, at least one of camerawork information, lighting information, sound information, color information, and overlay information used in the selected cut, at least one of motion data, background data, character data, and prop data held, and at least one of adjustable setting information.

[0021] Furthermore, content information is information relating to at least one of video data, animation, music data, and image data. Content information being edited is data included on the display screen of the content being edited. Support information is advice information for supporting editing in content production. Furthermore, for example, support information is evaluation information relating to the content information being edited.

[0022] The text answer generation unit 135 generates support information for content creation based on background information and content information being edited. The text answer generation unit 135 generates support information based on information about a user who creates content using the information processing system 1. The text answer generation unit 135 generates support information based on information input by the user. The text answer generation unit 135 generates support information based on a query from the user regarding the content information being edited. For example, the text answer generation unit 135 generates support information based on a query including a user's comment on the content information being edited.

[0023] The text answer generation unit 135 generates a prompt based on information about the user, and generates support information based on the prompt. The text answer generation unit 135 generates the support information using a machine learning agent that references background information about content production and content information being edited.

[0024] For example, the machine learning agent is an AI model (machine learning model) that outputs assistance information in response to a prompt input, such as model M1 described below. The machine learning agent is an AI model that has been fine-tuned from a pre-trained AI model based on prompts related to content creation. For example, the machine learning agent is fine-tuned based on information related to content creation created by a user. For example, the information related to content creation created by a user includes at least one of content information previously created by the user, material data created by the user, and text data created by the user. In this way, the information processing system 1 can appropriately assist content creation by using a fine-tuned AI model. Model M1, an example of a machine learning agent, outputs assistance information as text data.

[0025] The text answer generation unit 135 generates support information by inputting a prompt generated based on background information into an AI model such as model M1 and causing the AI ​​model such as model M1 to output support information. The text answer generation unit 135 generates support information by inputting a prompt generated based on information about a user who creates content and causing the AI ​​model to output support information. The text answer generation unit 135 generates support information by inputting a prompt generated based on a user's comment and causing the AI ​​model to output support information.

[0026] The example answer generation unit 136 outputs examples of content as support information based on background information related to content production and content information being edited. For example, the example answer generation unit 136 outputs examples of generated content as support information. For example, the example answer generation unit 136 inputs a prompt generated based on the background information into an AI model such as model M2, and causes the AI ​​model such as model M2 to output examples of content, thereby generating support information.

[0027] The example answer prompt generation unit 137 generates a prompt to be input to an AI model that outputs examples related to the content, based on background information related to content production and content information being edited. For example, the example answer prompt generation unit 137 outputs the generated prompt as an example answer prompt. For example, the example answer prompt generation unit 137 inputs the prompt generated based on the background information to an AI model such as model M3 and causes the AI ​​model such as model M3 to output examples related to the content, thereby generating an example answer prompt. Then, the example answer prompt generation unit 137 inputs the generated example answer prompt to an AI model such as model M4 and causes the AI ​​model such as model M4 to output examples related to the content, thereby generating support information.

[0028] The video generation unit 140 executes processing related to video generation. The video generation unit 140 acquires video data based on the code. The video generation unit 140 generates video using various information generated by the prompt generation unit 130. For example, the video generation unit 140 generates video data based on the code. Note that the video generation unit 140 may acquire video data in any manner. For example, the video generation unit 140 may acquire video data by transmitting data used to generate the video data to an external service providing device (such as a vendor) that provides a video data generation service, and receiving the video data generated by the service providing device from the service providing device. In FIG. 1 , the video generation unit 140 includes a USD generation unit 141, a rendering unit 142, and a video refinement unit 143.

[0029] The USD generation unit 141 generates various information related to a Universal Scene Description (USD). For example, the USD generation unit 141 generates USD-Python or the like using an AI model such as a Large Language Model (hereinafter also referred to as "LLM"), using a prompt obtained by video prompt generation in the video generation unit 132.

[0030] The rendering unit 142 executes various processes related to rendering, such as rendering the USD generated by the USD generation unit 141.

[0031] The image refinement unit 143 executes various processes for refining the image. The rendering unit 142 improves the quality of the generated image through image refinement processing. For example, the image refinement unit 143 executes image quality improvement processing to improve the image quality of the video data.

[0032] The sound generation unit 150 executes a process of generating sounds. The sound generation unit 150 generates sound information such as background music (BGM), sound effects (SE), narration, and dialogue using an AI model such as a contrastive learning model, using prompts obtained by the sound-oriented prompt generation in the sound-oriented generation unit 133.

[0033] The text / logo generator 160 executes a process for generating at least one of text and a logo. The text / logo generator 160 generates at least one of text and a logo using the information generated by the text / logo generator 134.

[0034] The image generation module 100, with the above-described configuration, generates a prompt for generating a scenario by combining a pre-stored prompt with a user's input. The image generation module 100 generates a scenario by inputting the generated prompt into an AI model such as an LLM. The image generation module 100 also generates prompts for generating images, sounds, and text / logos from the generated scenario. The image generation module 100 performs image generation, sound generation, and text / logo generation using prompts, scenarios, etc. for generating images, sounds, and text / logos.

[0035] The composite editing unit 170 executes processes related to editing, such as combining (combining) the generated video, sound, and text / logo into one video.

[0036] The evaluation unit 180 executes an evaluation process for evaluating various targets. The evaluation unit 180 evaluates information generated by the above-described configuration. For example, the evaluation unit 180 generates information indicating an evaluation of at least one of scenario data and video data. For example, the evaluation unit 180 generates information indicating an evaluation of content information being edited.

[0037] The client UI module 190 executes processing related to output on a UI (User Interface) on the client side. For example, the client UI module 190 generates various information related to output on the UI on the client side. In this case, the client UI module 190 executes processing to generate a UI to be displayed on the user side. The client UI module 190 generates various information to be displayed on the client UI display unit 400.

[0038] Furthermore, the information acquisition module 200 acquires various types of information. The information acquisition module 200 includes an information acquisition unit 210, a sensor acquisition unit 220, and the like. The information acquisition unit 210 is an acquisition unit that acquires information. For example, the information acquisition unit 210 acquires information used in processing from a storage unit (such as the memory / storage 14) of the information processing system 1. The information acquisition unit 210 acquires background information related to content production. The information acquisition unit 210 acquires content information currently being edited. The information acquisition unit 210 acquires text information input via the keyboard 320 or the microphone 330. For example, the information acquisition unit 210 acquires text information input by a user via the keyboard 320 or the microphone 330. For example, the information acquisition unit 210 acquires an input query related to video generation from a user.

[0039] The sensor acquisition unit 220 acquires information (also referred to as "sensor information") detected by a sensor such as a camera 340 or a motion capture device. The information acquisition module 200 provides (transmits) the acquired various pieces of information to the image generation module 100. Note that the information acquisition module 200 may be integrated with the image generation module 100.

[0040] The sensor unit 300 has various sensors. The sensor unit 300 senses user input. The sensor unit 300 accepts user operations. For example, the sensor unit 300 is a reception unit that accepts video editing operations from the user. For example, the sensor unit 300 has a mouse 310, a keyboard 320, a microphone 330, a camera 340, an IMU 350 which is an inertial measurement unit, and the like. In this way, the sensor unit 300 includes, in addition to the mouse 310 and the keyboard 320, a user terminal (such as a smartphone) equipped with the microphone 330, the camera 340, and the IMU 350, and sensors such as motion capture, and senses user input.

[0041] The client UI display unit 400 displays various information to be presented to the client (user). The client UI display unit 400 displays the UI generated by the client UI module 190 on a display (display device) of the client. For example, the client UI display unit 400 is a display control unit that displays a storyboard based on scenario data. The storyboard is configured to display video data for each cut of the video.

[0042] The information processing system 1 may have a hardware configuration as shown in Fig. 2. Fig. 2 is a diagram showing an example of a hardware configuration of the information processing system of the present disclosure. In Fig. 2, the information processing system 1 has, as its hardware configuration, a cloud-side computer 10, a client-side computer 20, etc. The information processing system 1 may also include an information providing device (computer) that provides information resources 30 such as learning data and an AI model 40 to the computer 10.

[0043] 2 is merely an example, and any hardware configuration can be adopted for the information processing system 1 as long as it can execute the desired processing. For example, the computer 10 and the computer 20 may be integrated. Furthermore, the information resource 30 and the AI ​​model 40 may be stored inside the computer 10.

[0044] The computer 10 includes a CPU (Central Processing Unit) 11, a GPU (Graphics Processing Unit) 12, a communication device 13, and a memory / storage 14. For example, the computer 10 corresponds to the image generation module 100 and the information acquisition module 200 in FIG. 1 . The computer 10 may be a service providing device (server device) that provides an image generation service. The CPU 11 and the GPU 12 are so-called processors, and execute calculations (arithmetic operations) related to various processes such as image generation.

[0045] The above is merely an example, and the computer 10 can have any configuration as long as it can perform the desired processing. For example, the computer 10 may use circuits such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array) to perform calculations (arithmetic processing) related to various processes, such as video display. Furthermore, the computer 10 may be configured so that programs are directly embedded in the processor circuitry instead of storing programs in memory (such as the memory / storage 14). In this case, the processor realizes its function by reading and executing the program embedded in the circuitry. Note that each processor in this embodiment is not limited to being configured as a single circuit, but may also be configured as a single processor by combining multiple independent circuits to realize its function. Similarly to the computer 10, the computer 20 can also have any configuration as long as it can perform the desired processing.

[0046] The communication device 13 is a communication device having a communication function for transmitting and receiving information to and from the computer 20, an information providing device, etc., and may be, for example, a communication circuit, a NIC (Network Interface Card), etc. The communication device 13 communicates with other devices such as the computer 20 and the information providing device via a predetermined network (such as the Internet). For example, the communication device 13 is connected to the predetermined network via a wired or wireless connection, and transmits and receives information to and from other devices such as the computer 20 and the information providing device.

[0047] The memory / storage 14 is a storage device that stores various types of information. The memory / storage 14 is, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 14 stores various types of information used for processing by processors such as the CPU 11 and the GPU 12. The memory / storage 14 may also store information resources 30, AI models 40, etc.

[0048] The computer 20 includes a CPU 21, a GPU 22, a communication device 23, and a memory / storage 24. For example, the computer 20 corresponds to the client UI display unit 400 in Fig. 1. The computer 20 may be a terminal device (such as a personal computer (PC) or a mobile device such as a smartphone) used by a user who uses the video generation service. The CPU 21 and the GPU 22 are so-called processors, and execute calculation processes (arithmetic processing) related to various processes such as video display.

[0049] The communication device 23 is a communication device having a communication function for transmitting and receiving information to and from the computer 10, etc., and may be, for example, a communication circuit, a NIC, etc. The communication device 23 communicates with other devices such as the computer 10 via a predetermined network (such as the Internet). For example, the communication device 23 is connected to the predetermined network via a wired or wireless connection, and transmits and receives information to and from other devices such as the computer 10.

[0050] The memory / storage 24 is a storage device that stores various types of information. The memory / storage 24 is, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. The memory / storage 24 stores various types of information that are used for processing by processors such as the CPU 21 and the GPU 22.

[0051] Furthermore, the information resource 30 includes various information such as training data. For example, the information resource 30 includes training data used for training various AI models such as LLM. The AI ​​model 40 includes information on AI models used for processing related to image generation such as LLM. For example, the AI ​​model 40 includes information on various AI models such as models M1 to M3 described below.

[0052] As described above, the information processing system 1 may have a configuration other than that shown in Fig. 2. For example, the information processing system 1 may have a hardware configuration (also referred to as a "sensor device") corresponding to the sensor unit 300 in Fig. 1.

[0053] The sensor device senses user input. The sensor device accepts user operations. Furthermore, for example, the computer 20 may have an IO interface, which is an input / output interface device. In this case, the computer 20 receives input from the sensor device via the IO interface. For example, the computer 20 receives input from an input device such as a keyboard or a mouse via the IO interface. Furthermore, the computer 20 outputs information from a display (display device) and a speaker (audio output device) via the IO interface. For example, the computer 20 plays video on a display and a speaker via the IO interface.

[0054] <1-2. First Example> A first example, which is an example of processing executed by the information processing system 1, will now be described. Note that the processing described with the information processing system 1 as the processing subject may be performed by any device capable of executing that processing, depending on the device configuration included in the information processing system 1. First, prior to describing the processing executed by the information processing system 1, an example of a user interface (UI) that provides support information to a user is shown in FIG. 3. Note that explanations of points similar to those described above will be omitted as appropriate. FIG. 3 is a diagram showing an example of a user interface.

[0055] 3, the information processing system 1 provides the user with content CT11. The content CT11 is a display screen for the user to use the information processing system 1 to create content such as video data.

[0056] For example, the information processing system 1 accepts user input information via a display screen such as that shown in content CT11. For example, the client UI display unit 400 displays a display screen such as that shown in content CT11. The content CT11 shows a state in which a comment CM12 containing support information generated by the information processing system 1 is displayed for a video MV11, which is content being produced by a user. FIG. 3 shows a state in which a cut selected by the user from the video MV11 is displayed. Also, in FIG. 3, a comment list TH11 including the comment CM12 shows a list of comments corresponding to a location (place) in the video MV11 to which an icon IC11 is attached. For example, the comment list TH11 is a chat thread (also simply referred to as a "thread") provided by a chat function.

[0057] Comment CM12 is advice information for supporting editing of video MV11 by a fictional creator who is given a predetermined personality as an AI Co-Creator (hereinafter also referred to as "AI creator"). In comment list TH11, comments by creators corresponding to creator names prefixed with "[AI]", that is, comment CM12 and comment CM14 correspond to comments by AI creators.

[0058] Assigning a personality to an AI creator will be described later. For example, the personality of an AI creator can be assigned by adding information corresponding to the AI ​​creator's personality (also referred to as "personality information") to the prompt input to the AI ​​model. Note that the assignment of a personality to an AI creator is not limited to adding personality information to the prompt, and any manner can be adopted as long as it is possible to assign a desired personality. For example, the personality of an AI creator can be assigned by using an AI model corresponding to each AI creator.

[0059] Thus, Figure 3 shows a case in which "Toru Kikuchi," an AI creator with a specific personality, gives advice to a user (hereinafter also referred to as "target user") who is editing video MV11, represented in Figure 3 as "Uesr X."

[0060] <1-2-1. First Processing> <1-2-1-1. Procedure of the First Processing> Hereinafter, an example of information processing (also referred to as "first processing") executed by the information processing system 1 in a first example will be described. First, the procedure of the first processing will be described using Fig. 4. Fig. 4 is a flowchart showing the procedure of the first processing.

[0061] 4, in the information processing system 1, a user selects a time from the timeline and a location within the video and starts commenting (step S101). For example, a target user who edits a video MV11 using the information processing system 1 selects an arbitrary location within the video MV11 and starts commenting, as shown in Fig. 5. For example, in the information processing system 1, a user selects a location to comment on by right-clicking and selecting a menu, and starts commenting.

[0062] In the information processing system 1, a user makes a comment (step S102). For example, a target user who edits a video MV11 using the information processing system 1 selects a location in the video MV11 marked with an icon IC11 to make a comment, as shown in FIG. 5. FIG. 5 is a diagram showing an example of a user interface. As shown in FIG. 5, the information processing system 1 provides the user with content CT12. The content CT12 is a display screen in a state in which the first comment from the target user has been received, i.e., a state prior to the content CT11.

[0063] Content CT12 shows a state in which a target user editing a video MV11 selects a location in the video MV11 marked with icon IC11, and a comment list TH11 containing the comment CM11 entered is displayed. In FIG. 5 , the target user designates the AI ​​creator "Toru Kikuchi" and enters a comment CM11, which is a question such as "Should this person's movements be made more natural?" In this way, the target user makes a comment by mentioning the AI ​​creator "Toru Kikuchi." The target user enters the comment CM11 by entering it into an input field located at the bottom of the comment list TH11.

[0064] As a result, the information processing system 1 accepts the comment CM11 from the user who mentioned the AI ​​creator. Note that the information processing system 1 may accept the user's comment in any manner. For example, the manner of accepting input from the user is not limited to the above, and the information processing system 1 may accept the user's comment by voice input. Furthermore, for example, the information processing system 1 may accept the user's comment by the user specifying a specific time and correction location on the timeline and asking a question in text.

[0065] 5, a state in which an @ symbol is added to the beginning of a creator name indicates that the creator corresponding to that creator name has been mentioned, but mentioning is not limited to adding an @ symbol and may be performed in any manner. The information processing system 1 waits until a user makes a comment (step S102). That is, in the processing of the information processing system 1, step S102 may be a conditional branch based on whether or not a user comment has been accepted. If a user comment has been accepted, the processing of step S103 is executed. If a user comment has not been accepted, the processing waits until a user comment is accepted.

[0066] The information processing system 1 that has received the user's comment determines whether the user's comment mentions the AI ​​Co-Creators (step S103). If the user's comment does not mention the AI ​​Co-Creators (step S103: No), the information processing system 1 returns to step S102 and repeats the process.

[0067] If the user's comment mentions the AI ​​Co-Creators (step S103: Yes), the AI ​​Co-Creators tuned to the specified characters provide a text response based on the video information (step S104). For example, if the user's comment mentions at least one AI creator, the information processing system 1 generates a text response to the user's comment based on the personality of the mentioned AI creator. As a result, the information processing system 1 generates a text response according to the AI ​​creator specified by the target user.

[0068] For example, as shown in Fig. 6 , the information processing system 1 generates and outputs a comment CM12 that is a text response to the target user's comment CM11 based on the personality of "Toru Kikuchi," an AI creator mentioned by the target user. Fig. 6 is a diagram showing an example of a user interface. As shown in Fig. 6 , the information processing system 1 provides the user with content CT13. The content CT13 is a display screen in a state in which the comment CM12 including support information from the AI ​​creator in response to the target user's first comment is output (displayed).

[0069] Content CT13 shows a state in which a comment list TH11 including a comment CM12 including support information in the personality of the AI ​​creator "Toru Kikuchi" in response to a comment CM11 of the target user is displayed. In Fig. 6, the information processing system 1 mentions and outputs support information based on the personality of the AI ​​creator "Toru Kikuchi" to the target user. The information processing system 1 inputs a prompt generated based on the comment CM11 of the target user to the model M1 and causes the model M1 to output a text response including the support information, thereby generating a text response including the support information.

[0070] The comment CM12 includes support information, which is a response from "Toru Kikuchi," an AI creator mentioned by the target user. For example, the comment CM12 includes support information, which is advice information for supporting editing of the video MV11, such as, "Your movements are a little stiff. If you turn around more naturally, viewers will be able to watch it more smoothly without being distracted." This allows the information processing system 1 to display the comment CM12, which includes support information, which is advice information for supporting editing of the video MV11, to the target user. The comment CM12 also includes information, such as "Shall I try making one?", which asks the target user whether or not to generate a sample answer (also referred to as "sample answer necessity confirmation information"). This allows the information processing system 1 to confirm with the target user whether or not to generate a sample answer corresponding to the support information included in the comment CM12, which is advice information for supporting editing of the video MV11.

[0071] The information processing system 1 determines whether or not the user has requested creation of an example answer for the question (step S105). If the user has not requested creation of an example answer (step S105: No), the information processing system 1 returns to step S102 and repeats the process.

[0072] If the user has requested the creation of an example answer (step S105: Yes), the information processing system 1 has the AI ​​Co-Creators tuned to the specified character create an example answer based on the video information (step S106), and then returns to step S102 to repeat the process. For example, if the user's comment includes information requesting the creation of an example answer, the information processing system 1 creates an example answer for the user's comment. Furthermore, if the information processing system 1 has provided the target user with advice information to assist editing and confirmed whether or not an example answer corresponding to the advice information needs to be created, and the user has responded to the confirmation with a request, the information processing system 1 creates an example answer.

[0073] 3, the information processing system 1 receives a comment CM13 from the target user requesting an example answer such as "Yes, please" in response to a comment CM12 including confirmation information for confirming whether an example answer is required such as "Shall I try making one?" In this way, the information processing system 1 receives the comment CM13 from the target user requesting the generation of an example answer corresponding to advice information for supporting the editing of the video MV11.

[0074] The information processing system 1 then executes a generation process to generate an example answer based on the personality of "Toru Kikuchi," the AI ​​creator mentioned by the target user. After completing the generation of the example answer, the information processing system 1 outputs a comment CM14 such as "I made it. Please check it out." Then, as shown in FIG. 7 , the information processing system 1 outputs an example answer EA11 including multiple videos. FIG. 7 is a diagram illustrating an example of a user interface. As shown in FIG. 7 , the information processing system 1 provides content CT14 to the user. The content CT14 is a display screen on which the example answer EA11 including multiple videos edited to have a natural turning-around movement based on the support information in the comment CM12 is output (displayed). For example, the information processing system 1 inputs a prompt generated based on the support information in the comment CM12 into the model M2 and causes the model M2 to output the example answer EA11, thereby generating the example answer EA11. In this manner, the information processing system 1 generates an example answer when a user requests an example answer. As shown in FIG. 7, the example response EA11 is displayed in a different area from the image MV11 (to the side of the image MV11 in FIG. 7).

[0075] As described above, the information processing system 1 provides the target user with support information that reflects the personality of the AI ​​creator. For example, if the target user mentions the AI ​​creator, the information processing system 1 automatically generates a text response. For example, if the target user requests the creation of a video, motion, or the like, the information processing system 1 generates an example response. For example, when generating a response, the information processing system 1 generates an example response based on various background information, including video information, etc. Note that if the target user mentions another user who is a human creator rather than the AI ​​creator, the information processing system 1 does not generate a text response or an example response based on the personality of the AI ​​creator.

[0076] <1-2-1-2. Overview of the First Processing> The overview of the first processing described above will be explained using Fig. 8. Fig. 8 is a diagram showing an overview of the first processing.

[0077] The information processing system 1 generates information used to generate prompts using data DT11 including user comments, data DT12 including comments from past AI creators, data DT13 including preset prompts, and data set DT14. For example, data DT11 includes past user comments in the same thread. For example, data DT12 includes past comments from AI creators in the same thread. For example, data DT13 includes data that serves as a template for prompts to be input into an AI model. For example, data DT14 includes at least one of content creation information BIN11 to BIN19, etc.

[0078] Content creation information BIN11 is video creation purpose information that indicates the purpose of video creation. For example, content creation information BIN11 is video creation purpose information that indicates the purpose of the target user creating content. Content creation information BIN12 is scenario information of the entire scenario and selected cuts. For example, content creation information BIN12 is scenario information of the entire scenario of the video for which the target user is creating content and selected cuts.

[0079] The content creation information BIN13 is information about the entire video. For example, the content creation information BIN13 is video data being edited by the target user. The content creation information BIN14 is information about the selected cut and the video (images) before and after it. For example, the content creation information BIN14 is information about the cut selected by the target user and the cuts before and after that cut.

[0080] The content creation information BIN 15 is information such as the name and location information of a 3DCG model that appears in the selected cut. For example, the content creation information BIN 15 is information such as the name and location information of a 3DCG model such as a USD file of a person that appears in the cut selected by the target user. The content creation information BIN 16 is information such as camerawork information, lighting, sound, color, overlay, etc. that is used in the selected cut. For example, the content creation information BIN 16 is information such as camerawork information, lighting information, sound information, color information, overlay information, etc. that is used in the cut selected by the target user.

[0081] The content creation information BIN 17 is screen information currently being edited, such as cuts, characters, sounds, etc. For example, the content creation information BIN 17 is information on cuts, characters, sounds, etc. that the target user is editing. The content creation information BIN 18 is motion data, background data, character data, prop data, etc. that are held. For example, the content creation information BIN 18 is motion data, background data, character data, prop data, etc. that are held by the information processing system 1.

[0082] Content creation information BIN19 is information on adjustable settings such as camera work. For example, content creation information BIN19 is information on adjustable settings such as camera work related to a video created by a target user. Note that content creation information BIN11 to BIN19 are merely examples, and the content creation information may include various information other than content creation information BIN11 to BIN19. For example, the content creation information includes AI creator personality information that indicates the personality settings of the AI ​​creator.

[0083] Note that data DT14 uses at least one piece of content production information such as content production information BIN11 to BIN19. For example, the information processing system 1 selects, as data DT14, the most appropriate piece of content production information such as content production information BIN11 to BIN19, depending on the screen information being edited. Also, for example, the information processing system 1 selects, as data DT14, the most appropriate piece of content production information such as content production information BIN11 to BIN19, depending on the content of the comment made by the target user. Also, for example, the information processing system 1 selects, as data DT14, the most appropriate piece of content production information such as content production information BIN11 to BIN19, depending on the temporal and physical location of the video selected by the target user.

[0084] The information processing system 1 uses at least one of the data DT11 to DT14 to generate a prompt PT11, which is a prompt for generating a text answer. The information processing system 1 uses at least one of the data DT11 to DT14 to generate a prompt PT11, which is a prompt for generating a text answer that is used as input for the model M1. For example, the model M1 is an AI model that outputs information including support information, which is advice information for supporting content editing, in response to the input of a prompt for generating a text answer. Any AI model, such as an LLM (large-scale language model), can be used for the model M1 as long as it is capable of producing the desired output in response to the input.

[0085] 8 , the information processing system 1 generates a text response OT11 including support information by inputting a prompt PT11 into a model M1 and causing the model M1 to output a text response OT11 that includes support information. For example, the information processing system 1 inputs a prompt PT11 including AI creator personality information of an AI creator mentioned by a user into the model M1, thereby generating a text response OT11 including support information that reflects the personality of the AI ​​creator mentioned by the user. The information processing system 1 then displays the generated text response OT11. For example, the information processing system 1 arranges the text response OT11 on a display screen such as that shown in content CT11, and displays it on the client UI display unit 400.

[0086] The information processing system 1 uses at least one of the data DT11 to DT14 to generate a prompt PT12, which is a prompt for generating an example answer. The information processing system 1 uses at least one of the data DT11 to DT14 to generate a prompt PT12, which is a prompt for generating an example answer, to be used as input for the model M2. For example, the model M2 is an AI model such as a motion generation model that outputs an example answer such as a video in response to the input of a prompt for generating an example answer. Any AI model, such as an LLM (large-scale language model), can be used for the model M2 as long as it is capable of producing the desired output in response to the input.

[0087] 8 , the information processing system 1 generates the example answer OT12 including images, etc., by inputting a prompt PT12 into a model M2 and causing the model M2 to output an example answer OT12 that is information including images, etc. For example, the information processing system 1 inputs the prompt PT12 including a comment from an AI creator mentioned by a user into the model M2, thereby generating the example answer OT12 including images, etc. that reflect the personality of the AI ​​creator mentioned by the user. The information processing system 1 then displays the generated example answer OT12. For example, the information processing system 1 arranges the example answer OT12 on a display screen such as that shown in content CT14, etc., and displays it on the client UI display unit 400.

[0088] The information processing system 1 uses at least one of the data DT11 to DT14 to generate a prompt PT13, which is a prompt for generating a prompt for generating an example answer. The information processing system 1 uses at least one of the data DT11 to DT14 to generate a prompt PT13, which is a prompt for generating a prompt for generating an example answer, to be used as input for the model M3. For example, the model M3 is an AI model that outputs a prompt for generating an example answer in response to input of a prompt for generating a prompt for generating a prompt for generating an example answer. Any AI model, such as an LLM (large-scale language model), can be used for the model M3 as long as it is capable of producing a desired output in response to the input.

[0089] 8, the information processing system 1 generates the prompt PT14 by inputting the prompt PT13 to the model M3 and causing the model M3 to output the prompt PT14, which is a prompt for generating an example answer. For example, the information processing system 1 generates the prompt PT14 that reflects the personality of the AI ​​creator mentioned by the user by inputting the prompt PT13, which includes the AI ​​creator personality information of the AI ​​creator mentioned by the user, to the model M3.

[0090] The information processing system 1 then uses the generated prompt PT14, which is a prompt for generating an example answer, as input for model M4. For example, model M4 is an AI model such as a motion generation model that outputs an example answer such as a video in response to the input of a prompt for generating an example answer. Any AI model such as an LLM (large-scale language model) can be used for model M4 as long as it is capable of producing the desired output in response to the input. For example, model M4 may be the same AI model as model M2.

[0091] 8 , the information processing system 1 generates the example answer OT13 including images, etc., by inputting a prompt PT14 into a model M4 and causing the model M4 to output an example answer OT13 that is information including images, etc. For example, the information processing system 1 inputs the prompt PT14 that reflects the personality of an AI creator into the model M4, thereby generating the example answer OT13 including images, etc. that reflect the personality of the AI ​​creator mentioned by the user. The information processing system 1 then displays the generated example answer OT13. For example, the information processing system 1 arranges the example answer OT13 on a display screen such as that shown in content CT14, etc., and displays it on the client UI display unit 400.

[0092] Note that when providing an example answer to a user, the information processing system 1 may generate only one of the example answer OT12 and the example answer OT13. For example, when providing the example answer OT12 to a user, the information processing system 1 may not perform processing related to the generation of the example answer OT13, and may not generate, for example, prompts PT13, PT14, and the example answer OT13. For example, when providing the example answer OT13 to a user, the information processing system 1 may not perform processing related to the generation of the example answer OT12, and may not generate, for example, prompt PT12 and the example answer OT12.

[0093] Here, an example of a prompt generation process executed by the information processing system 1 will be described with reference to FIG. 9 . FIG. 9 is a diagram illustrating an example of a prompt generation process. For example, FIG. 9 is a diagram illustrating an example of a prompt generation process that reflects the personality of an AI creator. For example, as shown in FIG. 9 , the information processing system 1 uses first information INF1 and second information INF2 to generate a prompt PT15 that reflects the personality of the AI ​​creator "Toru Kikuchi."

[0094] The first information INF1 includes AI creator personality information indicating that the AI ​​creator "Toru Kikuchi" is a film director. The first information INF1 also includes information requesting a user-based evaluation of a video, reflecting the personality of the AI ​​creator "Toru Kikuchi." The first information INF1 also includes information requesting output of a comment on the video, reflecting the personality of the AI ​​creator "Toru Kikuchi." The first information INF1 also includes constraint information indicating a character limit for a text response. The first information INF1 also includes constraint information indicating, for example, that an example response should be provided if instructed to provide an example response. For example, the first information INF1 is a tuning prompt. The information processing system 1 may use the first information INF1 as data DT13 including a preset prompt.

[0095] The second information INF2 includes information on copyrighted works such as videos, scripts, scenarios, and music for video works that the AI ​​creator "Toru Kikuchi" has produced to date. The second information INF2 also includes information on materials such as articles, PDFs, and presentation (slide) data that the AI ​​creator "Toru Kikuchi" has created to date. The second information INF2 also includes information on emails and chats that the AI ​​creator "Toru Kikuchi" has exchanged to date. The second information INF2 also includes history information such as comment data that the AI ​​creator "Toru Kikuchi" has made to date on other people's projects.

[0096] In this way, the second information INF2 is a dataset input as information indicating the personality assigned to the AI ​​creator. Note that the second information INF2 includes various information related to the personality assigned to the AI ​​creator, including but not limited to the above. For example, the second information INF2 may include name, age, gender, occupation, experience, educational background, awards, technical skills, personality traits, work strengths, daily life, career goals, challenges faced, representative works, industry reputation, works and directors that influenced the AI ​​creator, current activities, speaking characteristics, speaking examples, etc.

[0097] 9 , the information processing system 1 generates a prompt PT15 by adding the second information INF2 to the first information INF1. For example, the information processing system 1 generates the prompt PT15 as a prompt for generating a text answer. The information processing system 1 generates the prompt PT15 as a prompt for generating an example answer. In this case, the information processing system 1 generates the prompt for generating an example answer by, for example, changing the first line of the constraint in the first information INF1 from "Please output a comment for the INPUT comment" to "Please output an example answer for the INPUT comment."

[0098] Next, an example of the support information generation process executed by the information processing system 1 will be described with reference to Fig. 10 . Fig. 10 is a diagram showing an example of the support information generation process. For example, Fig. 10 is a diagram showing an example of the support information generation process for a video selected and commented on by a user. For example, as shown in Fig. 10 , the information processing system 1 generates a prompt PT21 using a dataset DT21 including video information such as comments CM21 from the target user on the selected video (cut), the purpose of the video, scenario information, etc.

[0099] For example, the information processing system 1 generates a prompt PT21 using a cut image CN21, which is an image of a cut selected by the target user (the second cut in FIG. 10 ), and a template TP21, which is a prompt template (preset prompt). For example, the template TP21 may be selected based on a mention. Although not shown in FIG. 10 , the prompt PT21 may include AI creator personality information indicating the personality of the AI ​​creator. For example, the prompt PT21 is a tuned prompt that reflects the personality of the AI ​​creator. The prompt PT21 may also include instruction information, constraint information, and the like, which indicate specific instructions regarding output, as shown in FIG. 9 .

[0100] 10 , the information processing system 1 generates a text response OT21 including support information by inputting a prompt PT21 into a model M1 and causing the model M1 to output a text response OT21 that is information including support information. For example, the information processing system 1 inputs a prompt PT21 including a user's comment into the model M1, thereby generating a text response OT21 that includes support information that serves as advice in response to the user's comment. The information processing system 1 then displays the generated text response OT21. For example, the information processing system 1 arranges the text response OT21 on a display screen such as that shown in content CT11, and displays it on the client UI display unit 400.

[0101] <1-3. Second Example> Note that the above-described first example is merely one example of processing executed by the information processing system 1, and the information processing system 1 may execute various processing, not limited to the processing shown in the first example. For example, the AI ​​creator may make a comment without receiving a mention (instruction) from the user. An example of this point will be described below as a second example. Note that the information processing system 1 may execute processing that combines the first processing and the second processing. Furthermore, explanations of points similar to those described in the first example and the like will be omitted as appropriate.

[0102] <1-3-1. Second Processing> <1-3-2-1. Procedure of Second Processing> Hereinafter, an example of information processing (also referred to as "second processing") executed by the information processing system 1 in a second example will be described. First, the procedure of the second processing will be described using FIG. 11 and FIG. 12. FIG. 11 and FIG. 12 are flowcharts showing the procedure of the second processing. For example, FIG. 11 is a flowchart showing the procedure of processing to output a comment by an AI creator not based on a mention (instruction) from a user. For example, FIG. 12 is a flowchart showing the procedure of processing to output a comment by an AI creator based on a mention (instruction) from a user. Note that explanations of points similar to those in the first processing described above will be omitted as appropriate. For example, the processing shown in FIG. 11 and the processing shown in FIG. 12 are executed in parallel.

[0103] 11, the information processing system 1 determines whether it is time to post a comment (step S201). If it is not time to post a comment (step S201: No), the information processing system 1 repeats the process of step S201.

[0104] If it is time to make a comment (step S201: Yes), the information processing system 1 has the AI ​​Co-Creators make a comment (text response) (step S202), and the process returns to step S201 to repeat. For example, if the information processing system 1 determines that there is a point in the content being edited by the user that should be commented on, it generates a comment based on the personality of the AI ​​creator.

[0105] In addition, when an AI creator has multiple personalities, the information processing system 1 generates a comment based on the personality of the AI ​​creator selected based on predetermined criteria. For example, when an AI creator has multiple personalities, the information processing system 1 generates a comment based on the personality of an AI creator that the user has previously mentioned. For example, the information processing system 1 generates a comment including support information that reflects the personality of the AI ​​creator.

[0106] The information processing system 1 outputs the generated comment. For example, the information processing system 1 arranges icons indicating locations corresponding to the generated comment on the display screen and displays them on the client UI display unit 400. For example, for video MV21, the information processing system 1 makes comments corresponding to locations in video MV21 marked with icon IC21 and locations marked with icon IC22, as shown in Fig. 13. Fig. 13 is a diagram showing an example of content.

[0107] 13 , if there are multiple places in the content where comments should be made, the information processing system 1 may make multiple comments. For example, when a user selects icon IC21 in video MV21, the information processing system 1 displays a comment list corresponding to icon IC21 on the client UI display unit 400. For example, when a user selects icon IC22 in video MV21, the information processing system 1 displays a comment list corresponding to icon IC22 on the client UI display unit 400.

[0108] In the process shown in Fig. 12, the user makes a comment in the information processing system 1 (step S301). The user may select an arbitrary part in the content and make a comment, or may make a comment in response to a comment made by the information processing system 1 in the process shown in Fig. 11. In the process shown in Fig. 12, the information processing system 1 waits until the user makes a comment (step S301). That is, in the process of the information processing system 1, step S301 may be a conditional branch based on whether or not a user comment has been accepted. If a user comment has been accepted, the process of step S302 is executed, and if a user comment has not been accepted, the process waits until a user comment is accepted.

[0109] The information processing system 1 that has received the user's comment determines whether the user's comment mentions the AI ​​Co-Creators (step S302). If the user's comment does not mention the AI ​​Co-Creators (step S302: No), the information processing system 1 returns to step S301 and repeats the process.

[0110] If the user's comment mentions the AI ​​Co-Creators (step S302: Yes), the AI ​​Co-Creators tuned to the specified characters provide a text response based on the video information (step S303). For example, if the user's comment mentions at least one AI creator, the information processing system 1 generates a text response to the user's comment based on the personality of the mentioned AI creator. As a result, the information processing system 1 generates a text response according to the AI ​​creator specified by the target user.

[0111] The information processing system 1 determines whether or not the user has requested creation of an example answer for the question (step S304). If the user has not requested creation of an example answer (step S304: No), the information processing system 1 returns to step S301 and repeats the process.

[0112] If the user has requested the creation of an example answer (step S304: Yes), the information processing system 1 causes the AI ​​Co-Creators tuned to the specified character to create an example answer based on the video information (step S305), and the process returns to step S301 to repeat. For example, if the user's comment includes information requesting the creation of an example answer, the information processing system 1 creates an example answer for the user's comment. Furthermore, if the information processing system 1 confirms with the target user whether or not an example answer corresponding to the advice information needs to be created along with advice information to assist editing, and if the user has responded to the confirmation with a request, the information processing system 1 creates an example answer.

[0113] As described above, the information processing system 1 provides support information that reflects the personality of the AI ​​creator to the target user. For example, the information processing system 1 automatically evaluates a video and outputs a comment, reflecting the personality of the AI ​​creator, even without a mention. For example, the information processing system 1 evaluates content such as a video using scenario information and images. For example, the information processing system 1 evaluates at least one of the good and bad parts and displays a comment indicating the evaluation in the corresponding location.

[0114] For example, the information processing system 1 generates evaluation information such as, for the video music video 21 in Figure 13, "Since the target users of this product are Japanese women, it would be more natural for viewers to have at least one Asian performer in the cast." Furthermore, for the video music video 21 in Figure 13, the information processing system 1 generates evaluation information such as, "In this case, rather than turning around to pick up the speaker, it would be more natural to pick up the speaker that was originally placed on the sofa." For example, the information processing system 1 generates evaluation information using the above-described model M1. The information processing system 1 generates evaluation information by inputting a prompt including information indicating the evaluation target to the model M1 and causing the model M1 to output evaluation information.

[0115] <1-3-2-2. Overview of the second process> The second process described above is similar to the first process except that the information processing system 1 outputs comments regardless of user comments. Therefore, a detailed description of the overview will be omitted, but the process shown in Fig. 14 will be described below. Fig. 14 is a diagram showing an example of a process for generating support information. Fig. 14 shows an example of a prompt used when generating evaluation information, which is an example of support information.

[0116] For example, as shown in FIG. 14, the information processing system 1 generates a prompt PT31 using a dataset DT31 including a comment CM31, which is a preset comment instructing the evaluation of a cut in a video, and video information such as the purpose of the video and scenario information.

[0117] For example, the information processing system 1 generates a prompt PT31 using a cut image CN31, which is an image of a cut to be evaluated (the second cut in FIG. 14 ), and a template TP31, which is a prompt template (preset prompt). For example, the template TP31 may be selected based on a mention. Although not shown in FIG. 14 , the prompt PT31 may include AI creator personality information indicating the personality of the AI ​​creator. For example, the prompt PT31 is a tuned prompt that reflects the personality of the AI ​​creator. The prompt PT31 may also include instruction information, constraint information, and the like, which indicate specific instructions regarding output, as shown in FIG. 9 .

[0118] 14 , the information processing system 1 generates a text response OT31 including evaluation information by inputting a prompt PT31 into a model M1 and causing the model M1 to output a text response OT31, which is information including evaluation information. For example, the information processing system 1 generates a text response OT31 including evaluation information indicating an evaluation of the evaluation target by inputting a prompt PT31 including an evaluation target into the model M1. The information processing system 1 then displays the generated text response OT31. For example, the information processing system 1 arranges the text response OT31 on a display screen such as that shown in content CT11, and displays it on the client UI display unit 400.

[0119] The information processing system 1 outputs comments including evaluation information of various evaluation contents at any timing. For example, the information processing system 1 generates and outputs comments including evaluation information regarding video information for each editing screen. For example, in the case of a cut editing screen, the information processing system 1 generates and outputs comments including evaluation information regarding cuts in the video. For example, in the case of a sound editing screen, the information processing system 1 generates and outputs comments including evaluation information regarding sound.

[0120] Furthermore, the information processing system 1 generates and outputs a comment including evaluation information indicating whether the video meets the purpose. For example, the information processing system 1 generates and outputs a comment including evaluation information indicating whether the video has overall consistency. For example, the information processing system 1 generates and outputs a comment including evaluation information indicating whether the content is of high quality, such as composition, camerawork, lighting, and color.

[0121] Furthermore, the information processing system 1 outputs comments including evaluation information at timing according to the user's actions (operations). For example, if the user repeatedly watches the same cut, the information processing system 1 determines that the user feels uncomfortable with that cut, executes an evaluation process, and generates and outputs a comment including the evaluation information. For example, if the user edits the video themselves rather than correcting it based on a suggestion (comment) from an AI creator, the information processing system 1 executes an evaluation process, generates and outputs a comment including the evaluation information.

[0122] Furthermore, when a user is viewing a specific cut, the information processing system 1 executes an evaluation process for N cuts (N is any integer equal to or greater than 1) before and after the specific cut, and generates and outputs a comment including evaluation information. Furthermore, the information processing system 1 may execute the evaluation process, generate and output a comment including evaluation information when the user is not performing any operation. Furthermore, the information processing system 1 may execute the evaluation process, generate and output a comment including evaluation information when the user has logged out of a content editing application (software) and is not performing any operation.

[0123] Furthermore, when a video contains unnatural expressions, the information processing system 1 may execute an evaluation process and generate and output a comment including evaluation information. For example, when people collide with each other or with an object such as a wall in the 3DCG, the information processing system 1 may execute an evaluation process and generate and output a comment including evaluation information.

[0124] The information processing system 1 may also have a critical level (e.g., a threshold value for evaluation) and output the critical level (e.g., a threshold value for evaluation) at the same time as making an evaluation. The information processing system 1 may also express the critical level by changing the color of the comment or by emitting a different sound.

[0125] Furthermore, the information processing system 1 may display a different color, size, shape, or other aspect of the icon of a comment that the user has already confirmed from that of a comment that the user has not confirmed. For example, the information processing system 1 may make the icon of a comment that the user has already confirmed lighter in color, thereby making it easier for the user to determine whether the comment has been confirmed in the past.

[0126] The information processing system 1 may also receive feedback from the user regarding the evaluation information. In this case, the information processing system 1 may, for example, display buttons that allow the user to select whether the evaluation indicated by the evaluation information is good or bad, together with the evaluation information, and when the user selects a button, acquire information indicating the user's preferences (tastes) as user preference information. For example, the information processing system 1 may use the acquired user preference information to generate a prompt that reflects the user's preferences.

[0127] In this way, by using prompts that reflect the user's preferences, the information processing system 1 can provide appropriate information according to the user. For example, the information processing system 1 may use the acquired user preference information to train an AI model such as model M1. For example, the information processing system 1 may use the acquired user preference information to fine-tune model M1.

[0128] <1-4. Other Examples> The first and second examples described above are merely examples, and the information processing system 1 is not limited to the above and may perform various information processes and may provide any UI. In this regard, several examples are described below.

[0129] For example, the information processing system 1 may perform processing to adapt a specific portion of a generated video to a video in response to a user operation, as shown in Fig. 15. Fig. 15 is a diagram showing an example of a user interface. For example, Fig. 15 shows an example of a UI when adapting only motion to a main video by dragging and pasting. Note that the content CT21 shown in Fig. 15 is similar to the content CT14 shown in Fig. 7 except that it includes motion information MT21 for adapting a specific portion to a video, and a description of this point will be omitted as appropriate.

[0130] A user can drag a portion of a video, such as a specified motion or lighting, and paste it onto the main video to replace only that portion of the main video. Figure 15 shows a case in which a target user applies the motion of the top example answer (also referred to as the "source answer") of the example answers EA11 to video MV11. When the information processing system 1 drags the motion information MT21 of the source answer of the example answers EA11 and pastes it onto video MV11, the information processing system 1 applies the motion corresponding to the motion information MT21 of the source answer to video MV11. This changes (updates) the motion in video MV11 to the motion corresponding to the motion information MT21 of the source answer.

[0131] The information processing performed by the information processing system 1 may also be used on an actual site (such as a filming site). For example, the information processing system 1 may acquire information from the filming scene, the filmed video, a storyboard, and a schedule, and evaluate the video. The information processing system 1 may output the information as a comment in the rush edit during filming, or may utter it from a voice interface such as a voice service providing device.

[0132] Furthermore, in the information processing system 1, the user may communicate with the AI ​​creator not only through a chat function such as a comment list, but also through voice. For example, when the information processing system 1 inputs a comment prompt into the AI ​​model, the comment prompt may be input together with video information or separately. For example, when the user inputs a comment or when the user enters the editing screen, the information processing system 1 may first input the portion other than the user's input comment (also referred to as "INPUT comment") into the AI ​​model, and then input only the INPUT comment portion into the AI ​​model later, thereby shortening the processing time.

[0133] Furthermore, in the information processing system 1, learning of the AI ​​creator (personality) may be performed in any manner. For example, the AI ​​creator (personality) may learn through comments in a project, but may reset data learned through previous chats along the way. For example, the AI ​​creator (personality) may manage versions learned through comments, allowing the AI ​​creator (personality) to return to the stage at which it learned up to the expected comments.

[0134] Furthermore, in the information processing system 1, the response of the AI ​​creator may be changed depending on the level of understanding of the user (human creator). For example, the information processing system 1 may change the AI ​​model that inputs a prompt depending on a setting value. For example, the information processing system 1 may change the prompt that is input to the AI ​​model depending on a setting value.

[0135] Furthermore, the information processing system 1 may allow a comment to specify not only one point on the timeline and one location but also a change over time. For example, when a user selects the position of a hand, the information processing system 1 may continue to specify the position of the hand over time based on the position information of the 3DCG hand bones, even if the position changes over time.

[0136] Furthermore, the information processing system 1 may not only allow a user to specify a location by using a comment, but may also allow the user to change the range from point to finger to hand to arm to body. For example, the information processing system 1 may accept such a change in range by the user scrolling with a mouse. Furthermore, for example, the information processing system 1 may automatically select a range that has been segmented in advance.

[0137] Furthermore, the information processing system 1 may not only share all chat threads with all users (all human creators), but may also make them available in private mode for interaction between only one user (human creator) and the AI ​​creator. The information processing system 1 may also provide a function for users (human creators) and the AI ​​creator to take a majority vote on presented example answers.

[0138] The above-described UI is merely an example, and the information processing system 1 may provide a UI as shown in Fig. 16. Fig. 16 is a diagram showing an example of a user interface.

[0139] 16, the information processing system 1 provides the user with content CT31. The content CT31 is a display screen for the user to use the information processing system 1 to create content such as video data.

[0140] For example, the information processing system 1 accepts information input by the user via a display screen such as that shown in content CT31, etc. For example, the client UI display unit 400 displays a display screen such as that shown in content CT31, etc.

[0141] 16, the user selects the motion correction area LN31 in the content CT31 and selects the Activities tab in the motion correction area in the comment list TH31. The user then makes a comment mentioning the AI ​​creator "Kenta Nishimura." As a result, the information processing system 1 generates a prompt (also called a "prompt PTX") that corresponds to the user's comment and reflects the personality of the AI ​​creator "Kenta Nishimura."

[0142] The information processing system 1 then inputs the prompt PTX to the model M1 and causes the model M1 to output a comment (also referred to as a "comment CMX"), which is a text response containing information about the support information, thereby generating a comment CMX containing the support information. For example, the information processing system 1 inputs the prompt PTX containing the AI ​​creator personality information of "Nishimura Kenta," an AI creator mentioned by a user, to the model M1, thereby generating a comment CMX containing support information that reflects the personality of "Nishimura Kenta," the AI ​​creator mentioned by the user. The information processing system 1 then displays the generated comment CMX. For example, the information processing system 1 arranges the comment CMX on a display screen such as that shown in content CT11 and displays it on the client UI display unit 400. In this way, when a comment mentioning an AI creator is made, support information including evaluation information or (suggested) advice is output.

[0143] As described above, the information processing system 1 outputs support information based on text, image, and video information previously input by a user at a production site (including a real-world location or within a PC tool) for producing video, photos, music, etc., and on text, image, and video information generated based on the text, image, and video information previously input by the user. For example, the information processing system 1 generates text responses and suggested changes to video, photos, music, etc., in response to a user's text input, to improve the quality of the production, using a personalized (personalized) AI creator.

[0144] For example, the information processing system 1 allows a personalized (personalized) AI creator to automatically evaluate a production based on text, images, videos, and audio information previously input by the user, text, images, videos, and audio information generated based on these, and user behavior. For example, the information processing system 1 leaves comments in the form of text, videos, photos, music, etc. at the points of criticism, specifying the location and time.

[0145] As described above, the information processing system 1 provides a service for creating videos together with co-creators (co-creators) including AI creators. For example, the information processing system 1 executes information processing related to the creation of a vision, a story, and a video. For example, the information processing system 1 enables a user to be a co-creator, and allows the user to create a desired video by adjusting it themselves or through conversations with the co-creator, rather than by writing and modifying prompts themselves.

[0146] For example, the information processing system 1 enables co-producers, including the user, to create a vision together. For example, in the information processing system 1, once a vision is created, the co-producer automatically generates a story and video. For example, in the information processing system 1, the created story and video can be checked and revised together with the co-producer to complete the project.

[0147] For example, the information processing system 1 provides a video editing tool (application) for co-creation among co-creators, including an AI creator. For example, in the tool provided by the information processing system 1, the AI ​​creator and the user (human) participate in the project on an equal footing. For example, in the information processing system 1, the AI ​​creator is personalized and registered as a video director, a boss, yourself, etc., and each co-creator creates and comments on the content.

[0148] For example, in the information processing system 1, a user can specify co-producers by mentioning them, collect ideas, have them submit multiple ideas, and then decide on a policy or the like from among them. For example, the information processing system 1 outputs an evaluation of the content. In the information processing system 1, a user can request other co-producers to submit ideas. For example, in the information processing system 1, a request can be made in text along with specific revision locations and times.

[0149] For example, other co-creators can rate videos and leave comments in the information processing system 1. Note that, even when an AI creator makes a comment in the information processing system 1, the behavior of the sudden comment is the same as when another user (human) makes a comment, which is easy for users of the information processing system 1 to accept.

[0150] For example, the information processing system 1 accepts input of video creation purpose information, which is information indicating the purpose of the user's video creation. The video creation purpose information is information that indicates in text what kind of video the user wants to create. For example, video creation purpose information is useful because, unless it is translated into a linguistic level, the client cannot understand what the user wants to create, and neither can an AI creator or other users.

[0151] The video creation purpose information includes information such as the purpose of the video, the message to be conveyed through the video, the target audience, the features of the function / service, the length of the video, the aspect ratio, the content style (emotional, informational, etc.), and the visual style (cinematic, animated, etc.). For example, a user may create the video creation purpose information in consultation with other collaborators.

[0152] The storyboard described above contains information that allows users to understand the overall picture at a glance, including scenes and cuts, and is the central page for editing. The storyboard allows users to modify the flow of cuts / scenes, add scenes / cuts, and delete scenes / cuts, rather than focusing on the details of each cut. For example, in the information processing system 1, users can improve the quality of their storyboards by chatting with co-creators and receiving feedback.

[0153] Furthermore, for the above-mentioned cuts (scenes), for example, the user can closely examine each cut and modify the characters, background, props, camera, lighting, etc. For example, the user can place the created / modified objects in "Character" and "Prop" by marking the characters and necessary props they want to fix in the cut (scene) with "+". Furthermore, any changes the user wants to make directly within the cut can be made through chat with co-producers. Furthermore, if the user makes a change that affects the before and after the cut they are editing, the co-producer can add an evaluation comment. Furthermore, the user can improve the quality of the video of the cut by chatting with the co-producer and receiving evaluations.

[0154] Furthermore, for the above-mentioned characters, for example, the user creates a character image by importing many photos or using a generation AI, etc. In the information processing system 1, the output picture (image) is a picture (image) that has passed through a refiner. Furthermore, the user improves the quality of the character by chatting with co-creators and receiving evaluations.

[0155] Regarding the props mentioned above, for example, a user can import a 3D CG image of a product or other prop. In the information processing system 1, the output picture (image) is a picture (image) that has passed through a refiner. The user can improve the quality of the prop by chatting with co-creators and receiving evaluations.

[0156] Regarding the above-mentioned overlay, for example, users can add text, logos, etc. to the videos they create using overlays. Users can also improve the quality of their overlays by chatting with co-creators and receiving evaluations.

[0157] Regarding the sound mentioned above, for example, users can add audio information such as background music, sound effects, dialogue, and narration to the videos they create. For background music, users can modify the music by setting "Genre," "Instrument," "Mood / Theme," and "Intensity Level" along the timeline. Users can also improve the quality of the sound by chatting with co-creators and receiving feedback.

[0158] Furthermore, the information processing system 1 may provide a UI that allows direct interaction with various objects, not just with co-creators including AI creators. An example of this is shown in Fig. 17. Fig. 17 is a diagram illustrating an example of a user interface.

[0159] 17, the information processing system 1 provides the user with content CT41. The content CT41 is a display screen for the user to use the information processing system 1 to create content such as video data.

[0160] For example, the information processing system 1 accepts information input by the user via a display screen such as that shown in content CT41, etc. For example, the client UI display unit 400 displays a display screen such as that shown in content CT41, etc.

[0161] For example, the information processing system 1 may allow the user to have a conversation with a person (AI) related to the video MV41, such as a cast member appearing in the video MV41 in the content CT41 or a cameraman who shoots the video MV41. In this case, the information processing system 1 updates the video MV41 based on the conversation between the user and the person (AI) related to the video MV41.

[0162] For example, when a user selects an icon CR41 indicating a cameraman in content CT41, the information processing system 1 accepts input of a comment for the cameraman (also referred to as an "AI cameraman") corresponding to the icon CR41 selected by the user. In FIG. 17, when a user selects an icon CR41, input of a comment for the cameraman (AI cameraman) corresponding to that icon CR41 is accepted. FIG. 17 shows a case where the user inputs a comment to the cameraman, such as "Take the camera at a low angle." As a result, the information processing system 1 generates a prompt (also referred to as a "prompt PTP") that corresponds to the user's comment and reflects the personality of the AI ​​cameraman.

[0163] The information processing system 1 then inputs a prompt PTP to the model M1 and causes the model M1 to output a comment (also referred to as a "comment CMP"), which is a text response containing information including support information, thereby generating a comment CMP containing support information. For example, the information processing system 1 inputs a prompt PTP containing personality information of an AI cameraman selected (designated) by the user to the model M1, thereby generating a comment CMP containing support information reflecting the personality of the AI ​​cameraman selected (designated) by the user. The information processing system 1 then displays the generated comment CMP. For example, the information processing system 1 displays a comment CMP such as "It's a low-angle shot. How about something like this, with the room's space captured in the background?" along with the user's input, "The camera is at a low-angle angle." The information processing system 1 also updates the video MV41 to reflect the comment CMP. In FIG. 17 , the information processing system 1 updates the video MV41 to a low-angle shot capturing the room's space in the background.

[0164] In this way, when a user selects the icon CR41 representing the cameraman and makes a comment, the information processing system 1 outputs support information (such as a comment CMP) including evaluation information or (suggested) advice in accordance with the comment, and updates the video MV 41. This allows the user to work while communicating with the production staff (also called "AI staff") of the video MV 41, allowing the user to generate the video with the feeling of actually working together.

[0165] For example, when a user selects a cast member CS41 in a video MV41, the information processing system 1 accepts input of a comment for the cast member CS41 selected by the user. In FIG. 17, when a user selects a cast member CS41, input of a comment for that cast member CS41 is accepted. FIG. 17 shows a case where a user inputs a comment for the cast member CS41, saying, "Can you make a sadder expression?" As a result, the information processing system 1 updates the video MV41 to content that reflects the user's comment, "Can you make a sadder expression?" In FIG. 17, the information processing system 1 updates the video MV41 to content that corresponds to the user's comment by changing the expression of the cast member CS41 to a sadder expression.

[0166] The information processing system 1 may also generate a prompt (also referred to as a "prompt PTS") that corresponds to a user's comment and reflects the personality of the cast member CS41. In this case, the information processing system 1 inputs the prompt PTS to the model M1 and causes the model M1 to output a comment (also referred to as a "comment CMS"), which is a text response that includes support information, thereby generating a comment CMS that includes support information. For example, the information processing system 1 inputs a prompt PTS that includes personality information of the cast member CS41 selected (specified) by the user to the model M1, thereby generating a comment CMS that includes support information that reflects the personality of the cast member CS41 selected (specified) by the user. The information processing system 1 then displays the generated comment CMS. For example, the information processing system 1 displays the comment CMS "Is it like this?" along with the comment CMS input by the user, "Can you make a sadder expression?"

[0167] In this way, when a user selects a cast CS41 and comments on it, the information processing system 1 makes changes to the cast CS41 in response to the comment and updates the video MV41. Furthermore, the information processing system 1 may output support information (such as a comment CMS) including evaluation information or (suggested) advice in response to the user's comment. This allows the user to work while communicating with the production staff (AI staff) of the video MV41, allowing the user to create the video with the feeling of actually working together.

[0168] The content CT41 also includes an icon group CL41 indicating Co-Creators (joint creators) including an AI creator. Fig. 17 shows a case where an icon group CL41 including four icons corresponding to four Co-Creators (joint creators) is displayed.

[0169] For example, when a user selects one icon from the icon group CL41, the information processing system 1 accepts input of a comment for the Co-Creators selected by the user. In FIG. 17, when a user selects the icon of the AI ​​creator (also referred to as "AI creator Z") with the name "AAAA" at the top of the icon group CL41, input of a comment for AI creator Z is accepted. In FIG. 17, a case is shown in which the user inputs a comment to AI creator Z, such as "What do you think of this person's motion?" As a result, the information processing system 1 generates a prompt (also referred to as "prompt PTZ") that corresponds to the user's comment and reflects the personality of AI creator Z.

[0170] The information processing system 1 then inputs the prompt PTZ to the model M1 and causes the model M1 to output a comment (also referred to as a "comment CMZ"), which is a text response that is information including support information, thereby generating a comment CMZ including support information. For example, the information processing system 1 inputs the prompt PTZ, which includes AI creator personality information of the AI ​​creator Z mentioned by the user, to the model M1, thereby generating a comment CMZ including support information that reflects the personality of the AI ​​creator Z mentioned by the user. The information processing system 1 then displays the generated comment CMZ. For example, the information processing system 1 displays a comment CMZ such as "It would be easier to understand if you made the movements a little gentler" along with the question, "What do you think of this person's motion?" input by the user.

[0171] In this way, when a user selects an AI creator's icon and makes a comment, the information processing system 1 outputs support information (such as a comment CMZ) including evaluation information or (suggested) advice based on the personality of the AI ​​creator. This allows the user to work while being aware of the co-creators, and thus allows the user to generate video as if they were talking to the icons representing the co-creators.

[0172] Furthermore, the information processing system 1 may intervene with the user at any timing. An example of this point will be described with reference to FIG. 18 . FIG. 18 is a diagram showing an example of the timing of intervention. For example, the information processing system 1 may determine the timing of intervention with the user based on a condition such as the timing condition TL41 in FIG. 18 .

[0173] 18 is generated based on the user's answer to the following question, for example: For example, the information processing system 1 asks the user the following question #1, and generates the timing condition TL41 based on the answer from the user.

[0174] Question #1: "What are some things you would check to determine when to intervene with a human? However, you are in a position to review their work and provide suggestions for improvement as you go. These suggestions may come up repeatedly. If you were an AI, how would you determine when to intervene with a human?"

[0175] The timing condition TL41 in FIG. 18 includes a timing condition for intervention based on understanding the user's situation and context. The timing condition TL41 in FIG. 18 also includes a timing condition for intervention based on analysis of the user's behavioral patterns. The timing condition TL41 in FIG. 18 also includes a timing condition for intervention based on real-time feedback. The timing condition TL41 in FIG. 18 also includes a timing condition for intervention based on user preferences and settings. The timing condition TL41 in FIG. 18 also includes a timing condition for intervention based on determination of urgency and importance.

[0176] The information processing system 1 determines the timing of intervention based on the timing condition TL41 in Fig. 18. For example, the information processing system 1 determines the timing of issuing a comment in step S201 in Fig. 11 based on the timing condition TL41 in Fig. 18. Note that the timing condition TL41 in Fig. 18 is merely an example, and the timing condition is not limited to that shown in Fig. 18 and may include various conditions related to the timing of intervention.

[0177] Furthermore, the information processing system 1 may use an AI model such as LLM to solve concepts like Calm Technology. An example of this point will be described using FIG. 19 . FIG. 19 is a conceptual diagram showing an example of the application of Calm Technology. For example, the information processing system 1 may perform processing based on the Calm Technology concept as shown in FIG. 19 . Calm Technology here refers to a concept related to the design of information technology that blends into the user's life. For example, Calm Technology is information technology that effectively conveys information while minimizing the user's attention, allowing the user to naturally interact with technology. For example, Calm Technology is technology that allows users to use digital devices without being aware of their presence, while still achieving effective functionality.

[0178] For example, in FIG. 19 , information such as user operation information and current setting values ​​in a system as shown in the first element EL1 is shared with co-creators as shown in the second element EL2. For example, in FIG. 19 , the co-creators as shown in the second element EL2 determine whether to intervene. For example, the co-creators determine whether to intervene based on the information shared by the first element EL1 and the presence of a user-like entity, such as onLLM, as shown in the third element EL3. The third element EL3 includes information such as a definition of the user's personality (considered similarly to the co-creators) and timing trends learned from past user operations (e.g., whether or not attention was paid to suggestions from the co-creators).

[0179] The second element EL2 acquires the results of the intervention (e.g., whether or not the Co-Creators' suggestions were attended to). The second element EL2 reflects the acquired results of the intervention (e.g., whether or not the Co-Creators' suggestions were attended to) in the first element EL1. The second element EL2 updates the setting values ​​of the first element EL1 based on the acquired results of the intervention (e.g., whether or not the Co-Creators' suggestions were attended to). This allows the information processing system 1 to be updated to a state appropriate for the user's personality and past behavior. Therefore, the information processing system 1 can effectively convey information while minimizing the user's attention and allow the user to naturally engage with technology. In other words, the information processing system 1 can be used without the user being aware of the presence of a digital device, while still achieving effective functionality.

[0180] <1-5. Regarding AI Models> Note that the AI ​​models used in the above-described processes are not limited to the examples described in each section, and any internal structure can be adopted as long as desired information can be output in response to input. Any combination of input, output, and internal structure of the AI ​​model can be adopted as long as desired information can be output.

[0181] The input of the AI ​​model may be text, images, audio, 3D data, etc., or a combination thereof. The output of the AI ​​model may be text, images, audio, 3D data, etc. Note that the above-mentioned inputs and outputs are merely examples, and the above-mentioned AI model may have any inputs and outputs.

[0182] Furthermore, the internal structure of the AI ​​model can be any structure depending on the combination of input and output. In other words, the internal structure of the AI ​​model can be any structure as long as it can produce a desired output for the input.

[0183] For example, the AI ​​model may have a structure related to Transformer. For example, the AI ​​model may have a structure related to Transformer and perform processing taking into account context, such as context within data, such as text or time-series data. For example, the AI ​​model may have a self-attention mechanism. For example, the AI ​​model may have any attention mechanism, such as single-head attention or multi-head attention. Note that the AI ​​model does not necessarily have to have an attention mechanism.

[0184] The AI ​​model may have a mechanism for extracting features from an input. For example, the AI ​​model may have an encoder. The AI ​​model may have a mechanism for generating information based on the extracted features. For example, the AI ​​model may have a decoder.

[0185] The AI ​​model may have a structure related to a convolutional neural network (CNN). For example, when processing an image, the AI ​​model may have a structure related to a CNN. For example, the AI ​​model may have at least one of a convolution layer, a pooling layer, a fully connected layer, etc.

[0186] The above-described internal structure is merely an example, and the AI ​​model may have any internal structure. For example, the AI ​​model may have a skip connection. Furthermore, the AI ​​model may have a structure related to a diffusion model.

[0187] Furthermore, the above-described AI model may be generated (trained) by any learning process. The AI ​​model may be a machine learning model trained using any machine learning method. For example, the AI ​​model may be a model generated based on a so-called Foundation Model by fine-tuning the Foundation Model to apply it to a specific task (e.g., generating a text answer, generating an example answer, etc.). For example, an AI model such as the above-described LLM may be a model generated by fine-tuning the Foundation Model to apply it to a specific task.

[0188] The base model here is a model that has been trained to be applicable to various tasks, for example, to be able to perform a wide variety of tasks. For example, the base model is a neural network that has been pre-trained with a large amount of unlabeled data set. Note that the base model may have any structure, such as a Transformer-based architecture. For example, the base model is generated by self-supervised learning using data without correct answer labels. As described above, the base model is fine-tuned so that it can be adapted to a wide range of downstream tasks.

[0189] For example, when applied to a task of generating a text answer, the base model is fine-tuned so that it can be adapted to the task of generating a text answer, and an AI model (model M1, etc.) applied to the task of generating a text answer is generated. Also, when applied to a task of generating an example answer, the base model is fine-tuned so that it can be adapted to the task of generating an example answer, and an AI model (model M2, etc.) applied to the task of generating an example answer is generated.

[0190] For example, an AI model (such as model M1) applied to a task of generating a text answer is trained using training data including a combination of input information corresponding to the AI ​​model and a text answer (also referred to as "correct answer information") that is the correct output when the input information is input. The training data, such as the input information and correct answer information, may be data created by a person or data automatically generated by a computer that generates the training data. For example, the text answer that is the correct answer information may be data created by a person. Below, model M1 will be briefly described as an example. For example, model M1 is trained to output correct answer information corresponding to each piece of input information in the training data when that input information is input. For example, model M1 is trained by adjusting (correcting) parameters (connection coefficients) using a technique such as backpropagation (error backpropagation) so as to reduce the error between the output of model M1 when certain input information is input and the correct answer information corresponding to that input information. Other AI models, such as the AI ​​model (such as model M2) applied to the task of generating a text answer, may also be trained using a similar training process.

[0191] The above-described learning process is merely an example, and the above-described AI model may be trained by any learning process depending on the input, output, and internal structure of the AI ​​model. For example, the AI ​​model may be trained using an unsupervised learning method such as a generative adversarial network (GAN). The AI ​​model may also be trained in a distributed state without aggregating data, such as federated learning. In this case, each video generation service device (e.g., server) may generate local models collected by the service, and a server (aggregation server) that aggregates information (e.g., parameters) of the local models generated by each video generation service device (e.g., server) may generate a global model using the information on the local models. In this case, the information processing system 1 may receive the global model generated by the aggregation server from the aggregation server and use the received global model as an AI model for processing.

[0192] In this way, the above-described AI models may be generated (learned) by any computer. That is, the learning process for generating the AI ​​models may be performed by any device (computer, etc.) in the information processing system 1, or may be performed by a device outside the information processing system 1. For example, when a device outside the information processing system 1 generates at least one of the above-described AI models, the information processing system 1 acquires the AI ​​model from the device outside the information processing system 1 and performs processing using the acquired AI model.

[0193] 2. Other Embodiments The processing according to each of the above-described embodiments may be implemented in various different forms (modifications) other than the above-described embodiments and modifications.

[0194] <2-1. Other Configuration Examples> The above-described configuration of the information processing system 1 is merely an example, and any desired division of functions in the information processing system 1 may be adopted. In other words, the above-described configuration is merely an example, and the information processing system 1 may have any desired division of functions and any desired configuration as long as it can provide the above-described service related to image generation. For example, the information processing system 1 may be configured by a single device (e.g., a computer) that performs the above-described processing. In this case, one device of the information processing system 1 may have the functions of the image generation module 100, the information acquisition module 200, the sensor unit 300, and the client UI display unit 400. For example, the image generation service provided by the information processing system 1 may be provided to a user as a program such as a tool (AI Assist Creation Tool) that runs on a terminal device (e.g., the computer 20) used by the user.

[0195] <2-2. Others> Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using a known method. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.

[0196] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.

[0197] Furthermore, the above-described embodiments and modifications can be combined as appropriate within the scope of not causing any contradiction in the processing content.

[0198] Furthermore, the effects described in this specification are merely examples and are not limiting, and other effects may also be present.

[0199] 3. Hardware Configuration An information processing device (information equipment) having the image generation module 100, information acquisition module 200, client UI display unit 400, etc. according to each of the above-described embodiments is realized by, for example, a computer 1000 configured as shown in FIG. 20 . An information processing device having the image generation module 100 will be described as an example. FIG. 20 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device. The computer 1000 has a processing circuitry 1100, a RAM 1200, a ROM 1300, a secondary storage device 1400, a communication interface 1500, an input / output interface 1600, a display unit 1700, a camera unit 1800, a microphone 1900, and a speaker 2000. The components of the computer 1000 are connected by a bus 1050.

[0200] The processing circuit 1100 operates and controls each unit based on programs stored in the ROM 1300 or the secondary storage device 1400. For example, the processing circuit 1100 loads the programs stored in the ROM 1300 or the secondary storage device 1400 into the RAM 1200 and executes processing corresponding to the various programs.

[0201] The ROM 1300 stores boot programs such as a basic input output system (BIOS) that is executed by the processing circuit 1100 when the computer 1000 is started up, and programs that depend on the hardware of the computer 1000 .

[0202] The secondary storage device 1400 is a computer-readable recording medium that non-temporarily records programs executed by the processing circuit 1100 and data used by such programs. Specifically, the secondary storage device 1400 is a recording medium that records programs for each process of an information processing device having the image generation module 100 according to an embodiment, which are an example of program data 1450.

[0203] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550. The communication interface 1500 corresponds to a communication unit (communication device) provided in an information processing device having the image generation module 100. For example, the processing circuit 1100 receives data from other devices and transmits data generated by the processing circuit 1100 to other devices via the communication interface 1500.

[0204] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the processing circuit 1100 receives data from an input device such as a microphone 1900 or a touch panel via the input / output interface 1600. The processing circuit 1100 also transmits data to an output device such as a display unit 1700 or a speaker 2000 via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of the media include optical recording media such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), magneto-optical recording media such as an MO (Magneto-Optical Disk), tape media, magnetic recording media, and semiconductor memories.

[0205] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 is, for example, a liquid crystal display or an organic electroluminescence display (EL display). The display unit 1700 may also be a touch panel display device or a video projection device.

[0206] The camera unit 1800 is an interface through which the computer 1000 captures images. The microphone 1900 is an interface through which the computer 1000 captures audio. The speaker 2000 is an interface through which the computer 1000 outputs audio processed by the computer 1000. The components of the computer 1000 are connected by a bus 1050. The interfaces do not necessarily need to be provided inside the computer 1000, but may be provided outside the computer 1000 via a network or the like. Furthermore, the components constituting the computer 1000 may be controlled by a circuit different from the processing circuit 1100. For example, the display unit 1700 may be controlled not by the processing circuit 1100 but by a circuit dedicated to display processing provided in the display unit 1700.

[0207] For example, when the computer 1000 functions as an information processing device having an image generation module 100 according to an embodiment, the processing circuit 1100 of the computer 1000 functions as the image generation module 100 by executing a program loaded onto the RAM 1200. The secondary storage device 1400 stores the information processing program according to the present disclosure and various data stored in the storage unit of the information processing device. The processing circuit 1100 reads and executes program data 1450 from the secondary storage device 1400. Alternatively, the processing circuit 1100 may obtain these programs from another device via an external network 1550. That is, the secondary storage device 1400 does not need to be located inside the computer 1000, but may also be located outside the computer 1000. The processing circuit 1100 is an example of an integrated circuit, and a CPU, MPU, GPU, APU, ASIC, and FPGA can all be considered integrated circuits.

[0208] The present technology can also be configured as follows. (1) An information processing system including: an acquisition unit that acquires background information related to content creation and content information being edited; and a processing unit that outputs support information related to the content creation based on the background information and the content information being edited. (2) The information processing system described in (1), in which the processing unit outputs the support information based on information about a user who creates the content using the information processing system. (3) The information processing system described in (2), in which the processing unit outputs the support information based on information input by the user. (4) The information processing system described in (3), in which the processing unit outputs the support information based on a query from the user related to the content information being edited. (5) The information processing system described in any one of (2) to (4), in which the processing unit generates a prompt based on information about the user, and outputs the support information based on the prompt. (6) The information processing system according to any one of (1) to (5), wherein the processing unit outputs the support information using a machine learning agent that references background information related to the content production and information about the content being edited. (7) The information processing system according to (6), wherein the machine learning agent is a machine learning model that outputs the support information in response to a prompt input, and the processing unit receives a prompt generated based on the background information and causes the machine learning model to output the support information, thereby generating the support information. (8) The information processing system according to (7), wherein the processing unit receives a prompt generated based on information about a user who produces the content, and causes the machine learning model to output the support information, thereby generating the support information. (9) The information processing system according to (8), wherein the processing unit receives a prompt generated based on a comment from the user, and generates the support information by causing the machine learning model to output the support information.(10) The information processing system according to any one of (6) to (9), wherein the machine learning agent is a machine learning model obtained by fine-tuning a pre-trained machine learning model based on prompts related to content creation. (11) The information processing system according to (10), wherein the machine learning agent is fine-tuned based on information related to content creation created by a user. (12) The information processing system according to (11), wherein the information related to content creation created by the user includes at least one of content information created in the past by the user, material data created by the user, and text data by the user. (13) The information processing system according to any one of (6) to (12), wherein the machine learning agent outputs the support information as text data. (14) The information processing system according to any one of (1) to (13), wherein the content information is information related to video data, animation, music data, or image data. (15) The information processing system according to any one of (1) to (14), wherein the background information related to content production includes at least one of information on the purpose of content creation, scenario information related to the entire content and a selected cut, the content data itself in the middle of production, content information on the selected cut and the parts before and after it, 3DCG models appearing in the selected cut and their position information, at least one of camerawork information, lighting information, sound information, color information, and overlay information used in the selected cut, at least one of retained motion data, background data, character data, and prop data, and adjustable setting information. (16) The information processing system according to any one of (1) to (15), wherein the content information being edited is data included on a display screen of the content being edited. (17) The information processing system according to any one of (1) to (16), wherein the support information is advice information for supporting editing of the content production. (18) The information processing system according to any one of (1) to (17), wherein the support information is evaluation information related to the content information being edited.(19) An information processing method including: acquiring background information related to content production and content information currently being edited; and outputting support information related to the content production based on the background information and the content information currently being edited. (20) An information processing program causing a computer to execute the following operations: acquiring background information related to content production and content information currently being edited; and outputting support information related to the content production based on the background information and the content information currently being edited.

[0209] 1 Information processing system 100 Video generation module 110 Information analysis unit 120 Sensor analysis unit 130 Prompt etc. generation unit 131 Scenario generation unit 132 Video generation unit 133 Sound generation unit 134 Text / logo generation unit 135 Character answer generation unit 136 Example answer generation unit 137 Example answer prompt generation unit 140 Video generation unit 141 USD generation unit 142 Rendering unit 143 Video refinement unit 150 Sound generation unit 160 Text / logo generation unit 170 Composite editing unit 180 Evaluation unit 190 Client UI module 200 Information acquisition module 210 Information acquisition unit 220 Sensor acquisition unit 300 Sensor unit 400 Client UI display unit

Claims

1. An information processing system comprising: an acquisition unit that acquires background information related to content production and content information being edited; and a processing unit that outputs support information related to the content production based on the background information and the content information being edited.

2. The information processing system according to claim 1, wherein the processing unit outputs the support information based on information relating to a user who creates the content using the information processing system.

3. The information processing system according to claim 2, wherein the processing unit outputs the support information based on information input by the user.

4. The information processing system according to claim 3, wherein the processing unit outputs the support information based on a query from the user regarding the content information being edited.

5. The information processing system according to claim 2, wherein the processing unit generates a prompt based on information about the user and outputs the support information based on the prompt.

6. The information processing system according to claim 1, wherein the processing unit outputs the support information using a machine learning agent that references background information related to the content production and information about the content being edited.

7. The information processing system according to claim 6, wherein the machine learning agent is a machine learning model that outputs the support information in response to a prompt input, and the processing unit generates the support information by inputting a prompt generated based on the background information and causing the machine learning model to output the support information.

8. The information processing system of claim 7, wherein the processing unit generates the support information by inputting a prompt generated based on information about the user who creates the content and having the machine learning model output the support information.

9. The information processing system according to claim 8, wherein the processing unit generates the support information by inputting a prompt generated based on the user's comment and causing the machine learning model to output the support information.

10. The information processing system according to claim 6, wherein the machine learning agent is a machine learning model that is fine-tuned based on a pre-trained machine learning model and prompts related to content creation.

11. The information processing system of claim 10, wherein the machine learning agent is fine-tuned based on information about user-created content creation.

12. The information processing system according to claim 11, wherein the information relating to the creation of content created by the user includes at least one of content information created in the past by the user, material data created by the user, and text data by the user.

13. The information processing system according to claim 6, wherein the machine learning agent outputs the assistance information as text data.

14. The information processing system according to claim 1, wherein the content information is information relating to video data, animation, music data, and image data.

15. The information processing system of claim 1, wherein the background information regarding the content production includes information regarding the purpose of content creation, scenario information regarding the entire content and the selected cut, the content data itself in the process of being created, content information regarding the selected cut and the content before and after it, 3DCG models appearing in the selected cut and their position information, at least one of camerawork information, lighting information, sound information, color information, and overlay information used in the selected cut, at least one of motion data, background data, character data, and prop data held, and at least one of adjustable setting information.

16. The information processing system according to claim 1, wherein the content information being edited is data included on a display screen of the content being edited.

17. The information processing system according to claim 1, wherein the support information is advice information for supporting editing of the content production.

18. The information processing system according to claim 1, wherein the support information is evaluation information regarding the content information being edited.

19. An information processing method comprising: acquiring background information related to content production and content information currently being edited; and outputting support information related to the content production based on the background information and the content information currently being edited.

20. An information processing program that causes a computer to acquire background information related to content production and information about the content being edited, and output support information related to the content production based on the background information and the content information being edited.

Citation Information

Patent Citations

  • Method and system for content generation, computing device and storage medium

    CN117115303A

  • Image editing method and device, electronic equipment and storage medium

    CN117765117A

  • Video editing material generation method, client, server and system

    CN117835008A