Generation program, generation method, and information processing device
The information processing device enhances video production by using a parent agent and child agents to analyze and generate videos based on target audience personas, addressing inefficiencies in conventional video production systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2024-10-30
- Publication Date
- 2026-05-07
AI Technical Summary
Conventional technologies in video production face challenges in generating effective videos efficiently due to limited contextual understanding, making it difficult to propose specific improvements based on analysis results, especially when dealing with large amounts of video material or long production times.
An information processing device employs a parent agent and multiple child agents, including a persona-simulating agent and a video generation agent, to analyze user-created videos, evaluate them based on target audience personas, and generate improved videos and related text by integrating video analysis and generation modules.
The system effectively generates enhanced videos and accompanying text by simulating target audience perspectives, addressing the inefficiencies of conventional methods and improving video production quality and efficiency.
Smart Images

Figure JP2024038781_07052026_PF_FP_ABST
Abstract
Description
Generation program, generation method, and information processing device.
[0001] The present invention relates to a generation program, a generation method, and an information processing device.
[0002] In recent years, advancements in LMM (Large Multi-modal Model) technologies such as GPT®-4o and Gemini®-1.5 Pro have led to remarkable improvements in the image and video comprehension capabilities of information processing devices. This improved image and video comprehension capability enables information processing devices to perform practical tasks such as generating captions and performing visual question answering (VQA) related to input images and videos.
[0003] For example, one practical task involves the use of visual prompts for images. A visual prompt is a visual instruction written directly on an image by the user. By using visual prompts for images, LMM can perform image comprehension and visual quality analysis (VQA) under conditions that focus on specified areas.
[0004] Zeyi Sun, Ye Fang, Tong Wu, Pan Zhang, Yuhang Zang, Shu Kong, Yuanjun Xiong, Dahua Lin, Jiaqi Wang, “Alpha-CLIP: A CLIP Model Focusing on Wherever You Want”, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 13019-13029
[0005] In video production, when there is a large amount of video material or the production is long, editing and review require a great deal of time and effort. Therefore, there is a need to produce more effective videos within a limited time. However, the conventional technology described above has limited contextual understanding, making it difficult to propose specific improvements based on the analysis results. Consequently, the conventional technology described above cannot generate effective videos.
[0006] In one aspect, the objective is to provide a generation program, generation method, and information processing device that can generate effective images.
[0007] In the first proposal, the generation program is characterized in that it causes a computer to acquire a first video to be analyzed and a request for the first video, and when the first video and the request are input to a first agent that generates information according to the input information, it identifies an agent that simulates a persona and an agent that generates a second video from among a plurality of second agents that can cooperate with the first agent, the first video and the request are input to the agent that simulates a persona, which generates an evaluation result that evaluates the first video based on the request and the persona, the generated evaluation result and the first video are input to the agent that generates the second video, which generates the second video based on the generated evaluation result, and the first agent outputs the second video and text related to the second video as a response to the request based on the second video.
[0008] According to one embodiment, effective video can be generated.
[0009] Figure 1 is a diagram illustrating the information processing device according to Embodiment 1. Figure 2 is a functional block diagram showing the functional configuration of the information processing device according to Embodiment 1. Figure 3 is a flowchart illustrating the processing of the parent agent. Figure 4 is a diagram illustrating an example of a prompt given to the parent agent. Figure 5 is a diagram illustrating an example of prompt information set for the child agent. Figure 6 is a diagram illustrating an example of prompt information set for the child agent. Figure 7 is a diagram illustrating an example of persona information. Figure 8 is a diagram illustrating an example of video generation. Figure 9 is a diagram illustrating an example of evaluation generation. Figure 10 is a diagram illustrating an overview of the related area estimation process by the related area estimation unit. Figure 11 is a diagram illustrating video generation (before question input) according to Embodiment 2. Figure 12 is a diagram illustrating video generation (when question input) according to Embodiment 2. Figure 13 is a diagram illustrating an example of hardware configuration.
[0010] The following describes in detail, with reference to the drawings, embodiments of the generation program, generation method, and information processing device according to the present invention. However, the present invention is not limited by these embodiments. Each embodiment can be combined as appropriate within a non-consistent range.
[0011] <Description of Information Processing Device> Figure 1 is a diagram illustrating the information processing device 10 according to Embodiment 1. The information processing device 10 shown in Figure 1 is an example of a computer device that executes AI agents (hereinafter simply referred to as agents) and generates and outputs responses in cooperation with each agent in response to user requests. This embodiment will be explained as an example in which a video created by a user is evaluated and the video is regenerated (updated) based on the evaluation results.
[0012] An agent is a software program that collects data, uses that data to perform self-determined tasks, and achieves predetermined goals. For example, an agent might proactively generate responses using a pre-trained machine learning model or a large-scale language model. The agent independently selects and executes the optimal actions necessary to achieve the goals set by the administrator or other relevant parties.
[0013] As shown in Figure 1, the information processing device 10 executes a parent agent 330, which is an example of a first agent, and multiple child agents (child agent 330X, child agent 330Y), which are examples of a second agent. Note that each child agent may also have multiple instances, for example, child agent 330X may have child agent 330X1 and child agent 330X2.
[0014] The parent agent 330 obtains video requests and videos created by the user for analysis from the user, and requests each child agent to evaluate the video according to the requests. The parent agent 330 obtains the evaluations from the child agents and requests the child agents to recreate the video according to the evaluations. The parent agent 330 then outputs the video generated by the child agents, along with text indicating the intention and perspective from which the video was generated, to the user as a response to the requests.
[0015] Child agent 330X is an agent that analyzes video using a video analysis module or the like. Specifically, child agent 330X is an agent that simulates a persona, and it evaluates the video and requests from the parent agent 330 from the perspective of the simulated persona, and outputs the evaluation results to the parent agent 330. Multiple child agents 330X may be prepared for each persona.
[0016] Child agent 330Y is an agent that generates video using video modules such as multimodal. Specifically, child agent 330Y obtains the video created by the user and the evaluation results of child agent 330X from parent agent 330, and outputs the video and text generated using these to parent agent 330.
[0017] In this system configuration, the information processing device 10 acquires a first video to be analyzed and a request for the first video. When the information processing device 10 receives the first video and the request as input to the parent agent 330, which generates information according to the input information, it identifies two child agents from among several child agents that can cooperate with the parent agent: a child agent 330X that simulates a persona and a child agent 330Y that generates a second video. The information processing device 10 uses the first video and the request as input and causes the child agent 330X that simulates a persona to execute, thereby generating an evaluation result that evaluates the first video based on the request and the persona. The information processing device 10 uses the generated evaluation result and the first video as input and causes the child agent 330Y that generates the second video to execute, thereby generating a second video based on the generated evaluation result. Based on the second video, the information processing device 10 causes the parent agent 330 to output the second video and text related to the second video as a response to the request.
[0018] For example, the information processing device 10 causes the parent agent 330 to obtain from the user who is operating the video, the video created by the user, and requests including the setting of the target audience and suggestions for improvement to enhance the audience's evaluation (S1). Subsequently, the information processing device 10 causes the parent agent 330 to request video analysis processing from the child agent 330X, which simulates the persona of the target audience included in the request (S2).
[0019] Then, the information processing device 10 instructs the child agent 330X, which has been requested to perform video analysis processing, to execute the processing (S3), and outputs the analysis results, including a virtual evaluation result obtained by simulating the viewer demographic, to the parent agent 330 (S4). Subsequently, the information processing device 10 outputs the virtual evaluation result of child agent 330X and the video created by the user to the child agent 330Y, which performs video generation processing (S5), and instructs the child agent 330Y to perform video production (updating the input video to be analyzed) based on the evaluation result (S6).
[0020] After that, the information processing apparatus 300 causes the parent agent 330 to acquire the generated result including the video generated by the child agent 330Y and the text and analysis content in the video (S7), and causes the parent agent 330 to output the generated result to the user (S8).
[0021] In this way, since the information processing apparatus 300 can execute evaluation corresponding to the user's request and video production based on the evaluation, it can generate an effective video.
[0022] <Functional Configuration> FIG. 2 is a functional block diagram showing the functional configuration of the information processing apparatus 10 according to the first embodiment. As shown in FIG. 2, the information processing apparatus 10 includes a communication unit 11, a storage unit 12, and a control unit 20.
[0023] The communication unit 11 is a processing unit that controls communication with other devices, and is realized by, for example, a communication interface or the like. For example, the communication unit 11 receives video and requests from the user terminal and transmits responses to the user terminal.
[0024] The storage unit 12 is a processing unit that stores various data and various programs executed by the control unit 20, and is realized by, for example, a memory or a hard disk. For example, this storage unit 12 stores a domain knowledge DB 13 and video data DB 14. In addition to the above, the storage unit 12 stores, for example, various learned machine learning models and multimodals used by the control unit 20.
[0025] The domain knowledge DB 13 is a database that stores knowledge specialized in a certain field. Specifically, the domain knowledge DB 13 stores information about personas, setting information for simulating personas, information necessary for considering requests, and various prompts and information to be set in the prompts.
[0026] The video data DB 14 is a database that stores videos to be analyzed. Specifically, the video data DB 14 stores videos acquired from the user terminal, an example of a video preferred by each persona for each persona, and the like.
[0027] The control unit 20 is a processing unit that manages the information processing device 10 and is implemented by, for example, a processor. This control unit 20 executes the response control unit 30, the video analysis unit 40, and the video generation unit 50. The response control unit 311, the response control unit 30, the video analysis unit 40, and the video generation unit 50 are implemented by electronic circuits of the processor or processes executed by the processor.
[0028] (Response Control Unit 30: Parent Agent 330) The response control unit 30 is a processing unit that executes the parent agent 330 and causes the parent agent 330 to perform various controls. Specifically, the response control unit 30 acquires the first video to be analyzed and the request for the first video. When the first video and the request are input, the response control unit 30 identifies a child agent 330X that simulates a persona and an agent 330Y that generates a second video from among a plurality of child agents that can cooperate with the parent agent 330. The response control unit 30 takes the first video and the request as input and executes the persona-simulating agent 330X to acquire an evaluation result that evaluates the first video based on the request and the persona. The response control unit 30 takes the evaluation result and the first video as input and executes the child agent 330Y that generates the second video to generate a second video based on the generated evaluation result. The response control unit 30 outputs the second video and text related to the second video as an answer to the request.
[0029] For example, the response control unit 30 instructs the parent agent 330 to perform the following processing. Figure 3 is a flowchart illustrating the processing of the parent agent 330. As shown in Figure 3, the parent agent 330 obtains video (first video) and a request from the user (S101). Subsequently, the parent agent 330 instructs the video analysis child agent 330X to perform processing according to the planning information in which instructions and processing request conditions are predetermined (S102).
[0030] Then, the parent agent 330 obtains the evaluation result from the child agent 330X that requested the processing (S103). Here, the parent agent 330 may, if necessary, have another child agent that simulates a different persona perform the analysis processing, or it may determine whether the evaluation result obtained from the child agent contains enough information to generate an answer for the user, and if it is insufficient, it may request reprocessing (S104).
[0031] Subsequently, if the acquired evaluation result is appropriate, the parent agent 330 outputs the evaluation result and the first video to the video generation child agent 330Y and instructs it to generate the second video (S105).
[0032] Then, the parent agent 330 obtains the generation result (second video and text) from the child agent 330Y that requested the processing (S106). Here, the parent agent 330 determines whether the generated video is appropriate or not, and requests reprocessing depending on the result of the determination (S107). After that, the parent agent 330 outputs the generation result to the user as a response (S108).
[0033] Here, the planning information and other details set for the parent agent 330 as described above can be set as prompts. Figure 4 illustrates an example of a prompt given to the parent agent 330. As shown in Figure 4, various pieces of information such as "behavior," "instructions," "wording," and "response format" can be set for a prompt.
[0034] "Behavior" is information that defines the behavior of the parent agent 330, and can be set to things like "knowledgeable person," "gentleman," or "expert." "Instructions" is information that defines the planning information of the parent agent 330, and sets the correspondence between the content of the question and the process to be performed. For example, "Instructions" could be "If video and request input are received, perform video analysis" or "If evaluation result input is received, perform video generation." In addition, instructions can also be set to indicate priority, such as which condition to prioritize when multiple conditions apply.
[0035] In this way, when a question corresponding to an instruction with a specified combination or order is input, the parent agent 330 has the child agents perform processing according to the instruction and aggregates the results. On the other hand, even when a question that does not correspond to an instruction is input, the parent agent 330 interprets the information in the instruction as an example, autonomously determines an appropriate child agent, and requests processing.
[0036] Furthermore, "language usage" is information that defines the language used when the parent agent 330 outputs a response, and can be set to, for example, "expert". "Response format" is information that defines the format in which the parent agent 330 provides a response to the user, and can be set to, for example, "text format", "text format and image", or "audio".
[0037] Depending on the framework used by the parent agent 330, the parent agent can also determine whether the processing result of the child agent is sufficient as an answer. For example, the parent agent 330 can determine whether the processing result of the child agent is appropriate as information to be used in the answer, and if it is determined to be inappropriate, it can request the child agent to reprocess. For example, the parent agent 330 can determine that the processing result of the child agent is insufficient as an answer if it does not include any of the pre-specified information such as "who does what and how" or "is it an evaluation by the persona of the request?". If the parent agent 330 determines that it is appropriate, it outputs an answer to the question.
[0038] (Video analysis: Child agent 330X) The video analysis unit 40 is a processing unit that executes the child agent 330X and causes the child agent 330X to perform various controls. Specifically, the video analysis unit 40 takes the first video and the request as input, and based on the simulated persona, generates an evaluation result that evaluates the first video from the perspective of the request and outputs it to the parent agent 330.
[0039] For example, the video analysis unit 40 has setting information about a persona and a multimodal model. The video analysis unit 40 inputs a prompt consisting of the setting information, a first video, and a request for the first video into the multimodal model, thereby generating an evaluation result that evaluates the video from the perspective of the request and the persona.
[0040] Next, an example of prompt information set on the child agent 330X will be described. Figure 5 shows an example of prompt information set on the child agent 330X. For example, the prompt information 52a set on the child agent 330X includes prompt information 52a-1, 52a-2, and 52a-3.
[0041] Prompt information 52a-1 sets an overview of CoT (Chain of Thought) prompting. For example, among the information included in prompt information 52a-1, "Identify elements from the interviewee's answers that contribute to deeper exploration" and "Set / update the direction of the questions" are information to set a framework to prevent questions from diverging in various directions. "Suggest questions that contribute to the interviewee's underlying needs that lead to their purchasing behavior, based on the direction of the questions. The trick here is to ask about actions and thoughts alternately" indicates a tip for interviews aimed at deeper exploration, and instructs that after hearing about thoughts, ask about related actions, and after hearing about actions, ask about the motivations and feelings at the time the action occurred.
[0042] For example, in Example 1, information used to set a framework to prevent the above questions from diverging in various directions would include "Identify the image, aversion, and positive aspects that the persona associates with the video." Information used to inquire about tips and feelings would include "Show the persona what makes the video popular," and "Show the persona what improves their perception of the video."
[0043] The prompt information 52a-2 sets the rules for creating questions. For example, "Adjust the questions to be as relevant to the interviewee as possible" is constraint information to prevent the discussion from continuing into irrelevant topics. For example, in Example 1, this would include "Narrow down the evaluation target to focus solely on the persona's perspective and the information shown in the first video." Alternatively, in Example 1, this would include information to narrow down the evaluation target, such as "If the user has defined an area within the video, evaluate the video within the defined area from the persona's perspective."
[0044] The prompt information 52a-3 includes multiple examples of follow-up questions. For example, in the case of Example 1, these include automatically narrowing down the area to be evaluated to a predetermined range, and broadening the range of personas to be evaluated to a predetermined age group (e.g., personas within ±10 years).
[0045] Next, another example of prompt information will be described. Figure 6 shows an example of prompt information set for child agent 330X. For example, prompt information 52b includes prompt information 52b-1, 52b-2, and 52b-3.
[0046] Prompt information 52b-1 contains an overview of CoT prompting. Among the information included in prompt information 52b-1, "Please imitate the person given as persona information" and "Please interpret the context, social situation, and general consumer trends from the persona's perspective" are instructions to respond from the perspective of the given persona. "Please generate a response in accordance with the rules, taking the conversation history into consideration" is information to constrain responses with rules to prevent duplicate answers.
[0047] Prompt information 52b-2 sets the rules for creating answers to questions. For example, "Do not repeat what you have said before" and "If you refer to similar information, mention the motivation or underlying desire related to the information" are pieces of information designed to generate more in-depth answers by allowing the use of similar content while requiring the addition of new information.
[0048] Prompt information 52b-3 is set with reference information. Reference information is intended to provide information that serves as the source or basis for the answer. For example, reference information may include persona information, information about the interviewee and related information about the interviewee, information about social conditions, general consumer trends, questions from the interviewer, chat history, etc.
[0049] For example, persona information is the information shown in Figure 7. Figure 7 is a diagram illustrating an example of persona information. As shown in Figure 7, persona information 53b includes name, gender, age, occupation, address, personality, hobbies, lifestyle, family structure, values, consumer behavior, trusted sources of information, desired future, etc.
[0050] (Video generation: Child agent 330Y) The video generation unit 50 is a processing unit that executes the child agent 330Y and causes the child agent 330Y to perform various controls. Specifically, the video analysis unit 40 has a multimodal model and generates a second video from the perspective of the evaluation results by inputting a prompt consisting of the evaluation results generated by the child agent 330X and the first video into the multimodal model.
[0051] (Process Overview) Figure 8 is a diagram illustrating an example of video generation. As shown in Figure 8, the video generation unit 50 is connected to the response control unit 30 and the user (user terminal device) 3. The response control unit 30 has been described above, so its explanation will be omitted.
[0052] User 3 is the user terminal device 3 used by the user, and is a device that outputs requests regarding the video created by the user. The user refers to the display screen of the user terminal device 3 and uses the user terminal device 3 to select a frame from the video acquired by the user terminal device 3 to specify a visual prompt indicating the object of their interest. Hereinafter, the frame selected by the user as the frame for specifying the visual prompt will be called the "selected frame". The user then uses the user terminal device 3 to specify the area in the selected frame that contains the object of their interest using a visual prompt. Furthermore, the user terminal device 3 receives input from the user in the form of questions related to the area of interest indicated by the visual prompt.
[0053] The user terminal device 3 then outputs information about the selected frame and a visual prompt indicating the object of the user's attention to the designated area extraction unit 104 of the video generation unit 50. The user terminal device 3 also outputs a text prompt containing a question about the object of the user's attention to the text conversion unit 111 of the video generation unit 50. Here, the object specified by the user using the visual prompt is an example of the "first object," and the frame in which the object of attention is specified using the visual prompt is an example of the "predetermined video frame." Furthermore, the text prompt containing a question about the object of the user's attention is an example of the "request (question) about the first object."
[0054] Furthermore, users can use the display screen of the user terminal device 3 to check the response from the video generation unit 50 to their requests regarding the object of their interest.
[0055] The video generation unit 50 includes a visual encoder 101, a spatiotemporal feature calculation unit 102, an overall projector 103, a specified region extraction unit 104, an ROI tracker 105, a related region estimation unit 106, a partial region feature calculation unit 107, a selection unit 108, and a projector 109. Furthermore, the video generation unit 50 includes an LLM (Large Language Models) decoder 110, a text conversion unit 111, and an embedding unit 112.
[0056] The visual encoder 101 receives evaluation results and video input from the response control unit 30. The visual encoder 101 then calculates the overall feature quantities for each frame of the video. Here, the picture represented by the entirety of each frame of the video is called an image. In other words, a video is a continuous collection of images, frame by frame. Furthermore, below, the overall feature quantities of an image will be called image features. The visual encoder 101 outputs the image features for each frame to the spatiotemporal feature calculation unit 102 and the subregion feature calculation unit 107.
[0057] The spatiotemporal feature calculation unit 102 receives image feature data for each frame of the video from the visual encoder 101. The spatiotemporal feature calculation unit 102 then calculates the spatial and temporal feature data for the entire video based on the temporal and spatial relationships of each object in the image for each frame. The spatiotemporal feature calculation unit 102 then outputs the spatial and temporal feature data for the entire video to the overall projector 103.
[0058] In this embodiment, the spatiotemporal feature calculation unit 102 calculates both spatial and temporal features of the entire video, but it may do so with either one or the other. That is, the spatiotemporal feature calculation unit 102 calculates either spatial or temporal image features of the video.
[0059] The overall projector 103 receives spatial and temporal features of the entire video as input from the spatiotemporal feature calculation unit 102. The overall projector 103 then performs embedding on the spatial and temporal features of the entire video to match the feature space of the LLM decoder 110. For example, the overall projector 103 performs processing such as matching the number of dimensions of the spatial and temporal features of the entire video to the number of dimensions of the feature space of the LLM decoder 110. After that, the overall projector 103 outputs the embedded data of the spatial and temporal features of the entire video to the LLM decoder 110.
[0060] The designated region extraction unit 104 receives information from the user terminal device 3 regarding the user's selected frame from the video output by the response control unit 30, and information regarding the visual prompt specified by the user for the image of the selected frame. The designated region extraction unit 104 then extracts a region on the image indicated by the visual prompt for the image of the selected frame as an ROI. For example, the designated region extraction unit 104 can set the X and Y axes for the image of the selected frame and represent the region using the X and Y coordinates representing each point in the image.
[0061] In this embodiment, the specified region extraction unit 104 extracts the ROI as a region called a BBox (Bounding Box). A BBox is a rectangular sub-region that encloses the area of the object of interest relative to the external region with the smallest possible rectangle and demarcates it with a boundary. For example, a BBox is represented as a rectangle that encloses a predetermined area on the image of the selected frame. The specified region extraction unit 104 can represent a BBox using the XY coordinates of two vertices on the diagonal, and the region enclosed by that BBox is considered the ROI. The specified region extraction unit 104 outputs the ROI information to the ROI tracker 105.
[0062] In this way, the designated area extraction unit 104 accepts the input of information for a visual prompt, that is, an operation to specify a first area in which a first object is located within a predetermined video frame displayed on the display screen of the user terminal device 3.
[0063] The ROI tracker 105 receives video input output by the response control unit 30. The ROI tracker 105 also receives BBox input from the designated area extraction unit 104, which indicates ROI information corresponding to the visual prompt specified by the user.
[0064] Next, the ROI tracker 105 searches for and tracks the subregion on the image of the selected frame that corresponds to the ROI for each frame of the video. This allows the ROI tracker 105 to extract the subregion corresponding to the visual prompt specified by the user for each frame of the video output by the response control unit 30. Subsequently, the ROI tracker 105 outputs information about the ROI and the subregions for each frame of the video to the related region estimation unit 106 and the subregion feature calculation unit 107. Hereafter, the ROI and the subregions for each frame of the video will be collectively referred to as the "ROI-corresponding subregion."
[0065] This ROI tracker 105 is an example of a "region identification unit." Furthermore, the ROI-corresponding region extracted by the ROI tracker 105 is an example of a "first region in which a first object is located within a predetermined video frame among multiple video frames included in the acquired video." In other words, the designated region extraction unit 104 acquires the video to be monitored and, based on the processing of the designated region extraction unit 104, identifies the first region in which a first object is located within a predetermined video frame among multiple video frames constituting the acquired video.
[0066] The related region estimation unit 106 has a machine learning model that estimates related regions associated with ROI-corresponding subregions in the image of each frame. The related region estimation unit 106 receives video input output by the response control unit 30. The related region estimation unit 106 also receives information about ROI-corresponding subregions from the ROI tracker 105.
[0067] The related region estimation unit 106 uses a machine learning model to estimate a predetermined number of related regions in descending order of relevance to the ROI-corresponding partial region in each frame, taking the overall image and the image of the ROI-corresponding partial region as input for each frame in which the ROI-corresponding partial region has been extracted. The related region estimation unit 106 then outputs information on the related regions that have a high degree of relevance to the estimated ROI-corresponding partial region as the estimation result.
[0068] The subregion feature calculation unit 107 receives image feature quantities for each frame of the video calculated by the visual encoder 101. The subregion feature calculation unit 107 also receives information on ROI-corresponding subregions output from the ROI tracker 105. Furthermore, the subregion feature calculation unit 107 receives related region information for each target frame estimated by the related region estimation unit 106.
[0069] Next, the subregion feature calculation unit 107 calculates the feature quantities of the ROI-corresponding subregion from the image feature quantities of each frame of the video. The subregion feature calculation unit 107 also calculates the feature quantities of the related region from the image feature quantities of each frame of the video. Finally, the subregion feature calculation unit 107 outputs the feature quantities of the ROI-corresponding subregion and the feature quantities of the related region to the selection unit 108.
[0070] The selection unit 108 receives feature quantities for the ROI-corresponding subregion and related regions as input from the subregion feature quantity calculation unit 107. The selection unit 108 selects the ROI-corresponding subregion features and related region features to be used to generate answers to questions by removing duplicate or unimportant features from among the ROI-corresponding subregion features and related region features. For example, the selection unit 108 can classify the ROI-corresponding subregion features and related region features using the K-means method, select groups considering the similarity of each group, and then select a predetermined number of features that match specific conditions from the selected groups. After that, the selection unit 108 outputs the selected ROI-corresponding subregion features and related region features to the projector 109.
[0071] The selection unit 108 selects features based on both the features of the ROI-corresponding subregion and the features of the related region, thereby taking into account the state of the related region as well as the ROI-corresponding subregion. For example, the selection unit 108 can select features even if there is little change in the ROI-targeted subregion, but there is a large change in the related region. This makes it possible to include important information about the related region in the question. In this way, the image generation unit 50 selects multiple image features from the image features of the first object and the second object using the K-means method.
[0072] The projector 109 receives input of feature quantities for the ROI-corresponding subregion and related regions selected by the selection unit 108. The projector 109 then performs an embedding process on the feature quantities for the ROI-corresponding subregion and related regions to match the feature quantity space of the LLM decoder 110. After that, the projector 109 outputs the embedded data of the ROI-corresponding subregion and related region feature quantities to the LLM decoder 110.
[0073] The text conversion unit 111 receives text prompt input containing questions about the video, including the ROI, from the answer control unit 30 and the user terminal device 3. For example, the text conversion unit 111 receives text prompts containing questions that include user requests or evaluation results of the video evaluated by the child agent. The text conversion unit 111 then performs text conversion processing according to the format of the question to the LLM decoder 110, such as dividing the text in the text prompt into vocabulary. This allows the text conversion unit 111 to identify what kind of question (request) was input to the LLM decoder 110. After that, the text conversion unit 111 outputs the converted text prompt to the embedding unit 112.
[0074] In this way, the text conversion unit 111 identifies a question about a first object specified by the user using a visual prompt. More specifically, the text conversion unit 111 receives a question document (for example, a request and evaluation result) from the user regarding a first object that exists in the first area, and identifies a question based on the question document.
[0075] The embedding unit 112 receives the text prompt input, which has undergone text conversion, from the text conversion unit 111. The embedding unit 112 then performs embedding processing, such as converting the text prompt into a vector, to convert it into a format that can be input to the LLM decoder 110. After that, the embedding unit 112 outputs the embedded text prompt to the LLM decoder 110.
[0076] The LLM decoder 110 is a machine learning model that receives input of image features and text prompts for questions about the image, and outputs a response to those questions. The LLM decoder 110 receives embedded data of spatial and temporal features of the entire video from the overall projector 103. The LLM decoder 110 also receives embedded data of features of ROI-corresponding subregions and related regions from the projector 109. Furthermore, the LLM decoder 110 receives embedded data of text prompts from the embedding unit 112.
[0077] The LLM decoder 110 then generates an answer to the question indicated by the text prompt, based on the embedded data of spatial and temporal features of the entire video, features of the ROI-corresponding subregion, and features of the related region. Subsequently, the LLM decoder 110 outputs the generated answer to the user terminal device 3.
[0078] In this way, the LLM decoder 110 generates an answer based on the spatial and temporal features of the entire video, as well as the features of the ROI-corresponding subregion, and the features of the related region. That is, the LLM decoder 110 can generate an answer to a question by considering events that occurred in the related region. The answer generated by the LLM decoder 110 is transmitted to the user terminal device 3 and displayed on the display screen.
[0079] In this way, the LLM decoder 110 generates an answer (video and text) to a question based on a question about a first object specified by the user using a visual prompt, which is an object specified by a visual prompt, and the image features of a second object present in the related region. More specifically, the LLM decoder 110 generates an answer based on a plurality of image features selected by the selection unit 108. The LLM decoder 110 is an example of a "multimodal model".
[0080] In other words, the video generation unit 50 generates an answer to a question by inputting a prompt containing the request, the image features of the first object, and the second object into a multimodal model. The video generation unit 50 also calculates spatial or temporal image features, a plurality of image features selected by the selection unit 108, and embeddings for each of the questions, and inputs the calculated embeddings into the multimodal model to generate an answer.
[0081] Figure 9 shows an example of evaluation generation. Next, referring to Figure 9, we will summarize the overall picture of the video generation process that takes into account the evaluation results from the video generation unit 501. Figure 9 also shows the data used in each process. Each piece of data will be explained using the name shown in Figure 9.
[0082] (Overall configuration and flow of processing) The response control unit 30 outputs video V. Video V contains a large number of consecutive frames. The user uses the user terminal device 3 to select a frame F from video V and sets a visual prompt P for the selected frame F.
[0083] The visual encoder 101 extracts image feature quantities f from the video V for each frame. t Calculate.
[0084] The spatiotemporal feature calculation unit 102 calculates the image feature f of each frame calculated by the visual encoder 101. t From this, the spatial features f of the video V spatial and time feature f temporal Calculate.
[0085] The overall projector 103 has spatial feature f spatialExecute an embedding process to match the feature space of the LLM decoder 110 for the spatial feature amount, and generate the embedded data e of the spatial feature amount. ν spatial Similarly, the overall projector 103 executes an embedding process to match the feature space of the LLM decoder 110 for the temporal feature amount f temporal and generates the embedded data e of the temporal feature amount. ν temporal
[0086] The specified region extraction unit 104 generates a BBox 21 indicating the ROI, which is the partial region specified by the visual prompt P, based on the visual prompt P for the selected frame F.
[0087] The ROI tracker 105 performs a search for each frame of the video V using the BBox 21 and generates a BBox 22 indicating the ROI corresponding partial region of each frame.
[0088] The related region estimation unit 106 estimates the related regions in each target frame from which the ROI corresponding partial region was extracted, based on the BBox 22 indicating the ROI corresponding partial region of each frame and the video V. Here, the related region estimation unit 106 estimates L related regions in descending order of the degree of relevance.
[0089] The partial region feature amount calculation unit 107 calculates the feature amount f Roi t,0 of the ROI corresponding partial region of each target frame from the BBox 22 indicating the ROI corresponding partial region of each target frame.
[0090] Also, the partial region feature amount calculation unit 107 calculates the feature amount f RRoi t,1 ~f of each related region of each target frame from the information indicating the related regions in each target frame. Here, since there are L related regions, the partial region feature amount calculation unit 107 calculates the feature amounts of each of the L related regions, namely, the feature amounts f RRoi t,L ~f RRoi t,1 ~f RRoi t,L .
[0091] The selection unit 108 selects the feature quantity f of the ROI-corresponding subregion of each target frame. Roi t,0 , and the feature quantities f of each related region RRoi t,1 ~f RRoi t,L Select the features to use in your question from the options provided.
[0092] The projector 109 performs embedding processing on the feature quantities selected by the selection unit 108 to embed data e related to the ROI-corresponding subregion and related regions. RoI 0 and embedded data e RoI 1 ~e RoI L Generates.
[0093] The text conversion unit 111 performs text conversion processing on the text prompt T according to the format of the question to the LLM decoder 110.
[0094] The embedding unit 112 performs an embedding process on the text prompt T that has undergone text conversion to embed data e t Generates.
[0095] The LLM decoder 110 uses the embedded spatial feature data e ν spatial , embedded data of time features e ν temporal , embedded data e related to ROI-compatible subregion and related regions RoI 0 and embedded data e RoI 1 ~e RoI L , and embedded data e indicating the question t The LLM decoder 110 receives the input. Based on the input data, it generates an answer A to the question regarding the subject specified in the video and visual prompt.
[0096] Figure 10 shows an overview of the related region estimation process performed by the related region estimation unit. Next, referring to Figure 10, we will summarize the overall picture of the related region estimation process performed by the related region estimation unit. Figure 10 also shows the data used in each process. Each data will be explained using the name shown in the figure.
[0097] The preprocessing unit 161 generates a cropped image 32 by cutting out the region indicated by the ROI-corresponding partial region R from the image 31 of each frame included in the video V.
[0098] The visual encoder 162 calculates the subregion feature quantities of the ROI-corresponding subregion from the cropped image 32. The visual encoder 162 also calculates the overall feature quantities of each frame from the image 31 of each frame.
[0099] The subregion projector 163 performs a conversion process on the ROI-compatible subregion to generate subregion feature quantities 33.
[0100] The overall projector 164 performs a conversion process on each frame to generate the overall feature vector 34.
[0101] The synthesis unit 165 generates a composite feature by performing matrix integration of the subregion feature 33 and the overall feature 34.
[0102] The normalization unit 166 performs a normalization process on the composite feature quantities.
[0103] The decoding unit 167 performs a decoding process on the normalized composite features to generate a relevance attention map 35.
[0104] The region generation unit 168 generates related region information 36, which indicates related regions for each frame, from the relatedness attention map 35.
[0105] As described above, the related region estimation unit 106 generates a relevance attention map 201 of the ROI-corresponding subregion 210 from a specific frame 200 and the ROI-corresponding subregion 210. For example, in the relevance attention map, darker colors indicate a higher degree of relevance. The related region estimation unit 106 then generates related region information indicating two related regions with a high degree of relevance from among the related regions 4 shown in the relevance attention map.
[0106] <Effects> The information processing device 10 uses an agent that simulates a persona to evaluate the video created by the user and identify areas for improvement. The information processing device 10 then modifies the video created by the user, taking the evaluation results into consideration. In this way, the information processing device 10 can rewrite the video created by the user into a video that receives the evaluation the user desires. Therefore, the information processing device 10 can streamline the planning, editing, and review processes in video production. Furthermore, the information processing device 10 can enable the creation of effective videos for viewers. In addition, the information processing device 10 can improve the quality of videos based on objective evaluations. Furthermore, the information processing device 10 can enable video production that takes into account diverse or specific personas. Furthermore, the information processing device 10 can provide a new video production workflow that supports user creativity.
[0107] Furthermore, the information processing device 10 can decompress the image while considering its relationship with other objects on the screen, thereby improving the accuracy of the answer. In addition, the information processing device 10 tracks the specified ROI for all frames to extract the ROI-corresponding region, and extracts related regions that are related to the ROI-corresponding region in each frame and have a high degree of relevance. Therefore, it can generate an answer using the spatial and temporal features of the entire video, the features of the ROI-corresponding region, and the features of the related regions.
[0108] Furthermore, the information processing device 10 can automatically acquire surrounding information related to the specified object and provide it to the LMM. This allows the information processing device 10 to consider not only the temporal and spatial changes in importance of the entire image and the important changes of the object of interest, but also the important changes of related objects such as people and objects that are highly relevant to the object of interest. Therefore, the information processing device 10 can improve its ability to understand images and videos. Moreover, by considering the important changes of related objects such as people and objects that are highly relevant to the object of interest when conducting question and answer sessions, it becomes possible to improve the accuracy of VQA responses.
[0109] By the way, the processing performed by the video generation unit 50 is not limited to the method described in Example 1. Therefore, Example 2 describes an example in which video and text related to the video are input to the VQA and video is generated. In this example, the text includes requests and evaluation results.
[0110] For example, the information processing device 10 according to Embodiment 2 takes a question sentence containing text such as requests as input and selects a topic of interest based on the content of the input question sentence. The information processing device 10 also takes a video as input and acquires the feature quantities of the input video. The information processing device 10 then compresses the feature quantities based on context. Next, the information processing device 10 samples feature quantities using the selected topic of interest and the compressed feature quantities. After that, the information processing device 10 generates an answer to the question sentence by inputting a prompt, which is composed of the sampled feature quantities and the input question sentence, into the LLM.
[0111] LLM is, for example, a pre-trained language model composed of a neural network. Furthermore, LLM is a language model that learns using large computational resources, data, and parameters, and then takes natural language as input, executes the learned processing, and returns a response.
[0112] Figure 11 is a diagram illustrating the video generation process (before question input) according to Example 2, and Figure 12 is a diagram illustrating the video generation process (during question input) according to Example 2.
[0113] As shown in Figure 11, the information processing device 10 inputs each frame of the video (video frame) into an encoder, extracts the visual features of each video frame, and stores each visual feature. Subsequently, the information processing device 10 inputs each visual feature into a compression mechanism, such as an autoencoder, and extracts contextual features from each visual feature.
[0114] The information processing device 10 then inputs the feature quantities of each context into a first topic extraction mechanism, which is a mechanism that predicts, extracts, and prioritizes objects and topics that may be the subject of questions, such as site characteristics and people involved, to generate topics of interest (hereinafter sometimes simply referred to as topics) and stores them in a topic bank. Subsequently, the information processing device 10 performs sampling to extract features corresponding to topics from the feature quantities of the contexts and stores the sampled feature quantities of the contexts in a memory bank.
[0115] In other words, during the initial video input stage, since no questions have been entered, the information processing device 10 extracts candidate topics that appear to be important from the video alone and generates the initial state of the topic bank using Topic extraction, which is an example of a first topic extraction mechanism. Furthermore, the information processing device 10 stores information that has a high degree of relevance to the feature quantities of the topic bank in the memory bank, such as when the number of frames exceeds the memory bank length.
[0116] Subsequently, when a question is input, as shown in Figure 12, the information processing device 10 inputs the question into an analysis mechanism to decompose it into morphemes. Next, the information processing device 10 inputs the obtained morphemes into a second topic extraction mechanism, which extracts the object or topic currently being questioned from the question and updates the topic bank, to extract topics. Then, the information processing device 10 inputs the extracted topics into a first conversion mechanism, which is an example of a projector that performs shape conversion to topics to be stored in the topic bank, and updates the topic bank with the shape-converted topics.
[0117] Subsequently, the information processing device 10 performs sampling to extract features corresponding to the topics in the updated topic bank, and stores the sampled context features in the memory bank. In other words, the information processing device 10 can update (regenerate) the memory bank using the stored (stocked) image features. The information processing device 10 can also establish update criteria for situations such as when a question on a previously unexplored topic is input. Furthermore, the information processing device 10 extracts context features that have a high degree of relevance to the top K topics (K is any number) in the topic bank.
[0118] The information processing device 10 then repeats the process shown in Figure 12 each time a question is input. When the question is finished, the information processing device 10 inputs the morphemes obtained from the question into a second conversion mechanism (Embedding) that converts it into an input format for LLM, and converts it into a numerical vector. Similarly, the information processing device 10 inputs the context features stored in the memory bank into a projector and converts (restores) them into features (visual embeddings) that the LLM can understand. After that, the information processing device 10 inputs the numerical vector of the question and the features (visual embeddings) into the LLM and outputs the answer obtained.
[0119] In this way, the information processing device 10 performs feature compression based on context and extracts candidate important topics from the video information. Subsequently, the information processing device 10 updates the topics of interest according to the content of the questions that are input each time, and performs information compression or extracts information from stored video features based on the topics of interest, updating the memory so that it contains more information related to the topics of interest. After that, the information processing device 10 restores the compressed features so that the LLM can understand them and inputs them into the LLM.
[0120] Therefore, the information processing device 10 can store important information in long-duration video based on the context of the video and the content of the questions, and can also compress the feature quantities, thereby improving the accuracy of the VQA output.
[0121] As described above, the information processing device 10 can realize VQA that appropriately recognizes and selects context by compressing video information using context as the basis for compression decisions. Furthermore, even if the monitored video is long-duration, the information processing device 10 can appropriately output answers to questions about events that occurred within the monitored video, thereby improving the accuracy of the answers.
[0122] Furthermore, the information processing device 10 according to Embodiment 2 acquires the video to be evaluated. Next, the information processing device 10 analyzes the video frames that make up the video and extracts first feature quantities related to the context from the video frames for each type of context that indicates the attributes of objects and / or relationships between objects that the video frames possess. The information processing device 10 also performs sampling processing on the extracted first feature quantities related to the context based on information on the topic of interest related to the video. Then, the information processing device 10 outputs a response to the request related to the video based on the first feature quantities on which the sampling processing has been performed. After that, the information processing device 10 displays the output result on the display screen.
[0123] Specifically, for example, the information processing device 10 according to Embodiment 2 acquires the video to be evaluated and questions about events that occurred in the monitored area within the video. Next, the information processing device 10 analyzes the video frames that make up the video and extracts first feature quantities related to the context from the video frames for each type of context that indicates the attributes of objects in the video frames and / or the relationships between objects in the video frames. The information processing device 10 also generates information on topics of interest using the content of the acquired questions about the events, and performs sampling processing on the extracted first feature quantities related to the context based on the generated information on topics of interest. Then, based on the first feature quantities from which the sampling processing has been performed and the questions about the events, the information processing device 10 generates information about events that occurred in the monitored area as answers to the questions.
[0124] Now, although embodiments of the present invention have been described, the present invention may be implemented in various other forms besides those described above.
[0125] (Numerical values, etc.) The machine learning model, context, topic, features, images, etc. used in the above example are merely examples and can be changed at will. Also, the processing flow described in each flowchart can be changed as appropriate within a range that does not contradict each other.
[0126] (System) Unless otherwise specified, the processing procedures, control procedures, specific names, and information including various data and parameters shown in the above documents and drawings may be changed at will.
[0127] Furthermore, the specific forms of distribution and integration of the components of each device are not limited to those shown in the diagram. For example, each agent may be run on a separate device. In other words, all or part of the components may be functionally or physically distributed and integrated in any unit depending on various loads and usage conditions. Moreover, each processing function of each device may be implemented, in whole or in any part, by a CPU and a program that is analyzed and executed by that CPU, or by hardware using wired logic.
[0128] Furthermore, each processing function performed by each device can be implemented, in whole or in part, by a CPU and a program executed for analysis by that CPU, or by hardware using wired logic.
[0129] (Hardware) Figure 13 is a diagram illustrating an example of hardware configuration. As shown in Figure 13, the information processing device 10 includes a communication device 10a, an HDD (Hard Disk Drive) 10b, memory 10c, and a processor 10d. Furthermore, each of the parts shown in Figure 13 is interconnected by a bus or the like.
[0130] The communication device 10a is a network interface card or the like, and communicates with other devices. The HDD 10b stores programs and databases that operate the functions shown in Figure 2.
[0131] The processor 10d operates a process that performs the functions described in Figure 2 by reading a program that performs the same processing as each processing unit shown in Figure 2 from the HDD 10b or the like and loading it into the memory 10c. For example, this process performs the same functions as each processing unit of the information processing device 10. Specifically, the processor 10d reads a program that has the same functions as the answer control unit 30, the video analysis unit 40, the video generation unit 50, etc., from the HDD 10b or the like. Then, the processor 10d executes a process that performs the same processing as the answer control unit 30, the video analysis unit 40, the video generation unit 50, etc.
[0132] Thus, the information processing device 10 operates as an information processing device that executes the generation method by reading and executing a program. Furthermore, the information processing device 10 can also achieve the same functionality as in the above-described embodiment by reading the program from the recording medium using a media reader and executing the read program. Note that the program referred to in this other embodiment is not limited to being executed by the information processing device 10. For example, the above embodiment may also be applied similarly when another computer or server executes the program, or when they collaborate to execute the program.
[0133] This program may be distributed via a network such as the Internet. Alternatively, this program may be recorded on a computer-readable recording medium such as a hard disk, flexible disk (FD), CD-ROM, MO (Magneto-Optical disk), or DVD (Digital Versatile Disc), and executed by being read from the recording medium by a computer.
[0134] 10 Information Processing Unit 11 Communication Unit 12 Storage Unit 13 Domain Knowledge Database 14 Video Data Database 20 Control Unit 30 Response Control Unit 40 Video Analysis Unit 50 Video Generation Unit
Claims
1. A generation program characterized by causing a computer to perform the following processes: acquire a first video to be analyzed and a request for the first video; when the first video and the request are input to a first agent that generates information according to the input information, it identifies, among a plurality of second agents that can cooperate with the first agent, an agent that simulates a persona and an agent that generates a second video; execute the agent that simulates a persona with the first video and the request as input to generate an evaluation result that evaluates the first video based on the request and the persona; execute the agent that generates the second video with the generated evaluation result and the first video as input to generate the second video based on the generated evaluation result; and the first agent outputs the second video and text related to the second video as a response to the request based on the second video.
2. The generation program according to claim 1, wherein the agent simulating the persona has setting information relating to the persona and a multimodal model, and generates the evaluation result which evaluates the video from the perspective of the request and the persona by inputting a prompt consisting of the setting information, the first video, and the request for the first video into the multimodal model.
3. The generation program according to claim 2, characterized in that the agent that generates the second video has a multimodal model, and generates the second video from the perspective of the evaluation result by inputting a prompt consisting of the generated evaluation result and the first video into the multimodal model.
4. The generation program according to claim 1, characterized in that the first agent is instructed by the computer to determine whether the information generated by the second agent is appropriate to be used in the answer result, and if it is determined to be inappropriate, the first agent is instructed by the computer to request the second agent to regenerate the information, and if it is determined to be appropriate, the first agent outputs the answer result based on the information generated by the second agent as the answer to the question.
5. The generation program according to claim 1, characterized in that it causes the computer to perform the following processing: identify a first region in which a first object is located within a predetermined video frame among a plurality of video frames contained in the first video, and a question concerning the first object present in the first region; analyze the first video to identify a second object related to the first object present in the first region among a plurality of objects present in each of the plurality of video frames; and generate an evaluation result that evaluates the first video based on the requests and persona concerning the first object, and the image features of the first object and the second object.
6. The generation program according to claim 5, wherein the process for identifying the first region and the request includes an operation to specify the first region in which the first object is located within the predetermined video frame displayed on the display screen, and a process for receiving a document relating to the first object in the first region from the user and identifying the request based on the document, and the process for generating the response includes a process for inputting a prompt including the request and persona, the image features of the first object and the second object into a large-scale multimodal model to generate an evaluation result that evaluates the first video from the perspective of the request and persona as an evaluation result of evaluating the first video.
7. The generation program according to claim 5, characterized in that it generates information about the topic of interest in the above request, analyzes the video frames constituting the first video to extract a first feature quantity related to the context from the video frames for each type of context indicating the relationship between objects in the video frames, performs a sampling process on the extracted first feature quantity related to the context based on the generated information about the topic of interest, and outputs the evaluation result of evaluating the first video based on the first feature quantity on which the sampling process has been performed.
8. A generation program characterized by causing a computer to perform the following processes: acquire a first video to be analyzed and a request for the first video; use the first video and the request as input to run an agent simulating a persona to generate an evaluation result that evaluates the first video based on the request and the persona; use the generated evaluation result and the first video as input to run an agent that generates a second video to generate the second video based on the generated evaluation result; and output the second video and text related to the second video as a response to the request.
9. A generation method characterized by the following processes: a computer acquires a first video to be analyzed and a request for the first video; when the first video and the request are input to a first agent that generates information according to the input information, the computer identifies an agent that simulates a persona and an agent that generates a second video from among a plurality of second agents that can cooperate with the first agent; the computer executes the agent that simulates a persona with the first video and the request as input to generate an evaluation result that evaluates the first video based on the request and the persona; the computer executes the agent that generates the second video with the generated evaluation result and the first video as input to generate the second video based on the generated evaluation result; and the first agent outputs the second video and text related to the second video as a response to the request based on the second video.
10. An information processing device comprising a control unit which acquires a first video to be analyzed and a request for the first video, and when the first video and the request are input to a first agent that generates information according to the input information, it identifies an agent that simulates a persona and an agent that generates a second video from among a plurality of second agents that can cooperate with the first agent, the first video and the request are input to the agent that simulates the persona, thereby generating an evaluation result that evaluates the first video based on the request and the persona, the generated evaluation result and the first video are input to the agent that generates the second video, thereby generating the second video based on the generated evaluation result, and the first agent outputs the second video and text related to the second video as a response to the request based on the second video.
Citation Information
Patent Citations
System and method for attention-based configurable convolutional neural network (abc-CNN) for visual question answering
JP2017091525A
Information processing device, information processing method, and information processing program
JP2024025997A
Information processing program, information processing method and information processing device
JP2024082634A