Information processing device, information processing method, and program
The information processing apparatus addresses the challenges of representing complex timelines and parallel actions in time-series content generation by converting text input into event sequence information, enhancing user efficiency and content creation.
Patent Information
- Application Number
- JP2023200501
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-09
AI Technical Summary
Conventional methods for generating time-series content from text struggle to accurately represent the desired structure and timing, leading to inefficiencies and user frustration.
An information processing apparatus that acquires text input and generates event sequence information, visualizing the structure of events in the order they occur, allowing for easier representation of complex timelines and parallel actions.
Enables users to more easily generate desired time-series content by providing a visual representation of event sequences, reducing the need for repeated text inputs and improving the efficiency of content creation.
Smart Images

Figure 2025086494000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] Conventionally, a technique for generating time-series data content (hereinafter, may be referred to as "time-series content") such as moving images and audio data from text has been known. For example, news and comments on the news are extracted from a website, the sentiment or subjectivity in input data that is text such as the extracted comments is analyzed, and based on the dynamic feature amounts of the sentiment or subjectivity included in the analyzed input data, a technique for generating a time animation is known.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, in the above conventional technology, since it only analyzes the sentiment or subjectivity in input data that is text and generates a time animation based on the dynamic feature amounts of the sentiment or subjectivity included in the analyzed input data, it is not always possible to make it easier to generate the time-series content desired by the user.
[0005] Therefore, the present disclosure proposes an information processing apparatus, an information processing method, and a program that can make it easier to generate the time-series content desired by the user.
Means for Solving the Problems
[0006] The information processing apparatus of the present disclosure includes an acquisition unit that acquires text input for generating time-series content, and a sequence generation unit that generates event sequence information, which is information visualizing a structure in which event information corresponding to events in the time-series content is arranged according to the order in which the events occur, based on the text.
Brief Description of Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Embodiments for Carrying Out the Invention
[0008] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each of the following embodiments, the same parts are denoted by the same reference numerals, and redundant explanations are omitted.
[0009] (Embodiment) (1. Introduction) In recent years, there has been an increasing demand for content creation by users such as UGC (User Generated Contents). It is desirable that anyone can easily create content, but creating content of time-series data such as moving images and audio data (hereinafter sometimes referred to as "time-series content") takes time corresponding to the length of the content. Note that the moving image may be an animation or a video.
[0010] Therefore, attention has been focused on generative AI (also referred to as generative AI), which is a machine learning model trained to output time-series content such as moving images from simple inputs such as text. Here, generative AI refers to a program or algorithm that generates various contents such as moving images and audio data based on learning data. For example, generative AI can autonomously learn the pattern (feature amount, etc.) of the content data to be output and reflect the learned content in the newly generated content.
[0011] However, in the current method for creating time-series content using generative AI, it is not possible to confirm what kind of content it will be until the time-series content is output over a long period of time. Therefore, a lot of time and effort are wasted until the user completes the time-series content desired. This point will be described in detail with reference to FIGS. 1 and 2.
[0012] FIG. 1 is a diagram for explaining an example of a technical problem when generating time-series content from text. In FIG. 1, when a user attempts to generate time-series content from text using generative AI and edit the generated content, it is explained that it may be difficult to sufficiently represent the structure in the time-axis direction in the time-series content only with text.
[0013] In FIG. 1, as an example of the text input by the user, it shows the structure in the time axis direction when generating time-series content based on the input sentence "The son who was reading a comic in the room was scolded by his mother to study and left home in disgust." In FIG. 1, it shows a visualization of a structure in which the elements indicating the respective actions of the son and the mother are arranged horizontally in the order in which each action occurs. In FIG. 1, the horizontal direction corresponds to the time axis direction. Specifically, in FIG. 1, it shows the state where time flows from left to right. That is, it shows that the action corresponding to the element located on the left side of FIG. 1 occurs earlier in time, and the action corresponding to the element located on the right side of FIG. 1 occurs later in time. Also, the length of the element corresponding to each action indicates the duration of each action.
[0014] It is inferred from the input sentence that the actions of the son "reading a comic" and the mother "getting angry" proceed simultaneously in time. Therefore, in FIG. 1, the element indicating the action of the son "reading a comic" and the element indicating the action of the mother "getting angry" are arranged in parallel in a direction orthogonal to the time axis direction. Here, the fact that each of the elements indicating two or more actions is arranged in parallel in a direction orthogonal to the time axis direction corresponds to the fact that two or more actions proceed simultaneously in time. Hereinafter, two or more actions that proceed simultaneously may be referred to as parallel actions. For example, when the number of action subjects appearing in the time-series content increases and the number of parallel actions increases to three or more, there is a limit to expressing three or more parallel actions using only the text. Thus, it may be difficult to appropriately express parallel actions using only text.
[0015] In addition, the input text does not specifically describe the situation where the son is performing the action of "reading a comic". Therefore, in Figure 1, since it is inferred that the son is lying on the bed reading a comic, elements indicating the action of the son "lying on the bed" are supplemented. Specifically, elements indicating the action of the son "lying on the bed" are supplemented at positions parallel to the direction orthogonal to the time axis direction and in parallel with the elements indicating the action of the son "reading a comic". Also, although not specified in the input text, it is inferred that before the son gets angry and leaves the house, an action of placing the comic the son is reading is required. Therefore, elements indicating the action of the son "putting down the book" he is reading are supplemented at a position behind the elements indicating the action of the son "reading a comic". Also, although not specified in the input text, it is inferred that before the son gets angry and leaves the house, an action of getting up from the bed where he was lying is required. Therefore, elements indicating the action of the son "getting up from the bed" where he was lying are supplemented at a position behind the elements indicating the action of the son "lying on the bed". Also, although not specified in the input text, it is inferred that before the mother performs the action of "getting angry", an action of the mother "entering the room" is required. Therefore, elements indicating the action of the mother "entering the room" are supplemented at a position before the elements indicating the action of the mother "getting angry". Thus, when generating time-series content from text, the generative AI may need to supplement actions not specified in the text. Also, it may be difficult for the generative AI to infer and appropriately supplement all actions not specified in the text from the text.
[0016] FIG. 2 is a diagram for explaining an example of a technical problem in generating time-series content from text. In FIG. 2, when a user attempts to generate time-series content from text using a generative AI and edit the generated content, it explains that the generation time is proportional to the length of the content. The upper part of FIG. 2 shows a visualization of the structure in the time-axis direction when generating time-series content from the same sentence as in FIG. 1: "The son who was reading a comic in the room was scolded by his mother to study and left home in disgust." The lower part of FIG. 2 shows the case where the time-series content corresponding to the sentence "The son who was reading a comic in the room was scolded by his mother to study and left home in disgust." is a moving image. As shown in the lower part of FIG. 2, since the moving image has a chronological relationship in the time-axis direction, the situation at a specific moment cannot be made into a moving image. Therefore, when wanting to check the frame at a specific timing in the moving image, one has to wait for the content from the beginning of the moving image to the specific timing to be generated. Specifically, the user cannot check the frame at a specific timing without waiting for the content to be generated for a time t obtained by multiplying the number of frames included from the beginning of the moving image to the specific timing by the generation time per frame.
[0017] In contrast, the information processing apparatus according to an embodiment of the present disclosure acquires text input for generating time-series content, and based on the text, generates event sequence information which is information visualizing a structure in which event information corresponding to events in the time-series content is arranged according to the order in which the events occur. In this way, the information processing apparatus enables the visual grasp of the structure in the time-axis direction in the time-series content by the event sequence information. Thereby, the information processing apparatus can save the trouble of the user having to repeatedly re-enter the text in order to appropriately represent the structure in the time-axis direction in the time-series content until the time-series content desired by the user is completed. Therefore, the information processing apparatus can make it easier for the user to generate the time-series content desired by the user.
[0018] In addition, the information processing apparatus generates predicted content corresponding to a predetermined timing in a scene including one or more events based on event sequence information and the like. In this way, the information processing apparatus enables the user to confirm the output prediction of the content for each scene including one or more events. Thereby, the information processing apparatus can save the trouble of waiting for the user until the content from the beginning to a specific timing of the moving image is generated. Therefore, the information processing apparatus can more easily generate the time-series content desired by the user.
[0019] (2. Configuration of Information Processing Apparatus) FIG. 3 is a diagram showing a configuration example of an information processing apparatus according to an embodiment of the present disclosure. As shown in FIG. 3, the information processing apparatus 1 includes an input unit 10, an output unit 20, a communication unit 30, a storage unit 40, and a control unit 50.
[0020] The input unit 10 is an input device that receives various inputs from the outside. The input unit 10 includes an operating device that receives the input operations of the user. The operating device is a device for the user to perform various operations, such as a keyboard, a mouse, and operation keys. When a touch panel is adopted in the information processing apparatus 1, the touch panel is also included in the operating device. In this case, the user performs various operations by touching the screen with a finger or a stylus. For example, the input unit 10 receives an input operation of text input by the user to generate time-series content.
[0021] The output unit 20 is a display device that displays various types of information. The output unit 20 is, for example, a liquid crystal display, an organic EL (Electro Luminescence) display, or the like. Further, the output unit 20 may include an audio output device such as a speaker that outputs audio. When a touch panel is adopted in the information processing apparatus 1, the output unit 20 may be an integrated device with the operating device of the input unit 10. The output unit 20 displays various types of information and provides it to the user. Also, in the following description, the output unit 20 may be described as a screen. For example, the output unit 20 displays the event sequence information generated by the sequence generation unit 52. Further, the output unit 20 displays the predicted content generated by the prediction generation unit 53. Also, when the time-series content is a moving image, the output unit 20 displays the moving image generated by the content generation unit 54. Also, when the time-series content is audio data, the output unit 20 outputs the audio data generated by the content generation unit 54 as audio.
[0022] The communication unit 30 is realized, for example, by a NIC (Network Interface Card) or the like. And the communication unit 30 is connected to a network by wire or wirelessly, and may transmit and receive information, for example, with a terminal device used by a user.
[0023] The storage unit 40 is realized, for example, by a semiconductor memory element such as a RAM (Random Access Memory), a flash memory, or a storage device such as a hard disk or an optical disk. For example, the storage unit 40 stores information regarding various programs (for example, the program according to the embodiment).
[0024] The control unit 50 is a controller, which is realized, for example, by a CPU (Central Processing Unit), MPU (Micro Processing Unit), ASIC (Application Specific Integrated Circuit), FPGA (Field Programmable Gate Array), etc., when various programs stored in the internal storage device of the information processing apparatus 1 are executed with a storage area such as a RAM as a work area. In the example shown in FIG. 3, the control unit 50 includes an acquisition unit 51, a sequence generation unit 52, a prediction generation unit 53, and a content generation unit 54.
[0025] FIG. 4 is a flowchart showing a processing procedure by the information processing apparatus according to an embodiment of the present disclosure. In FIG. 4, the acquisition unit 51 acquires the text input for generating time-series content (step S101). Specifically, the acquisition unit 51 acquires the text input by the user for generating time-series content via the input unit 10. In other words, the text is a character string. The acquisition unit 51 may acquire a character string in any form such as characters, words, phrases, or sentences. Also, the acquisition unit 51 may acquire text described in any language such as English or Japanese.
[0026] Also, in FIG. 4, when the text is acquired by the acquisition unit 51, the sequence generation unit 52 generates event sequence information and global information based on the text acquired by the acquisition unit 51 (step S102). Specifically, the sequence generation unit 52 generates event sequence information, which is information visualizing a structure in which event information corresponding to events in time-series content is arranged according to the order in which the events occur, based on the text acquired by the acquisition unit 51. Also, the sequence generation unit 52 further generates global information corresponding to the scene where the event is occurring based on the text acquired by the acquisition unit 51.
[0027] More specifically, the sequence generation unit 52 includes a text parser that analyzes text. The text parser may be a rule-based mechanism using natural language processing technology or a machine learning model using a neural network. The text parser arranges information in chronological order based on the input text and generates event sequence information and global information corresponding to the text. For example, the sequence generation unit 52 includes a first machine learning model that has been pre-trained to output event sequence information and global information corresponding to text when the text is input as a text parser.
[0028] Also, the sequence generation unit 52 uses the text parser to generate each of the event sequence information and the global information from the text input by the user to generate time-series content. For example, the sequence generation unit 52 uses a first machine learning model that has been pre-trained to generate event sequence information and global information from the text input by the user to generate time-series content. More specifically, the sequence generation unit 52 inputs the text input by the user to generate time-series content into the first machine learning model that has been pre-trained, and generates the event sequence information and the global information output from the first machine learning model.
[0029] The event sequence information is information that visualizes the flow of events occurring in the time-series content. Specifically, the event sequence information is information that visualizes a structure in which event information corresponding to each of various events occurring in the time-series content is arranged in the order in which the events occur. More specifically, the event sequence information is information that visualizes the event information arranged in chronological order according to the order in which the events occur. Specific examples of the event sequence information will be described in detail with reference to FIGS. 6 to 8 and FIG. 14 described later.
[0030] Global information is information related to the entire scene where an event such as location or weather occurs. Specific examples of global information will be described in detail with reference to FIGS. 9 and 10 described later. Also, event information and global information are expressed in a data format that describes structured data such as JSON (JavaScript Object Notation) or YAML (YAML Ain't a Markup Language) as a string.
[0031] Also, in FIG. 4, when the event sequence information and the global information are generated by the sequence generation unit 52, the prediction generation unit 53 generates prediction content based on the event sequence information and the global information generated by the sequence generation unit 52 (step S103). The prediction generation unit 53 generates prediction content corresponding to a predetermined timing in a scene including one or more events based on the event sequence information and the global information generated by the sequence generation unit 52.
[0032] A scene refers to a constituent unit of time-series content including one or more events. A scene is determined by a user dividing event sequence information in units of one or more events. Hereinafter, each of the event sequence information divided by the user may be referred to as "scene information". Scene information includes one or more pieces of event information.
[0033] Prediction content is content corresponding to a predetermined timing of each scene in the time-series content to be generated. In other words, the prediction content is content at the moment representing each scene in the time-series content to be generated. For example, when the time-series content is a moving image, the prediction content may be a still image corresponding to a predetermined timing of each scene in the moving image. Specific examples of the prediction content will be described in detail with reference to FIGS. 12 to 13 and FIGS. 15 to 21 described later.
[0034] Specifically, the prediction generation unit 53 includes a prediction generator that generates prediction content. The prediction generator may be a program created based on rules or a machine learning model using a neural network. For example, the prediction generation unit 53 includes, as a prediction generator, a second machine learning model that has been pre-trained to output second prediction content when scene information, global information, and first prediction content corresponding to the scene to be processed are input.
[0035] Also, the prediction generation unit 53 generates second prediction content from scene information, global information, and first prediction content corresponding to the scene to be processed using the prediction generator. For example, the prediction generation unit 53 generates second prediction content from scene information, global information, and first prediction content corresponding to the scene to be processed using a second machine learning model that has been pre-trained. More specifically, the prediction generation unit 53 generates the second prediction content output from the second machine learning model by inputting the scene information, global information, and first prediction content corresponding to the scene to be processed into the second machine learning model that has been pre-trained.
[0036] Also, in FIG. 4, when the prediction content is generated by the prediction generation unit 53, the content generation unit 54 generates time-series content based on the event sequence information and global information generated by the sequence generation unit 52 and the prediction content generated by the prediction generation unit 53 (step S104). Specifically, the content generation unit 54 includes a content generator that generates time-series content. Hereinafter, a case where the time-series content is a moving image and the content generator is a video generator that generates a moving image will be described. The video generator may be a machine learning model using a neural network (for example, generative AI) or a program that controls a three-dimensional space. For example, when event sequence information, global information, and prediction content are input as the video generator, the content generation unit 54 includes a third machine learning model that has been previously trained to output a moving image corresponding to the event sequence information, global information, and prediction content.
[0037] Also, the content generation unit 54 generates a moving image from the event sequence information, global information, and prediction content using the video generator. For example, the content generation unit 54 generates a moving image from the event sequence information, global information, and prediction content using a third machine learning model that has been previously trained. More specifically, the content generation unit 54 generates a moving image output from the third machine learning model by inputting the event sequence information, global information, and prediction content into the third machine learning model that has been previously trained.
[0038] FIG. 5 is a diagram showing an overview of information processing according to an embodiment of the present disclosure. In FIG. 5, a case where the time-series content is a moving image will be described. In FIG. 5, an acquisition unit 51 acquires an input sentence input by a user U for generating a moving image. In FIG. 5, four events #1 to #4 occur in the moving image to be generated. Among the moving images to be generated, event #1 occurs first, event #2 occurs after event #1, event #3 occurs after event #2, and event #4 occurs after event #3.
[0039] In FIG. 5, a sequence generation unit 52 uses a text parser to generate event sequence information and global information, which are information visualizing a structure in which each of event information #1 to #4 corresponding to each of events #1 to #4 is arranged along a time axis in the order in which events #1 to #4 occur. Specifically, the sequence generation unit 52 generates event sequence information visualizing a structure in which event information #1 is arranged at the earliest position in time along the time axis, event information #2 is arranged at the next earliest position in time after event information #1, event information #3 is arranged at the next earliest position in time after event information #2, and event information #4 is arranged at the next earliest position in time after event information #3.
[0040] Also, an output unit 20 displays a user interface for the user U to edit event information and event sequence information. The user U can edit event information and event sequence information via the user interface. Also, the output unit 20 displays a user interface for the user U to edit global information. The user U can edit global information via the user interface. Note that the user U can also edit the input sentence.
[0041] Also, in FIG. 5, the user U determines scenes #1 to #4 corresponding to each of events #1 to #4 by separating the event sequence information for each event information. That is, in FIG. 5, the scene information #1 corresponding to scene #1 is the event information #1. Also, the scene information #2 corresponding to scene #2 is the event information #2. Also, the scene information #3 corresponding to scene #3 is the event information #3. Also, the scene information #4 corresponding to scene #4 is the event information #4. Also, in FIG. 5, the global information is information common to each scene of events #1 to #4.
[0042] In FIG. 5, the prediction generation unit 53 uses a prediction generator to generate prediction content #1 corresponding to a predetermined timing in scene #1 from the scene information #1 and the global information. Subsequently, the prediction generation unit 53 uses the prediction generator to generate prediction content #2 corresponding to a predetermined timing in scene #2 from the scene information #2, the global information, and the prediction content #1. Subsequently, the prediction generation unit 53 uses the prediction generator to generate prediction content #3 corresponding to a predetermined timing in scene #3 from the scene information #3, the global information, and the prediction content #2. Subsequently, the prediction generation unit 53 uses the prediction generator to generate prediction content #4 corresponding to a predetermined timing in scene #4 from the scene information #4, the global information, and the prediction content #3. FIG. 5 shows a case where each of the prediction contents #1 to #4 is a still image. The output unit 20 displays the prediction content generated by the prediction generation unit 53. Also, the user U can edit the displayed prediction content.
[0043] Also, in FIG. 5, the content generation unit 54 uses a video generator to generate a moving image from the event sequence information, the global information, and the prediction contents #1 to #4. The output unit 20 displays the moving image generated by the content generation unit 54.
[0044] Next, with reference to FIG. 6, event information will be described in detail. FIG. 6 is a diagram showing an example of event information according to an embodiment of the present disclosure. Here, an event refers to a situation in which an acting entity such as a person or character appearing in time-series content performs a predetermined action. In the following, the acting entity appearing in time-series content may be referred to as a "subject object". Also, the predetermined action performed by the acting entity may be referred to as an "action".
[0045] Event information includes object information corresponding to the objects appearing in the event and action information corresponding to the actions of the objects. In FIG. 6, the event information 31 includes, as object information, subject object information 32 corresponding to the subject object. Specifically, the event information 31 includes, as the subject object information 32, information such as "id:0", which is identification information for identifying the subject object, and "name:obj0", which indicates the name of the subject object. Here, both the identification information for identifying the subject object and the name of the subject object are information capable of identifying the object. Thus, event information includes information capable of identifying the object as object information.
[0046] In addition, the event information 31 includes action information 33 corresponding to the actions of the subject object. Specifically, the event information 31 includes, as the action information 33, "id:0" which is the identification information for identifying the subject object, "type:walk" indicating that the type of action is the walking motion, and "state:start" indicating that the action has just started. In this way, the event information may include, as the action information, information that can identify the object that is the subject of the action, information indicating the type of the action, and information indicating the state of the action. Note that, in addition to "start" indicating that the action has just started, the information indicating the state of the action includes "playing" indicating that the action is in progress and "finish" indicating that the action is about to end, etc.
[0047] In addition, the event information includes relationship information indicating the relative positional relationship between a plurality of objects that appear in the event. In FIG. 6, the event information 31 includes, as the relationship information 34, "obj1:0" which is the identification information for identifying the first object that appears in the event, "obj2:1" which is the identification information for identifying the second object, and "near" indicating the relative positional relationship between the two objects that appear in the event. In this way, the event information may include, as the relationship information, information that can identify each of the plurality of objects that appear in the event, and information indicating the relative positional relationship between the plurality of objects. Note that, in addition to "near" indicating that the relative position between the plurality of objects is close, the information indicating the relative positional relationship between the plurality of objects includes "on" indicating that another object is in contact with the upper surface of one object or "above" indicating that another object is located above one object, etc.
[0048] FIG. 7 is a diagram showing an example of a screen of a user interface for receiving an editing operation of scene information by a user according to an embodiment of the present disclosure. In FIG. 7, the output unit 20 displays a screen of a user interface UI1 for enabling an editing operation of scene information by the user. The output unit 20 displays a screen of the user interface UI1 for enabling an editing operation for each of one or more pieces of event information included in the scene information. In FIG. 7, a portion of the screen of the user interface UI1 where one piece of event information is displayed is shown. Specifically, the output unit 20 displays a screen of the user interface UI1 for enabling the user to edit object information, action information, and relationship information. The input unit 10 receives an editing operation for editing scene information from the user via the screen of the user interface UI1 displayed by the output unit 20.
[0049] For example, the output unit 20 displays a screen of the user interface UI1 including an input field F11 for editing information that can identify an object appearing in an event and an edit button B11 for adding an object appearing in the event. Further, the output unit 20 displays a screen of the user interface UI1 including an input field F12 for editing information indicating the type of action and an edit button B12 for adding the type of action.
[0050] Further, the output unit 20 displays a screen of the user interface UI1 including input fields F131 and F132 for editing information that can identify each of a plurality of objects appearing in an event, and an input field F133 for editing information indicating the relative positional relationship between the plurality of objects. Further, the output unit 20 displays a screen of the user interface UI1 including an edit button B13 for adding the relative positional relationship between the plurality of objects.
[0051] FIG. 8 is a diagram showing an example of event sequence information according to an embodiment of the present disclosure. In FIG. 8, a sequence generation unit 52 generates event sequence information 100 which is a node graph including event nodes indicating event information and first edges connecting the event nodes to each other. Specifically, the sequence generation unit 52 includes a first machine learning model M11 that has been pre-trained as a text parser to output a node graph corresponding to the text when the text is input. Further, the sequence generation unit 52 generates event sequence information 100 which is a node graph output from the first machine learning model M11 by inputting the text input by the user for generating a moving image into the pre-trained first machine learning model M11. An output unit 20 displays the event sequence information 100 which is the node graph generated by the sequence generation unit 52. In FIG. 8, events are associated with event nodes, and the order of occurrence of the events is represented by the connections (edges) between the event nodes.
[0052] In FIG. 8, the user can view and edit the event sequence information 100 which is a node graph via a GUI (Graphical User Interface) tool. The user can perform editing operations equivalent to those of a general node graph editor (for example, editing properties, adding nodes, or deleting nodes, etc.) on the node graph via the GUI tool. Further, the user can change the scene division by drag and drop.
[0053] In FIG. 8, the sequence generation unit 52 generates event sequence information 100, which is a node graph visualizing a structure in which event nodes are arranged vertically according to the order in which each event occurs. In FIG. 8, the vertical direction corresponds to the time axis direction. Specifically, in FIG. 8, it shows how time flows from top to bottom. That is, it indicates that the event corresponding to the event node located on the upper side of FIG. 8 occurs earlier in time, and the event corresponding to the event node located on the lower side of FIG. 8 occurs later in time. As shown in FIG. 8, when event sequence information is represented by a node graph, since the event sequence information does not include specific time information, the user can confirm the general flow of events even without knowing the duration of each event.
[0054] Note that the sequence generation unit 52 may generate a node graph visualizing a structure in which event nodes are arranged horizontally according to the order in which each event occurs. In this case, the horizontal direction corresponds to the time axis direction. Specifically, the node graph shows how time flows from left to right. That is, it indicates that the event corresponding to the event node located on the left side of the node graph occurs earlier in time, and the event corresponding to the event node located on the right side of the node graph occurs later in time.
[0055] In FIG. 8, the sequence generation unit 52 generates a node graph in which five event nodes N1 to N5 are connected by edges. The event node N1 is an event element corresponding to an event in which character A appearing in the time-series content makes a phone call. The event node N2 is an event element corresponding to an event in which character B appearing in the time-series content answers a phone call from character A. The event node N3 is an event element corresponding to an event in which character B talks on the phone. The event node N4 is an event element corresponding to an event in which character A talks on the phone. The event node N5 is an event element corresponding to an event in which character B leaves home. The sequence generation unit 52 generates a node graph in which the event node N2 is connected by an edge below the event node N1, the event node N3 is connected by an edge below the event node N2, the event node N4 is connected by an edge below the event node N3, and the event node N5 is connected by an edge below the event node N4.
[0056] In addition, the user can create a layer structure for the node graph generated by the sequence generation unit 52. In FIG. 8, the input unit 10 receives a selection operation from the user to select the event node N5 displayed on the screen. When the sequence generation unit 52 receives a selection operation for the event node N5 via the input unit 10, it generates a layer structure in which the event node N5B corresponding to the event that character B heads towards the entrance is connected by an edge below the event node N5A corresponding to the event that character B takes a bag. The output unit 20 displays the layer structure generated by the sequence generation unit 52. The layer structure generated by the sequence generation unit 52 corresponds to the event corresponding to the event node N5 selected by the user being subdivided into more detailed events. The user can refer to the layer structure displayed by the output unit 20 and edit the node graph. For example, when the sequence generation unit 52 receives a selection operation from the user to select a layer structure via the input unit 10, it deletes the event node N5 from the node graph and adds the selected layer structure to the node graph. The sequence generation unit 52 generates event sequence information 100, which is a new node graph to which the layer structure selected by the user is added. In this way, the sequence generation unit 52 generates event sequence information, which is information that can be edited by the user.
[0057] In addition, the user can expand the events of the node graph generated by the sequence generation unit 52. In FIG. 8, when the sequence generation unit 52 receives an operation to add an event node N6 below the event node N4 displayed on the screen via the input unit 10, the sequence generation unit 52 adds the event node N6 below the event node N4. The event node N6 is an event element corresponding to the event that character B hangs up the phone. The sequence generation unit 52 generates event sequence information 100, which is a new node graph in which the event node N6 has been added by the user. As shown in FIG. 8, both the event node N5 and the event node N6 are arranged directly below the event node N4. That is, the event node N5 and the event node N6 are arranged in parallel in a direction orthogonal to the time axis direction. This corresponds to the fact that two events corresponding to the event node N5 and the event node N6 respectively proceed simultaneously in the time-series content to be generated. The sequence generation unit 52 generates event sequence information 100, which is a new node graph having a structure in which the event node N5 and the event node N6 are arranged in parallel in a direction orthogonal to the time axis direction. In this way, the sequence generation unit 52 generates event sequence information, which is information visualizing a structure in which event information corresponding to each of a plurality of events proceeding simultaneously is arranged in parallel in a direction orthogonal to the time axis direction.
[0058] Also, in FIG. 8, when the sequence generation unit 52 receives an operation to add an event node N7 below the event node N5 displayed on the screen via the input unit 10, the sequence generation unit 52 adds the event node N7 below the event node N5. The event node N7 is an event element corresponding to an event in which the character A greets. The sequence generation unit 52 generates event sequence information 100, which is a new node graph in which the event node N7 is added by the user. In this way, the sequence generation unit 52 generates event sequence information 100, which is a new node graph having a structure in which the event node N7 is arranged in series below the event node N5. In this way, the sequence generation unit 52 generates event sequence information, which is information visualizing a structure in which event elements corresponding to each of a predetermined event and another event occurring subsequently to the predetermined event are arranged in series in the time axis direction.
[0059] Next, the global information will be described in detail with reference to FIG. 9. FIG. 9 is a diagram showing an example of the global information according to the embodiment of the present disclosure. The global information includes information indicating a location, weather, or time corresponding to the scene where the event is occurring. Here, in the scene corresponding to the event, objects existing in the environment surrounding the acting entity, other characters or characters different from the acting entity appear. Hereinafter, objects existing in the environment surrounding the acting entity, other characters or characters different from the acting entity may be described as "object objects".
[0060] In FIG. 9, the global information 41 includes location information 42 indicating the location corresponding to the scene where the event occurs, weather information 43 indicating the weather corresponding to the scene where the event occurs, time information 44 indicating the time corresponding to the scene where the event occurs, and object object information 45 corresponding to the object objects appearing in the scene where the event occurs. Specifically, the global information 41 includes, as the location information 42, information "location:Shinagawa" indicating the name of the location corresponding to the scene where the event occurs. In addition to "Shinagawa" indicating Shinagawa, the location information includes information indicating all place names.
[0061] Also, the global information 41 includes, as the weather information 43, information "weather:sunny" indicating the weather corresponding to the scene where the event occurs. In addition to "sunny" indicating sunny weather, the weather information includes "rainy" indicating rainy weather, "cloudy" indicating cloudy weather, and so on.
[0062] Also, the global information 41 includes, as the time information 44, information "time:DAY" indicating the time zone corresponding to the scene where the event occurs, information "season:summer" indicating the season, and information "era:Edo" indicating the era. Thus, the global information includes, as the time information, information indicating the time zone corresponding to the scene where the event occurs, information indicating the season, and information indicating the era. In addition to "DAY" indicating daytime, the information indicating the time zone includes "NIGHT" indicating nighttime, and so on. In addition to "summer" indicating summer, the information indicating the season includes "autumn" indicating autumn, "winter" indicating winter, and "spring" indicating spring. In addition to "Edo" indicating the Edo period, the information indicating the era can include information indicating all eras.
[0063] In addition, the global information 41 includes, as object object information 45, the identification information "id:0" that identifies the object object, the "name:obj0" that indicates the name of the object object, the "category:vehicle" that indicates the category of the object object, and the tag information "tags:“car”,“red”" that indicates the characteristics of the object object. Further, the global information includes, as object object information, information that can identify the object object and information that indicates the category of the object object.
[0064] FIG. 10 is a diagram showing an example of a screen of a user interface for receiving an editing operation of global information by a user according to an embodiment of the present disclosure. In FIG. 10, the output unit 20 displays a screen of a user interface UI2 for enabling the user to edit the global information. Specifically, the output unit 20 displays a screen of a user interface UI2 for enabling the user to edit location information, weather information, and object object information. The input unit 10 receives an editing operation for editing the global information from the user via the screen of the user interface UI2 displayed by the output unit 20.
[0065] For example, the output unit 20 displays a screen of a user interface UI1 including an input field F21 for editing location information. Further, the output unit 20 displays a screen of a user interface UI1 including an input field F22 for editing weather information. Further, the output unit 20 displays a screen of a user interface UI2 including an input field F23 for editing object object information and an edit button B23 for adding an object object.
[0066] FIG. 11 is a diagram for explaining an example of the generation process of prediction content according to an embodiment of the present disclosure. In FIG. 11, the prediction generation unit 53 inputs scene information corresponding to the scene to be processed, global information, and first prediction content corresponding to the scene one before the scene to be processed into a second machine learning model that has been learned in advance, and generates second prediction content corresponding to the scene to be processed output from the second machine learning model. In this way, the prediction generation unit 53 generates second prediction content corresponding to the scene to be processed based on the first prediction content corresponding to the scene one before the scene to be processed.
[0067] FIG. 12 is a diagram for explaining the generation process of the integrated scene graph according to an embodiment of the present disclosure. In FIG. 12, the prediction generation unit 53 generates a scene graph including an object node indicating object information, a relationship node indicating relationship information, and a second edge connecting the object node and the relationship node based on object information corresponding to the objects appearing in the scene, action information corresponding to the actions of the objects, and relationship information indicating the relative positional relationship between a plurality of objects appearing in the scene, and generates prediction content that is an integrated scene graph 57 obtained by integrating a first scene graph 55 corresponding to the scene one before the scene to be processed and a second scene graph 56 corresponding to the scene to be processed.
[0068] In FIG. 12, the prediction generation unit 53 includes three modules: a relational analogy device, a scene graph generator, and a scene graph integrator. Specifically, the prediction generation unit 53 newly generates relational information from the object information and action information included in the scene information using the relational analogy device, and adds the newly generated relational information to the scene information. For example, when an event such as "Character A picks up a book." is included in a scene, it is presumed that Character A is near the book in order for Character A to pick up the book. In this case, the prediction generation unit 53 generates relational information indicating that Character A and the book are near each other, and adds the relational information indicating that Character A and the book are near each other to the scene information. The relational analogy device may be a program created based on rules or a machine learning model using a neural network. For example, the relational analogy device is a fourth machine learning model that has been previously learned to output relational information when object information and action information included in scene information are input.
[0069] Also, the prediction generation unit 53 uses the scene graph generator to generate a scene graph including an object node indicating object information, a relational node indicating relational information, and a second edge connecting the object node and the relational node from the object information corresponding to the objects appearing in the scene, the action information corresponding to the actions of the objects, and the relational information indicating the relative positional relationship between the plurality of objects appearing in the scene. The scene graph generator may be a program created based on rules or a machine learning model using a neural network. For example, the scene graph generator is a fifth machine learning model that has been previously learned to output a scene graph including an object node, a relational node, and a second edge when object information, action information, and relational information included in scene information are input.
[0070] For example, the prediction generation unit 53 uses a scene graph generator to generate a first scene graph corresponding to the scene one before the scene to be processed. In FIG. 12, the prediction generation unit 53 generates a first scene graph 55 in which an object node indicating character A, an object node indicating a door, and a relationship node indicating that character A and the door are close to each other are connected by edges. For example, the prediction generation unit 53 inputs the subject object information, action information, and relationship information included in the scene information of the previous scene into the fifth machine learning model, thereby generating the first scene graph 55 output from the fifth machine learning model.
[0071] In addition, the prediction generation unit 53 uses a scene graph generator to generate a second scene graph corresponding to the scene to be processed. In FIG. 12, the prediction generation unit 53 generates a second scene graph 56 in which an object node indicating character A, an object node indicating a desk, and a relationship node indicating that character A and the desk are close to each other are connected by edges. For example, the prediction generation unit 53 inputs the subject object information 32A, action information 33A, and relationship information 34A included in the scene information 31A of the scene to be processed into the fifth machine learning model, thereby generating the second scene graph 56 output from the fifth machine learning model.
[0072] Further, the prediction generation unit 53 generates prediction content, which is an integrated scene graph obtained by integrating a first scene graph corresponding to a scene one before the scene to be processed and a second scene graph corresponding to the scene to be processed, using a scene graph integrator. The scene graph integrator may be a program created based on rules or a machine learning model using a neural network. For example, the scene graph integrator is a sixth machine learning model that has been pre-trained to output an integrated scene graph obtained by integrating a first scene graph and a second scene graph when the first scene graph corresponding to a scene one before the scene to be processed and the second scene graph corresponding to the scene to be processed are input. In FIG. 12, the prediction generation unit 53 generates an integrated scene graph 57 by integrating the object node indicating character A included in the first scene graph 55 and the object node indicating character A included in the second scene graph 56 into one object node. For example, the prediction generation unit 53 inputs the first scene graph 55 and the second scene graph 56 into the sixth machine learning model to generate the integrated scene graph 57 output from the sixth machine learning model.
[0073] FIG. 13 is a diagram for explaining an editing operation of an integrated scene graph by a user according to an embodiment of the present disclosure. In FIG. 13, the output unit 20 displays the integrated scene graph 57 generated by the prediction generation unit 53. The user can view and edit the integrated scene graph 57 via the GUI tool. The user can perform editing operations equivalent to those of a general node graph editor (for example, editing properties, adding nodes, or deleting nodes, etc.) on the integrated scene graph via the GUI tool. For example, the user can edit the content of the relationship information indicated by the relationship nodes included in the integrated scene graph 57 generated by the prediction generation unit 53 (Example 1). Also, the user can add new object nodes and relationship nodes to the integrated scene graph 57 generated by the prediction generation unit 53 (Example 2). Further, the user can delete the object nodes and relationship nodes included in the integrated scene graph 57 generated by the prediction generation unit 53 (Example 3).
[0074] (3. Modification Example) Incidentally, several modification examples can be given for the embodiments of the present disclosure described above.
[0075] In the above-described embodiment, the case where the sequence generation unit 52 generates the event sequence information 100 which is a node graph has been described with reference to FIG. 8. In the modification example, with reference to FIG. 14, the case where the sequence generation unit 52 generates the event sequence information 200 which is information visualizing a structure in which event information corresponding to events in time-series content is arranged along the time axis will be described.
[0076] FIG. 14 is a diagram showing an example of event sequence information according to a modification of the embodiment of the present disclosure. In FIG. 14, the sequence generation unit 52 generates event sequence information 200, which is information visualizing a structure in which event information corresponding to events in time-series content is arranged along the time axis. Specifically, the sequence generation unit 52 generates event sequence information 200, which is information visualizing a structure in which the length of each event element indicating event information in the direction of the time axis corresponds to the duration of the event, and each of the event elements is arranged at a position on the time axis corresponding to the start time of each event.
[0077] Each of the event elements M1 to M7 shown in FIG. 14 corresponds to each of the event nodes N1 to N7 shown in FIG. 8. In FIG. 14, the main objects (character A and character B), which are the actors of the events, are arranged vertically, and the events corresponding to each of the actors are arranged horizontally, which is different from FIG. 8. For example, in FIG. 14, the event elements M1, M4, and M7 in which character A is the actor are arranged in the band of character A. Also, the event elements M2, M3, M5, and M6 in which character B is the actor are arranged in the band of character B. In FIG. 14, the horizontal direction corresponds to the time axis direction. Specifically, in FIG. 14, it shows that time flows from left to right. That is, it shows that the events corresponding to the event elements located on the left side of FIG. 14 occur earlier in time, and the events corresponding to the event elements located on the right side of FIG. 14 occur later in time. In FIG. 14, it is different from FIG. 8 in that each of the event elements M1 to M7 is arranged along the time axis. More specifically, in FIG. 14, it is different from FIG. 8 in that each of the event elements M1 to M7 is arranged at a position on the time axis corresponding to the start time of the event corresponding to each of the event elements M1 to M7. Also, in FIG. 14, it is different from FIG. 8 in that the lengths L1 to L7 of each of the event elements M1 to M7 in the direction of the time axis correspond to the durations of the events corresponding to each of the event elements M1 to M7. Also, in FIG. 14, it is different from FIG. 8 in that each of the event elements M1 to M7 is not connected by an edge.
[0078] As described above, in FIG. 14, time management is performed using the start time of an event and the duration of the event. For this reason, in FIG. 14, the sequence generation unit 52 includes an event parser that generates the start time of an event and the duration of the event, separately from the text parser. For example, the sequence generation unit 52 includes, as an event parser, a seventh machine learning model that has been pre-trained to output information indicating the start time and duration of each event in a moving image when text is input. The sequence generation unit 52 inputs the text input by the user for generating the moving image into the seventh machine learning model that has been pre-trained, and thereby obtains information indicating the start time and duration of each event output from the seventh machine learning model.
[0079] In addition, the sequence generation unit 52 includes, as a text parser, a first machine learning model M2 that has been pre-trained to output information visualizing a structure in which event information corresponding to events in a moving image is arranged along a time axis when text and information indicating the start time and duration of each event are input. More specifically, the sequence generation unit 52 includes, as a text parser, a first machine learning model M2 that has been pre-trained to output information visualizing a structure in which the length of an event element indicating event information in the direction of the time axis corresponds to the duration of the event and each of the event elements is arranged at a position on the time axis corresponding to the start time of each event when text and information indicating the start time and duration of each event are input. Further, the sequence generation unit 52 inputs the text input by the user for generating the moving image into the first machine learning model M2 that has been pre-trained, and thereby generates event sequence information 200, which is information visualizing a structure in which event information corresponding to events in the moving image output from the first machine learning model M2 is arranged along a time axis. Further, the output unit 20 displays the event sequence information 200 generated by the sequence generation unit 52.
[0080] In FIG. 14, the user can view and edit the event sequence information 200 via the GUI tool. The user can add, delete, or edit event elements with respect to the event sequence information 200 via the GUI tool. Also, the user can change the scene separation by drag and drop.
[0081] Also, similar to FIG. 8, the user can create a layer structure for the event sequence information 200 generated by the sequence generation unit 52. In FIG. 14, the input unit 10 receives a selection operation from the user to select the event element M5 displayed on the screen. When the sequence generation unit 52 receives a selection operation for the event element M5 via the input unit 10, it generates a layer structure in which an event element M5B (not shown) corresponding to the event that character B heads towards the entrance is arranged to the right of an event element M5A (not shown) corresponding to the event that character B takes a bag. The output unit 20 displays the layer structure generated by the sequence generation unit 52. The layer structure generated by the sequence generation unit 52 corresponds to the events corresponding to the event element M5 selected by the user being subdivided into more detailed events. The user can refer to the layer structure displayed by the output unit 20 and edit the event sequence information #2. For example, when the sequence generation unit 52 receives a selection operation from the user to select a layer structure via the input unit 10, it deletes the event element M5 from the displayed event sequence information 200 and adds the selected layer structure to the event sequence information 200. The sequence generation unit 52 generates new event sequence information 200 with the layer structure selected by the user added thereto. In this way, the sequence generation unit 52 generates event sequence information that can be edited by the user.
[0082] In addition, the user can expand the events of the event sequence information 200 generated by the sequence generation unit 52. In FIG. 14, when the sequence generation unit 52 receives an operation to add an event element M6 below the event element M5 displayed on the screen via the input unit 10, the sequence generation unit 52 adds the event element M6 below the event element M5. The sequence generation unit 52 generates new event sequence information 200 in which the event element M6 has been added by the user. As shown in FIG. 14, both the event element M5 and the event element M6 have the same start time t5 at the start of the event. That is, the event element M5 and the event element M6 are arranged in parallel in a direction orthogonal to the time axis direction. This corresponds to the situation where two events corresponding to the event element M5 and the event element M6 respectively proceed simultaneously in the time-series content to be generated. The sequence generation unit 52 generates new event sequence information 200 having a structure in which the event element M5 and the event element M6 are arranged in parallel in a direction orthogonal to the time axis direction. In this way, the sequence generation unit 52 generates event sequence information, which is information visualizing a structure in which event information corresponding to each of a plurality of events proceeding simultaneously is arranged in parallel in a direction orthogonal to the time axis direction.
[0083] In addition, in the above-described embodiment, the case where the prediction generation unit 53 generates prediction content that is an integrated scene graph has been described with reference to FIG. 12. In a modified example, the case where the prediction generation unit 53 generates prediction content other than the integrated scene graph will be described with reference to FIGS. 15 to 21.
[0084] FIG. 15 is a diagram showing an example of a map according to a modification of an embodiment of the present disclosure. FIG. 15 shows a top view 61 of a map generated by the prediction generation unit 53. In FIG. 15, the prediction generation unit 53 generates prediction content, which is a three-dimensional map in which objects appearing in the scene are arranged in a virtual space of a scene corresponding to the scene, based on the integrated scene graph. The output unit 20 displays the three-dimensional map generated by the prediction generation unit 53. The output unit 20 displays a two-dimensional map of the three-dimensional map generated by the prediction generation unit 53 viewed from an arbitrary viewpoint. For example, the output unit 20 displays a top view or a side view of the three-dimensional map generated by the prediction generation unit 53. In FIG. 15, the prediction generation unit 53 generates prediction content, which is a three-dimensional map in which the object O1 of the character A, the object O2 of the door, and the object O3 of the desk appearing in the scene to be processed are arranged in a virtual space in which the room R1 corresponding to the scene to be processed is visualized. The output unit 20 displays the top view 61 of the three-dimensional map generated by the prediction generation unit 53.
[0085] Specifically, the prediction generation unit 53 generates a three-dimensional map using a program created based on rules or a machine learning model using a neural network. For example, the prediction generation unit 53 includes an eighth machine learning model that has been pre-trained to output a three-dimensional map in which objects appearing in the scene are arranged in a virtual space of a scene corresponding to the scene when the integrated scene graph is input. Further, the prediction generation unit 53 generates a three-dimensional map output from the eighth machine learning model by inputting the integrated scene graph into the eighth machine learning model that has been pre-trained. Note that the prediction generation unit 53 may generate a two-dimensional map of the three-dimensional map viewed from an arbitrary viewpoint instead of generating the three-dimensional map. For example, the prediction generation unit 53 may generate a top view or a side view of the three-dimensional map.
[0086] FIG. 16 is a diagram showing an example of an editing operation of a map by a user according to a modification of an embodiment of the present disclosure. FIG. 16 is a top view 61 of the map shown in FIG. 15. In FIG. 16, the user can view and edit the top view 61 of the map via the GUI tool. For example, the user can move or rotate the objects arranged on the map by performing a drag-and-drop operation on the objects arranged on the map via the GUI tool. In FIG. 16, the state where the user moves the object O1 of character A by performing a drag-and-drop operation on the object O1 of character A is shown.
[0087] FIG. 17 is a diagram showing an example of an editing operation of a map by a user according to a modification of an embodiment of the present disclosure. FIG. 17 is a side view 62 of the map shown in FIG. 15. In FIG. 17, the user can view and edit the side view 62 of the map via the GUI tool. For example, the user can check the height of each object arranged on the map from the side view 62 of the map. Similarly to the top view 61 of the map, the user can move or rotate the objects arranged on the map by performing a drag-and-drop operation on the objects arranged on the map via the GUI tool.
[0088] FIG. 18 is a diagram for explaining a generation process of an image including a composition corresponding to camera information according to a modification of an embodiment of the present disclosure. The prediction generation unit 53 generates second camera information corresponding to the camera work of the virtual camera when shooting the scene to be processed, based on the first scene information corresponding to the scene one before the scene to be processed and the first camera information corresponding to the camera work of the virtual camera when shooting the scene one before the scene to be processed with the virtual camera. Then, based on the generated second camera information, an image including a composition corresponding to the second camera information is searched for, and prediction content is generated based on the searched image.
[0089] In FIG. 18, the prediction generation unit 53 includes two modules: a camera work meta generator and a camera work template searcher. The camera work meta generator is a module that generates camera information corresponding to the camera work of a virtual camera when a predetermined scene is photographed with the virtual camera. The prediction generation unit 53 uses the camera work meta generator to generate camera information corresponding to the camera work of the virtual camera when a predetermined scene is photographed with the virtual camera. The camera information includes information indicating a shooting target (e.g., a subject object), a shot type, a shot size, a shot direction, and a shot angle when each scene is photographed with the virtual camera. For example, shot types include static, push-in, pan, etc. Shot sizes include extreme close-up shot, close-up shot, medium shot, cowboy shot, full shot, etc. Shot directions include front, over-the-shoulder, side, etc. Shot angles include high angle, eye level, shoulder level, hip level, etc.
[0090] The camera work meta generator may be a rule-based program or a machine learning model using a neural network. For example, the prediction generation unit 53, as a camera work meta generator, when the first scene information corresponding to the scene one before the scene to be processed and the first camera information corresponding to the camera work of the virtual camera when the scene one before the scene to be processed is photographed by the virtual camera are input, it has a ninth machine learning model pre-trained to output the second camera information corresponding to the camera work of the virtual camera when the scene to be processed is photographed by the virtual camera. For example, in FIG. 18, the prediction generation unit 53 inputs the first scene information 31B corresponding to the scene one before the scene to be processed, and the first camera information 71A and 71B corresponding to the camera work of the virtual camera when the scene one before the scene to be processed is photographed by the virtual camera into the ninth machine learning model pre-trained, and generates the second camera information 72 corresponding to the camera work of the virtual camera when the scene to be processed is photographed by the virtual camera output from the ninth machine learning model.
[0091] In addition, when the prediction generation unit 53 includes camera information 71B indicating an instruction sentence for the camera work from the user such as "camera_operation" shown in FIG. 18, it makes an estimation considering the instruction sentence in the camera work meta generator. Here, in "camera_operation", it assumes the input of the name of camera work such as zoom or pan and instructions by sentences such as "shoot character A in the background".
[0092] Also, the camera work template searcher is a module that searches for a template image prepared in advance from camera information. The template image is an image that includes a composition indicated by specific camera information and is associated with the specific camera information. The camera work template searcher may be a program created based on rules or a machine learning model using a neural network. The prediction generation unit 53 uses the camera work template searcher to search for a template image 73 that includes a composition corresponding to the generated second camera information 72 based on the generated second camera information 72. The prediction generation unit 53 generates prediction content based on the searched template image 73. For example, the prediction generation unit 53 may use the searched template image 73 as the prediction content.
[0093] FIG. 19 is a diagram for explaining a generation process of an image including a composition corresponding to camera information according to a modification of the embodiment of the present disclosure. The prediction generation unit 53 generates second camera information corresponding to the camera work of the virtual camera when the scene to be processed is photographed with the virtual camera based on the first scene information corresponding to the scene one before the scene to be processed and the first camera information corresponding to the camera work of the virtual camera when the scene one before the scene to be processed is photographed with the virtual camera, and generates prediction content, which is an image including a composition corresponding to the second camera information, based on the second scene information corresponding to the scene to be processed and the generated second camera information.
[0094] In FIG. 19, the prediction generation unit 53 includes two modules: a camera work meta generator and a composition graph generator. Similar to FIG. 18, the prediction generation unit 53 uses the camera work meta generator to generate camera information corresponding to the camera work of a virtual camera when shooting a predetermined scene with the virtual camera. For example, the prediction generation unit 53 includes the ninth machine learning model described in FIG. 18 as the camera work meta generator. The prediction generation unit 53 inputs the first scene information 31B and the first camera information 71A and 71B into the ninth machine learning model that has been learned in advance, thereby generating the second camera information 72 output from the ninth machine learning model.
[0095] Also, the composition graph generator is an image generation model generated using a neural network. For example, the composition graph generator is an image generation model that has been learned in advance to output a picture content image (for example, a rough image such as a sketch of a painting) including a composition corresponding to the camera information when scene information and camera information are input. The prediction generation unit 53 uses the composition graph generator to generate prediction content, which is a picture content image 74 including a composition corresponding to the second camera information 72, based on the second scene information 31A (not shown) corresponding to the scene to be processed and the generated second camera information 72.
[0096] FIG. 20 is a diagram for explaining the generation process of an image including an expression indicating an emotion corresponding to the emotion information of a character according to a modification of the embodiment of the present disclosure. The prediction generation unit 53 generates second emotion information corresponding to the emotion of the character appearing in the scene to be processed based on the first scene information corresponding to the scene one before the scene to be processed and the first emotion information corresponding to the emotion of the character appearing in the scene one before the scene to be processed, and generates prediction content, which is an image including an expression indicating the emotion corresponding to the second emotion information of the character appearing in the scene to be processed, based on the second scene information corresponding to the scene to be processed and the generated second emotion information.
[0097] In FIG. 20, the prediction generation unit 53 includes two modules: an emotion meta-generator and an emotion rough generator. The emotion meta-generator is a module that generates second emotion information corresponding to the emotion of the character appearing in the scene to be processed based on the first scene information corresponding to the scene one before the scene to be processed and the first emotion information corresponding to the emotion of the character appearing in the scene one before the scene to be processed. The emotion meta-generator may be a program created based on rules or a machine learning model using a neural network. For example, the prediction generation unit 53 includes a tenth machine learning model that has been pre-trained to output second emotion information when the first scene information and the first emotion information are input as the emotion meta-generator. In FIG. 20, the prediction generation unit 53 inputs the first scene information 31B corresponding to the scene one before the scene to be processed, and the first emotion information 81A and 81B corresponding to the emotion of the character appearing in the scene one before the scene to be processed into the tenth machine learning model that has been pre-trained, and generates the second emotion information 82 corresponding to the emotion of the character appearing in the scene to be processed output from the tenth machine learning model.
[0098] Note that when the prediction generation unit 53 includes the emotion information 81B indicating an instruction sentence for the emotion from the user such as "emotion_operation" shown in FIG. 20, the emotion meta-generator performs an estimation considering the instruction sentence.
[0099] The emotion rough generator is an image generation model generated using a neural network. For example, the emotion rough generator is an image generation model that has been pre-trained to output a picture content image including an expression indicating the emotion corresponding to the emotion information when the scene information and the emotion information are input. In FIG. 20, the prediction generation unit 53 uses the emotion rough generator to generate prediction content, which is a picture content image 83 including an expression indicating the emotion corresponding to the second emotion information 82, based on the second scene information 31A (not shown) corresponding to the scene to be processed and the generated second emotion information 82.
[0100] FIG. 21 is a diagram for explaining a generation process of an image including a pose corresponding to pose information of a character according to a modification of an embodiment of the present disclosure. The prediction generation unit 53 generates pose information corresponding to the pose of a character appearing in the scene to be processed based on the target scene information corresponding to the scene to be processed, and based on the generated pose information, generates prediction content that is an image including the pose corresponding to the pose information of the character appearing in the scene to be processed.
[0101] In FIG. 21, the prediction generation unit 53 includes two modules: a pose meta-generator and a pose generator. The pose meta-generator generates pose information corresponding to the pose of a character appearing in the scene from the object information and action information included in the scene information. The pose meta-generator may be a program created based on rules or a machine learning model using a neural network. For example, the prediction generation unit 53 includes an eleventh machine learning model that has been previously trained to output pose information corresponding to the pose of a character appearing in the scene when the scene information is input. In FIG. 21, the prediction generation unit 53 inputs the target scene information 31A corresponding to the scene to be processed into the eleventh machine learning model that has been previously trained, and generates pose information 91 corresponding to the pose of a character appearing in the scene to be processed, which is output from the eleventh machine learning model.
[0102] The pose generator generates a three-dimensional image in which a character appearing in a scene assumes a pose corresponding to pose information. The pose generator may be a program created based on rules or a machine learning model using a neural network. For example, when pose information is input, the prediction generation unit 53 includes a twelfth machine learning model that has been previously trained to output an image including a pose corresponding to the pose information of the character appearing in the scene. In FIG. 21, the prediction generation unit 53 inputs the generated pose information 91 into the eleventh machine learning model that has been previously trained, and generates prediction content that is a three-dimensional image 92 including a pose corresponding to the pose information of the character appearing in the scene to be processed, which is output from the eleventh machine learning model.
[0103] (4. Effect) As described above, the information processing apparatus 1 according to the embodiment of the present disclosure includes an acquisition unit 51 and a sequence generation unit 52. The acquisition unit 51 acquires text input for generating time-series content. The sequence generation unit 52 generates event sequence information, which is information visualizing a structure in which event information corresponding to events in the time-series content is arranged in the order in which the events occur, based on the text.
[0104] In this way, the information processing apparatus 1 enables the visual grasping of the structure in the time-axis direction in the time-series content by the event sequence information. As a result, the information processing apparatus 1 can save the trouble of the user repeatedly inputting the text to appropriately represent the structure in the time-axis direction in the time-series content until the time-series content desired by the user is completed. Therefore, the information processing apparatus 1 can make it easier for the user to generate the time-series content desired by the user.
[0105] In addition, the sequence generation unit 52 generates event sequence information, which is information visualizing a structure in which event information corresponding to each of a plurality of events proceeding simultaneously is arranged in parallel in a direction orthogonal to the time-axis direction.
[0106] As a result, the information processing apparatus 1 can visually grasp parallel operations that are difficult to appropriately represent only by text based on the event sequence information, so that the user can more easily generate time-series content as desired.
[0107] In addition, the sequence generation unit 52 generates event sequence information, which is information editable by the user.
[0108] As a result, the information processing apparatus 1 can be edited by the user event sequence information, so that the user can more easily generate time-series content as desired.
[0109] In addition, the event information includes object information corresponding to an object appearing in the event and action information corresponding to the action of the object.
[0110] As a result, the information processing apparatus 1 can visually grasp the object information and the action information, so that the user can more easily generate time-series content as desired.
[0111] In addition, the event information includes relationship information indicating the relative positional relationship between a plurality of objects appearing in the event.
[0112] As a result, the information processing apparatus 1 can visually grasp the relationship information, so that the user can more easily generate time-series content as desired.
[0113] In addition, the sequence generation unit 52 generates event sequence information, which is a node graph including event nodes indicating event information and first edges connecting the event nodes to each other.
[0114] As a result, the information processing apparatus 1 can visually grasp the structure in the time axis direction in the time-series content by means of the node graph.
[0115] In addition, the sequence generation unit 52 generates event sequence information which is information obtained by visualizing the structure in which the event information is arranged along the time axis.
[0116] As a result, the information processing apparatus 1 can visually grasp the structure in the time axis direction in the time-series content by means of the information obtained by visualizing the structure in which the event information is arranged along the time axis.
[0117] In addition, the sequence generation unit 52 generates event sequence information which is information obtained by visualizing the structure in which the lengths of the event elements indicating the event information in the direction of the time axis correspond to the durations of the events and the respective event elements are arranged at the positions on the time axis corresponding to the start times of the respective events.
[0118] As a result, the information processing apparatus 1 can visually grasp the duration of the event and the start time of the event.
[0119] In addition, the sequence generation unit 52 further generates global information corresponding to the scene where the event occurs based on the text.
[0120] As a result, the information processing apparatus 1 can visually grasp the global information, so that the user can more easily generate the time-series content desired by the user.
[0121] In addition, the global information includes information indicating the location, weather or time corresponding to the scene.
[0122] As a result, the information processing apparatus 1 can visually grasp the information indicating the location, weather or time corresponding to the scene.
[0123] Further, the information processing apparatus 1 further includes a prediction generation unit 53. The prediction generation unit 53 generates prediction content corresponding to a predetermined timing in a scene including one or more events based on the event sequence information and the global information.
[0124] In this way, the information processing apparatus 1 enables the user to confirm the output prediction of the content for each scene including one or more events. Thereby, the information processing apparatus 1 can save the trouble of waiting for the content from the beginning to a specific timing of the moving image to be generated. Therefore, the information processing apparatus 1 can more easily generate the time-series content desired by the user.
[0125] Also, the prediction generation unit 53 generates second prediction content corresponding to the scene to be processed based on the first prediction content corresponding to the scene one before the scene to be processed.
[0126] Thereby, the information processing apparatus 1 can appropriately generate the second prediction content corresponding to the scene to be processed in consideration of the first prediction content corresponding to the previous scene.
[0127] Also, the prediction generation unit 53 generates a scene graph including an object node indicating object information, a relationship node indicating relationship information, and a second edge connecting the object node and the relationship node based on the object information corresponding to the object appearing in the scene, the action information corresponding to the action of the object, and the relationship information indicating the relative positional relationship between a plurality of objects appearing in the scene, and generates prediction content that is an integrated scene graph obtained by integrating the first scene graph corresponding to the scene one before the scene to be processed and the second scene graph corresponding to the scene to be processed.
[0128] Thereby, the information processing apparatus 1 enables the user to confirm the output prediction of the content for each scene by the integrated scene graph.
[0129] In addition, the prediction generation unit 53 generates prediction content, which is a map in which objects appearing in the scene are arranged in the virtual space of the scene corresponding to the scene, based on the integrated scene graph.
[0130] Thereby, the information processing apparatus 1 enables the user to confirm the output prediction of the content for each scene by means of the map.
[0131] Also, the prediction generation unit 53 generates second camera information corresponding to the camera work of the virtual camera when shooting the scene to be processed, based on first scene information corresponding to the scene one before the scene to be processed and first camera information corresponding to the camera work of the virtual camera when shooting the scene one before the scene to be processed. Then, based on the generated second camera information, it searches for an image including the composition corresponding to the second camera information, and generates prediction content based on the searched image.
[0132] Thereby, the information processing apparatus 1 enables the user to confirm the output prediction of the content for each scene by means of the searched image.
[0133] Also, the prediction generation unit 53 generates second camera information corresponding to the camera work of the virtual camera when shooting the scene to be processed, based on first scene information corresponding to the scene one before the scene to be processed and first camera information corresponding to the camera work of the virtual camera when shooting the scene one before the scene to be processed. Then, based on the second scene information corresponding to the scene to be processed and the generated second camera information, it generates prediction content, which is an image including the composition corresponding to the second camera information.
[0134] Thereby, the information processing apparatus 1 enables the user to confirm the output prediction of the content for each scene by means of the image including the composition corresponding to the camera information.
[0135] Further, the prediction generation unit 53 generates second emotion information corresponding to the emotion of the character appearing in the scene to be processed based on the first scene information corresponding to the scene one before the scene to be processed and the first emotion information corresponding to the emotion of the character appearing in the scene one before the scene to be processed, and generates prediction content that is an image including an expression indicating the emotion corresponding to the second emotion information of the character appearing in the scene to be processed based on the second scene information corresponding to the scene to be processed and the generated second emotion information.
[0136] Thereby, the information processing apparatus 1 enables the user to confirm the output prediction of the content for each scene by means of an image including the expression of the character appearing in the scene.
[0137] Also, the prediction generation unit 53 generates pose information corresponding to the pose of the character appearing in the scene to be processed based on the target scene information corresponding to the scene to be processed, and generates prediction content that is an image including the pose corresponding to the pose information of the character appearing in the scene to be processed based on the generated pose information.
[0138] Thereby, the information processing apparatus 1 enables the user to confirm the output prediction of the content for each scene by means of an image including the pose corresponding to the pose information of the character appearing in the scene.
[0139] (5. Hardware Configuration) The information processing apparatus 1 according to the above-described embodiment is reproduced by a computer 1000 having a configuration as shown in FIG. 22, for example. FIG. 22 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing apparatus according to the present disclosure. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface 1500, and an input / output interface 1600. Each part of the computer 1000 is connected by a bus 1050.
[0140] The CPU 1100 operates based on programs stored in the ROM 1300 or the HDD 1400 and controls each part. For example, the CPU 1100 expands the programs stored in the ROM 1300 or the HDD 1400 in the RAM 1200 and executes processes corresponding to various programs.
[0141] The ROM 1300 stores boot programs such as BIOS (Basic Input Output System) executed by the CPU 1100 when the computer 1000 is started up, programs dependent on the hardware of the computer 1000, and the like.
[0142] The HDD 1400 is a computer-readable non-transitory recording medium that non-transitorily records programs executed by the CPU 1100 and data used by such programs. Specifically, the HDD 1400 is a non-transitory recording medium that records the program according to the present disclosure, which is an example of the program data 1450.
[0143] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550 (for example, the Internet). For example, the CPU 1100 receives data from other devices or transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0144] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a keyboard and a mouse via the input / output interface 1600. Also, the CPU 1100 transmits data to output devices such as a display, a speaker, and a printer via the input / output interface 1600. Further, the input / output interface 1600 may function as a media interface for reading a program or the like recorded on a predetermined recording medium (media). The media is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory or the like.
[0145] For example, when the computer 1000 functions as the information processing apparatus 1 according to the embodiment, the CPU 1100 of the computer 1000 reproduces functions such as the control unit 50 by executing the program loaded on the RAM 1200. Also, the HDD 1400 stores the program according to the present disclosure and various types of data. Note that the CPU 1100 reads and executes the program data 1450 from the HDD 1400, but as another example, these programs may be acquired from other devices via the external network 1550.
[0146] As described above, each embodiment of the present disclosure has been described. However, the technical scope of the present disclosure is not limited to the above-described embodiments as they are, and various modifications are possible without departing from the gist of the present disclosure. Also, components across different embodiments and modifications may be appropriately combined.
[0147] Also, the effects in each of the embodiments described in this specification are merely examples and are not limiting, and there may be other effects.
[0148] Note that the present technology can also adopt the following configurations. (1) An acquisition unit that acquires text input for generating time-series content, A sequence generation unit that generates event sequence information, which is information visualizing a structure in which event information corresponding to events in the time-series content is arranged in the order in which the events occur, based on the text. An information processing apparatus comprising the above. (2) The sequence generation unit generates the event sequence information, which is information visualizing a structure in which the event information corresponding to each of a plurality of events proceeding simultaneously is arranged in parallel in a direction orthogonal to the direction of the time axis, for the information processing apparatus according to (1) above. (3) The sequence generation unit generates the event sequence information, which is information editable by a user, for the information processing apparatus according to (1) or (2) above. (4) The event information includes object information corresponding to an object appearing in the event and action information corresponding to an action of the object. for the information processing apparatus according to any one of (1) to (3) above. (5) The event information includes relationship information indicating a relative positional relationship between a plurality of objects appearing in the event. for the information processing apparatus according to any one of (1) to (4) above. (6) The sequence generation unit generates the event sequence information, which is a node graph including event nodes indicating the event information and first edges connecting the event nodes. for the information processing apparatus according to any one of (1) to (5) above. (7) The sequence generation unit generates the event sequence information, which is information visualizing a structure in which the event information is arranged along a time axis. The information processing apparatus according to any one of (1) to (6) above. (8) The sequence generation unit generates the event sequence information, which is information visualizing a structure in which the length of each event element indicating the event information in the direction of the time axis corresponds to the duration of the event and each of the event elements is arranged at a position on the time axis corresponding to the start time of each event. The information processing apparatus according to (7) above. (9) The sequence generation unit further generates global information corresponding to a scene where the event occurs based on the text. The information processing apparatus according to any one of (1) to (8) above. (10) The global information includes information indicating a location, weather, or time corresponding to the scene. The information processing apparatus according to (9) above. (11) A prediction generation unit that generates prediction content corresponding to a predetermined timing in a scene including one or more of the events based on the event sequence information and the global information is further provided. The information processing apparatus according to (9) or (10) above. (12) The prediction generation unit generates second prediction content corresponding to the scene to be processed based on first prediction content corresponding to a scene one before the scene to be processed. The information processing apparatus according to (11) above. (13) The prediction generation unit Based on the object information corresponding to the objects appearing in the scene, the action information corresponding to the actions of the objects, and the relationship information indicating the relative positional relationship between a plurality of objects appearing in the scene, a scene graph including an object node indicating the object information, a relationship node indicating the relationship information, and a second edge connecting the object node and the relationship node is generated, and a first scene graph corresponding to a scene one before the scene to be processed and the prediction content which is an integrated scene graph obtained by integrating the second scene graph corresponding to the scene to be processed is generated. The information processing apparatus according to (11) or (12). (14) The prediction generation unit Based on the integrated scene graph, the prediction content which is a map in which the objects appearing in the scene are arranged in the virtual space of the scene corresponding to the scene is generated. The information processing apparatus according to (13). (15) The prediction generation unit Based on the first scene information corresponding to a scene one before the scene to be processed and the first camera information corresponding to the camera work of the virtual camera when shooting the scene one before the scene to be processed with the virtual camera, the second camera information corresponding to the camera work of the virtual camera when shooting the scene to be processed with the virtual camera is generated, an image including a composition corresponding to the generated second camera information is searched based on the generated second camera information, and the prediction content is generated based on the searched image. The information processing apparatus according to any one of (11) to (14). (16) The prediction generation unit Based on the first scene information corresponding to the scene one before the scene to be processed and the first camera information corresponding to the camera work of the virtual camera when shooting the scene one before the scene to be processed, generate the second camera information corresponding to the camera work of the virtual camera when shooting the scene to be processed, and based on the second scene information corresponding to the scene to be processed and the generated second camera information, generate the predicted content which is an image including the composition corresponding to the second camera information. The information processing apparatus according to any one of (11) to (15) above. (17) The prediction generation unit Based on the first scene information corresponding to the scene one before the scene to be processed and the first emotion information corresponding to the emotion of the character appearing in the scene one before the scene to be processed, generate the second emotion information corresponding to the emotion of the character appearing in the scene to be processed, and based on the second scene information corresponding to the scene to be processed and the generated second emotion information, generate the predicted content which is an image including the expression indicating the second emotion information of the character appearing in the scene to be processed. The information processing apparatus according to any one of (11) to (16) above. (18) The prediction generation unit Based on the target scene information corresponding to the scene to be processed, generate pose information corresponding to the pose of the character appearing in the scene to be processed, and based on the generated pose information, generate the predicted content which is an image including the pose corresponding to the pose information of the character appearing in the scene to be processed. The information processing apparatus according to any one of (11) to (17) above. (19) The computer Obtain the text input for generating time-series content Generating event sequence information, which is information visualizing a structure in which event information corresponding to events in the time-series content is arranged according to the order in which the events occur, based on the text; An information processing method including the above. (20) Obtaining text input for generating time-series content; Generating event sequence information, which is information visualizing a structure in which event information corresponding to events in the time-series content is arranged according to the order in which the events occur, based on the text; A program for causing a computer to execute the above.
Explanation of Signs
[0149] 1 Information processing apparatus 10 Input unit 20 Output unit 30 Communication unit 40 Storage unit 50 Control unit 51 Acquisition unit 52 Sequence generation unit 53 Prediction generation unit 54 Content generation unit
Claims
1. An acquisition unit that acquires text input for generating time-series content; A sequence generation unit that generates event sequence information, which is information visualizing a structure in which event information corresponding to events in the time-series content is arranged according to the order in which the events occur, based on the text; Comprising: An information processing apparatus.
2. The sequence generation unit: Generates the event sequence information, which is information visualizing a structure in which the event information corresponding to each of a plurality of events proceeding simultaneously is arranged in parallel in a direction orthogonal to the direction of the time axis. The information processing apparatus according to claim 1.
3. The sequence generation unit: Generates the event sequence information, which is information editable by a user. The information processing apparatus according to claim 1.
4. The event information includes object information corresponding to an object appearing in the event and action information corresponding to an action of the object. The information processing apparatus according to claim 1.
5. The event information includes relationship information indicating a relative positional relationship between a plurality of objects appearing in the event. The information processing apparatus according to claim 1.
6. The sequence generation unit: Generates the event sequence information, which is a node graph including event nodes indicating the event information and first edges connecting the event nodes. The information processing apparatus according to claim 1.
7. The sequence generation unit: Generates the event sequence information, which is information visualizing a structure in which the event information is arranged along the time axis. The information processing apparatus according to claim 1.
8. The sequence generation unit: Generates the event sequence information, which is information visualizing a structure in which the length of the event elements indicating the event information in the direction of the time axis corresponds to the duration of the event and each of the event elements is arranged at a position on the time axis corresponding to the start time of each of the events. The information processing apparatus according to claim 7.
9. The sequence generation unit: Further generates global information corresponding to the scene in which the event occurs, based on the text. The information processing apparatus according to claim 1.
10. The global information includes information indicating a location, weather, or time corresponding to the scene. The information processing apparatus according to claim 9.
11. A prediction generation unit that generates prediction content corresponding to a predetermined timing in a scene including one or more of the events based on the event sequence information and the global information further comprising The information processing apparatus according to claim 9
12. The prediction generation unit generates second prediction content corresponding to the scene to be processed based on first prediction content corresponding to a scene one before the scene to be processed The information processing apparatus according to claim 11
13. The prediction generation unit generates a scene graph including an object node indicating the object information, a relationship node indicating the relationship information, and a second edge connecting the object node and the relationship node based on the object information corresponding to the object appearing in the scene, the action information corresponding to the action of the object, and the relationship information indicating the relative positional relationship between a plurality of objects appearing in the scene, and generates the prediction content which is an integrated scene graph obtained by integrating a first scene graph corresponding to a scene one before the scene to be processed and a second scene graph corresponding to the scene to be processed The information processing apparatus according to claim 11
14. The prediction generation unit generates the prediction content which is a map in which the objects appearing in the scene are arranged in the virtual space of the scene corresponding to the scene based on the integrated scene graph The information processing apparatus according to claim 13
15. The prediction generation unit generates second camera information corresponding to the camera work of the virtual camera when the scene to be processed is photographed by the virtual camera based on first scene information corresponding to a scene one before the scene to be processed and first camera information corresponding to the camera work of the virtual camera when a scene one before the scene to be processed is photographed by the virtual camera, searches for an image including a composition corresponding to the generated second camera information based on the generated second camera information, and generates the prediction content based on the searched image The information processing apparatus according to claim 11
16. The prediction generation unit Based on the first scene information corresponding to the scene one before the scene to be processed and the first camera information corresponding to the camera work of the virtual camera when shooting the scene one before the scene to be processed, generate the second camera information corresponding to the camera work of the virtual camera when shooting the scene to be processed, and based on the second scene information corresponding to the scene to be processed and the generated second camera information, generate the predicted content which is an image including the composition corresponding to the second camera information. The information processing apparatus according to claim 11.
17. The prediction generation unit Based on the first scene information corresponding to the scene one before the scene to be processed and the first emotion information corresponding to the emotion of the character appearing in the scene one before the scene to be processed, generate the second emotion information corresponding to the emotion of the character appearing in the scene to be processed, and based on the second scene information corresponding to the scene to be processed and the generated second emotion information, generate the predicted content which is an image including the expression showing the emotion corresponding to the second emotion information of the character appearing in the scene to be processed. The information processing apparatus according to claim 11.
18. The prediction generation unit Based on the target scene information corresponding to the scene to be processed, generate the pose information corresponding to the pose of the character appearing in the scene to be processed, and based on the generated pose information, generate the predicted content which is an image including the pose corresponding to the pose information of the character appearing in the scene to be processed. The information processing apparatus according to claim 11.
19. A computer Obtains the text input for generating the time-series content, Generates event sequence information which is information visualizing the structure in which the event information corresponding to the events in the time-series content is arranged according to the order in which the events occur, based on the text. An information processing method including the above.
20. Obtains the text input for generating the time-series content, Generates event sequence information which is information visualizing the structure in which the event information corresponding to the events in the time-series content is arranged according to the order in which the events occur, based on the text. A program that causes a computer to execute.
Citation Information
Patent Citations
Animation generation device, animation generation method, and program
JP2015176592A