Content generation device, content generation method, program, and recording medium

The content generation device addresses the issue of insufficient data by incorporating additional information through analysis, ensuring accurate content generation without prolonged training, thus enhancing content quality.

WO2026070326A1PCT designated stage Publication Date: 2026-04-02FUJIFILM CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing content generation systems struggle to produce accurate content when insufficient learning data is available, often relying on estimation which can lead to inappropriate information.

Method used

A content generation device that receives additional information based on the analysis of existing information, using a trained model to generate content that incorporates third-party information when necessary, ensuring accuracy without excessive training time.

Benefits of technology

Enables the generation of appropriate content even with limited initial data by utilizing additional information, reducing the reliance on estimation and improving content quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025031587_02042026_PF_FP_ABST
    Figure JP2025031587_02042026_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a content generation device, a content generation method, a program, and a recording medium with which it is possible to generate appropriate content in cases where existing information alone is insufficient for generating content. A content generation device according to one embodiment of the present invention comprises a processor, wherein the processor receives input of first information and second information for generating target content, receives third information based on a result of analysis of one or both of the first information and the second information, and generates the target content from the third information and the other or both of the first information and the second information by using a content generation model that is trained to generate the target content.
Need to check novelty before this filing date? Find Prior Art

Description

Content generation device, content generation method, program, and recording medium

[0001] One embodiment of the present invention relates to a content generation device, a content generation method, a program, and a recording medium.

[0002] For example, a technique for generating content using artificial intelligence (AI: Artificial Intelligence) is already known, and an example thereof is the technique described in Patent Document 1. In the generation device described in Patent Document 1, a pseudo-camera image is output using a learned model from the camera image acquired by the camera image acquisition unit and the sensor information acquired by the sensor information acquisition unit, so that even when the state of the camera image is poor, an image (content) that pseudo-represents an object can be obtained.

[0003] Japanese Patent Application Laid-Open No. 2023-139901

[0004] In the generation device described in Patent Document 1, the more learning data is learned by the learned model, the more likely the accuracy of the learned model is to increase. And by increasing the accuracy of the learned model, a more appropriate pseudo-camera image (content) can be obtained. On the other hand, as the learning data increases, it takes time to create the learned model. Also, when existing information alone is insufficient to generate content, the learned model will estimate additional information (such as images) necessary to generate content. In this case, the more learning data is learned by the learned model, the higher the probability that the estimated information will be more appropriate information. However, since this is still only estimated information, there is also a possibility that it will be inappropriate information, and it is preferable that the information obtained without estimation be used as the additional information required.

[0005] One embodiment of the present invention has been made in view of the above circumstances, and an object thereof is to provide a content generation device, a content generation method, a program, and a recording medium that can generate appropriate content when existing information alone is insufficient to generate content.

[0006] The above objective is achieved by a content generation device described in any of the following [1] to

[18] . [1] A content generation device equipped with a processor, wherein the processor receives input of first information and second information for generating target content, receives third information based on the results of analyzing one or both of the first and second information, and generates target content from the other or both of the first and second information and the third information using a content generation model learned to generate target content. [2] The content generation device according to [1], wherein the processor generates at least one of a still image and a moving image as target content. [3] The content generation device according to [1] or [2], wherein the processor receives at least one of a still image and a moving image as first information and receives text information as second information. [4] The content generation device according to any of the following [1] to [3], wherein the processor receives at least one of a still image and a moving image as first information and receives a moving image as second information. [5] A content generation device according to any one of [1] to [4], wherein the processor accepts at least one of still images, moving images, and text information as third information. [6] A content generation device according to any one of [1] to [5], wherein the processor accepts additional information necessary to satisfy the criteria as third information if there is a part in the target content that does not satisfy the criteria when generated without using third information. [7] A content generation device according to any one of [1] to [6], wherein the third information is information that is different from the first and second information that has already been accepted as input, and relates to at least one of the first and second information that is additionally necessary. [8] A content generation device according to any one of [1] to [7], wherein the processor generates third information from at least one of the first and second information using an information generation model that has been learned to generate third information.[9] A content generation device according to any one of [1] to [8], wherein the processor generates a plurality of candidates from at least one of the first information and the second information using a candidate generation model trained to generate candidates for the third information, and accepts one or more target candidates selected from the plurality of candidates as the third information.

[10] A content generation device according to any one of [1] to [9], wherein the processor outputs fourth information necessary to obtain the third information based on the analysis results.

[11] A content generation device according to

[10] , wherein the processor accepts information regarding the characteristics of at least one of the first information and the second information output based on the fourth information as the third information.

[12] A content generation device according to

[11] , wherein the third information is information regarding the characteristics of an image when at least one of the first information and the second information is an image.

[13] A content generation device according to any one of

[10] to

[12] , wherein the processor outputs information regarding the shooting of at least one of a still image and a moving image as the fourth information.

[14] A content generation device according to any one of

[10] to

[13] , wherein the processor outputs text information as fourth information.

[15] A content generation device according to any one of

[10] to

[14] , wherein the processor outputs fourth information indicating the portion of the target content that requires the use of third information when generated without using third information.

[16] A content generation device according to any one of

[10] to

[15] , wherein the processor outputs fourth information indicating the portion of the image that requires the use of third information when at least one of the first information and the second information includes an image.

[17] A content generation device according to any one of

[10] to

[16] , wherein the processor outputs information indicating the reason why third information is necessary to generate the target content as fourth information.

[18] A content generation device according to any one of

[10] to

[17] , wherein the processor generates an image as fourth information from at least one of the first information and the second information using an image generation model that has been trained to generate an image as fourth information.

[0007] Furthermore, the above objective can also be achieved by the content generation method described in

[19] below.

[19] A content generation method in which a processor performs the following steps: a process of receiving input of first information and second information for generating target content; a process of receiving third information based on the results of analyzing one or both of the first and second information; and a process of generating target content from the other or both of the first and second information and the third information, using a content generation model trained to generate target content.

[0008] Furthermore, a program according to one embodiment of the present invention is a program that causes a computer to perform each step included in the content generation method described in

[19] above. Furthermore, a recording medium according to one embodiment of the present invention is a recording medium that is readable by a computer and on which a program that causes a computer to perform each step included in the content generation method described in

[19] above is recorded.

[0009] According to one embodiment of the present invention, a content generation device, a content generation method, a program, and a recording medium are provided that can generate appropriate content when existing information alone is insufficient for generating content.

[0010] This is a schematic diagram illustrating content generation according to a comparative example of the present invention. This is a schematic diagram illustrating the overview of content generation according to one embodiment of the present invention. This is a diagram showing an example of the use of a content generation device according to one embodiment of the present invention. This is a diagram showing the hardware configuration of a content generation device according to one embodiment of the present invention. This is a schematic diagram illustrating the functions of a content generation device according to one embodiment of the present invention. This is a schematic diagram illustrating an example of fourth information according to one embodiment of the present invention. This is a schematic diagram illustrating an example of fourth information according to one embodiment of the present invention. This is a schematic diagram illustrating an example of fourth information according to one embodiment of the present invention. This is a diagram showing an example of a generation flow according to one embodiment of the present invention. This is a diagram showing an example of a generation flow according to modification 1 of the present invention. This is a diagram showing an example of a generation flow according to modification 2 of the present invention. This is a schematic diagram illustrating content generation according to the appendix of the present invention.

[0011] The following describes specific embodiments of the present invention. For convenience of explanation, the following descriptions may sometimes be based on the perspective of a GUI (Graphical User Interface). Furthermore, the fundamental data processing technologies for realizing the present invention (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, image display technologies, and visualization technologies, etc.) are known technologies, and therefore, their descriptions will be omitted.

[0012] Furthermore, in this specification, the concept of "device" includes not only a single device that performs a specific function, but also a combination of multiple devices that exist independently and in a distributed manner while cooperating (linking) to perform a specific function.

[0013] Furthermore, in this specification, "user" refers to the user of the content generation device of the present invention. Specifically, a user is a person who uses the information obtained by the functions of the content generation device of the present invention (in detail, the target content Ct described later). Also, in this specification, unless otherwise specified, "image" includes both still images and moving images, and includes both two-dimensional images and three-dimensional images. Unless otherwise specified, "image" refers to digital image data (hereinafter referred to as "image data"), and defines the grayscale values ​​of each of the multiple pixels that constitute the image. Image data includes RAW image data before compression and image data after compression. Image data after compression includes image data with lossy compression such as the JPEG format, image data with lossless compression such as the GIF (Graphics Interchange Format) or PNG (Portable Network Graphics) format, etc. Image data with lossless compression such as the PNG (Portable Network Graphics) format, image data of three-dimensional images used on a web browser such as the VRML (Virtual Reality Modeling Language) format, and image data of moving images composed of multiple frame images (still images) may also be used. Furthermore, the image data may include supplementary information such as the file name, date and time of shooting, and location of shooting.

[0014] Furthermore, in this specification, "person" means an entity that performs a specific action, and includes individuals, groups, corporations and other legal entities, and organizations, and may even include computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) realizes intelligent functions such as reasoning, prediction, and judgment using hardware and software resources. The algorithm of artificial intelligence is arbitrary and may include, for example, expert systems, case-based reasoning (CBR), Bayesian networks, or inclusion architectures.

[0015] <<Overview of Content Generation>> Content generation (hereinafter referred to as "this content generation") carried out using a content generation device (hereinafter referred to as "content generation device 10") and a content generation method according to one embodiment of the present invention (hereinafter referred to as "this embodiment") will be described with reference to Figures 1 to 3. Figure 1 is an explanatory diagram of content generation according to a comparative example of the present invention for easy understanding of this content generation, and Figures 2 and 3 are explanatory diagrams of this content generation.

[0016] <Content Generation in the Comparative Example> In the content generation in the comparative example, as shown in Figure 1, a specific content Ci is generated from first information D1 and second information D2 using, for example, a content generation model L0. The first information D1 is, for example, a still image Pi of person J taken from the front, and the second information D2 is, for example, a video Mi of person K dancing.

[0017] The content generation model L0 is a trained model that has been trained to generate specific content Ci when first information D1 and second information D2 are input. It may be constructed by performing machine learning using the first information (still image) and second information (moving image) for training, and the specific content for training, as training data. In the specific content Ci, only person K is replaced with person J displayed in the still image Pi in the moving image Mi, and the area other than person K is the same moving image as the moving image Mi. As a result, the specific content Ci is a moving image that is generated to make it appear as if person J is dancing the dance that person K was dancing. However, only information on the front side of person J can be obtained from the still image Pi, and information on the back side of person J cannot be obtained. On the other hand, the moving image Mi contains information on the back side of person K, so in order to generate the specific content Ci, it is necessary for the content generation model L0 to infer the information on the back side of person J. For this reason, in the specific content Ci, the area on the back side of person J is the area inferred by the content generation model L0.

[0018] <Generating this content> In contrast, in this content generation, as shown in Figure 2, for example, a content generation model L1 is used to generate target content Ct from first information D1, second information D2, and third information D3. The first information D1 and second information D2 were explained in the content generation related to the comparative example, so their explanation is omitted here. The third information D3 is, for example, a still image Pa of person J taken from the back, and this content generation differs from the content generation related to the comparative example in that it is generated using this third information D3. The content generation model L1 is a trained model that has been trained to generate target content Ct when first information D1, second information D2, and third information D3 are input, and may be constructed by performing machine learning using first information (still image), second information (moving image), and third information (still image) for training, and target content for training as training data.

[0019] The target content Ct, like the specific content Ci, is a video in which only person K in the video Mi is replaced with person J displayed in the still image Pi, and the dance that person K was performing is made to appear as if person J is performing it. On the other hand, unlike the specific content Ci, the target content Ct does not generate the area on the back of person J by inference, but rather generates it based on the information of the back of person J obtained from the third information D3 (still image Pa).

[0020] To explain in more detail, as shown in Figure 1, in the specific content Ci, the area on the back of person J is an estimated area, so the actual appearance of the back of person J is unknown. On the other hand, as shown in Figure 2, there is text (in Figure 2, "ENJOY") written on the back of person J's jacket as shown in the still image Pa, so in the target content Ct, the same text as in the still image Pa will be written on the back of person J's jacket.

[0021] The third piece of information D3 is input by the user to the content generation device 10 based on the fourth piece of information D4 obtained by analyzing the first piece of information D1 and the second piece of information D2. Examples of the fourth piece of information D4 include text information Tx, "Please add information about the back of person J," as shown in Figure 3. In the example shown in Figure 3, the user inputs a still image Pa of person J taken from the back to the content generation device 10, referring to the text information Tx. The content generation device 10 receives the input still image Pa as the third piece of information D3 and uses the content generation model L0 to generate the target content Ct from the first piece of information D1, the second piece of information D2, and the third piece of information D3.

[0022] When generating content, the more training data a content generation model receives, the higher the potential accuracy of the model, and this increased accuracy can lead to more appropriate content. On the other hand, increasing the amount of training data increases the time required to create the content generation model. Furthermore, if existing information is insufficient to generate content, the content generation model will estimate additional information (such as images) necessary for content generation. In this case, the more training data the content generation model receives, the higher the probability that the estimated information will be more appropriate. However, since this is only estimated information, there is a possibility that it may be inappropriate, and it is preferable for the additional information to be obtained without estimation. In this content generation process, the third piece of information D3, which is additionally necessary for content generation, can be appropriately obtained without estimation. By obtaining the third piece of information D3, appropriate content can be generated even when existing information (i.e., the first piece of information D1 and the second piece of information D2) is insufficient to generate content.

[0023] In the above description, the first information D1 is assumed to be a still image (still image Pi), but it is not limited to this and may be a moving image, for example. Also, in the above description, the second information D2 is assumed to be a moving image (moving image Mi), but it is not limited to this and may be text information, for example. Also, in the above description, the third information D3 is assumed to be a still image (still image Pa), but it is not limited to this and may be a moving image or text information, for example. Also, in the above description, the fourth information D4 is assumed to be text information (text information Tx), but it is not limited to this and may be an image, for example (see Figures 6 to 8 described later). Also, the target content Ct is assumed to be a moving image, but it is not limited to this and may be a still image, for example.

[0024] Let me explain with a few examples. As an example, in this content generation, the input of a current face image (still image) of a person (subject) taken from the front is accepted as the first piece of information D1, and the input of text information "Make this a face image of the person 50 years from now" is accepted as the second piece of information D2. Furthermore, as the third piece of information D3, the input of a current face image (still image) of the person taken from the side is accepted. In this case, the content generation device 10 may generate a face image (still image) of the person 50 years from now as the target content. Note that "subject" means a person, animal, object, background, etc. included in the image. As another example, in this content generation, the input of a two-dimensional video image of a person taken from a first viewpoint is accepted as the first piece of information D1, and the input of text information "Make this a three-dimensional video image" is accepted as the second piece of information D2. Furthermore, as the third piece of information D3, the input of a two-dimensional video image of a person taken from a second viewpoint different from the first viewpoint is accepted. In this case, the content generation device 10 may generate a three-dimensional video as the target content, in which the person can be viewed from any viewpoint. Alternatively, in this content generation, the first information D1 is an input of a full-body image (still image) of a person taken from the front, and the second information D2 is an input of text information stating, "Make a video of a person dancing." Furthermore, the third information D3 is an input of a full-body image (still image) of a person taken from the back. In this case, the content generation device 10 may generate a video of a person dancing as the target content.

[0025] Furthermore, in the above explanation, as shown in Figure 3, the user inputs an image (still image Pa) as the third piece of information D3. However, for example, text information may also be input as the third piece of information D3. Specifically, information about the back of person J may be expressed as text information, such as "The word ENJOY is written on the back of person J's jacket."

[0026] Furthermore, although the above description states that the content generation device 10 analyzes both the first information D1 and the second information D2, it is not limited to this, and for example, it may analyze only one of the first information D1 and the second information D2. Also, although the above description states that the content generation device 10 generates the target content Ct from both the first information D1 and the second information D2 and the third information D3 using the content generation model L1, it is not limited to this, and for example, if only one of the first information D1 and the second information D2 is used in the analysis described above, the target content may be generated from only the other of the first information D1 and the second information D2 and the third information D3.

[0027] To illustrate with an example, let's assume that the content generation device 10 requests a still image of the subject's entire body as the first piece of information D1. However, suppose the still image actually input by the user as the first piece of information D1 is a still image showing only the subject's face. In this case, the content generation device 10 analyzes only the first piece of information D1 and concludes that information about areas other than the subject's face is missing. It may then output, for example, text information such as "Please add a still image of the subject's entire body" as the fourth piece of information D4. In response, the user inputs a still image of the subject's entire body as the third piece of information D3 into the content generation device 10. The content generation device 10 then accepts the input still image as the third piece of information D3. Furthermore, if the second information D2 is text information "Create a video of the subject dancing", the content generation device 10 may generate a video of the subject dancing as the target content from only the second information D2 (the other of the first information D1 and the second information D2) and the third information D3.

[0028] <<Example Configuration of Content Generation Device According to This Embodiment>> Next, an example configuration of the content generation device 10 according to this embodiment will be described with reference to Figure 4. The content generation device 10 consists of a computer used by the user, specifically a client terminal, and is composed of, for example, a smartphone, a tablet terminal, and a PC (Personal Computer) such as a notebook or desktop computer. The content generation device 10 is not limited to a computer owned by the user, and may be composed of a terminal installed in a store, etc., which the user does not own but can use when visiting a store, etc., by entering a PIN or password or making a deposit. In the following, we will explain using the case in which the content generation device 10 is composed of a computer owned by the user, specifically a PC, as an example.

[0029] As shown in Figure 4, the computer comprising the content generation device 10 includes a processor 10a, memory 10b, communication interface 10c, storage 10d, input device 10e, and output device 10f.

[0030] The processor 10a is composed of, for example, a CPU (Central Processing Unit). The memory 10b is composed of, for example, semiconductor memory such as ROM (Read Only Memory) and RAM (Random Access Memory).

[0031] The communication interface 10c may be configured, for example, as a network interface card or a communication interface board. The computer constituting the content generation device 10 can communicate with other devices connected to a communication network such as the Internet and mobile communication lines via the communication interface 10c.

[0032] The storage 10d may consist of, for example, flash memory, HDD (Hard Disc Drive), SSD (Solid State Drive), FD (Flexible Disc), MO disk (Magneto-Optical Disc), CD (Compact Disc), DVD (Digital Versatile Disc), SD card (Secure Digital card), or USB memory (Universal Serial Bus memory). The storage 10d may be built into the computer main body that constitutes the content generation device 10, or it may be attached to the computer main body as an external device. Alternatively, the storage 10d may consist of a NAS (Network Attached Storage) or the like. Furthermore, the storage 10d may be an external device that can communicate with one of the computers that constitutes the content generation device 10 via a communication network, such as an online storage or database server.

[0033] The input device 10e is a device that accepts user input and is composed of, for example, a touch panel and a keyboard. The input device 10e may also include a camera, such as one built into a PC or smartphone, and a microphone for sound collection. The output device 10f is composed of, for example, a display and a speaker, and may also include, for example, a hologram displayed in three dimensions.

[0034] Furthermore, the computer constituting the content generation device 10 has software installed, including an operating system (OS) program and an application program for executing content generation (hereinafter referred to as the content generation app). When these programs are read and executed by the processor 10a, the computer constituting the content generation device 10 performs its functions, specifically executing a series of processes related to this content generation. The content generation app may be obtained by reading it from a recording medium that the computer can read, or by downloading it via a communication network such as the Internet or an intranet.

[0035] <<Functions of the Content Generation Device According to This Embodiment>> The configuration of the content generation device 10 will be explained again in terms of its functions with reference to Figure 5. As shown in Figure 5, the content generation device 10 has a first receiving unit 21, an analysis unit 22, a first output unit 23, a second receiving unit 24, a generation unit 25, and a second output unit 26. These functional units are realized by the processor 10a of the content generation device 10 executing the aforementioned content generation application and cooperating with other hardware devices of the content generation device 10. Furthermore, at least some of the functions are realized using artificial intelligence (AI). The following describes each functional unit.

[0036] <First Reception Unit> The first reception unit 21 receives input from the user of first information D1 and second information D2 for generating the target content Ct. "First information D1" may be at least one of a still image and a moving image. This allows the user to broaden the range of images to input. The still image may be, for example, a frame image selected by the user from one or more moving images input by the user.

[0037] "Second Information D2" may be at least one of a still image, a moving image, and text information. By making Second Information D2 an image (still image and moving image), i.e., information that is easily recognizable visually, the user's intent can be appropriately reflected in the generated target content Ct. Alternatively, by making Second Information D2 text information, the user's intent can be appropriately reflected even when there is no suitable moving image for Second Information D2, or when Second Information D2 is abstract and difficult to represent with an image.

[0038] The image as the first information D1 or second information D2 may be, for example, a photograph taken by the user using a camera or other photographic device, an image stored in storage 10d, or an image stored in an external device that can communicate via the communication interface 10c, such as online storage or a database server. The same applies to the image as the third information D3 described below. The text information as the second information D2 may be, for example, characters and strings entered by the user through the input device 10e. The same applies to the text information as the third information D3 described below.

[0039] Furthermore, one of the first information D1 and the second information D2 may be, for example, main information for determining the main subject of the target content Ct, specifically a particular person (person J in Figure 3), while the other of the first information D1 and the second information D2 may be, for example, sub-information (reference information) for determining auxiliary parts of the target content Ct, such as the actions of that person or other subjects (landscape, etc.).

[0040] <Analysis Unit> The analysis unit 22 performs analysis based on one or both of the first information D1 and the second information D2. Specifically, first, the analysis unit 22 extracts the features of the first information D1 and the second information D2, more specifically, the image features of the still image Pi and the moving image Mi. "Image features" include information about the image quality of each region of the image, the grayscale values ​​of the pixels included in each region, and information about the subject estimated from this information, such as the size of the subject region relative to the image and the sharpness of the image. "Subject information" may include the type of subject, the state of the subject, the position of the subject in the image, and, if the subject is a person, the facial expression. It is desirable that the image features can be digitized, vectorized, or tensorized. In that case, the image features become digitized, vectorized, or tensorized image features, i.e., feature quantities.

[0041] In the example shown in Figure 3, the analysis unit 22 acquires region information for each of the subjects, persons J and K, based on the characteristics of the still image Pi and the moving image Mi, such as information about the face, chest, abdomen, back, arms, and legs of persons J and K. The analysis unit 22 then compares the region information of persons J and K and identifies region information that is present in one subject but not in the other. More specifically, the analysis unit 22 identifies region information that is present in person K but not in person J, and obtains an analysis result indicating that information on the back side of person J is insufficient compared to person K.

[0042] <First Output Unit>The first output unit 23 outputs the fourth information D4 to, for example, the output device 10f based on the analysis result by the analysis unit 22. "The fourth information D4" is information necessary to obtain the third information D3 and is information regarding the analysis result in the analysis unit 22. As described above, the "analysis result" is information indicating a fact such as "the information on the back side of person J is insufficient compared to person K". "Information regarding the analysis result" is, for example, a reformulation of the analysis result so as to prompt the user to input the third information D3. For example, text information such as "Please add information on the back side of person J" corresponds to this. In this way, in the content generation device 10, by using the fourth information D4 related to the analysis result, it is possible to prompt the user to appropriately input the third information D3.

[0043] As a display method of the fourth information D4, as long as it prompts the user to input the third information D3, it is not particularly limited. For example, it may be a mere presentation of the analysis result (insufficient information), a question regarding the insufficient information, and a proposal, advice, and instruction regarding the acquisition method of the insufficient information. Further, the fourth information D4 may be information related to shooting. For example, it may be an instruction regarding the shooting conditions of an image as the third information D3. Examples of the shooting conditions include an angle of view, a shooting direction, a focal length, a focus position, an aperture (f-number), a depth of field, a white balance, a shutter speed, brightness, contrast, and an ISO sensitivity. Thereby, the user can smoothly shoot an image as the third information D3.

[0044] Further, the fourth information D4 may be text information such as the text information Tx shown in FIG. 3. Thereby, for example, even when the fourth information D4 has an abstract content and is difficult to represent by an image, if the fourth information D4 is text information, it can be appropriately represented.

[0045] Furthermore, the fourth piece of information D4 may be information indicating the reason why the third piece of information D3 is necessary to generate the target content Ct. For example, it may be text information such as "Information about the back of person J included in the target content Ct is insufficient, therefore additional information is needed" or "Information about the back of person J included in the target content Ct is insufficient, therefore the back of person J will be generated by estimation." By presenting the reason in this way, the user's understanding can be deepened, and the user can be prompted to input more appropriate third piece of information D3.

[0046] Furthermore, the fourth information D4 may be an image used to facilitate the creation and selection of images as the third information D3, in other words, a sample image that serves as a model for the image to be input as the third information D3. For example, the first output unit 23 may generate an image as the fourth information D4 from at least one of the first information D1 and the second information D2 using an image generation model that has been trained to generate an image as the fourth information D4. Referring more specifically to Figure 6, the first output unit 23 may generate a still image Pg as the fourth information D4 from both a still image Pi and a moving image Mi using an image generation model. The image generation model is a trained model that has been trained to generate an image as the fourth information D4 when the first information D1 and the second information D2 are input, and may be constructed by performing machine learning using the first and second information for training and the training image as training data. The generated still image Pg may be, for example, a sample image of the still image Pa as the third information D3. This allows the user to smoothly capture a still image Pa as third information D3 by, for example, setting the field of view and shooting direction of the shooting device based on the generated still image Pg. The content generation device 10 may also generate multiple still images Pg as fourth information D4 and present multiple still images Pg to the user as candidate images. The content generation device 10 may then accept the still image Pg selected by the user from among the candidate images as third information D3.

[0047] Furthermore, the fourth information D4 may be information indicating the portion of the image where the third information D3 needs to be used. Specifically, as shown in Figure 7, the first output unit 23 may output the fourth information D4 (corresponding to the emphasis section F1 in Figure 7) indicating the portion of the image where the third information D3 needs to be used within the target content Ct (corresponding to the study content Ce in Figure 7) when generated without using the third information D3. More specifically, the first output unit 23 first generates a study content Ce corresponding to the specific content Ci in Figure 1 from the first information D1 and the second information D2 using a study content model corresponding to the content generation model L0 in Figure 1. Next, based on the analysis result in the analysis unit 22 that there is insufficient information on the back side of person J, the first output unit 23 displays the emphasis section F1 as the fourth information D4 in the portion where the third information D3 needs to be used, that is, in the area on the back side of person J within the study content Ce. The first output unit 23 then outputs the review content Ce with the emphasis unit F1 displayed to the output device 10f. This allows the content generation device 10 to prompt the user to input appropriate third information D3, using the review content Ce with the emphasis unit F1 displayed as a reference. The emphasis method by the emphasis unit F1 is not particularly limited; for example, the target area may be masked, or the target area may be enclosed with a line. The same applies to the emphasis method by the emphasis unit F2 described below.

[0048] Further, as shown in FIG. 8, when at least one of the first information D1 and the second information D2 includes an image, the first output unit 23 may output fourth information D4 (corresponding to the highlighted portion F2 in FIG. 8) indicating a portion in the image (corresponding to the moving image Mi in FIG. 8) where it is necessary to use the third information D3. More specifically described, based on the analysis result in the analysis unit 22 indicating that the information on the back side of the person J is insufficient, the first output unit 23 displays the highlighted portion F2 as the fourth information D4 in the area on the back side of the person K in the moving image Mi received by the first reception unit 21. Then, the first output unit 23 outputs the moving image Mi with the highlighted portion F2 displayed to the output device 10f. As a result, in the content generation device 10, as in the example shown in FIG. 7, the fourth information D4 (highlighted portion F2) can be displayed using the existing moving image Mi without newly generating the consideration content Ce. That is, in the example shown in FIG. 8, by using the existing moving image Mi, the processing time (processing man-hours) in the content generation device 10 can be reduced compared to the case of generating new consideration content Ce as shown in FIG. 7. Thus, in the example shown in FIG. 8, while reducing the processing time in the content generation device 10, the user can be made to input appropriate third information D3 by referring to the moving image Mi with the highlighted portion F2 displayed. Further, as described above, a consideration content model is required to generate the consideration content Ce shown in FIG. 7, but in the example shown in FIG. 8, it is not necessary to prepare such a consideration content model.

[0049] <Second Reception Unit> The second reception unit 24 receives the third information D3 (in other words, the third information D3 based on the analysis results) input by the user, with reference to the fourth information D4 output based on the analysis results by the analysis unit 22. The "third information D3" may be one of still images, moving images, and text information. In other words, the AI ​​may be designed to accept any of the information input as the third information D3, whether it is a still image, a moving image, or text information. This allows the user to input only text information without having to take additional still or moving images and input them into the content generation device 10, thereby reducing the burden on the user. On the other hand, in cases where it is difficult to express something in text information (for example, when it is difficult to explain the pattern on the back of the jacket worn by person J in text information), the user only needs to input an image, thereby reducing the burden on the user.

[0050] More specifically, the third information D3 is information that is different from the first information D1 and second information D2 that have already been received as input, and is additionally required and related to at least one of the first information D1 and second information D2. Referring to Figure 3, the still image Pa as the third information D3 includes a person J, which is a common subject with the still image Pi of the first information D1, and is related to the first information D1 in that it includes a common subject. On the other hand, the still image Pa as the third information D3 is different from the still image Pi of the first information D1 in that it displays the person J as a subject in a different direction (direction of shooting), and is therefore a different image from the still image Pi. Also, the still image Pa of the third information D3 is a still image and is a different image from the second information D2, which is a moving image Mi. In this way, the content generation device 10 can avoid receiving the same information as the first information D1 and second information D2 as the third information D3, and can also avoid receiving information unrelated to the first information D1 and second information D2 as the third information D3. This allows the content generation device 10 to efficiently acquire the third information D3.

[0051] Furthermore, the third information D3 is additional information necessary to satisfy the criteria when there are parts in the target content Ct, i.e., the content Ce for consideration in Figure 7 (or the specific content Ci in Figure 1), that do not meet the criteria when the third information D3 is not used. The "criteria" here may include, for example, "generated based on facts." In the content Ce for consideration in Figure 7, the area on the back side of the subject, person J, is a part that was generated based on speculation and therefore does not meet the above criteria. For this reason, the information on the back side of person J becomes the third information D3. In this way, the content generation device 10 can acquire appropriate third information D3 according to the criteria.

[0052] The "criteria" are not limited to those described above, and may include, for example, "resolution being above a certain threshold." Furthermore, different threshold values ​​may be set depending on the type of subject in the image, the state of the subject, the position of the subject in the image, and the facial expression. For example, the threshold may be set lower for the area behind a person and higher for the face area of ​​the person. This means that, for example, in the content Ce for consideration, even if the area behind person J is a part generated by estimation, it may be considered to satisfy the criteria because the resolution is above a predetermined value. Conversely, in the content Ce for consideration, even if the face area of ​​person J is a part based on a still image Pi (a part that has not been estimated), it may be considered not to satisfy the criteria because the resolution is below a predetermined value. In this case, the additional information required to satisfy the criteria (third information D3) would be the information of the face area of ​​person J.

[0053] Furthermore, the second receiving unit 24 may receive information regarding the characteristics of at least one of the first information D1 and second information D2 output based on the fourth information D4 as third information D3. Here, third information D3 corresponds to information regarding the characteristics of the image when at least one of the first information D1 and second information D2 is an image. "Information regarding the characteristics of the image" corresponds to, for example, information about the subject, and more specifically, if the subject is a person, it corresponds to information such as the person's posture, movement and orientation (facing forward, etc.), the display range of the person in the image (whole body or face), and the position of the person in the image. In the example shown in Figure 3, information regarding the orientation of person J, more specifically, information about the back of person J, becomes third information D3. The information about the back of person J that becomes third information D3 may be an image of the back of person J, or it may be text information about the back of person J. The text information may be, for example, "The word ENJOY is written on the back of person J's jacket." This allows the content generation device 10 to appropriately acquire the third information D3 necessary for generating the target content Ct.

[0054] <Generation Unit> The generation unit 25 generates target content Ct from the other or both of the first information D1 and the second information D2, and the third information D3, using a content generation model L1 that has been trained to generate target content Ct. In the example shown in Figure 3, the generation unit 25 generates target content Ct from a still image Pi, a moving image Mi, and a still image Pa, using the content generation model L1. Note that the target content Ct may be at least one of a still image and a moving image. This allows the user to obtain output of a type according to their preference.

[0055] <Second Output Unit> The second output unit 26 outputs the generated target content Ct to, for example, the output device 10f.

[0056] <<Example of Content Generation Method According to This Embodiment>> Next, as an example of the operation of the content generation device 10 according to this embodiment, a generation flow using the device will be described. In the generation flow described below, the content generation method of the present invention is used. In other words, each step in the generation flow described below corresponds to a component of the content generation method of the present invention. Note that the generation flow below is merely an example, and new steps may be added to the flow without departing from the spirit of this embodiment.

[0057] Each step in the generation flow according to this embodiment is performed by the processor 10a of the content generation device 10 in the order shown in Figure 9. In other words, in each step of the generation flow, the processor 10a executes the data processing specified in the content generation application that corresponds to each step in Figure 9.

[0058] First, the content generation flow is initiated when the user launches a content generation application installed on the content generation device 10, which in turn generates a control signal to start the generation flow. After the content generation application is launched, the processor 10a transitions the screen of the display (output device 10f) to an input screen (not shown) that prompts the user to input first information D1 and second information D2, based on a predetermined operation by the user. Then, when the user inputs first information D1 and second information D2, the processor 10a receives first information D1 and second information D2 (S001). In the example shown in Figure 3, a still image Pi is input as first information D1 and a moving image Mi is input as second information D2.

[0059] Next, the processor 10a performs analysis based on one or both of the first information D1 and the second information D2 (S002), and outputs the fourth information D4 based on the analysis results (S003). In the example shown in Figure 3, for example, the processor 10a outputs the text information Tx "Please add information about the back of person J" as the fourth information D4. The user uses the fourth information D4 as a reference and inputs the third information D3 corresponding to the fourth information D4 into the input device 10e. As a result, the processor 10a receives the information input by the user into the input device 10e as the third information D3 (S004). In the example shown in Figure 3, for example, the processor 10a receives the additional image (still image Pa) as the third information D3.

[0060] Subsequently, the processor 10a generates target content Ct from the other or both of the first information D1 and the second information D2, and the third information D3, using the content generation model L1 (S005). In the example shown in Figure 3, for example, the processor 10a generates target content Ct from a still image Pi, a moving image Mi, and a still image Pa, using the content generation model L1.

[0061] The generation flow according to this embodiment ends when the above series of processes is completed. In the generation flow, each time the user inputs the first information D1 and the second information D2, the series of content generation processes shown in Figure 9 are repeatedly executed by the processor 10a.

[0062] <<Other Embodiments>> Although specific embodiments of the present invention have been described above, the embodiments described above are merely examples given to facilitate understanding of the present invention and do not limit it. That is, the present invention can be modified or improved from the embodiments described below, without departing from its spirit. Furthermore, the present invention includes equivalents thereof. Modifications will be described below. In the following, the differences between the modifications and the embodiments described above will be the main points of explanation, and for convenience, the same names and reference numerals as those in the embodiments described above will be used for the modifications that correspond to those in the embodiments described above.

[0063] <Modification 1> In the above embodiment, the third information D3 is input by the user, but it is not limited to this, and the processor 10a may, for example, generate the third information D3 from at least one of the first information D1 and the second information D2 using an information generation model. The information generation model is a trained model that has been trained to generate the third information D3 when the first information D1 and the second information D2 are input, and may be constructed by performing machine learning using the first and second information for training and the third information as training data. An example of the operation of the content generation device according to this modification will be described with reference to Figure 10. Note that the processing that is common to the above embodiment will be omitted from the explanation. In the generation flow according to this modification, first, the processing of steps S101 to S102, which is common to steps S001 to S002 of the above embodiment, is executed.

[0064] Next, the processor 10a outputs selection information to allow the user to choose whether or not to automatically generate the third information D3, along with the fourth information D4 (S103). If the user chooses to automatically generate the third information D3 (S104), the processor 10a automatically generates the third information D3 (S105). Specifically, the processor 10a generates the third information D3 from at least one of the first information D1 and the second information D2 using an information generation model. After that, the processor 10a receives the generated third information D3 (S106). On the other hand, if the user does not choose to automatically generate the third information D3 (S104), the processor 10a receives the third information D3 from the user, similar to step S004 in the above embodiment (S106). After that, the processor 10a generates the target content Ct, similar to step S005 in the above embodiment (S107).

[0065] As explained above, the content generation device 10 according to this modified example can acquire the third information D3 even when it is difficult for the user to input an image as the third information D3 (for example, when the person J in the still image Pi has already passed away).

[0066] <Modification 2> In Modification 1 described above, the processor 10a generates the third information D3 if the user selects to automatically generate the third information D3 (S104) (S105). However, it is not limited to this, and for example, the processor 10a may use a candidate generation model trained to generate candidates for the third information D3 to generate multiple candidates for the third information D3 from at least one of the first information D1 and the second information D2, and accept one or more target candidates selected from the multiple candidates as the third information D3. The candidate generation model is a trained model trained to generate candidates for the third information D3 when the first information D1 and the second information D2 are input, and may be constructed by performing machine learning using the first and second information for training and the third information as training data.

[0067] An example of the operation of the content generation device according to this modified example will be described with reference to Figure 11. Note that the processing common to the above embodiment and Modification Example 1 will not be explained. In the generation flow according to this modified example, first, the processing of steps S201 to S203, which are common to steps S101 to S103 of Modification Example 1, is executed.

[0068] Next, if the user selects to automatically generate the third information D3 (S204), the processor 10a uses an information generation model to generate a plurality of candidates for the third information D3 from at least one of the first information D1 and the second information D2 (S205), and outputs the generated plurality of candidates to the output device 10f (S206). Then, the processor 10a accepts one or more target candidates selected by the user from the plurality of candidates as the third information D3 (S207). After that, the processor 10a generates the target content Ct in the same manner as in step S207 of the modified example 1 above (S208).

[0069] As explained above, the content generation device 10 according to the modified example 2 automatically generates the third information D3, while also being able to reflect the user's intentions regarding the third information D3.

[0070] <Modification 3> In Modifications 1 and 2 described above, the processor 10a outputs selection information along with the fourth information D4 in step S103 or S203. However, it is not limited to this, and for example, steps S103 and S104 may be deleted from the generation flow of Figure 10, or steps S203 and S204 may be deleted from the generation flow of Figure 11. Specifically, in the generation flow of Figure 10, after performing the analysis (S102), the processor 10a may skip steps S103 and S104, automatically generate the third information D3 without allowing the user to choose whether or not to automatically generate the third information D3 (S105), and accept the generated third information D3 (S106). Similarly, in the generation flow of Figure 11, after the analysis is performed (S202), the processor 10a may skip steps S203 and S204 and automatically generate multiple candidates for the third information D3 without allowing the user to choose whether or not to automatically generate the third information D3 (S205). Then, after outputting the multiple candidates to the output device 10f (S206), the processor 10a may accept one or more target candidates selected by the user from among the multiple candidates as the third information D3 (S207).

[0071] <Other> In the above embodiment, the content generation device of the present invention is configured using a computer directly used by the user, such as a user-owned terminal (client terminal). However, it is not limited to this, and the content generation device of the present invention may be configured using a computer that the user can use indirectly, such as a server computer. Here, the server computer may be, for example, a server computer for cloud services, specifically a server computer for ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service). In this case, when the necessary information is input on the client terminal, the server computer performs various processes (calculations) based on the input information, and the calculation results are output on the client terminal side. In other words, the functions of the server computer that constitutes the content generation device of the present invention can be used on the client terminal side.

[0072] Furthermore, in the above embodiments, each process is executed on any computer. Alternatively, any computer may execute these processes using a processor as hardware, a program as software, or a combination thereof. In this case, the processor is configured to work in cooperation with the program to execute the various processes in the above embodiments, and can function as a unit or means in the above embodiments. The execution order of the processes by the processor is not limited to the order described and may be changed as appropriate. Any computer may be a general-purpose computer, a computer designed for a specific purpose, a workstation, or any other hardware element capable of executing a program.

[0073] The processor may consist of one or more hardware components, and the type of hardware is not limited. For example, the processor may consist of programmable logic devices such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), FPGA (Field Programmable Gate Array), dedicated circuits for executing specific processes such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphic Processing Unit), or an NPU (Neural Processing Unit). The processor also has various units or means that execute the various processes in this embodiment. Furthermore, the type of hardware may be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processes of a processor, these components may reside in physically separate devices or in the same device. Furthermore, in any embodiment, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware components are composed of electrical circuits (circuits) and the like, which are combinations of circuit elements such as semiconductor elements.

[0074] Furthermore, the program may be firmware or software such as microcode. Alternatively, the program may be, for example, a set of program modules, each function of which may be implemented by a processor configured to perform its respective function. The program may be program code or multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage media or other storage). The program may be divided and stored on multiple non-temporary computer-readable media located in physically separate devices. Program code, or code segments, may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. Program code, or code segments, may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.

[0075] The following additional information is disclosed regarding the embodiments described above.

[0076] (Note 1) A content generation device equipped with a processor, wherein the processor receives input of first information and second information for generating content, generates the content from the first information and the second information using a content generation model learned to generate the content, generates fourth information for obtaining third information based on the analysis result of at least one of the first information and the second information, and adds the fourth information to the content.

[0077] In the content generation device described in Appendix 1, as shown in Figure 12, a content generation model L0 (corresponding to the content generation model L0 in Figure 1) is used to generate specific content Ci (corresponding to specific content Ci in Figure 1) from first information D1 and second information D2. Furthermore, by analyzing at least one of the first information D1 and second information D2, a fourth piece of information D4, which is additionally necessary to obtain third information D3, is generated. The fourth piece of information D4 is then added to the specific content Ci and output to the output device 10f. Note that the fourth piece of information D4 may be added as supplementary information to the content (image), for example (Exif (Exchangeable Image File Format)). In other words, the fourth piece of information D4 is not normally displayed when the user is viewing the content, but may be displayed when the user intentionally performs an operation to display the supplementary information. Furthermore, similar to the embodiments described above, the fourth information D4 may be text information (such as the text information Tx in Figure 12), an image (such as the still image Pg in Figure 6), or information indicating the part of an image where the third information D3 needs to be used (see the highlighted parts F1 and F2 in Figures 7 and 8). As explained above, the content generation device in Appendix 1 adds the fourth information D4 to existing content. Therefore, for example, if a user wants to improve existing content, they can obtain the third information D3 by referring to the fourth information D4, and then generate content again using the third information D3 to obtain more appropriate content.

[0078] 10 Content generation device 10a Processor 10b Memory 10c Communication interface 10d Storage 10e Input device 10f Output device 21 First reception unit 22 Analysis unit 23 First output unit 24 Second reception unit 25 Generation unit 26 Second output unit Ci Specific content Ct Target content Ce Content for consideration D1 First information D2 Second information D3 Third information D4 Fourth information F1, F2 Emphasis unit J, K Person L0, L1 Content generation model Mi Moving image Pi, Pa, Pg Still image Tx Text information

Claims

1. A content generation device equipped with a processor, wherein the processor receives input of first information and second information for generating target content, receives third information based on the results of analyzing one or both of the first and second information, and generates the target content from the other or both of the first and second information and the third information using a content generation model that has been learned to generate the target content.

2. The content generation apparatus according to claim 1, wherein the processor generates at least one of a still image and a moving image as the target content.

3. The content generation apparatus according to claim 1, wherein the processor receives at least one of a still image and a moving image as first information and receives text information as second information.

4. The content generation apparatus according to claim 1, wherein the processor receives at least one of a still image and a moving image as first information and receives a moving image as second information.

5. The content generation apparatus according to claim 1, wherein the processor receives at least one of still images, moving images, and text information as the third information.

6. The content generation apparatus according to claim 1, wherein the processor, when the target content is generated without using the third information, accepts additional information necessary to satisfy the criteria as the third information.

7. The content generation apparatus according to claim 1, wherein the third information is information that is different from the first and second information that has already been input, and is additionally required, relating to at least one of the first and second information.

8. The content generation apparatus according to claim 1, wherein the processor generates the third information from at least one of the first information and the second information using an information generation model that has been trained to generate the third information.

9. The content generation apparatus according to claim 1, wherein the processor generates a plurality of candidates from at least one of the first information and the second information using a candidate generation model trained to generate candidates for the third information, and accepts one or more target candidates selected from the plurality of candidates as the third information.

10. The content generation apparatus according to any one of claims 1 to 9, wherein the processor outputs fourth information necessary to obtain the third information based on the analysis results.

11. The content generation apparatus according to claim 10, wherein the processor receives information relating to at least one of the features of the first information and the second information output based on the fourth information as the third information.

12. The content generating apparatus according to claim 11, wherein the third information is information relating to the characteristics of an image, when at least one of the first information and the second information is an image.

13. The content generation apparatus according to claim 10, wherein the processor outputs information relating to the capture of at least one of a still image and a moving image as the fourth information.

14. The content generation device according to claim 10, wherein the processor outputs text information as the fourth information.

15. The content generation apparatus according to claim 10, wherein the processor outputs the fourth information indicating the portion in the target content that requires the use of the third information when generated without using the third information.

16. The content generation apparatus according to claim 10, wherein the processor outputs a fourth information indicating a portion of the image in which the third information is to be used, when at least one of the first information and the second information includes an image.

17. The content generation apparatus according to claim 10, wherein the processor outputs information indicating the reason why the third information is necessary to generate the target content as the fourth information.

18. The content generation apparatus according to claim 10, wherein the processor generates an image as the fourth information from at least one of the first information and the second information using an image generation model that has been trained to generate an image as the fourth information.

19. A content generation method in which a processor performs the following steps: a process of receiving input of first information and second information for generating target content; a process of receiving third information based on the analysis results of one or both of the first and second information; and a process of generating the target content from the other or both of the first and second information and the third information, using a content generation model trained to generate the target content.

20. A program for causing a computer to perform each of the processes included in the content generation method described in claim 19.

21. A computer-readable recording medium on which a program is recorded causing a computer to perform each of the processes included in the content generation method described in claim 19.

Citation Information

Patent Citations

  • Image collection system and guiding apparatus for collecting image

    JP2003047026A

  • Generation device, generation method, and generation program

    JP2017130748A

  • Image capturing support device and image capturing support method

    JP2020113843A

  • Video generation program, video generation device, and video generation method

    JP2021033961A

  • Information processing device, information processing method and program

    WO2019026397A1