Content generation device, content generation method, program, and recording medium
The content generation device addresses the challenge of verifying and correcting inconsistencies in AI-generated content by using automated evaluation and regeneration based on predetermined rules, ensuring high-quality output through user-specified corrections.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2026-04-02
AI Technical Summary
The challenge of accurately verifying and correcting inconsistencies in large or high-resolution content generated by generative AI, such as images and videos, is cumbersome and resource-intensive, often leading to suboptimal quality due to insufficient personnel and time for thorough verification.
A content generation device and method that includes a processor for generating and evaluating content based on predetermined rules, regenerating portions that do not meet standards, and allowing user specification of areas for regeneration, while considering the relative importance of objects and viewer/intent in the content.
Enables accurate responses to errors in generated content by regenerating problematic areas, ensuring higher quality and user satisfaction through automated evaluation and correction processes.
Smart Images

Figure JP2025028603_02042026_PF_FP_ABST
Abstract
Description
Content generation device, content generation method, program, and recording medium
[0001] One embodiment of the present invention relates to a content generation device, a content generation method, a program, and a recording medium that enable accurate responses to errors in content generated by a generation AI.
[0002] In recent years, so-called generative AI (Artificial Intelligence) has become available in internet services and specific applications, making it easily accessible in a variety of situations. In such generative AI, various types of content, such as videos and still images, are generated by inputting appropriate information such as text, images, and audio into the user interface.
[0003] On the other hand, there is no guarantee that the content generated in this way will be free from problems from various perspectives. Therefore, users will need to perform a verification process before using the content. Technologies to support content verification and appropriate responses based on the results are already known. One example of such a technology is the technology described in Patent Document 1. Patent Document 1 discloses a technology relating to an image generation device, an image generation method, and a program that can generate 3D characters that give a natural impression while suppressing the occurrence of visual differences and unnecessary configurations.
[0004] This technology is an image generation device that applies a 2D image having 2D-direction information to a 3D template as an image template of a 3D model as an image model having 3D-direction information to generate a 3D image of an organism, comprising: an image acquisition means for acquiring the recorded 3D template and the 2D image; a biological 3D model generation means for adding 3D-direction position information to the 2D image of the organism to generate a biological 3D model as a 3D image; a fitting means for using the position information of the biological 3D model to generate a fitted image in which the biological 3D model is fitted to the 3D template; and an image correction means for generating a corrected image in which, when a predetermined inconsistency that causes a visual discomfort occurs in the fitted image, at least a part of the fitted image is corrected by changing the 3D-direction position information.
[0005] Also, in Patent Document 2, in response to a situation where, when the vector used for interpolation is incorrect, edges that do not originally exist are generated and the generated image breaks down, a technique is disclosed for making the breakdown of the generated image less noticeable.
[0006] This technology includes: a determination means for determining whether the image is corrupted in the pixels constituting the second image generated by frame interpolation, which detects motion vectors on the frame of the first image, assigns the detected motion vectors to pixels on the frame of the second image, and generates pixel values for the pixels on the frame of the second image based on the assigned motion vectors; a cluster-up extraction means for extracting a predetermined number of pixels from the second image, including the pixel of interest, as a cluster-up, which are used for classifying the generated pixels of a third image with a higher resolution than the second image, located at a position corresponding to the pixel of interest in the second image; a class classification means for classifying the generated pixels using the cluster-up; and the generated image The present invention relates to an image processing apparatus comprising: a prediction tap extraction means for extracting a predetermined number of pixels from the second image, including the pixel of interest, as prediction taps, which are pixels used for prediction; and a prediction calculation means for generating a third image, which, for the pixel of interest determined by the determination means to be image-degraded, uses a prediction variable for the class of the generated pixel according to the class classification means in a first prediction coefficient group predetermined using the pixel of image degradation, and uses the prediction taps to predict the pixel value of the generated pixel; and for the pixel of interest determined by the determination means not to be image-degraded, uses a prediction variable for the class of the generated pixel according to the class classification means in a second prediction coefficient group predetermined using the pixel of image degradation, and uses the prediction taps to predict the pixel value of the generated pixel.
[0007] Japanese Patent Publication No. 2022-164390 Japanese Patent Publication No. 2009-135878
[0008] On the other hand, the functions of the aforementioned generation AI continue to evolve, and the size and quality of the content that can be generated (for example, the size and resolution of images, and the length of videos) are also expanding. However, it is cumbersome to thoroughly check the generated large or high-resolution content and visually verify that there are no inconsistencies, inconsistencies, or unintended object rendering or inclusion (hereinafter referred to as content errors). If sufficient personnel and time resources are not available for such verification, it becomes difficult to maintain the quality of the content appropriately.
[0009] One embodiment of the present invention has been made in view of the above circumstances, and aims to provide a content generation device, a content generation method, a program, and a recording medium that enable accurate responses to errors in content generated by generation AI.
[0010] The above objective is achieved by a content generation device described in any of the following [1] to [9]. [1] A content generation device comprising a processor that generates content, wherein the processor performs a process of generating content and a process of evaluating at least a portion of the generated content based on predetermined rules.
[0011] [2] The content generation device according to [1], wherein the processor regenerates the portion of the result that does not meet the standard based on the evaluation results.
[0012] [3] The content generation device described in [2], wherein the processor also regenerates the portion of the result that meets the standard.
[0013] [4] The content generation device described in [3], wherein the processor performs regeneration if the event indicated by the result satisfies the requirements.
[0014] [5] The content generating device described in [1], wherein the processor displays information related to the evaluation results.
[0015] [6] The content generation device according to [5], wherein the processor accepts user specification of the area to be regenerated in the content.
[0016] [7] The content generation device according to [1], wherein the processor modifies predetermined rules based on at least one of the information input for content generation and the generated content.
[0017] [8] The content generating apparatus according to [1], wherein the processor evaluates at least a portion of the content based on at least one of the following predetermined rules: the relative importance of objects in the content, the intent of the content creator, and the intent of the content viewer.
[0018] [9] The content generating apparatus according to [8], wherein the processor evaluates at least a portion of the content based on input relating to at least one of the relative importance of objects in the content, the intent of the content creator, and the intent of the content viewer.
[0019] Furthermore, the above objective can also be achieved by the content generation method described in
[10] below.
[10] A content generation method comprising the steps of: generating content with a processor; and evaluating at least a portion of the generated content with the processor based on predetermined rules.
[0020] Furthermore, the above objectives can also be achieved by the program described in
[11] below.
[11] A program for causing a computer to perform each step included in the content generation method described in
[10] .
[0021] Furthermore, the above objectives can also be achieved by the recording medium described in
[12] below.
[12] A computer-readable recording medium on which a program is recorded causing a computer to perform each step of the content generation method described in
[10] .
[0022] According to one embodiment of the present invention, a content generation device, a content generation method, a program, and a recording medium are provided that enable accurate responses to errors in content generated by generation AI.
[0023] This figure shows an example of a system configuration including the content generation device according to this embodiment. This figure shows an example of the hardware configuration of the content generation device according to this embodiment. This figure shows the functional part of the content generation device according to this embodiment. This figure shows an example of content generation in this embodiment. This figure shows an example of a text information table according to this embodiment. This figure shows an example of an input image table according to this embodiment. This figure shows an example of a generated content DB according to this embodiment. This figure shows an example of a generated image according to this embodiment. This figure shows an example of a generated image according to this embodiment. This figure shows an example of an evaluation rule table according to this embodiment. This figure shows an example of a regeneration rule table according to this embodiment. This figure shows an example of an evaluation information output according to this embodiment. This figure shows an example of a content generation method according to this embodiment. This figure shows an example of a generation screen according to this embodiment. This figure shows an example of an evaluation value according to this embodiment. This figure shows an example of a judgment result according to this embodiment. This figure shows an example of a screen based on the evaluation result according to this embodiment. This figure shows an example of a modified image (output image) according to this embodiment.
[0024] The following describes specific embodiments of the present invention. For convenience of explanation, the following descriptions may sometimes be based on the perspective of a GUI (Graphical User Interface). Furthermore, since the fundamental data processing technologies for realizing the present invention (communication / transmission technologies, data acquisition technologies, data recording technologies, data processing / analysis technologies, machine learning technologies, image processing technologies, and visualization technologies, etc.) are known technologies, their descriptions will be omitted.
[0025] Furthermore, in this specification, the concept of "device" includes not only a single device that performs a specific function, but also a combination of multiple devices that exist independently and in a distributed manner while cooperating (linking) to perform a specific function.
[0026] Furthermore, in this invention, "user" refers to a user of the content generation device of the present invention, and specifically, for example, a person who, using the functions of the content generation device of the present invention, checks the content generated by the generation AI and takes appropriate action (e.g., correction) based on the results.
[0027] Furthermore, in this specification, "person" means an entity that performs a specific action, and includes individuals, groups, corporations and other legal entities, and organizations, and may also include computers and devices that constitute artificial intelligence (AI). Artificial intelligence (AI) realizes intelligent functions such as reasoning, prediction, and judgment using hardware and software resources. The algorithm of artificial intelligence is arbitrary and includes, for example, expert systems, case-based reasoning (CBR), convolutional neural networks (CNN), deep neural networks (DNN), Bayesian networks, or inclusion architectures.
[0028] <<About one embodiment of the present invention>> [Configuration of the content generation system] In one embodiment of the present invention (hereinafter, this embodiment), a content generation system 5 is configured with a content generation device 10, a user terminal 30, a shooting device 40, and a server computer 50 connected to a network 1 as shown in Figure 1. The user terminal 30 is an information processing device for which the user inputs prompts (hereinafter, text information) necessary for content generation by the generation AI and uploads them to the content generation device 10. The user terminal 30 and the shooting device 40 also accept user input of images, audio, etc. (hereinafter, base content) that will serve as the basis for content generation by the generation AI, or acquire them by shooting or recording, and upload them to the content generation device 10.
[0029] Furthermore, the user (content creator) may, without using the user terminal 30 or the shooting device 40, read the base content and input text information as prompts using the UI (User Interface) of the content generation device 10. In this case, the UI of the content generation device 10 will read the base content by shooting or recording images from the user's photo prints, the screen of the user terminal 30 (displaying still images or videos), or the audio being played on the user terminal 30. Alternatively, the UI will read such base content data from a recording medium presented by the user. In this case, the user may also directly input the text information using the UI, such as a keyboard, mouse, or touch panel.
[0030] The content generation device 10 stores the base content uploaded from the user terminal 30 and the shooting device 40 in the input image table 212 (described later in Figures 3 and 6), and the text information in the text information table 211 (described later in Figures 3 and 5), in preparation for content generation.
[0031] The content generation device 10 is an information processing device that uses a generation AI 221 (see Figure 3), which is either provided by the device itself or by the server computer 50, to perform image generation and modification (including the concept of regeneration; the same applies hereinafter) based on the base content and text information. Such processing is performed autonomously by the content generation device 10 in accordance with the content generation method of the present invention, or in response to user instructions received through a predetermined UI.
[0032] Furthermore, as described above, the server computer 50 is a server device that provides the content generation device 10 with the same functions as the generation AI 221 and evaluation AI 230 (described later; held by the evaluation unit 23) via the network 1. Of course, if the content generation device 10 is configured in a way that does not require the provision of the generation AI 221 and evaluation AI 230 functions, the server computer 50 does not need to be included in the content generation system 5.
[0033] In addition to the configuration in which the content generation device 10 and the shooting device 40 are connected by a network 1, there is also a configuration in which the internal bus wiring of the content generation device 10 and the interface of the shooting device 40 are directly connected.
[0034] [Example Configuration of Content Generation Device] Next, an example configuration of the content generation device 10 according to this embodiment will be described with reference to Figures 2 and 3. Specifically, the content generation device 10 is composed of a PC (Personal Computer), a smartphone, a tablet terminal, a notebook PC, or a server device, etc. Note that the content generation device 10 is not limited to a computer owned by the user or accessible via network 1. For example, it may be composed of a terminal that is not owned by the user but is available when visiting a store or facility, such as a store-installed terminal. In the following, we will explain using the case in which the content generation device 10 is composed of a computer owned by the user, specifically a PC, as an example.
[0035] As shown in Figure 2, the computer comprising the content generation device 10 includes a processor 11, an auxiliary storage device 12, a main storage device 13, an input device 14, an output device 15, and a communication device 16.
[0036] The processor 11 is composed of, for example, a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), an MCU (Micro Controller Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), a TPU (Tensor Processing Unit), or an ASIC (Application Specific Integrated Circuit).
[0037] The auxiliary storage device 12 is composed of, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), FD (Flexible Disk), MO disk (Magneto-Optical Disk), CD (Compact Disk), DVD (Digital Versatile Disk), SD card (Secure Digital card), or USB memory (Universal Serial Bus memory).
[0038] Such auxiliary storage devices 12 may be built into the computer main body that constitutes the content generation device 10, or they may be attached to the computer main body as external devices. Alternatively, the auxiliary storage device 12 may be configured as a NAS (Network Attached Storage) or the like. Furthermore, the auxiliary storage device 12 may be an external device that can communicate with one of the computers that constitutes the content generation device 10 via a communication network, such as an online storage or database server.
[0039] Furthermore, the auxiliary storage device 12 holds the operating system (OS) and programs 121 such as applications for executing content generation processing. When the programs 121 are read and executed by the processor 11, the computer constituting the content generation device 10 performs the functions of the reception unit 21, generation unit 22, evaluation unit 23, regeneration unit 24, and output unit 25 shown in Figure 3. Specifically, it performs a series of processes related to the generation of content based on text information and base content, and the regeneration of that content based on the evaluation results.
[0040] The main memory 13 is composed of semiconductor memory such as ROM (Read Only Memory) and RAM (Random Access Memory). The processor 11 loads the program 121 onto the main memory 13 and executes it there.
[0041] The input device 14 is a device that receives input operations from a user, and is configured by, for example, a keyboard, a mouse, a touch panel, or a camera unit. Further, the input device 14 may include a photographing device implemented by a digital camera, a microphone for sound collection, and the like. The output device 15 is configured by, for example, a display, a speaker, or the like.
[0042] The communication device 16 may be configured by, for example, a network interface card or a communication interface board. The computer constituting the content generation device 10 can communicate with other devices connected to the network 1 such as the Internet and a mobile communication line via the communication device 16.
[0043] As shown in FIG. 3, the content generation device 10 includes a reception unit 21, a generation unit 22, an evaluation unit 23, a regeneration unit 24, and an output unit 25. These functional units are realized by the processor 11 of the content generation device 10 executing the above program 121 and cooperating with other hardware devices of the content generation device 10. Further, the generation AI 221 in the generation unit 22 and the evaluation AI 230 in the evaluation unit 23 may be called and executed from the server computer 50.
[0044] [Various Definitions and the Like Regarding Content] In the present embodiment, the content generation device 10 of the present invention provides, for example, as shown in FIG. 4, text information Tp (prompt) obtained from the user via the user interface U1 distributed to the user terminal 30 or displayed on the output device 15, or base content Bc such as an image provided by the user (or an image previously generated by the generation AI 221 of the content generation device 10 may also be included) to the generation AI 221 to generate an image P1 (hereinafter referred to as a generated image), which is one type of content.
[0045] The generated image P1 is an image that includes a background image Pb and one or more subject images Ps placed on the background image Pb. In the generated image P1 shown in Figure 4, the background image Pb is an image that constitutes a landscape corresponding to the configuration indicated by the text information Tp (prompt), including an autumn-colored forest Pb1, the sky Pb2, the sea including the coastline Pb3, a railway bridge Pb4, and leaves of trees reflected in the foreground Pb5. Furthermore, the subject image Ps is a train speeding across the railway bridge Pb4 that extends from the background to the foreground in the space of this background image Pb. Of course, these background image Pb and subject image Ps are merely examples, and various things can be used depending on the text information Tp (prompt) and base content Bc provided by the user.
[0046] The background image Pb described above is an image on which the subject image Ps is placed. This background image Pb is the same size as or larger than the subject image Ps, and can be a landscape consisting of natural objects such as mountains, forests, lakes, and seas, or artificial objects such as cities, buildings, and rooms, or it can be a specific monochrome image such as white or black. Furthermore, the background image Pb may be the base content Bc described above. In that respect, the subject image Ps may also be the base content Bc described above.
[0047] The content generation device 10 evaluates the generated image P1 produced by the generation AI 221 from predetermined viewpoints, for example, using the evaluation AI 230, and regenerates a specific area or all of the subject image Ps or background image Pb according to the evaluation result. It is preferable that such regeneration be performed repeatedly (as far as processing time and computer resources allow) so that the evaluation of the image obtained through regeneration meets the criteria. The image that meets the criteria through the above regeneration, i.e., the image that the user is satisfied with, is output as an output image Po to the output device 15 or user terminal 30.
[0048] Note that the "image" in the present invention is composed of a plurality of pixels, represented by the gradation values of each of the plurality of pixels, and includes at least one subject image Ps. Further, digital image data (hereinafter referred to as image data) that defines an image at a set resolution is generated by compressing data in which gradation values for each pixel are recorded using a predetermined compression method. Examples of the types of image data include non-reversible compression image data such as JPEG (Joint Photographic Experts Group) format, and reversible compression image data such as GIF (Graphics Interchange Format) or PNG (Portable Network Graphics) format, etc.
[0049] [Functional Units in Content Generation Device] Subsequently, the processing executed by each functional unit provided in the processor 11 (which is a CPU) of the content generation device 10, namely the reception unit 21, generation unit 22, evaluation unit 23, regeneration unit 24, and output unit 25, will be described.
[0050] (Reception Unit) The processing of the reception unit 21 includes receiving base content Bc such as text information Tp serving as a prompt and various images from the user terminal 30 or the imaging device 40 in the input device 14 or via the communication device 16, and storing these, for example, in the auxiliary storage device 12. Further, the above-described storage processing in the reception unit 21 includes storing the received text information Tp in the text information table 211 and storing the image that is the base content Bc in the input image table 212.
[0051] FIG. 5 shows a configuration example of the text information table 211 of the present embodiment. The text information table 211 is an aggregate of records including values of user ID, date and time, and text information. Among these, the user ID is identification information of the user who input the text information Tp, and may be identification information of the user terminal 30 (the same applies hereinafter). Also, when the content generation device 10 is operated by a single user, this user ID becomes unnecessary. The date and time is the date and time when the text information Tp was input.
[0052] The text information Tp is the description of the prompt that the user provides to the generation AI 221 for content generation. This text information Tp stores a sentence or group of words received by the UI when the user inputs using the input device 14 or user terminal 30. Alternatively, the reception unit 21 may provide the input sentence to a predetermined morphological analysis engine to extract one or more words and store them as text information Tp. The morphological analysis engine may be held by the reception unit 21 or made callable from an external system via the network 1.
[0053] Figure 6 shows an example of the configuration of the input image table 212 in this embodiment. The input image table 212 is a collection of records containing the values of User ID, Date and Time, and Image Data. The User ID is the identification information of the user who input the image which is the base content Bc, and may also be the identification information of the user terminal 30. Furthermore, if the content generation device 10 is operated by a single user, this User ID is not required. The Date and Time is the date and time when the base content Bc was input.
[0054] The method for receiving text information Tp and base content Bc in the reception unit 21 is not particularly limited, but includes methods of acquisition by performing photo printing of images taken by the user terminal 30 or the shooting device 40, or by reading playback videos with the camera unit. In addition, the reception unit 21 may acquire images by downloading image data from external devices or a web server via the network 1.
[0055] (Generation Unit) The processing of the generation unit 22 includes the process of generating a generated image P1 using, for example, the generation AI 221, based on the text information Tp and base content Bc obtained in the processing of the reception unit 21. The generation AI 221 is an artificial intelligence that automatically generates, for example, an image as content desired by the user, based on the text information Tp and base content Bc entered by the user.
[0056] In the processing of the generation unit 22, the generated image P1 is stored in the generated content DB 222. As shown in Figure 7, this generated content DB 222 is a collection of records consisting of a user ID, date and time, and image data. The user ID is the ID of the user who gave the instruction to generate the generated image P1. The date and time is the date and time the generated image P1 was generated. The image data is the data file of the generated image P1. The base content Bc may also include what the generation unit 22 generates by attaching text information Tp and base content Bc to the generation AI 221. Furthermore, the text information Tp may also include what the generation unit 22 generates by attaching base content Bc to an appropriate generation AI.
[0057] An example of a generation AI that generates text information Tp is a caption generation AI that outputs keywords that appropriately describe an input image for that input image. Therefore, the caption generation AI has a model that has been machine-learned using sets of various input images and keywords that appropriately describe those input images as training data. The generation AI 221 (including the caption generation AI) can be used through a service provided on network 1, but the content generation device 10 may also keep it itself.
[0058] In this embodiment, the processing of the generation unit 22 includes generating a subject image Ps corresponding to text information Tp or base content Bc using the generation AI 221, and compositing it with another base content Bc which is a background image Pb, when generating the generated image P1. Alternatively, the processing of the generation unit 22 includes generating a background image Pb corresponding to text information Tp or base content Bc using the generation AI 221, and compositing it with another base content Bc which is a subject image Ps. Alternatively, the processing of the generation unit 22 includes generating a subject image Ps corresponding to text information Tp or base content Bc using the generation AI 221, and compositing it with a background image Pb generated by the generation AI 221 based on another text information Tp or another base content Bc, when generating the generated image P1.
[0059] (Evaluation Unit) The evaluation unit 23 also performs a process of evaluating at least a portion of the generated image P1 generated by the generation unit 22 based on predetermined rules. These "predetermined rules" may include those determined immediately before the evaluation. The basic concept of "evaluation" in the evaluation unit 23 is, for example, to determine whether or not there is any sense of discomfort that a person would feel when viewing the image. The region of the generated image P1 that is subject to such evaluation is at least one of the subject image Ps and the background image Pb, and may be all or part of the subject image Ps or the background image Pb.
[0060] The above evaluation itself is performed by an evaluation AI 230 that has a judgment model for events that can lead to the above-mentioned "discomfort," such as differences in the standard shape, number, and size of an object or the parts that make it up, the distance between the part and the surrounding parts, and discontinuities in contours and colors. The judgment model of this evaluation AI 230 is a model that uses deep learning to determine the correspondence between events such as the general shape, number, size, color, distance from the surroundings and continuity of each object and its parts, and the state in which a person feels discomfort regarding those events, and outputs a higher evaluation value the less discomfort there is. For example, a judgment model that judges discomfort with a person's face in an image would be a model that judges the lack of events such as differences in the standard shape, number, and size of facial parts (e.g., eyes, nose, mouth, ears, chin, etc.), differences in the standard distance between the part and the surrounding parts, and discontinuities in contours and colors.
[0061] For example, if the subject image Ps of the generated image P1 includes a "human face," viewers of that generated image P1 (content viewers) are likely to focus on the face. Therefore, as in image G1 shown in Figure 8A, if the shape or color of the face image Gf, which is the subject image Ps in the generated image P1, is abnormal or discontinuous with its surroundings, such as the eyes Ge, the viewer will immediately feel something is off. On the other hand, as in image G2 shown in Figure 8B, even if the leaves and branches Gt of the trees in the background image Pb are in an unrealistic state (e.g., needle-like with no branches at all, or multiple types of plant leaves on one tree), viewers are less likely to notice this phenomenon. Even if they do notice such a phenomenon, viewers often don't pay it much attention. In other words, people unconsciously tend to focus on the "face" in the generated image P1 and judge whether there is anything unusual about it.
[0062] Therefore, in the processing of the evaluation unit 23, regarding the evaluation value by the evaluation AI 230, if the generated image P1 contains an image (object) of a "face", the evaluation AI 230 sets a standard such that the face region must meet a relatively stricter standard (which can also be called importance or weighting) compared to the other regions. Regions that do not meet this standard will be subject to regeneration. For example, if the standard for the face region is set to "80", and the evaluation value output by the evaluation AI 230 for the "face" region in a given generated image P1 is "70", then that "face" region will be subject to regeneration.
[0063] On the other hand, for elements included in the background image Pb of the "face" region, such as "tree leaves," a relatively looser standard will be set compared to the face region. For example, if the standard for the "tree leaves" region is set at "60," and the evaluation value output by the evaluation AI 230 for the "tree leaves" region in a generated image P1 is "65," then that "tree leaves" region will not be included in the regeneration process.
[0064] Figure 9 shows an example of the configuration of the evaluation rule table 231 in this embodiment. This evaluation rule table 231 is a table that defines the above criteria according to the conditions of the subject image Ps and background image Pb. The evaluation rule table 231 is a collection of records that link each viewpoint and each value of the criteria, with an ID that identifies the rule as the key. The viewpoints describe the definitions of events such as subject, background, user input, area ratio (subject / background), and inappropriate elements. The criteria are thresholds that serve as the criteria for determining whether the evaluation value (determined by the evaluation AI 230) for the target area of the generated image P1 is subject to regeneration, that is, whether a person feels something is wrong with it.
[0065] The evaluation unit 23 applies each rule from the evaluation rule table 231 or a predetermined rule to the generated image P1 to be evaluated, and determines whether each region in the generated image P1 (regions corresponding to images of people or landscapes, or images of parts that make them up) is natural or not, that is, whether or not it should be regenerated, based on the evaluation value (determined by the evaluation AI 230).
[0066] Among the rules in the evaluation rule table 231, for example, ID "R00" indicates that if the generated image P1 contains an inappropriate element, the image of the area containing that inappropriate element will be subject to regeneration according to criterion "101," that is, even if the evaluation value for the absence of incongruity by the evaluation AI 230 is 100. These inappropriate elements include, for example, images, sounds, and text that violate public order and morals or the standards of personal information protection. Therefore, the evaluation AI 230 has already learned about such inappropriate elements and is capable of identifying inappropriate elements in the generated image P1.
[0067] Furthermore, ID "R01" specifies that if the subject in the generated image P1 is a "person", and the ratio of the area of the person to the area of the rest of the background image Pb is 90% or more, that is, if the generated image P1 is almost entirely occupied by an image of a person, then it is determined whether or not to regenerate the image based on the following criteria: the face area of the person's image is based on a standard of "80", the body area on a standard of "70", and the rest of the background image Pb on a standard of "50".
[0068] Furthermore, ID "R01-1" specifies that if the subject in the generated image P1 includes a "face," and the ratio of the area of the face to the area of the rest of the background image Pb is 60% or more, that is, if the human face image occupies more than half of the generated image P1, then the human face image is judged to be subject to regeneration based on a standard of "80," and the rest of the background image Pb is judged to be subject to regeneration based on a standard of "50."
[0069] Furthermore, ID "R01-2" specifies that if the subject in the generated image P1 includes "hands and feet," the images of those hands and feet are judged to be subject to regeneration based on a standard of "70," while the rest of the background image Pb is judged to be subject to regeneration based on a standard of "65." Additionally, ID "R01-3" specifies that if the subject in the generated image P1 includes a "body," the images of that body are judged to be subject to regeneration based on a standard of "70," while the rest of the background image Pb is judged to be subject to regeneration based on a standard of "65."
[0070] Furthermore, ID "R02" specifies that if the subject of the generated image P1 includes a "person" and the background image Pb is a landscape, the person's image is judged to be subject to regeneration based on a criterion of "80," while the rest of the background image Pb is judged to be subject to regeneration based on a criterion of "60." On the other hand, ID "R02A" specifies that if the subject of the generated image P1 includes a "person," the background image Pb is a landscape, and the user input (text information Tp (prompt) or base content Bc) is a landscape, that is, if it can be inferred that the user who instructed the generation of the generated image P1 places importance on the "background," then the person's image is judged to be subject to regeneration based on a criterion of "60," while the rest of the background image Pb is judged to be subject to regeneration based on a criterion of "80." Since these areas that are subject to user input can be considered areas that should be given importance based on the user's intentions and preferences, the evaluation unit 23 modifies rule "R02" to specify in rule "R02A" that the person's image is judged to be subject to a criterion of "60," while the rest of the background image Pb is judged to be subject to a criterion of "80."
[0071] Furthermore, ID "R02B" specifies that if the subject of the generated image P1 includes a "person", the background image Pb is a landscape, and the ratio of the area of the person to the area of the rest of the background image Pb is 60% or more, and the user input (text information Tp (prompt) or base content Bc) is a landscape, that is, even though the area of the person in the subject image Ps is larger than the area of the landscape in the background image Pb, it can be inferred that the user who instructed the generation of generated image P1 prioritized the "background", then it is determined whether or not to regenerate the image based on the criterion "60" for the area of the rest of the background image Pb. This rule "R02B", like rule "R02A", is a modification of rule "R03" by the evaluation unit 23.
[0072] Furthermore, ID "R03" specifies that if the subject in the generated image P1 includes something other than a person (e.g., a building or equipment), and the background image Pb is a landscape, and the ratio of the area of the non-person subject to the area of the rest of the background image Pb is 15% or less, then the criteria for whether or not to regenerate the image of the non-person is set to "70", and the area of the rest of the background image Pb is set to "80".
[0073] On the other hand, ID "R03A" specifies that if the generated image P1 includes something other than a person, the background image Pb is a landscape, and the user input (text information Tp (prompt) or base content Bc) is a landscape, that is, if it can be inferred that the user who instructed the generation of generated image P1 prioritized the "background", then the image of the non-person object is judged to be subject to regeneration based on criterion "60", and the rest of the background image Pb is judged to be subject to regeneration based on criterion "80". In this case as well, the evaluation unit 23 modifies rule "R04" and generates rule "R03A".
[0074] Furthermore, the concept of modifying the rules as described above may include modifications based not only on the intention of the person who issued the instruction to generate the generated image P1, but also on the intentions of the viewers of such generated image P1. One example of a viewer's intention is receiving specifications regarding the viewer's preferences from the user terminal 30 of the person who will view (or is expected to view) the generated image P1, and modifying the values of the above criteria according to those specifications. These specifications may include, for example, the attributes of the viewer, their associates, or preferred subjects.
[0075] Furthermore, the evaluation unit 23 may perform the evaluation using a rule specified by the user terminal 30 of the person who gave the instruction to create the generated image P1, or a viewer, from among the rules included in the evaluation rule table 231.
[0076] (Regeneration Unit) The regeneration unit 24 processes the generated image P1 to the extent that the evaluation value from the evaluation unit 23 for the generated image P1 satisfies the requirements, and regenerates the region of the generated image P1 that includes at least the region to be evaluated. The regeneration process in the regeneration unit 24 is executed by instructing the generation unit 22 to regenerate, or by calling and using the generation AI 221. The image regenerated by the regeneration unit 24 (the generated image P1 or an image of a part of its region) is stored in the generation content DB 222 or held in the main memory 13.
[0077] Figure 10 shows the regeneration rule table 232 that defines the above requirements. This regeneration rule table 232 is a collection of records that associate an ID that identifies a regeneration rule with values for perspective and regeneration policy. The perspective defines the evaluation values and conditions related to various situations concerning the subject image Ps, background image Pb, or each region that constitutes them in the generated image P1. The regeneration policy defines the content of the regeneration to be performed when the conditions, i.e., requirements defined in the perspective are met.
[0078] In the regeneration rule table 232 of Figure 10, for example, ID "R51" specifies the regeneration policy "regenerate the entire area of the target subject" for the criterion "the evaluation of two or more areas among the face, hands, and feet is below the standard." Also, ID "R52" specifies the regeneration policy "regenerate only the target area of the target subject" for the criterion "the evaluation of only one area among the face, hands, and feet is below the standard." Also, ID "R53" specifies the regeneration policy "regenerate only the face of the target subject" for the criterion "the evaluation of only the face among the face, hands, and feet is below the standard." Also, ID "R54" specifies the regeneration policy "regenerate the entire area of the target subject" for the criterion "the evaluation of two or more areas including the face among the face, hands, and feet is below the standard." Also, ID "R55" specifies the regeneration policy "regenerate only the target area of the target subject" for the criterion "the evaluation of areas other than the face among the face, hands, and feet is below the standard." Furthermore, for ID "R56," the perspective "Expected playback time is below the acceptable standard" specifies the regeneration policy "Regenerate the area where the expected playback time is below the acceptable standard." Also, for ID "R57," the perspective "Content is a still image" specifies the regeneration policy "Regenerate the entire area." And for ID "R58," the perspective "Content is a video" specifies the regeneration policy "Regenerate the area where the evaluation is below the standard."
[0079] As shown in the configuration of the regeneration rule table 232 above, based on the perspective of unnaturalness already mentioned, the main policy is to regenerate the image by giving weight to the main parts of the subject image Ps, such as the face, hands, and feet, when a person is the subject. Furthermore, the lower the overall evaluation value of these main parts, the more likely the subject image Ps is to be regenerated as a whole. Therefore, the regeneration unit 24 applies the evaluation result from the evaluation unit 23 (whether the evaluation value exceeds the standard) to each rule in the regeneration rule table 232 to determine the regeneration policy, and then executes the regeneration process according to that regeneration policy.
[0080] The regeneration unit 24 may, instead of determining the regeneration policy using the regeneration rule table 232, identify the area to be regenerated in the generated image P1 in response to instructions from the person who issued the generation instruction for the generated image P1 or from the viewer (received via the user terminal 30), and then perform regeneration on that area.
[0081] Furthermore, as exemplified by rules "R57" and "R58" in the regeneration rule table 232 of Figure 10, the regeneration policy may be determined based on the attributes of the generated image P1 itself, such as whether the generated image P1 is a still image or a video, and then regeneration may be performed. In addition, a rule may be adopted to determine the regeneration policy according to the length of time required for regeneration. In this case, for example, for areas where the evaluation value is below a standard, the time required to regenerate only that area and the time required to regenerate the entire image including that area (e.g., the entire subject) are predicted, and the regeneration policy that is within the acceptable time or has a shorter time is adopted. This method forms the basis of the concept of deciding to "regenerate the entire area" when the content is a still image (i.e., the time required is shorter than for a video) as described in rule "R57", and to "regenerate the area where the evaluation is below a standard" when the content is a video (i.e., the time required is longer than for a still image) as described in rule "R58".
[0082] (Output Unit) The output unit 25 processes the generated image P1 generated by the generation unit 22, or the image of the region regenerated by the regeneration unit 24, or the modified image P2 including the region, on an output device 15 such as a display or on a user terminal 30. These generated image P1 and modified image P2 can become the output image Po. If the generated image P1 and modified image P2 are images that are in a state where there are no problems with the evaluation value by the evaluation unit 23, they may be displayed as the final output image Po. For the display of such images, the output unit 25 appropriately retrieves image data from the generated content DB 222 (see Figure 7) in response to requests from the user terminal 30 or the arrival of a predetermined timing, and outputs it.
[0083] The method of outputting the output image Po is not particularly limited, but includes, for example, the output unit 25 displaying the output image Po on the display or monitor of the output device 15 or user terminal 30, printing the output image Po, transmitting the output image Po to other users, and providing the output image Po as a commercial product. The output image Po as a commercial product may include media consisting of one or more pages or cards on which images are posted, such as albums, photobooks, postcards, message cards, electronic albums, and bromide prints.
[0084] Furthermore, the output unit 25 may display information related to the evaluation value from the evaluation unit 23 on the user terminal 30 or output device 15. In the display example shown in Figure 11, on screen G3A, the areas of the subject image Ps, which is the image of a person Gk, the background image Pb, which is the image of a dog Gd, and the image of trees Gt are each demarcated by line segments, and a marker Gm is attached to the area in each region where the evaluation value is below the standard. In other words, information related to the evaluation value is arranged on screen G3A as a marker Gm.
[0085] In the example shown in Figure 11, the evaluation value for the image Gt of trees is below the standard, and a marker Gm is placed there. The person who issued the generation instruction for the generated image P1, i.e., screen G3A in this case, or the viewer, can take action by operating the user terminal 30 to notify the regeneration unit 24 of a click operation using the cursor Cs on the area containing the marker Gm and to instruct it to regenerate.
[0086] In the example shown in Figure 11, when the user clicks the area of the dog image Gd near the center of screen G3A with the cursor Cs, the color of the area of the dog image Gd inverts to black, as shown in screen G3B, and the regenerate button Br is clicked. Of course, the form illustrated in Figure 11 is just one example, and a form can be adopted in which different colors or patterns are assigned to each area depending on the magnitude of the evaluation value. Alternatively, a form can be adopted in which the numerical value of the evaluation value is assigned to each area.
[0087] [Example of Content Generation Method Flow] Next, as an example of the operation of the content generation device 10 in this embodiment, an image processing flow using the device will be described. The content generation method of the present invention is used in the image processing flow described below. In other words, each step in the image processing flow described below corresponds to a component of the content generation method of the present invention. Note that the following flow is merely an example, and some steps in the flow may be deleted, new steps added to the flow, or the execution order of two steps in the flow may be changed without departing from the spirit of this embodiment.
[0088] Each step in the content generation flow according to this embodiment is performed by the processor 11 of the content generation device 10 in the order shown in Figure 12. In other words, in each process of the content generation flow, the processor 11 executes the data processing defined in the application program for content generation that corresponds to each step in Figure 12.
[0089] To explain in more detail, in the content generation flow according to this embodiment, first, the receiving unit 21 receives at least one of either text information Tp or base content Bc from, for example, the user terminal 30 (S1). During this process, the receiving unit 21 delivers a receiving screen, including U1 as shown in Figure 4, to the user terminal 30. The receiving unit 21 also obtains the text information Tp via the prompt input field of U1 on the receiving screen. The receiving unit 21 also obtains the base content Bc file specified by the user in the file specification field of the receiving screen.
[0090] The acquisition of this text information Tp and base content Bc is necessary for the generation of the generated image P1 by the generation AI 221 in the generation unit 22. If the receiving unit 21 receives only the text information Tp, the value of the text information Tp is transmitted to the generation unit 22, and the generation image P1 is generated by the generation AI 221 (S2). Figure 13 shows an example of the generated image P1. In this generated image P1, the subject image Ps of the woman is placed on the background image Pb which includes the forest image Gt and the sky image Ga. The subject image Ps includes the image Gf of the woman's face and the image Gh of her hands.
[0091] If the reception unit 21 receives only the base content Bc, the base content Bc is transmitted to the generation unit 22, and the base content Bc is used as the subject image Ps or background image Pb when the generation AI 221 generates the generated image P1. On the other hand, if the reception screen receives both text information Tp and the base content Bc, they are transmitted to the generation unit 22, and the generation AI 221 of the generation unit 22 executes, for example, the process of generating a subject image Ps based on the text information Tp, and the process of arranging the subject image Ps as a background image Pb on the base content Bc to generate the generated image P1.
[0092] Next, the evaluation unit 23 extracts the generated image P1 generated in S2 from the main memory 13 or the generated content DB 222 and assigns it to the evaluation AI 230, thereby evaluating at least a portion of the generated image P1 based on the rules (S3). As already mentioned, this evaluation involves the evaluation AI 230 determining whether or not there is any sense of discomfort that a human would feel when viewing the image. The region of the generated image P1 that is subject to such evaluation is at least one of the subject image Ps and the background image Pb, and may be all or part of the subject image Ps or the background image Pb.
[0093] Figure 14 shows an example of evaluation values obtained in the evaluation in S3 for the generated image P1 (Figure 13) described above. In this example of evaluation values, the generated image P1 in Figure 13 is the identification information "GP01", and the area ratio, presence or absence of inappropriate elements, and evaluation value are set for each part of the subject image Ps and background image Pb. Among the subject image Ps and background image Pb in the generated image P1, the subject image Ps of the "hand" has an evaluation value of "48", and if the criterion for regeneration is, for example, "70", then this subject image Ps of the "hand" will be determined to be subject to regeneration. Such evaluation value information is stored, for example, in the main memory 13.
[0094] Furthermore, when performing the above evaluation in the evaluation unit 23, if, for example, user input is received prior to the generation of the generated image P1 (text information Tp or base content Bc is input from the user terminal 30 or input device 14), or if the area ratio between the subject image Ps and the background image Pb in the generated image P1 satisfies the requirements, the relevant rules in the evaluation rule table 231 may be modified as already mentioned. Of course, modifications based on the viewer's intentions for the generated image P1 (for example, the preferences of those who will view the generated image P1 (including prospective viewers)) may also be included.
[0095] Next, the evaluation unit 23 applies each rule from the evaluation rule table 231 (see Figure 9; previously mentioned) or a predetermined rule to the generated image P1 (generated in S2) that is to be evaluated, and determines whether each region in the generated image P1 (regions corresponding to images of people or landscapes, or images of parts that make them up) is natural or not, that is, whether or not it is subject to regeneration, based on the evaluation value (determined by the evaluation AI 230) (S4).
[0096] In the generated image P1 produced in S2 above, the evaluation value (by the evaluation AI 230) for the human "hand" in the subject image Ps was "48". In this case, based on the criterion "Subject: 70" of rule "R01-2" in the evaluation rule table 231, the area of the "hand" is determined to be an image that causes discomfort and is subject to regeneration by the regeneration unit 24. The result of such a determination is stored, for example, in the main memory 13.
[0097] Figure 15 shows an example of the results of the above judgment stored in the main memory 13. The judgment result table shown in Figure 15 has the same evaluation value table structure as shown in Figure 14, but differs in that the judgment result column has either an OK or NG value set. As described above, in the generated image P1 obtained in S2, the "hand" is subject to regeneration, so the judgment result is set to "NG".
[0098] The evaluation unit 23 obtains the determination result for whether or not various subjects and their parts included in the subject image Ps and background image Pb in the generated image P1 are targets for regeneration (i.e., whether or not the evaluation value by the evaluation AI 230 is below the standard). The evaluation unit 23 then applies the determination results obtained for each subject and its parts to the rules of the regeneration rule table 232 to determine the regeneration policy (S5). For example, if the evaluation value for the "hand" region among the "face," "hand," and "foot" regions of the subject image in a given generated image P1 is lower than the standard and it is determined to be a target for regeneration, the regeneration policy is determined to "regenerate only the target region of the target subject image" based on rule "R52".
[0099] The determination of such a regeneration policy may be made based on the attributes of the generated image P1 itself, such as whether the generated image P1 is a still image or a video, as explained with respect to the regeneration rule table 232 in Figure 10, as exemplified by rules "R57" and "R58," and then the regeneration may be performed. Alternatively, a rule may be adopted to determine the regeneration policy according to the length of time required for regeneration. In this case, for example, for areas where the evaluation value is below a standard, the time required to regenerate only that area and the time required to regenerate the entire image including that area (e.g., the entire subject) are predicted, and the regeneration policy that is within the acceptable time or has a shorter time is adopted.
[0100] Regarding the time required, for example, the evaluation unit 23 may employ a method of calculation based on information such as the regeneration time per unit area of still images and the regeneration time per unit duration of video in the generating AI 221, as well as the area of the region to be regenerated and the length of the video.
[0101] Next, the regeneration unit 24 performs regeneration of the target area according to the regeneration policy determined in S5 (S6). At this time, it is preferable to highlight the area determined to be regenerated (hand image Gh) by dividing it with line segments and inverting it with a specific color, as shown in Figure 16. The person who issued the generation instruction for the generated image P1 or the viewer can view the screen in Figure 16 and visually recognize the location of the target to be generated and its validity.
[0102] The image regenerated in S6, i.e., the corrected image P2 (see Figure 17), is stored, for example, in the main memory 13. The evaluation unit 23 also performs an evaluation of the corrected image P2 obtained in S6 in the same manner as in S3 (S7). The evaluation unit 23 also determines, in the same manner as in S4, whether the corrected image P2 contains the object to be regenerated. If it contains the object to be regenerated (S8:Y), the process returns to S5. On the other hand, if the evaluation unit 23 determines that the corrected image P2 does not contain the object to be regenerated (S8:N), it passes the corrected image P2 to the output unit 25. The output unit 25 displays this corrected image P2 as output image Po on the user terminal 30 or output device 15 (S9), and this flow ends.
[0103] In the example above, we described a case where the content generated and regenerated by the content generation device 10 is an image or video, but the content is not limited to this. For example, it could be audio or music. In that case, the base content Bc could be spoken audio, singing voices, instrument sounds, or musical compositions.
[0104] Although specific embodiments of the present invention have been described above, these embodiments are merely examples given to facilitate understanding of the present invention and do not limit it. That is, the present invention can be modified or improved from the embodiments described below, without departing from its spirit. Furthermore, the present invention includes equivalents thereof. Moreover, embodiments of the present invention may include forms that combine the above embodiments with one or more of the following modifications.
[0105] (Regarding the computer constituting the content generation device) In the above embodiment, the content generation device 10 of the present invention is configured with a computer used directly by the user, such as a user-owned PC (Personal Computer). However, it is not limited to this, and the content generation device 10 of the present invention may be configured with a computer that the user can use indirectly, for example, a server computer 50. Here, the server computer 50 may be, for example, a server computer for cloud services, specifically a server computer for ASP (Application Service Provider), SaaS (Software as a Service), PaaS (Platform as a Service), or IaaS (Infrastructure as a Service). In this case, when the user inputs the necessary information on the user terminal 30 owned by the user, the server computer 50 performs various processes (calculations), including image generation, based on the input information, and the calculation results are output on the user terminal 30. In other words, the functions of the server computer 50 that constitute the content generation device 10 of the present invention can be used on the user terminal 30. Alternatively, the content generation device 10 may be configured using a shooting device 40 or a smartphone (a type of user terminal 30) used by the user, in addition to the PC or server computer mentioned above.
[0106] (Regarding the processor configuration) In this embodiment, each process is executed on any computer. Furthermore, any computer may execute these processes using a processor as hardware, a program as software, or a combination thereof. In that case, the processor is configured to work in cooperation with the program to execute the various processes in this embodiment, and can function as a unit or means in this embodiment. Also, the execution order of the processes by the processor is not limited to the order described and may be changed as appropriate. Any computer may be a general-purpose computer, a computer designed for a specific purpose, a workstation, or any other system capable of executing each process.
[0107] A processor may consist of one or more hardware components, and the type of hardware is not limited. For example, a processor may consist of programmable logic devices such as a CPU (Central Processing Unit), MPU (Micro Processing Unit), FPGA (Field Programmable Gate Array), dedicated circuits for performing specific processing such as an ASIC (Application Specific Integrated Circuit), a GPU (Graphic Processing Unit), or an NPU (Neural Processing Unit).
[0108] Furthermore, the processor has various units or means that execute the various processes in this embodiment. The hardware may also be a combination of different types of hardware. When multiple hardware components are configured to execute one or more processes of a processor, these components may reside in physically separate devices or in the same device. In any embodiment, the order of the processes performed by the processor is not limited to the order described above and may be changed as appropriate. The hardware is composed of electrical circuits (circuitry) and the like, which are combinations of circuit elements such as semiconductor elements.
[0109] Furthermore, this embodiment may be implemented by hardware, software, firmware, microcode, or a combination thereof. Software, firmware, and microcode are composed of a program. The program may also be, for example, a group of program modules, and each function may be implemented by a processor configured to perform its respective function. The program may be program code or multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage media or other storage). Furthermore, the program may be firmware or software such as microcode. The program may also be, for example, a group of program modules, and each function may be implemented by a processor configured to perform its respective function. The program may be program code or multiple code segments stored on one or more non-temporary computer-readable media (e.g., storage media or other storage). The program may be divided and stored on multiple non-temporary computer-readable media located on devices that are physically separated from each other. The program code or code segments may represent any combination of procedures, functions, subprograms, routines, subroutines, modules, software packages, classes, or instructions, data structures, or program statements. Program code or code segments may be connected to other code segments or hardware circuits by sending and receiving information, data, arguments, parameters, or memory contents.
[0110] 1 Network 5 Content Generation System 10 Content Generation Device 11 Processor 12 Auxiliary Storage Device 121 Program 13 Main Storage Device 14 Input Device 15 Output Device 16 Communication Device 21 Reception Unit 211 Text Information Table 212 Input Image Table 22 Generation Unit 221 Generation AI 222 Generation Content DB 23 Evaluation Unit 230 Evaluation AI 231 Evaluation Rule Table 232 Regeneration Rule Table 24 Regeneration Unit 25 Output Unit 30 User Terminal 40 Shooting Device 50 Server Computer U1 User Interface Tp Prompt (Text Information) Bc Base Content P1 Generated Image P2 Modified Image Pb Background Image Ps Subject Image Po Output Image
Claims
1. A content generation device comprising a processor that generates content, wherein the processor performs a process of generating content and a process of evaluating at least a portion of the generated content based on predetermined rules.
2. The content generation apparatus according to claim 1, wherein the processor regenerates the portion of the result that does not meet the standard, based on the result of the evaluation.
3. The content generation apparatus according to claim 2, wherein the processor also regenerates the portion of the result that meets the standard.
4. The content generation apparatus according to claim 3, wherein the processor performs the regeneration when the event indicated by the result satisfies the requirements.
5. The content generation apparatus according to claim 1, wherein the processor displays information related to the results of the evaluation.
6. The content generation apparatus according to claim 2, wherein the processor accepts a user's designation of the area to be regenerated in the content.
7. The content generation apparatus according to claim 1, wherein the processor modifies the predetermined rules based on at least one of the information input for generating the content and the generated content.
8. The content generation apparatus according to claim 1, wherein the processor evaluates at least a portion of the content based on at least one of the predetermined rules: the relative importance of objects in the content, the intent of the content creator, and the intent of the content viewer.
9. The content generating apparatus according to claim 8, wherein the processor evaluates at least a portion of the content based on input relating to at least one of the relative importance of objects in the content, the intent of the content creator, and the intent of the content viewer.
10. A content generation method comprising: a step of generating content using a processor; and a step of evaluating at least a portion of the generated content using a processor based on predetermined rules.
11. A program for causing a computer to perform each step included in the content generation method described in claim 10.
12. A computer-readable recording medium on which a program is recorded causing a computer to perform each step included in the content generation method described in claim 10.
Citation Information
Patent Citations
Information processing system, information processing method, and program
JP7396762B1
Information processing system, program, and information processing method
JP7448271B1
Method for providing image generation service by sharing image generation ai model trained in specific style and apparatus therefor
KR102692283B1