Information processing method, program, and information processing apparatus
The method addresses the issue of restricting content generated by generative models that are similar to copyrighted works by using a system of models like prompt extraction and similarity calculation to determine whether or not to allow the use of the generated content, ensuring compliance with copyright laws and ethical use of generated content.
Patent Information
- Application Number
- JP2025068991
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-28
- Filing Date
- 2025-04-18
- Publication Date
- 2025-12-10
AI Technical Summary
Existing technologies do not adequately address the restriction of content generated by generative models when it resembles copyrighted works.
An information processing method that generates content using a generative model and determines its use based on similarity to copyrighted works, employing models like prompt extraction, edge detection, and similarity calculation to control the output.
Enables the restriction of content similar to copyrighted works, ensuring compliance with copyright laws and ethical use of generated content.
Smart Images

Figure 2025179807000001_ABST
Abstract
Description
[Technical Field]
[0001] The present technology relates to an information processing method, a program, and an information processing device. [Background technology]
[0002] Conventionally, a technique for displaying an alter ego of a user in a virtual space has been proposed (for example, Patent Document 1). [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2024-021529 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology described in Patent Document 1 does not take into consideration restricting the use of content generated by a generative model when the content is similar to a copyrighted work.
[0005] The present disclosure has been made in consideration of such circumstances, and aims to provide an information processing method, etc., that can restrict the use of content when the content generated by a generative model is similar to a copyrighted work. [Means for solving the problem]
[0006] An information processing method according to one embodiment of the present disclosure generates first content using a generation model that generates content, and determines whether or not to allow use of the first content based on the similarity between the first content and second content, which is a copyrighted work. [Effects of the Invention]
[0007] In an information processing method according to an embodiment of the present disclosure, if content generated by a generative model is similar to a copyrighted work, it is possible to restrict the use of the content. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is an explanatory diagram illustrating an overview of an information processing system. [Figure 2] FIG. 1 is a block diagram illustrating an example of the configuration of an information processing device. [Figure 3] FIG. 2 is a block diagram illustrating an example of the configuration of a user terminal. [Figure 4] FIG. 10 is an explanatory diagram showing an example of image conversion specification information (prompt) settings. [Figure 5] FIG. 2 is an explanatory diagram illustrating an example of an edge detection module. [Figure 6] FIG. 1 is an explanatory diagram showing an example of a learning method for an image generation model. [Figure 7] FIG. 1 is an explanatory diagram illustrating an example of an image generation model. [Figure 8] FIG. 10 is an explanatory diagram illustrating an example of a similarity calculation model. [Figure 9] FIG. 10 is an explanatory diagram illustrating an example of a work table. [Figure 10] FIG. 10 is an explanatory diagram illustrating an example of a similarity table. [Figure 11] FIG. 10 is an explanatory diagram showing an example of a work selection screen. [Figure 12] FIG. 10 is an explanatory diagram showing an example of a user image capturing screen. [Figure 13] FIG. 10 is an explanatory diagram showing an example of a work level selection screen. [Figure 14] 10 is a flowchart illustrating an example of processing by a processing unit of the information processing device. [Figure 15] FIG. 10 is an explanatory diagram showing an example of a learning method for an image generation model according to the second embodiment. [Figure 16] FIG. 10 is an explanatory diagram showing an example of a similarity calculation model according to the second embodiment. [Figure 17] FIG. 10 is an explanatory diagram showing an example of a work table according to the second embodiment. [Figure 18] FIG. 10 is an explanatory diagram illustrating an example of a character table. [Figure 19] FIG. 10 is an explanatory diagram showing an example of a character selection screen. [Figure 20] FIG. 10 is an explanatory diagram showing an example of a work level selection screen according to the second embodiment. [Figure 21] 10 is a flowchart illustrating an example of processing by a processing unit of an information processing device according to a second embodiment. [Figure 22] FIG. 10 is an explanatory diagram showing an example of a generated image use screen. DETAILED DESCRIPTION OF THE INVENTION
[0009] (Embodiment 1) An information processing device 1 according to the first embodiment generates a generated image (first content) that resembles the artistic style of a work based on an acquired user image. The information processing device 1 calculates the similarity between the generated image and an image (second content) of a character appearing in the copyrighted work, and determines whether or not to permit use of the generated image based on the calculated similarity. The information processing device 1 transmits the generated image whose use is permitted to a user terminal 2 and provides it to the user. Hereinafter, the first embodiment will be described with reference to the drawings.
[0010] FIG. 1 is an explanatory diagram showing an overview of an information processing system S. The information processing system S includes an information processing device 1 and a user terminal 2. The information processing device 1 is, for example, a cloud server, and the user terminal 2 is, for example, a smartphone, tablet terminal, or personal computer owned by a user, or a passport photo machine or print sticker machine such as Purikura (registered trademark) operated by the user. In this embodiment, the user terminal 2 is described as a smartphone. The information processing device 1 and the user terminal 2 can communicate with each other via a network N. The user terminal 2 transmits a user image captured of the user to the information processing device 1. The information processing device 1 generates a generated image based on the acquired user image and transmits the generated image to the user terminal 2. The user terminal 2 displays the generated image received from the information processing device 1. Note that if the user terminal 2 is a passport photo machine or print sticker machine, the user terminal 2 may have a user interface function and communicate with the information processing device 1 via the network N, and may execute some or all of the processing executed by the information processing device 1.
[0011] FIG. 2 is a block diagram showing an example configuration of information processing device 1. Information processing device 1 generates a generated image (first content) that resembles the style of a work based on an acquired user image. Information processing device 1 calculates the similarity between the generated image and an image (second content) of a character appearing in the copyrighted work, and determines whether or not to permit use of the generated image based on the calculated similarity. Information processing device 1 transmits the generated image whose use is permitted to user terminal 2 and provides it to the user. Information processing device 1 includes processing unit 11, storage unit 12, and communication unit 13. Note that information processing device 1 may have its functions realized by multiple server devices or computers, or may be a device corresponding to a node on a blockchain. Information processing device 1 may also execute some or all of the processes executed by user terminal 2.
[0012] The processing unit 11 of the information processing device 1 is composed of a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), an NPU (Neural Processing Unit), or a TPU (Tensor Processing Unit), and performs various control processes, arithmetic processes, etc. by reading and executing a program P (program product) etc. that is pre-stored in the memory unit 12.
[0013] The storage unit 12 of the information processing device 1 is, for example, a volatile memory and a non-volatile memory. The storage unit 12 stores a program P, a prompt extraction model M1, an edge detection module M2, an image generation model M3, a similarity calculation model M4, a work table 121, a similarity table 122, and a character table 123. The image generation model M3 generates content. The prompt extraction model M1 and the edge detection module M2 may be used to generate content. The similarity calculation model M4 determines whether or not to permit use of the first content based on the similarity between the first content and the second content. The prompt extraction model M1 and the edge detection module M2 may be used to generate content. The program P may be provided to the information processing device 1 using a computer-readable storage medium 12a. The storage medium 12a may be, for example, a portable memory. Examples of the portable memory include a CD-ROM, a USB (Universal Serial Bus) memory, an SD card, a microSD card, and a Compact Flash Memory (registered trademark). If the recording medium 12a is a portable memory, the processing element of the processing unit 11 may read the program P from the recording medium 12a using a reading device (not shown). The read program P is written to the storage unit 12. Furthermore, the program P may be provided to the information processing device 1 by the communication unit 13 communicating with an external device. Details of the prompt extraction model M1, edge detection module M2, image generation model M3, similarity calculation model M4, work table 121, and similarity table 122 will be described later. Details of the character table 123 will be described in embodiment 2. Note that the image generation model M3 may be stored in another server device and executed on that server device.
[0014] The communication unit 13 of the information processing device 1 is a communication module or communication interface for communicating with other devices such as the user terminal 2 via wired or wireless means, and is, for example, a wide-area wireless communication module such as LTE (registered trademark), 4G, or 5G. The processing unit 11 communicates with the user terminal 2 via the communication unit 13 through an external network N such as the Internet.
[0015] 3 is a block diagram showing an example configuration of the user terminal 2. The user terminal 2 includes a device processing unit 21, a storage unit 22, a communication unit 23, a display unit 24, an input unit 25, and an image capturing unit 26. The functions of the user terminal 2 may be realized by a plurality of devices. For example, arithmetic processing may be performed by a smartphone, and input reception or screen display may be performed by wearable glasses. The user terminal 2 may also perform some or all of the processing performed by the information processing device 1. The device processing unit 21 is configured with a CPU, MPU, GPU, NPU, TPU, or the like, and performs various control processing, arithmetic processing, and the like.
[0016] The storage unit 22 of the user terminal 2 stores an application program Pa that captures a user image, accepts a selection from the user of a work to which the style of the user image is to be adapted, and displays the generated image received from the information processing device 1. The application program Pa is provided to the user terminal 2 using, for example, a storage medium 22a. Note that the device processing unit 21 of the user terminal 2 may obtain the application program Pa using the Internet and store it in the storage unit 22.
[0017] The communication unit 23 of the user terminal 2 is a communication module or a communication interface for wirelessly communicating with the information processing device 1. The terminal processing unit 21 communicates with the information processing device 1 via the communication unit 23 and an external network N.
[0018] The display unit 24 of the user terminal 2 displays the generated image received from the information processing device 1. The input unit 25 of the information processing device 1 receives input from the user, such as a selection of a work to which the style of the user image is to be adapted. In this embodiment, the user terminal 2 is a smartphone, and the display unit 24 and the input unit 25 are integrated into a touch panel.
[0019] The photographing unit 26 of the user terminal 2 photographs a user image with the user as the photographing subject. In this embodiment, the user terminal 2 is a smartphone, and the photographing unit 26 is, for example, a camera built into the smartphone. The device processing unit 21 of the user terminal 2 transmits the user image photographed by the photographing unit 26 to the information processing device 1. The device processing unit 21 may also transmit a user image acquired from another photographing device or a user image pre-stored in the storage unit 22 to the information processing device 1.
[0020] FIG. 4 is an explanatory diagram showing an example of image conversion specification information (prompt) settings. The image conversion specification information (prompt) is information specifying the conditions for a converted image when converting a user image, and includes, for example, the gender, age, hair color, eye color, accessories, and clothing of the converted image. The image conversion specification information (prompt) uses features extracted from the user image by the prompt extraction model M1. Note that the image conversion specification information (prompt) may be set by a user selecting and instructing on an operation terminal screen. The prompt extraction model M1 may be configured, for example, by a model that detects objects contained in an input image and outputs the class of the detected object, such as a convolutional neural network (CNN), a You Only Look at Once (YOLO), a vision transformer, or a single-shot multibox detector (SSD); a large language model (LLM) that verbalizes the features of an input image, such as a generative pre-trained transformer (GPT); or an image classification model such as contrastive language-image pre-training (CLIP). When the prompt extraction model M1 is configured using a neural network including a CNN, the input layer of the prompt extraction model M1 has multiple neurons that receive input pixel values of the user image and passes the input pixel values to the middle layer. The middle layer has multiple neurons that extract image features of the user image and passes the extracted image features to the output layer. The output layer outputs the user's class (feature) contained in the user image based on the image features. In the example shown in FIG. 4, the user classes (features) include "male," "20s," "black hair," "black eyes," "glasses," and "shirt." Note that the classes output by the prompt extraction model M1 are not limited to these, and other classes may also be output, such as facial expression, gaze, or background. Furthermore, the classes output by the prompt extraction model M1 are not limited to words, and may be sentences such as "he wears glasses and has black hair that is long enough so that his bangs do not cover his glasses."
[0021] FIG. 5 is an explanatory diagram showing an example of the edge detection module M2. The edge detection module M2 is, for example, a Canny edge detector using a Sobel filter. Note that the edge detection module M2 may also be an edge detector using a Gaussian filter or a Laplacian filter. The edge detection module M2 has a function of detecting the boundaries (edges) between bright and dark parts of an input image based on the brightness of each pixel in the input image. As a result, when a user image is input, the edge detection module M2 detects the user's characteristic features (hair, eyes, nose, mouth, etc.) and contours, and outputs the positions of the characteristic features and contours (edge information). In the output image shown in FIG. 5, the positions of the characteristic features and contours are indicated by dashed lines.
[0022] FIG. 6 is an explanatory diagram showing an example of a training method for the image generation model M3. The image generation model M3 may be configured with either or both of an image-to-image model including stable diffusion and a text-to-image model. The image generation model M3 may be a diffusion model such as DALL-E3 or Imagene, a CNN, a variational auto-encoder (VAE), a generative adversarial network (GAN), or a style GAN. The image generation model M3 is trained to output an image generated based on a prompt and edge information indicating the characteristics of the image to be generated. That is, the input data for the image generation model M3 includes the prompt and edge information. It may also include facial images, facial feature information, weights, and training data. The image generation model M3 is trained to output an image that matches the artistic style of each work by fine-tuning (additional learning) for each work based on character images (first content) of multiple characters appearing in each work. When the image generation model M3 is configured using stable diffusion, the image generation model M3 is trained (pre-trained) using images obtained from a copyright holder or publicly available images. The processing unit 11 of the information processing device 1 fine-tunes (additionally trains) the pre-trained image generation model M3 for each work using character images of multiple characters included in each work as training data so as to output images that are closer to the artistic style of each work. Note that the image generation model M3 may be fine-tuned in a server device external to the information processing device 1. Furthermore, content that may be included in the training data for the image generation model M3 according to this embodiment is not limited to character images related to character images appearing in the work, but may also be an image related to a scene in the work (background image) or an image related to a scene in the work in which a character appears (background image including a character image).
[0023] The image generation model M3 is fine-tuned, for example, by Low-Rank Adaptation (LORA). Specifically, by adding a low-rank decomposition matrix based on character images and captions (character image features) extracted from the character images to each layer of the image generation model M3, the image generation model M3 is fine-tuned to output images that reflect the artistic style of the work. That is, the image generation model M3 is fine-tuned to learn the artistic style of the work based on character images of multiple characters. Fine-tuning (additional learning) of the image generation model M3 is performed using a dataset of character images of multiple characters for each work, and multiple image generation models M3 (M3A, M3B, M3C, etc.) that have learned the artistic style of each work are generated. The image generation model M3 may also be trained using, for example, Dream Booth, Textual Inversion Prefix Tuning, or Prompt Tuning. Furthermore, the dataset according to this example may include multiple character images of the same character.
[0024] In the example shown in FIG. 6, an image generation model M3A is fine-tuned using character images of characters appearing in work AAA, an image generation model M3B is fine-tuned using character images of characters appearing in work BBB, and an image generation model M3C is fine-tuned using character images of characters appearing in work CCC. The training data dataset for fine-tuning the image generation model M3 includes multiple character images and captions extracted from each character image. The image generation model M3 is fine-tuned by inputting a dataset related to, for example, 20 character images per work. Note that the number of datasets input to the image generation model M3 may be 19 or less or 21 or more. The storage unit 12 of the information processing device 1 stores the fine-tuned (additionally learned) image generation models M3 (M3A, M3B, M3C, etc.) related to each work.
[0025] FIG. 7 is an explanatory diagram showing an example of the image generation model M3. As described above, the image generation model M3 is configured by either or both of an image-to-image model including stable diffusion and a text-to-image model. When a user image, a prompt, and edge information are input, the image generation model M3 outputs a generated image based on a character image used for fine tuning. The processing unit 11 of the information processing device 1 reads out the image generation model M3 (one of M3A, M3B, M3C, etc.) related to the work selected by the user (see FIG. 11), and inputs the user image, the prompt (class) extracted from the user image by the prompt extraction model M1, and the edge information extracted by the edge detection module M2 into the read image generation model M3. Note that all or part of the prompt input to the image generation model M3 may be instructions input by the user. The prompt input to the image generation model M3 may also include a designation of the use of the generated image or a designation of the number of types of generated images output by the image generation model M3. When the image generation model M3 is a stable diffusion model, the encoder of the VAE (Variational AutoEncoder) of the image generation model M3 converts the input user image into a latent representation and passes it to the diffusion model. The text encoder of the image generation model M3 vector-converts the input prompt and passes it to the diffusion model. The diffusion model generates a pixel-by-pixel probability map of the generated image based on the latent representation of the user image, the vector of the acquired prompt, and the input edge information, and passes it to the VAE decoder. The VAE decoder constructs and outputs the generated image based on the received probability map. In other words, the user image is directly used as input data for the image generation model M3 to output the generated image. The image generation model M3 (Stable Diffusion) may be provided with additional functions (e.g., a control net), and the processing unit 11 may input control information related to image generation, such as pose or composition designation, to the image generation model M3.
[0026] In this embodiment, the image generation model M3 outputs multiple generated images with different degrees of artistic quality. The artistic quality indicates the degree to which the generated image reflects the artistic style of the artwork, and varies depending on the degree of reflection (weight) of edge information when generating the generated image. The weight varies, for example, within a range of 0 to 1.0, with a higher value indicating a higher degree of reflection of edge information and a lower value indicating a lower degree of reflection of edge information. The artistic quality is, for example, a value obtained by subtracting the degree of reflection of edge information from 1.0. In other words, the higher the degree of reflection of edge information, the lower the artistic quality of the generated image, and the lower the degree of reflection of edge information, the higher the artistic quality of the generated image. In the example shown in FIG. 7, the image generation model M3 outputs four levels of generated images with edge information weights of 0.9, 0.7, 0.3, and 0.1. The processing unit 11 of the information processing device 1 may input information specifying the weight of edge information into the prompt and output a generated image in which the edge information is reflected based on the specified weight. Furthermore, the processing unit 11 may output multiple types of generated images with different degrees of quality by varying the applicability of the model learned by the LoRA when generating generated images using the image generation model M3. Furthermore, the image generation model M3 may generate multiple different images by changing the seed value in image generation.
[0027] The data input to the image generation model M3 may include at least one of a user image, a prompt, or edge information. For example, if the data input to the image generation model M3 includes a prompt and edge information, information extracted from the user image (prompt and edge information) is input to the image generation model M3. That is, the user image is indirectly used as input data for the image generation model M3 to output a generated image. The data input to the image generation model M3 may also include a character image of a character selected by the user (see FIG. 13). If the data input to the image generation model M3 includes a user image and a character image, the prompt input to the image generation model M3 may include an instruction to output a generated image by making the user image resemble the style of the character image. In this case, the character image is used as input data for the image generation model M3 to output a generated image. The data input to the image generation model M3 includes at least one of a user image, a character image, edge information, and a prompt. A user image of the user's entire body may also be input to the image generation model M3. In this case, image generation model M3 may output a generated image of the whole body. Alternatively, image generation model M3 may be stored in a server device different from information processing device 1, and may be trained or output generated images on that server device.
[0028] FIG. 8 is an explanatory diagram showing an example of a similarity calculation model M4. The similarity calculation model M4 is a learning model capable of extracting image features, such as a CNN or a Vision Transformer. One of the generated images generated by the image generation model M3 and one of the character images used in training the image generation model M3 that output the generated image are input to the similarity calculation model M4. When the similarity calculation model M4 is configured using a neural network including a CNN, the input layer of the similarity calculation model M4 has multiple neurons that receive input pixel values of the generated image and the character image, and passes the input pixel values to the intermediate layer. The intermediate layer has multiple neurons that extract image features of the generated image and the character image, and passes the extracted image features to the output layer. The output layer calculates and outputs the similarity between the input generated image and the character image based on the image features of the generated image and the character image. The similarity is represented by a value between 0 and 1, with a higher value indicating a greater similarity between the two images.
[0029] The processing unit 11 of the information processing device 1 inputs each of the multiple generated images, each with a different degree of reflection of edge information, into a similarity calculation model M4 along with all of the character images used in training the image generation model M3, and calculates the similarity between each of the generated images and the character images. The processing unit 11 may also extract feature points or feature quantities from the images and output the similarity between the generated images and the character images using a processing module that performs pattern matching based on the feature points or feature quantities extracted from the multiple images. The similarity calculation model M4 may also have a segmentation function that extracts the positions of people (characters or users) in the images as segments, and output the similarity between the extracted segments in the generated images and the character images.
[0030] 9 is an explanatory diagram showing an example of the work table 121. The work table 121 records the titles of works such as manga, anime, movies, or games in which characters of character images that are the subject of learning for the image generation model M3 appear. The management items (fields) of the work table 121 include a work ID field, a work name field, a work logo field, an image generation model field, and an allowable similarity field.
[0031] The work ID field stores an ID assigned to the work. The work name field stores the work name. The work logo field stores the work's logo, for example, in file format. The image generation model field stores the type of image generation model M3 (one of M3A, M3B, M3C, etc.) that has been fine-tuned (additionally trained) using character images of multiple characters appearing in the corresponding work. The allowable similarity field stores a threshold value for permitting use of the similarity of a generated image to a character image of the work, specified by, for example, the author of the work or a person who holds the rights to the work, such as the copyright holder. In other words, the threshold value corresponds to a criterion set as a permission condition for determining whether or not to permit use of the first content. Note that the work table 121 may store multiple different threshold values depending on conditions such as whether the user's use of the system is paid or free, or whether the user is a member of the system.
[0032] 10 is an explanatory diagram showing an example of the similarity table 122. The similarity table 122 stores a generated image (first content) generated by the image generation model M3 and the similarity of the generated image to a character image (second content) included in the training data of the image generation model M3 in association with each other. The management items of the similarity table 122 include a user ID field, a work ID field, an edge information weight field, multiple character image similarity fields, and a permission information field.
[0033] The user ID field of the similarity table 122 stores a user ID assigned to the user who owns the user terminal 2. The user ID stored in the user ID field is the user ID of the user associated with the user image whose prompt and edge information were input to the image generation model M3 when the generated image was generated. The work ID field stores the work ID of the work of the character image included in the training data of the image generation model M3 that generated the generated image. The edge information weight field stores the weight (reflection degree) of the edge information when the generated image was generated. The generated image is identified by the user ID, work ID, and edge information weight. The similarity table 122 also stores information about additional generated images (see FIG. 13) in the same way as the generated image (for example, the fifth record from the top of the similarity table 122 shown in FIG. 10).
[0034] The multiple character image similarity fields (character image (1) similarity, character image (2) similarity, character image (3) similarity, etc.) in similarity table 122 store the similarity of the generated image to each character image output by similarity calculation model M4. The character images of the characters appearing in each work (character images included in the dataset of training data for each image generation model M3) are assigned character numbers ((1), (2), (3), etc.) in each work. In other words, character images are identified by the work ID and character number. For example, the character image (1) similarity field of the top record stores the similarity of the generated image to the character image with character number (1) of work ID: WT001. The permission information field stores whether permission information has been assigned to the generated image of the corresponding record. The processing unit 11 of the information processing device 1 assigns permission information to a generated image when the similarity to all character images used in training the image generation model M3 that output the generated image of the corresponding record is less than the threshold of allowable similarity for the work in which the character image appears (see FIG. 9). When permission information is assigned to a generated image, the generated image becomes an unrestricted image, and "permitted" is stored in the permission information field. When the similarity to at least one character image used in training the image generation model M3 that output the generated image of the corresponding record is equal to or greater than the threshold of allowable similarity for the work in which the character image appears (see FIG. 9), the generated image becomes a restricted image, and "not permitted" is stored in the permission information field. That is, the processing unit 11 reads a threshold corresponding to the conditions of the selected work from the work table 121, from among multiple allowable similarity thresholds corresponding to the conditions of the work, and restricts the use of the generated image (first content) when the similarity is equal to or greater than the read threshold.
[0035] FIG. 11 is an explanatory diagram showing an example of a work selection screen. For example, when the device processing unit 21 of the user terminal 2 receives an input from the user instructing it to start generating a generated image, it displays the work selection screen on the display unit 24. The work selection screen displays a search field, multiple work display fields, and a start shooting button. The input unit 25 receives an input of a work name to search for a work from the user in the search field displayed on the display unit 24. The device processing unit 21 transmits the content entered in the search field to the information processing device 1. Based on the input content in the search field received from the user terminal 2, the processing unit 11 of the information processing device 1 identifies a work whose work name is highly similar to the input content from among the works stored in the work table 121. The processing unit 11 of the information processing device 1 transmits the work name and logo mark of the identified work to the user terminal 2.
[0036] The work display field displays the work name or logo mark of the work identified by the processing unit 11 of the information processing device 1 from among the works stored in the work table 121 (see FIG. 7) based on the input content in the search field. The work display field also displays a check box that allows the user to select the work name. Note that the work display field may also display, for example, character images of characters appearing in the work. In the example shown in FIG. 11, three works are identified based on the input content in the search field, and three work display fields are displayed on the work selection screen, but the number of works identified based on the input content in the search field and displayed in the work display field may be two or less or four or more.
[0037] When a check box is pressed in any of the work display fields among the multiple work display fields displayed on the work selection screen, the start shooting button becomes pressable. When the start shooting button is pressed, the device processing unit 21 of the user terminal 2 starts the shooting unit 26 and displays a user image shooting screen (see FIG. 12) on the display unit 24.
[0038] FIG. 12 is an explanatory diagram showing an example of a user image capture screen. When the capture start button is pressed on the work selection screen, the device processing unit 21 of the user terminal 2 displays the user image capture screen on the display unit 24. A captured image field is displayed on the user image capture screen. The captured image field displays images being captured by the capture unit 26 in real time. The device processing unit 21 determines, for example, an image captured by the capture unit 26 when a predetermined time (e.g., 10 seconds) has elapsed since the user image capture screen was displayed on the display unit 24 as the user image. Note that a capture button may be displayed on the user image capture screen, and an image captured when a predetermined time has elapsed since the capture button was pressed may be determined as the user image. Alternatively, the device processing unit 21 may display, on the display unit 24, multiple images captured before the predetermined time has elapsed, and cause the input unit 25 to accept input for selecting an image to be used as the user image from the multiple images. When a user image is captured, the device processing unit 21 transmits, to the information processing device 1, a work ID associated with the work name selected on the work display screen (see FIG. 11 ) and the captured user image. The device processing unit 21 may read out the user image stored in the storage unit 22 of the user terminal 2 and transmit it to the information processing device 1.
[0039] 13 is an explanatory diagram showing an example of a work quality selection screen. The processing unit 11 of the information processing device 1 receives the work ID and user image of the selected work from the user terminal 2. The processing unit 11 reads the image generation model M3 related to the selected work based on the work table 121, and inputs the captured user image, the prompt extracted from the user image, and the edge information extracted from the user image into the read image generation model M3. The processing unit 11 transmits the generated image output by the image generation model M3 to the user terminal 2. The device processing unit 21 of the user terminal 2 displays the generated image received from the information processing device 1 on the display unit 24.
[0040] The artwork level selection screen displays the captured user image, multiple types of generated images with different artwork levels output by the image generation model M3, and multiple function buttons corresponding to how the generated images will be used. In the example shown in FIG. 13, each generated image is displayed with an artwork level (a value obtained by subtracting the weight of the edge information from 1.0). Note that each generated image may also be displayed with the weight of the edge information itself. The display arrangement of the user image and multiple generated images on the artwork level selection screen is not limited to that shown in FIG. 13. The device processing unit 21 of the user terminal 2 may display the generated images such that, for example, generated images with lower artwork levels are positioned closer to the user image.
[0041] The device processing unit 21 of the user terminal 2 displays generated images (non-restricted images) to which permission information has been assigned, among the generated images generated by the image generation model M3, with a checkbox attached for accepting selection of the generated image to be used by the user. Furthermore, the device processing unit 21 displays generated images (restricted images) to which permission information has not been assigned, among the generated images generated by the image generation model M3, with, for example, a diagonal line attached. That is, output of generated images whose similarity to the character image is equal to or greater than a threshold value to the display unit 24 is restricted. Restricted images are not provided with a checkbox for accepting selection from the user, and the user cannot select and use the restricted images. That is, use of generated images whose similarity to the character image is equal to or greater than a threshold value is restricted. Generated images whose similarity to the character image is equal to or greater than a threshold value may be displayed in a manner that restricts some of their usage, for example, by disallowing output to another system but allowing printing. Furthermore, the device processing unit 21 may not display generated images (restricted images) whose similarity to the character image is equal to or greater than a threshold value on the artwork level selection screen.
[0042] If there are fewer than a predetermined number (e.g., three) of generated images (non-restricted images) among the generated images generated by the image generation model M3 whose similarity to all character images used in training the image generation model M3 is less than a threshold, the processing unit 11 of the information processing device 1 outputs, using the image generation model M3, a generated image (additional generated image) that reflects the edge information of the user image to a higher degree than generated images (restricted images) whose similarity to the character images is greater than or equal to the threshold.
[0043] In the example shown in FIG. 13, generated images with a degree of art work of 0.9 (degree of edge information reflection: 0.1) and a degree of art work of 0.7 (degree of edge information reflection: 0.3) have similarities to the character image equal to or greater than the threshold, and there are two generated images whose similarities to all character images used in training the image generation model M3 are less than the threshold. Therefore, the processing unit 11 generates, for example, one additional generated image using the image generation model M3. The degree of art work of the user image for the additional generated image is lower than the degree of art work of the generated image with the lowest degree of art work among the generated images whose similarities to the character images are equal to or greater than the threshold. In other words, the degree of reflection of the edge information of the user image for the additional generated image is higher than the degree of reflection of the generated image with the highest degree of edge information reflection among the generated images whose similarities to the character images are equal to or greater than the threshold. In the example shown in FIG. 13, the processing unit 11 causes the image generation model M3 to generate an additional generated image with a degree of edge information reflection: 0.6 (degree of art work: 0.4), and displays the generated additional generated image on the art work selection screen. If the similarity of the additional generated image with all the character images used in training the image generation model M3 is less than a threshold, permission information is assigned to the additional generated image (see FIG. 10 ), and the additional generated image becomes a non-restricted image and is displayed on the artistic quality selection screen. If the similarity between the additional generated image and at least one of the character images used in training the image generation model M3 is equal to or greater than a threshold, the additional generated image may be displayed as a restricted image on the artistic quality selection screen, or it may not be displayed on the artistic quality selection screen. In this case, the processing unit 11 may further increase the degree of reflection of edge information (reduce the artistic quality), or, if the similarity is dissimilarity, may decrease the degree of reflection of edge information (increase the artistic quality) to cause the image generation model M3 to generate the additional generated image. If the similarity of the additional generated image with all the character images used in training the image generation model M3 is less than a threshold, the additional generated image becomes a non-restricted image and is displayed with a checkbox attached for accepting the user's selection of the generated image to be used.
[0044] When one of the checkboxes for the generated image on the artwork quality selection screen is checked, one of the multiple function buttons can be pressed. The multiple function buttons include, for example, a print button, an output button, and an upload button. When the print button is pressed, the user terminal 2 transmits the generated image to a printer, for example, via communication, and causes the printer to print the generated image. When the output button is pressed, the user terminal 2 outputs the generated image to, for example, another terminal device, a projector (projection device), or a display (display device). When the upload button is pressed, the user terminal 2 outputs (uploads) the generated image to another system via the network N. The other system may, for example, be a web page where the generated image is shared by multiple users, or a metaverse space where avatars based on the generated image appear. The generated image is used by pressing these function buttons.
[0045] The processing unit 11 of the information processing device 1 may generate a corrected generated image based on the setting values of the setting items selected by the user. The processing unit 11 may also generate a body combination image by combining a body image with a generated image selected on the user terminal 2. For example, the body image may be an image below the neck trimmed from a character image used for training the image generation model M3. The processing unit 11 may also generate a background combination image by combining the body combination image with a background image (frame image). The processing unit 11 may also generate a user avatar by inputting the body combination image into an avatar generation model capable of outputting a three-dimensional object when a two-dimensional planar image is input, such as NeRF (Neural Radiance Fields) or 3D Gaussian Splatting. The generated avatar is displayed, for example, in a metaverse space (virtual space). The generation of the corrected generated image, body combination image, background combination image, or avatar may be executed by the device processing unit 21 of the user terminal 2 using the function of the application program Pa.
[0046] 14 is a flowchart showing an example of processing by the processing unit 11 of the information processing device 1. The processing unit 11 of the information processing device 1 receives (acquires) a work ID of a selected work title and a user image from the user terminal 2 (S1). The processing unit 11 inputs the received user image into a prompt extraction model M1 (S2) and outputs a prompt (S3). The processing unit 11 also inputs the user image into an edge detection module M2 (S4) and outputs edge information (S5). The processing unit 11 reads out an image generation model M3 related to the work selected by the user from the work table 121 based on the received work ID (S6). The processing unit 11 inputs the user image, the prompt, and the edge information into the image generation model M3 (S7) and outputs multiple generated images with different degrees of quality (S8).
[0047] The processing unit 11 of the information processing device 1 inputs all combinations of each generated image output in S8 and each character image used in training the image generation model M3 into the similarity calculation model M4 (S9), and outputs the similarity for each combination of the generated image and the character image (S10). The processing unit 11 stores the similarity for each combination of the generated image and the character image in the similarity table 122, in association with the work ID, the user ID, and the weight of the edge information (S11). The processing unit 11 reads out the threshold for the allowable similarity for the work associated with the work ID, using the work ID associated with the similarity between the generated image and the character image as a key (S12). The processing unit 11 identifies, from among the generated images output in S8, generated images for which at least one similarity to each character image stored in the similarity table 122 is equal to or greater than the threshold read out in S12 as restricted images (S13). The similarity threshold may be a predetermined value regardless of the work, or may be different for each character image. The threshold value may also vary depending on conditions such as whether the user is using the system for a fee or free of charge, or whether the user is a member of the system. In this case, the storage unit 12 of the information processing device 1 may store, for example, a threshold value for the acceptable similarity for each work corresponding to a user ID, and the processing unit 11 may read out the threshold value. The storage unit 12 may also store a correction degree for correcting the acceptable similarity threshold for each work corresponding to a user ID, and the processing unit 11 may calculate the correction threshold value based on the threshold value read out from the work table 121 and the correction degree corresponding to the user ID, and identify restricted images based on the correction threshold value and the similarity of the generated images. The processing unit 11 assigns permission information to generated images that are not restricted images (non-restricted images) among the generated images (storing the permission information in the similarity table 122) (S14). The processing unit 11 determines whether the number of non-restricted images among the multiple generated images is equal to or greater than a predetermined number (S15). If the number of non-restricted images is greater than or equal to a predetermined number (S15: YES), the processing unit 11 transmits (outputs) the non-restricted images to which permission information has been assigned and the generated image including the restricted image to the user terminal 2 (S16), and terminates the processing.
[0048] If the number of non-restricted images is less than the predetermined number (S15: NO), the processing unit 11 of the information processing device 1 increases the degree of reflection of edge information on the generated image output by the image generation model M3 (S17), and inputs the user image, prompt, and edge information to the image generation model M3 (S18). The processing unit 11 outputs the generated image (S19) and returns the process to S9. Note that in S9 after the process is returned, all combinations of the generated images output in S19 and the character images used in training the image generation model M3 are input to the similarity calculation model M4. Also, in S17 and S18, the processing unit 11 may reduce the applicability of the model trained by the LOA and output generated images by the image generation model M3.
[0049] According to the above configuration and processing, the processing unit 11 of the information processing device 1 can restrict the use of the first content if the generated image (first content) is similar to a copyrighted work. The information processing device 1 may be a photo booth or a print sticker machine such as a Purikura (registered trademark). In this case, the information processing device 1 may include a display unit, an input unit, and a photographing unit, and may execute the processing executed by the user terminal 2 in this embodiment. Some of the processing executed by the information processing device 1 may be executed by another device. For example, the generation of a generated image by the image generation model M3 may be executed by the user terminal 3 or another server device, and the calculation of similarity by the similarity calculation model M4 may be executed by the information processing device 1. Alternatively, the generation of a generated image by the image generation model M3 may be executed by the information processing device 1 or the user terminal 3, and the calculation of similarity by the similarity calculation model M4 may be executed by another server device. The content that can be included in the training data for the image generation model M3 according to this embodiment is not limited to character images of characters appearing in a work, but may also be an image (background image) of a scene in the work in which a character appears (background image including a character image). In this case, the generated image output by the image generation model M3 is not limited to a user image that matches the style of the work, but may be a background image that matches the style of the work, for example.
[0050] (Embodiment 2) The image generation model M3 according to the second embodiment is trained for each character using multiple character images of characters appearing in the work. The device processing unit 21 of the user device 2 accepts a character image selection from the user, identifies a character to which the user image's style is to be adapted, and transmits the identified character to the information processing device 1. The information processing device 1 outputs multiple types of generated images that are different from each other using the image generation model M3, which has been fine-tuned (additionally trained) based on the character image of the identified character. That is, the image generation model M3 has learned the style of the character image, and when a user image is input, it outputs a generated image in which the input user image is adapted to the style of the character image. Furthermore, the similarity calculation model M4 according to the second embodiment extracts character features based on multiple character images of one character and outputs the similarity between the character features and the user image features. The present invention according to the second embodiment will be described below with reference to the drawings. Among the components according to the second embodiment, components similar to those of the first embodiment are designated by the same reference numerals, and detailed description thereof will be omitted.
[0051] FIG. 15 is an explanatory diagram showing an example of a training method for the image generation model M3 according to the second embodiment. When the image generation model M3 is configured using stable diffusion, the image generation model M3 is trained (pre-trained) using images obtained from a copyright holder or publicly available images. The processing unit 11 of the information processing device 1 according to this embodiment trains the pre-trained image generation model M3 to output images that resemble the artistic style of each character image by fine-tuning (additional training) the pre-trained image generation model M3 for each character using character images of characters included in a work as training data. Note that the image generation model M3 may be fine-tuned in a server device external to the information processing device 1.
[0052] The image generation model M3 is fine-tuned, for example, by Low-Rank Adaptation (LORA). Specifically, by adding a low-rank decomposition matrix based on the character image and a caption (character image feature) extracted from the character image to each layer of the image generation model M3, the image generation model M3 is fine-tuned to output an image that resembles the style of the character image. In other words, the image generation model M3 is fine-tuned to learn the style of the character image. Fine-tuning (additional learning) of the image generation model M3 is performed using a dataset of multiple character images for each character, and multiple image generation models M3 (M3a, M3b, M3c, etc.) that have learned the style of each character are generated. Note that the image generation model M3 may also be trained by, for example, prefix tuning or prompt tuning.
[0053] In the example shown in FIG. 15, an image generation model M3a fine-tuned using a character image of character A, an image generation model M3b fine-tuned using a character image of character B, and an image generation model M3c fine-tuned using a character image of character C are shown. The training data dataset for fine-tuning the image generation model M3 includes one character image and a caption extracted from the character image. The image generation model M3 is fine-tuned by inputting a dataset related to, for example, 20 character images per character. Note that the number of datasets input to the image generation model M3 may be 19 or less or 21 or more. In the example shown in FIG. 15, one dataset per character image is illustrated, and other datasets are omitted. The storage unit 12 of the information processing device 1 stores the fine-tuned (additionally learned) image generation models M3 (M3a, M3b, M3c, etc.) related to each character.
[0054] FIG. 16 is an explanatory diagram showing an example of a similarity calculation model M4 according to the second embodiment. The similarity calculation model M4 according to the second embodiment receives as input one of the generated images generated by the image generation model M3 and multiple character images used in training the image generation model M3 that output the generated image. When the similarity calculation model M4 is configured using a neural network including a CNN, the input layer of the similarity calculation model M4 has multiple neurons that receive input of pixel values of the generated image and the character image, and passes the input pixel values to the intermediate layer. The intermediate layer has multiple neurons that extract image features of the generated image and the multiple character images, and passes the extracted image features to the output layer. The output layer calculates and outputs the similarity between the input generated image and the character based on the image features of the generated image and the multiple character images. The similarity is expressed as a value between 0 and 1, with a higher value indicating a greater similarity between the two images. That is, the similarity calculation model M4 according to this embodiment outputs the similarity between the generated image and the character features based on the features of the multiple character images. The feature amount of a character is, for example, an average value of the feature amounts of multiple character images of the character. Note that the information processing device 1 may input a character image of a character different from the character related to the character image used for training the image generation model M3 that output the generated image, and the generated image, and output the similarity between the generated image and a character that is not the target to which the style of the user image is to be applied.
[0055] 17 is an explanatory diagram showing an example of a work table 121 according to embodiment 2. The work table 121 according to embodiment 2 includes a character table ID field. The character table ID field stores an ID for identifying a character table 123 (see FIG. 18) that stores information about characters appearing in each work.
[0056] FIG. 18 is an explanatory diagram showing an example of the character table 123. The character table 123 records information about characters appearing in the work. The management items (fields) of the character table 123 include a character ID field, a character name field, a character image field, and an image generation model field. The character ID field stores an ID assigned to a character. The character name field stores the character name of the character. The character image field stores, for example, an image file of one of the character images used to fine-tune the image generation model M3 related to the corresponding character. The image generation model field stores the type of image generation model M3 (one of M3a, M3b, M3c, etc.) fine-tuned by the character image of the corresponding character.
[0057] In addition, the character table 123 is assigned attribute information for distinguishing between multiple character tables 123. The attribute information includes a character table ID stored in the work table 121 and a work name read from the work table 121 using the character table ID as a key.
[0058] FIG. 19 is an explanatory diagram showing an example of a character selection screen. The character selection screen displays character images of multiple characters appearing in the work selected on the work selection screen (see FIG. 11) and a generation start button. The device processing unit 21 of the user terminal 2 transmits the work ID of the work name selected on the work selection screen to the information processing device 1. The processing unit 11 of the information processing device 1 identifies the character table 123 for the selected work based on the work ID and work table 121 of the work selected on the work selection screen. The processing unit 11 transmits the character images stored in the identified character table 123, as well as the character names and character IDs of each character, to the user terminal 2. The device processing unit 21 of the user terminal 2 displays the received character images, as well as the character names and character IDs of each character, on the character selection screen. Each character image is displayed with a check box attached for accepting the user's selection of a character image to which the user image's style is to be influenced. When any one of the check boxes for the multiple character images is checked, the generation start button becomes pressable.
[0059] When the start generation button is pressed on the character selection screen, the device processing unit 21 of the user terminal 2 transmits the character ID of the selected character image to the information processing device 1, along with the work ID of the work name selected on the work selection screen (see FIG. 11) and the user image captured on the user image capture screen (see FIG. 12). The processing unit 11 of the information processing device 1 reads out an image generation model M3 related to the character of the selected character image based on the character ID and the character table 123, and inputs the prompt and edge information extracted from the captured user image to the read image generation model M3. When a generated image is output by the image generation model M3, the processing unit 11 transmits the generated image to the user terminal 2. The device processing unit 21 of the user terminal 2 displays a work level selection screen (see FIG. 20) on the display unit 24.
[0060] FIG. 20 is an explanatory diagram showing an example of a quality level selection screen according to the second embodiment. The quality level selection screen according to the second embodiment displays a captured user image, a character image selected on the character selection screen (see FIG. 13), multiple types of generated images with different quality levels output by the image generation model M3, and multiple function buttons. The display layout of the user image, character image, and multiple generated images on the quality level selection screen is not limited to that shown in FIG. 20. The processing unit 11 of the information processing device 1 may, for example, arrange and display each image in the order of user image, generated image, and character image. Furthermore, the processing unit 11 may display the generated images so that the higher the quality level of the generated image is located closer to the character image, and the lower the quality level of the generated image is located closer to the user image.
[0061] The processing unit 11 of the information processing device 1 identifies restricted images in which the similarity between each generated image output by the similarity calculation model M4 and a character related to a character image used in training the image generation model M3 is equal to or greater than the threshold of the allowable similarity for the work in which the character appears. Furthermore, if there are fewer than a predetermined number (e.g., three) of generated images (non-restricted images) in which the similarity to the character related to the character image used in training the image generation model M3 is less than the threshold, the processing unit 11 of the information processing device 1 outputs, by the image generation model M3, a generated image (additional generated image) in which the edge information of the user image is more highly reflected than in generated images (restricted images) in which the similarity to the character is equal to or greater than the threshold, and transmits the generated image to the user terminal 2. The device processing unit 21 of the user terminal 2 displays the non-restricted images and the additional generated images that are non-restricted images with check boxes attached to them for receiving the user's selection of the generated image to be used. Furthermore, the device processing unit 21 displays, for example, diagonal lines on generated images (restricted images) whose similarity to the characters related to the character images used in training the image generation model M3 is equal to or greater than a threshold value.
[0062] 21 is a flowchart showing an example of processing by the processing unit 11 of the information processing device 1 according to the second embodiment. The processing unit 11 of the information processing device 1 receives (acquires) a work ID, a character ID, and a user image of a selected work title from the user terminal 2 (S21). The processing unit 11 inputs the received user image into a prompt extraction model M1 (S22) and outputs a prompt (S23). The processing unit 11 also inputs the user image into an edge detection module M2 (S24) and outputs edge information (S25). The processing unit 11 reads out an image generation model M3 related to a character selected by the user from the character table 123 based on the received character ID (S26). The processing unit 11 inputs the user image, the prompt, and the edge information into the image generation model M3 (S27), and outputs a plurality of generated images with different degrees of creativity (S28).
[0063] The processing unit 11 of the information processing device 1 inputs each generated image output in S28 and the multiple character images used in training the image generation model M3 into the similarity calculation model M4 (S29), and outputs the similarity between the generated image and the character (S30). The processing unit 11 uses the work ID associated with the similarity between the generated image and the character image as a key to read out the threshold of the allowable similarity for the work associated with the work ID (S31). The processing unit 11 identifies, as restricted images, generated images whose similarity output by the similarity calculation model M4 in S30 is equal to or greater than the threshold read out in S31 (S32). The similarity threshold may be a predetermined value regardless of the work, or may be different for each character. The threshold may also be a different value depending on conditions such as whether the user's use of the system is paid or free, or whether the user is a member of the system. Furthermore, the processing unit 11 may output the similarity between the generated image and each character image using the similarity calculation model M4, and identify restricted images based on the similarity between the generated image and the character and the similarity between the generated image and each character image. The processing unit 11 assigns permission information to images that are not restricted images (non-restricted images) among the generated images (S33). The processing unit 11 determines whether the number of non-restricted images among the multiple generated images is equal to or greater than a predetermined number (S34). If the number of non-restricted images is equal to or greater than the predetermined number (S34: YES), the processing unit 11 transmits (outputs) the non-restricted images to which permission information has been assigned and the generated images including the restricted images to the user terminal 2 (S35).
[0064] If the number of non-restricted images is less than the predetermined number (S34: NO), the processing unit 11 of the information processing device 1 increases the degree of reflection of edge information on the generated images output by the image generation model M3 (S36) and inputs the prompt and edge information to the image generation model M3 (S37). The processing unit 11 outputs the generated images (S38) and returns the process to S29. Note that in S29 after the process is returned, each generated image output in S38 and multiple character images used in training the image generation model M3 are input to the similarity calculation model M4. Also, in S36 and S37, the processing unit 11 may reduce the applicability of the model trained by the LOA and output generated images by the image generation model M3.
[0065] (Variation) The device processing unit 21 of the user terminal 2 may store, in association with (associate) with (associate) the specified unrestricted image (generated image), generation information indicating that use of the unrestricted image is permitted (permission information), the work (original work) in which the character of the character image used in training the image generation model M3 that generated the unrestricted image appears, that the unrestricted image is a secondary use image (secondary use work), or that the unrestricted image was generated by the image generation model M3, and later display the generation information together with the unrestricted image (generated image). Furthermore, regardless of whether the generated image is a restricted or unrestricted image, if the generated image has been licensed for secondary use by the copyright holder of the original work, the device processing unit 21 may store, in association with (associate) with (associate) the generated image with generation information indicating that secondary use has been licensed (secondary use license information), and later display the generation information together with the generated image. Figure 22 is an explanatory diagram showing an example of a generated image usage screen. The generated image usage screen displays the generated image (unrestricted image) to which permission information has been assigned, along with the permission information, including the permission information, the title of the original work, information indicating that the unrestricted image is a secondary use image, information indicating that the unrestricted image was generated by the image generation model M3, and secondary use permission information. That is, the generated information is recorded as invisible information in the user terminal 2 and displayed as visible information. Similarly to the level of creation selection screen (see FIG. 13), the generated image usage screen also displays multiple function buttons, including a print button, an output button, and an upload button. The generated information may be recorded together with the generated image in a portable memory or the like by the information processing device 1 or the user terminal 2, read from the portable memory by another device, and displayed on that device. The generated information may also be transmitted together with the generated image to another device via the network N by the information processing device 1 or the user terminal 2. The generated image may be displayed with a recognition image indicating permission information, such as a secondary use permission mark or a C (Copyright) mark, or permission information or identification information may be added to image data (properties, etc.) indicating the generated image.
[0066] In each of the above-described embodiments, the processing unit 11 of the information processing device 1 identifies a character image or a generated image whose similarity to a character is less than a threshold as an unrestricted image, but this is not limited to this. For example, if the user is the copyright holder (including the copyright holder, author, neighboring rights holder, and those authorized by these holders) of a work related to a character image used in training the image generation model M3, the processing unit 11 may identify a character image or a generated image whose similarity to a character is greater than or equal to a threshold as an unrestricted image. In this case, the user who is the copyright holder of the work can select and use generated images similar to the style of the work from among the generated images output by the image generation model M3.
[0067] The generated image output by the image generation model M3 is not limited to one that resembles the artistic style of a work or character image of a user image. The image generation model M3 may output a generated painting, diagram, photograph, video, etc. (generated painting, etc.) based on, for example, an input prompt. In this case, the image generation model M3 may be trained using copyrighted paintings, diagrams, photographs, videos, etc. (paintings, etc.) as training data, and may output a painting, etc. that resembles the artistic style or flavor of the learned painting, etc. In this case, the generated painting, etc. output by the image generation model M3 corresponds to the first content, and the painting, etc. that serves as training data for the image generation model M3 corresponds to the second content. Note that when the image generation model M3 outputs a generated video, the image generation model M3 may be, for example, Phenaki or Emu Video. The image generation model M3 may also be a super-resolution, style transfer, inpainting, or other model that converts an input image and outputs a generated image. The processing unit 11 of the information processing device 1 may output the similarity between the generated painting or the like (first content) output by the image generation model M3 and the painting or the like (second content) that served as training data for the image generation model M3 using the similarity calculation model M4, and may determine whether or not to permit use of the generated painting or the like (first content) output by the image generation model M3 based on the output similarity. Furthermore, when the first content is a generated video and the second content is a video, the similarity calculation model M4 may output the similarity for each frame included in the video, or may output the similarity of the change between frames in a time series.
[0068] The processing unit 11 of the information processing device 1 may generate generated sentences or lyrics (generated sentences, etc.) using a sentence generation model that generates generated sentences or lyrics, etc. (generated sentences, etc.) based on an input prompt. The sentence generation model is configured, for example, with a Generative Pretrained Transformer (GPT), Bidirectional Encoder Representations from Transformers (BERT), Language Large Models Meta AI (LLaMa), Claude, or Pathways Language Model (PaLM), etc. In this case, the sentence generation model may be trained using sentences or lyrics (sentences, etc.) contained in literary works such as essays, novels, screenplays, poems, and lecture manuscripts as training data, and may output generated sentences, etc. that are in keeping with the style of the learned sentences, etc. In this case, the generated sentences, etc. output by the sentence generation model correspond to the first content, and the sentences, etc. that serve as training data for the sentence generation model correspond to the second content. The sentence generation model may be a model such as Whisper or SpeechGPT that outputs generated sentences based on input speech, a model such as CLIP (Contrastive Language-Image Pre-training) or DALL-E that outputs generated sentences based on input images, a model such as WaveNet or Tacotron that outputs speech of generated sentences based on input sentences, or a model such as Stable Diffusion or DALL-E that outputs program code based on input sentences. The processing unit 11 of the information processing device 1 may output a similarity between generated sentences (first content) output by the sentence generation model and sentences (second content) that served as training data for the sentence generation model using a sentence similarity calculation model capable of extracting features of sentences, and determine whether or not to permit use of the sentences (first content) output by the sentence generation model based on the output similarity. The sentence similarity calculation model may be configured, for example, by an RNN (Recurrent Neural Network) or a Transformer.
[0069] The processing unit 11 of the information processing device 1 may generate a generated song or music (generated song or the like) using a song generation model that generates a generated song or music (generated song or the like) based on an input prompt. The song generation model may be configured, for example, with Stable Audio, Jukebox, AudioLM, or MusicGen. In this case, the song generation model may be trained using copyrighted songs or music (song or the like) as training data, and may output a generated song or the like that is close to the style of the learned song or the like. In this case, the generated song or the like output by the song generation model corresponds to the first content, and the song or the like that serves as training data for the song generation model corresponds to the second content. Note that the song generation model may also be a model such as Soundify that outputs a generated song based on input video. The processing unit 11 of the information processing device 1 may use a music similarity calculation model capable of extracting feature quantities of music, etc. to output a similarity between the generated music, etc. (first content) output by the music generation model and the music, etc. (second content) that served as training data for the music generation model, and determine whether or not to permit use of the music, etc. (first content) output by the music generation model based on the output similarity. The music similarity calculation model is configured, for example, by a learning model capable of extracting feature quantities from time-series data such as an RNN. The music similarity calculation model extracts feature quantities of the music, etc. or the generated music, etc. based on melody spectrogram data or sheet music data of the music, etc. or the generated music, etc., and outputs the similarity by comparing the extracted feature quantities.
[0070] The processing unit 11 of the information processing device 1 determines that the first content is a restricted image when the similarity between the first content and the second content is equal to or greater than a threshold, and determines that the first content is an unrestricted image when the similarity between the first content and the second content is less than the threshold, but this is not limited to this. The processing unit 11 may determine that the first content is a restricted image when the similarity between the first content and the second content is greater than the threshold (exceeds the threshold), and may determine that the first content is an unrestricted image when the similarity between the first content and the second content is equal to or less than the threshold. The processing unit 11 may also determine whether the first content is a restricted image or an unrestricted image based on the dissimilarity indicating the degree of difference between the first content and the second content. In this case, the processing unit 11 may determine that the first content is a restricted image if the dissimilarity between the first content and the second content is less than a threshold, and may determine that the first content is an unrestricted image if the similarity between the first content and the second content is equal to or greater than the threshold, or may determine that the first content is a restricted image if the dissimilarity between the first content and the second content is equal to or less than the threshold, and may determine that the first content is an unrestricted image if the similarity between the first content and the second content is greater than the threshold (exceeding the threshold). In this case, the dissimilarity may be, for example, a value obtained by subtracting the similarity from 1 (for example, if the similarity is 0.7, the value of the dissimilarity is 0.3). Furthermore, the processing unit of the information processing device 1 may acquire the first content generated by a generative model in a device other than the information processing device 1.
[0071] The embodiments disclosed herein are illustrative in all respects and should not be considered limiting. The technical features described in each embodiment may be combined with one another, and the scope of the present invention is intended to include all modifications within the scope of the claims and equivalents thereto. Furthermore, independent and dependent claims described in the claims may be combined with one another in any and all combinations, regardless of the reference format. Furthermore, while the claims use a format in which a claim references two or more other claims (multiple claim format), this is not limiting. Multiple claims (multiple multiple claims) that reference at least one other multiple claim may also be used. [Explanation of symbols]
[0072] 1. Information processing equipment 11 Processing section 12 Storage section 121 Works Table 122 Similarity Table 123 Character Table 12a Recording Media 13 Communications Department 2. User terminal 21 Terminal processing section 22 Memory section 22a Storage medium 23 Communications Department 24 Display 25 Input section 26 Photography Department M1 Prompt Extraction Model M2 Edge Detection Module M3 image generation model M4 Similarity calculation model N Network P Program Pa App Program S Information Processing System
Claims
1. generating first content using a generative model that generates content; The permission or denial of use of the first content is determined according to the degree of similarity or dissimilarity between the first content and the second content, which is a copyrighted work. An information processing method that causes a computer to execute a process.
2. the second content is a copyrighted work whose use is permitted, A decision is made as to whether or not to permit use of the first content based on criteria set as conditions for permission. The information processing method according to claim 1 .
3. If the degree of similarity is equal to or greater than a threshold, or if the degree of dissimilarity is less than or equal to a threshold, the use of the first content is restricted. The information processing method according to claim 1 .
4. A threshold value corresponding to the acquired condition is read from a plurality of threshold values corresponding to the condition; If the degree of similarity is equal to or greater than a threshold, or if the degree of dissimilarity is less than or equal to a threshold, the use of the first content is restricted. The information processing method according to claim 3 .
5. When the degree of similarity is less than or equal to a threshold, the first content is displayed on a display unit, and when the degree of similarity is equal to or greater than a threshold, output of the first content to the display unit is restricted. or When the dissimilarity is equal to or greater than a threshold, the first content is displayed on a display unit, and when the dissimilarity is less than or equal to a threshold, output of the first content to the display unit is restricted.
4. The information processing method according to claim 1.
6. The generated first content and the similarity of the first content to the second content are stored in association with each other.
4. The information processing method according to claim 1.
7. Permission information is assigned to the first content for which the similarity is less than or equal to a threshold value, or the dissimilarity is greater than or equal to a threshold value.
4. The information processing method according to claim 1.
8. generating the first content by inputting a prompt to the generative model trained using the second content as training data; 4. The information processing method according to claim 1.
9. generating a plurality of first contents each having a different degree of reflection of input data on the generative model using the generative model; Among the generated plurality of first contents, outputting the first contents whose similarity is less than or equal to a threshold value or whose dissimilarity is equal to or greater than a threshold value.
4. The information processing method according to claim 1.
10. If the number of first contents whose similarity is less than or equal to a threshold or whose dissimilarity is greater than or equal to a threshold is less than or equal to a predetermined number among the generated plurality of first contents, the generation model is caused to generate the first contents whose reflection degree of the input data is increased or decreased. The information processing method according to claim 9.
11. the generative model is trained using the training data including a plurality of the second contents; outputting the first content whose similarity to all of the plurality of second contents included in the training data is less than or equal to a threshold, or whose dissimilarity is greater than or equal to a threshold; The information processing method according to claim 8.
12. Associating the first content with generation information including permission information, an original work, a fact that the first content is a secondary use work, a fact that the first content is generated by a generative model, or secondary use permission information; storing the corresponding first content and the corresponding creation information; 4. The information processing method according to claim 1.
13. Adding to the first content generation information including permission information, an original work, a statement that the first content is a secondary use product, a statement that the first content is generated by a generative model, or secondary use permission information; Displaying the first content together with the assigned generation information 4. The information processing method according to claim 1.
14. Obtaining first content generated by a generative model that generates content; The permission or denial of use of the first content is determined according to the degree of similarity or dissimilarity between the first content and the second content, which is a copyrighted work. An information processing method that causes a computer to execute a process.
15. Obtaining first content generated by a generative model that generates content; The permission or denial of use of the first content is determined according to the degree of similarity or dissimilarity between the first content and the second content, which is a copyrighted work. A program that causes a computer to perform a process.
16. generating first content using a generative model that generates content; The permission or denial of use of the first content is determined according to the degree of similarity or dissimilarity between the first content and the second content, which is a copyrighted work. Processing section An information processing device comprising:
Citation Information
Patent Citations
Program, information processing device and information processing method
JP2024021529A