Information processing apparatus, control method, medium, and program product
Patent Information
- Application Number
- CN202610195021.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-17
- Filing Date
- 2026-02-11
- Publication Date
- 2026-08-18
AI Technical Summary
特别地,在作为AI内容的图像包括大量的对象的情况下,有可能发生对具有相对小尺寸的对象的侵权的风险的评估的疏忽
[0008] According to another aspect of the present invention, a control method for an information processing apparatus includes: obtaining content generated by generative AI; identifying objects with infringement risks among objects included in the content; and recording the location information of the objects in the content as risk information in a storage unit.
Smart Images

Figure CN122596620A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to information processing apparatus, control methods, media, and program products, and particularly to technologies for supporting the assessment of risks of infringement of rights to content. Background Technology
[0002] With the proliferation of trained models (generative AI) that learn for data generation purposes, environments are being developed where a vast amount of diverse content (text, images, moving images, audio, 3D models, etc.) can be generated. On the other hand, content generated by generative AI (AI content) may include content that risks infringing on the rights of third parties (trademark rights, copyrights, portrait rights, etc.). In particular, some generative AI arbitrarily generates details not included in user-specified content (hints). Therefore, when using AI content, it is necessary to properly understand the risks of rights infringement.
[0003] Japanese Patent No. 7448271 (Patent Document 1) discloses a technique for obtaining relevant information (such as copyright license terms and portrait rights license terms) of image sets used for learning in image generation AI, and for providing this information as metadata to the AI content generated by the model. The Japanese Agency for Cultural Affairs' "Checklist & Guidance on AI and Copyright" (page 26) (Non-Patent Document 1) from July 2024 proposes using internet searches (text search and image search) as a method to confirm whether AI content is similar to existing works.
[0004] However, without access to relevant information about the image set used for generative AI learning, the method in Patent Document 1 cannot provide appropriate metadata for the AI content. As a result, the risk of infringement of rights by the AI content cannot be assessed.
[0005] When AI content includes a large number of objects, the method in Non-Patent Document 1 has the potential to overlook objects with a high risk of infringement. In particular, when an image as AI content includes a large number of objects, there is a possibility of neglecting to assess the risk of infringement for objects with relatively small sizes. Summary of the Invention
[0006] This disclosure provides techniques for supporting the assessment of the risk of infringement of rights to content.
[0007] According to one aspect of the present invention, an information processing apparatus includes: an acquisition unit that acquires content generated by generative artificial intelligence (AI); an identification unit that identifies objects among the objects in the content that pose a risk of infringement; and a recording unit that records location information of the objects in the content as risk information.
[0008] According to another aspect of the present invention, a control method for an information processing apparatus includes: obtaining content generated by generative AI; identifying objects with infringement risks among objects included in the content; and recording the location information of the objects in the content as risk information in a storage unit.
[0009] The features of this disclosure will become clear from the following description of embodiments with reference to the accompanying drawings. The following description of the embodiments is by way of example. Attached Figure Description
[0010] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the embodiments.
[0011] Figure 1 This is a diagram illustrating the hardware configuration of an information processing device.
[0012] Figure 2 It is a diagram depicting a usage scenario of an information processing device.
[0013] Figure 3 This is a diagram illustrating the functional configuration of an information processing device.
[0014] Figure 4 It is a flowchart used to calculate and record risk values.
[0015] Figure 5 This is a diagram showing the table used to manage category names and select objects.
[0016] Figure 6 This is a diagram showing the risk values of the objects being managed (modified Example 1-3).
[0017] Figure 7 This is a diagram depicting a usage scenario of the information processing device (second embodiment).
[0018] Figure 8 This is a diagram illustrating the functional configuration of the information processing device (second embodiment).
[0019] Figure 9 This is a flowchart for calculating and recording risk values (second embodiment).
[0020] Figure 10This is a diagram depicting a usage scenario of the information processing device (third embodiment).
[0021] Figure 11 This is a diagram illustrating the functional configuration of the information processing device (third embodiment).
[0022] Figure 12 This is a flowchart used to evaluate a trained model. Detailed Implementation
[0023] In the following description, embodiments will be presented in detail with reference to the accompanying drawings. Note that the following embodiments are not intended to limit the scope of the claims. Several features are described in the embodiments, but not all such features are required, and several such features may be appropriately combined. Furthermore, in the drawings, the same or similar configurations are given the same reference numerals, and redundant descriptions thereof are omitted.
[0024] First Embodiment As a first embodiment of the information processing apparatus according to this disclosure, the following will describe, as an example, an information processing apparatus for assessing the risk of infringement of images generated using image generation AI.
[0025] summary In this embodiment, when assessing the infringement risk of an image generated using image generation AI, image regions with a high risk of infringement are presented to the user. Specifically, objects included in the image are extracted, a risk value for infringement is calculated for each object, and the location information and risk value in the image are recorded as image metadata. The risk value as metadata is not essential. For example, the information may relate to the presence or absence of risk. When the image is displayed, the metadata content is also displayed. For example, the regions recorded in the metadata are overlaid on the image.
[0026] Figure 2 This diagram describes the usage scenario of the information processing device. The calculation of location information and risk values, as well as their recording in metadata, are performed by the information processing device 201. Figure 2 In this process, location information and risk values are calculated for objects 203 to 205 included in image 202. The calculated location information and risk values are recorded as metadata 206 in the image.
[0027] Users can check the content recorded in the metadata 206 at any time on the display screen 207 displayed via the display unit 107. For example, a user can check it when considering the commercial use of image 202. By checking the content recorded in the metadata 206, users can easily identify the location of objects at risk of infringement and determine the level of infringement risk. This can reduce the burden associated with confirming the risk of infringement of image 202 for the user.
[0028] the term Image generation AI is an example of an image processing model that artificially generates new images based on specific input data (also known as prompts) such as text and images.
[0029] Objects at risk of infringement include, for example, objects included in an image, such as characters (like people or animals), buildings, symbols, or logos. There are no particular restrictions on the occupancy (area ratio, coverage) and number of these objects in the image. That is, an object can be drawn in a very small area that is part of the image, or it can occupy a large portion of the image. An image may include only one object or multiple objects.
[0030] The location information indicates the area in the image where the object is drawn. The shape of the area is not particularly restricted, as long as it includes the drawing area of the object. For example, the shape can be a rectangle or a shape that follows the outline of the object.
[0031] Risk value is an arbitrary indicator of the degree of likelihood that an object is subject to infringement of the rights (trademark rights, copyrights, portrait rights, etc.) of existing works protected by copyright.
[0032] Metadata is supplementary information about the image, and in this embodiment, it includes the aforementioned location information and risk value. Metadata can be embedded, for example, as an Exif (Exchangeable Image File Format) header of the image file. Note that the specific method of recording the metadata is not particularly limited, as long as the user can check the recorded content at any time.
[0033] Device configuration Figure 1 This diagram illustrates the hardware configuration of the information processing apparatus. CPU 101 controls various devices connected to bus 102 and performs information processing as disclosed herein. CPU is an abbreviation for Central Processing Unit. ROM 103 stores the Basic Input / Output (BIOS) program and the boot program. ROM is an abbreviation for Read-Only Memory. RAM 104 is used as the main storage device for CPU 101. RAM is an abbreviation for Random Access Memory.
[0034] External storage device 105 is a storage device such as an HDD or SSD that stores programs to be processed by information processing device 201. HDD is an abbreviation for hard disk drive, and SSD is an abbreviation for solid state drive.
[0035] Input unit 106 is a keyboard or mouse, and is hardware that receives user input. Display unit 107 is hardware that outputs the calculation results of information processing device 201 to a display device according to instructions from CPU 101. Note that the display device can be of any type, such as a liquid crystal display, projector, or LED indicator. LED is an abbreviation for light-emitting diode. I / O 108 is an interface for communication connection with a management server (not shown) storing a given set of protected data. I / O is an abbreviation for input / output. Note that the communication connection can be wired or wireless. Bus 102 is a bus that connects the above units in a communicable manner.
[0036] Figure 3 This is a diagram illustrating the functional configuration of the information processing device. Note that details of the processing operations for each functional unit will be provided later. Figure 4 The information processing device 301 includes a first acquisition unit 302, an identification unit 303, a first feature calculation unit 304, a second acquisition unit 305, a second feature calculation unit 306, a risk value calculation unit 307, and a recording unit 308.
[0037] Note that it is assumed that each functional unit included in the information processing device 301 is implemented by software (by executing a program via the CPU 101), but some or all of them can be implemented by hardware such as an application-specific integrated circuit (ASIC). Furthermore, it can be implemented by a single information processing device, or by the collaboration of multiple information processing devices.
[0038] The first obtaining unit 302 obtains content generated by generative AI (such as text, images, and audio). In this embodiment, it obtains an image 202 generated using image generation AI.
[0039] The identification unit 303 identifies one or more objects (objects at risk of infringement) included in the content obtained by the first acquisition unit 302. In this embodiment, objects 203 to 205 are identified as objects included in the image 202, and the position information (coordinates) of each of the objects in the image are calculated.
[0040] The first feature calculation unit 304 calculates the feature representation of the object identified by the recognition unit 303. In this embodiment, the feature representation of each of objects 203 to 205 is calculated.
[0041] The second acquisition unit 305 obtains data (images) from a management server (not shown) storing rights-protected data, which serves as the basis for comparisons (feature representations) to be performed by the risk value calculation unit 307, described later. In this embodiment, the rights-protected data is a set of images known to be protected by copyright, trademark rights, portrait rights, etc.
[0042] The second feature calculation unit 306 calculates the feature representation of the protected data obtained by the second acquisition unit 305. In this embodiment, the second feature calculation unit 306 calculates the feature representation of the protected image group obtained by the second acquisition unit 305. Note that the second feature calculation unit 306 calculates the feature representation using the same method as the first feature calculation unit 304. Note that the second feature calculation unit 306 and the first feature calculation unit 304 can be configured as a single shared functional unit.
[0043] The risk value calculation unit 307 calculates the risk value by comparing the feature representations obtained by the first feature calculation unit 304 and the second feature calculation unit 306. In this embodiment, the risk value calculation unit 307 calculates the risk value by comparing the feature representation of each of the objects 203 to 205 calculated by the first feature calculation unit 304 with the feature representation of each protected image calculated by the second feature calculation unit 306. Details of the risk value calculation method will be described later.
[0044] The recording unit 308 records the content obtained by the first obtaining unit 302, the location information of the object obtained by the identification unit 303, and the risk value calculated by the risk value calculation unit 307 in a mutually related manner as risk information. In this embodiment, the recording unit 308 records the risk information (the location information and risk value of each of the objects 203 to 205) in the metadata portion of the image 202 obtained by the first obtaining unit 302.
[0045] Device operation Figure 4 This is a flowchart for calculating and recording risk values for content. Here, it is assumed that the content is an image generated by generative AI based on instructions (prompts, etc.) from the user, and the following processes begin at a timed indicative of image acquisition. However, the processes in S406 and S407 can be performed before the image acquisition in S401.
[0046] In S401, the first obtaining unit 302 obtains an image generated by generative AI. In this embodiment, an image 202 generated using image generation AI is obtained. The data to be obtained is not particularly limited, as long as it is content generated using generative AI, and examples include moving images, audio, text, and 3D models. Multiple data items can be obtained.
[0047] In S402, the identification unit 303 identifies (detects) objects included in the generated data obtained in S401. For example, the identification unit 303 identifies objects 203 to 205 included in image 202. Note that the process can be simplified by identifying (detecting) only objects of a specific category (type) for which infringement needs to be checked.
[0048] Figure 5 This is a diagram showing a table used to manage category names and select objects. This table records the category name indicating the category of the object and the type of object to be selected. Users can select one or more categories from the object management table and narrow down the objects to be identified in the generated data.
[0049] The specific method for object identification is not particularly limited, as long as it is a method that can extract the feature regions of the object. For example, by calculating the feature representation of an arbitrary range in an image and comparing it with the general feature representation of each object recorded in an object management table, it is possible to determine whether the range corresponds to an object. The comparison of feature representations can be performed, for example, by inputting the feature representation data of an arbitrary range in the image into a pre-learned classification model.
[0050] Specific examples of classification models include models using neural networks, models using decision trees such as random forests or gradient boosting, and models using the k-nearest neighbor algorithm. Note that the comparison results of feature performance can be output numerically to determine whether to set it as a feature region of an object using an arbitrary threshold determined by the user.
[0051] Examples of image feature representations used to extract feature regions of an object include global feature representations such as color histograms and color distributions. Examples include feature representations using Scale Invariant Feature Transform (SIFT) and Speed-Up Robust Feature Transform (SURF), as well as feature representations generated by neural networks. Examples of methods for extracting feature regions of an object based on image feature representations include bounding boxes (BB), instance segmentation, and semantic segmentation. Note that users can manually set the feature regions.
[0052] This embodiment assumes that features are generated and compared by a neural network, and that feature regions of the object are extracted by black-and-white (BB).
[0053] In step S403, the recognition unit 303 obtains the location information of the feature region of the object obtained in step S402. The location information of the object's feature region is information used to indicate the location and extent of the feature region in the image of the object. In this embodiment, the recognition unit 303 obtains the vertex coordinates of the rectangular feature region in the image corresponding to the object.
[0054] In S404, the risk value calculation unit 307 determines whether one or more objects were identified in S402. If one or more objects were identified, the process proceeds to S405; otherwise, if no objects were identified... Figure 4 The entire process is now complete.
[0055] In S405, the first feature calculation unit 304 calculates the feature representation of the region for each object. Examples of methods for calculating feature representation include, for example, methods using neural networks and methods using principal component analysis. This embodiment assumes that feature representation calculation using a convolutional neural network is performed.
[0056] In S406, the second obtaining unit 305 connects to the management server storing the protected data set and obtains the protected data set. The management server storing the protected data set is not particularly limited, as long as it can store the protected data set, and can be an external storage device 105 or a network server such as a cloud server communicating via network using I / O 108.
[0057] In S407, the second feature calculation unit 306 calculates the feature representation of the protected data set obtained in S406. The method for calculating the feature representation is not particularly limited, as long as it allows for comparison with the feature representation of the object calculated in S405. This embodiment assumes the use of a method similar to the feature representation calculation method used in S405.
[0058] Note that the protected data set to be stored can be pre-converted into a characteristic representation, and the obtaining unit 305 can obtain the data converted into a characteristic representation. If multiple data items with similar characteristic representations exist within the protected data, only a portion of them can be obtained.
[0059] In S408, the risk value calculation unit 307 compares the characteristic performance of the object obtained in S405 with the characteristic performance of each feature in the protected data set obtained in S407 to calculate the risk value of the object. Various methods can be used to compare the characteristic performance of the object with that of the protected data. Examples include Euclidean distance and cosine similarity. The method of expressing the risk value is not particularly limited, as long as it is an expression that allows for comparison of the magnitude of risk. For example, it can be a numerical value, or it can be indicated by categories such as high, medium, or low. This embodiment assumes that the maximum value of the cosine similarity used in metadata 206 to compare the object's risk value with that of each feature in the protected data set is used. It is assumed that the object's risk value is a real value in the range of 0 to 1, and that the closer it is to 1, the higher the risk of infringement.
[0060] In S409, the recording unit 308 records the location information of the characteristic region of each object obtained in S403 and the risk value calculated in S408 in the external storage device 105. It can also compare the risk value with a threshold to record risk values exceeding the threshold as indicating the presence of risk and other risk values as indicating the absence of risk. The user can then determine whether a risk exists or not for an object.
[0061] In S410, the risk value calculation unit 307 confirms whether risk values have been calculated for all objects identified in S402. If there are still objects for which risk values have not yet been calculated, the processing steps S405 and S408 to S409 are performed on the unprocessed objects to calculate and record risk values. If risk values have been calculated and recorded for all objects, Figure 4 The entire process is now complete.
[0062] As described above, according to the first embodiment, one or more objects included in an image are identified, and the location information and risk value (or the presence or absence of risk) of each object are calculated and recorded as metadata. Then, when a user assesses the infringement risk of an image (such as content generated by image-generating AI), the target image for assessment is displayed, and image regions within the image with a high risk of infringement are shown. This makes it possible to reduce the possibility of ignoring objects with a high risk of infringement, even when the target image for assessment includes a large number of objects or small objects. Image regions that are displayed as a result of infringement checks to distinguish them from other areas also have the risk of infringement, even if the risk value and the presence or absence of risk are not explicitly indicated.
[0063] By recording location information and risk values (the presence or absence of risk) in the metadata for objects with a high risk of infringement, users can easily confirm information related to the risk of infringement at any time.
[0064] (Modified Example 1-1) As a method for specifying the categories of objects included in the content (which are the identification criteria used to identify objects), the above description employs a method whereby the user selects from a table ( Figure 5 The configuration allows selection of object categories. However, the invention is not limited to this; any method can determine the object category. For example, the object category can be determined from user input (prompts) when content is generated from generative AI. Scene analysis of the content can be performed, and the object category can be determined. This saves users time and effort in determining the object category based on the content.
[0065] (Modified Example 1-2) In the above description, the risk value is calculated from the maximum value of the cosine similarity between the characteristic features of the calculated object and each characteristic feature of the protected data set. However, as long as the value indicates the magnitude of the risk of infringement, it can be calculated using another derived method. For example, the risk value could be a value in which the comparison between the characteristic features of the object and each characteristic feature of the protected data set is weighted using information related to each protected data set or information related to the object. Examples of information related to protected data include the "remaining term of protection" and "existence or non-existence of a license" for the protected data. Examples of information related to the object include the "occupancy rate" and "center coordinates" of the object in the content. This allows users to obtain an assessment of the risk of infringement using composite information.
[0066] (Modify Example 1-3) In the above description, the location information of the object's characteristic regions and the risk value are recorded in the object's metadata (Exif, etc.). On the other hand, as long as the content, the location information of the object's characteristic regions, and the object's risk value can be recorded in a way that is correlated with each other, they can be recorded in another form. For example, for content channels (e.g., RGB color channels), channels indicating the distribution of risk values can be additionally recorded. It can be configured not to be recorded along with the content, but rather managed as a separate database (table).
[0067] Figure 6 It is a diagram showing the risk values of the objects being managed. For example, as in... Figure 6 As shown in the management table, content identification information (ID), location information of the object's characteristic regions (maximum and minimum XY coordinates), and the object's risk value can be recorded in relation to each other. Note that the items to be recorded in the management table are not limited to these items. The storage location (storage unit) of the data in which information such as the management table is recorded is not particularly limited, as long as the user can check it at any time. That is, it can be an external storage device 105 or a network server such as a cloud server that can communicate via I / O 108.
[0068] (Modify Examples 1-4) In the above description, the risk of infringement is assessed when the content is an image. On the other hand, the type of data in the content is not particularly limited, as long as a specific area within the content can be expressed as location information. For example, it could be a moving image, audio, text, etc. Specific examples of obtaining location information of an object when it is a moving image, audio, or text will be described.
[0069] When the content is a moving image, the object's location information includes, for example, location information specifying the position and range within a particular frame of the moving image. When the content is audio, the object's location information includes, for example, time information based on the elapsed time in the audio from the start time. The object's position and range can be specified by specifying the start and end times of the target range. When the content is text, the object's location information includes, for example, information based on a number obtained by counting the number of words in the text data starting from the beginning of the document. The object's position and range can be specified by specifying the start and end word positions of the target range. Note that another method of location expression can be used, as long as it indicates the position of the object within the content.
[0070] Second Embodiment In the second embodiment, the form in which additional information regarding the protection of rights is presented when a user assesses the risk of infringement of an image will be described.
[0071] summary In this embodiment, in addition to the screen display in the first embodiment ( Figure 2 In addition to the above, the system also presents relevant protected data and supplementary information, as well as alternative images, related to objects whose risk of infringement exceeds any threshold. Here, the relevant protected data is data identified as having a high probability of infringement from protected data managed by a management server or similar entity. Supplementary information is rights-related information concerning the protected data, specifically including information about the rights holder, the publication date of the protected data, and the existence or non-existence of a license. As described in detail later, the alternative images presented are images that can replace the object and have a low risk of infringement.
[0072] Figure 7 This is a diagram illustrating a usage scenario of the information processing device in the second embodiment. Figure 7 The image is a picture where the information processing device 701 calculates the location information and risk value of objects included in the image as content, and presents the image of objects that exceed the risk value threshold set by the user.
[0073] Specifically, the information processing device 701 displays an evaluation target image (in which images of three people's faces are drawn) in display area 702. Highlighting 703 is overlaid on objects exceeding a threshold specified by the threshold setting user interface (UI) 707. An image of rights-protected data similar to the object displayed in display area 704, along with supplementary information, is displayed in display area 705. An alternative image (an image with a low risk of infringement) of the object displayed in display area 704 is displayed in the alternative image UI 706.
[0074] By confirming display area 702, users can identify the location of objects in the image that pose a high risk of infringement. By confirming display area 704, users can confirm detailed information about objects that pose a high risk of infringement. Furthermore, by confirming display area 705, users can confirm the corresponding protected data and supplementary information about the protected data.
[0075] In addition, UI 706 presents one or more alternative images that are proposed as substitutes for objects with a high risk of infringement. By selecting one of these alternative images to replace the object, the user can perform object replacement using the alternative image.
[0076] In the threshold setting UI 707, users can arbitrarily set the threshold for the risk value, and users can arbitrarily determine the notification sensitivity of the risk of infringement according to the circumstances. Display area 708 presents the generation conditions used when generating the evaluation target image displayed in display area 702 (e.g., information about the generated model, prompts set during generation, etc.).
[0077] This allows users to easily identify objects in a target image that pose a risk of infringement, as well as similar protected data and supplementary information. It also makes it easy to generate images in which objects at risk of infringement are replaced with alternative images.
[0078] Device configuration Figure 8 This is a diagram illustrating the functional configuration of the information processing device in the second embodiment. Note that the hardware configuration is different from that in the first embodiment ( Figure 1 The hardware configurations of the two are similar, so the description will be omitted.
[0079] Information processing device 801 includes the first embodiment ( Figure 3 The information processing device 801 includes a first acquisition unit 302, an identification unit 303, a first feature calculation unit 304, a second feature calculation unit 306, and a risk value calculation unit 307, as described in [reference to document 801]. The information processing device 801 also includes an input unit 802, a second acquisition unit 803, a condition acquisition unit 804, a risk determination unit 805, a substitute acquisition unit 806, and a display unit 807.
[0080] The input unit 802 performs the selection of content to be obtained by the first obtaining unit 302, and the acquisition of thresholds to be used by the risk determination unit 805. In this embodiment, the input unit 802 selects an image to be displayed in the display area 702 and obtains a threshold input by the user in the threshold setting UI 707.
[0081] The second acquisition unit 803 obtains data from the management server storing the data protected by rights, which is used as comparison information and supplementary information. In this embodiment, if the risk determination unit 805 determines that the risk value exceeds a threshold, the obtained data is displayed in the display area 705.
[0082] The condition acquisition unit 804 acquires the generation conditions used when generating the content acquired by the first acquisition unit 302. In this embodiment, the acquired generation conditions are displayed in the display area 708.
[0083] The risk determination unit 805 determines the risk of infringement by comparing the risk value calculated by the risk value calculation unit 307 with the threshold obtained by the input unit 802. In this embodiment, the risk determination unit 805 makes the determination by comparing the threshold input through the threshold setting UI 707 with the risk value calculated by the risk value calculation unit 307.
[0084] The substitute acquisition unit 806 acquires an image of a substitute for an object determined by the risk determination unit 805 to have a high risk of infringement. For example, the image can be acquired from an external storage device 105, or from a web server such as a cloud server that can communicate via I / O 108. In this embodiment, the acquired substitute image is displayed in the substitute image UI 706. Note that by operating the substitute image UI 706, the user can arbitrarily determine which image to use as a substitute for the object.
[0085] Display unit 807 generates and displays Figure 7 The screen shown is an example of a display unit 807 that controls the display of information obtained by the first acquisition unit 302, the identification unit 303, the second acquisition unit 803, the condition acquisition unit 804, and the substitute acquisition unit 806, as well as the determination result determined by the risk determination unit 805.
[0086] Device operation Figure 9 This is a flowchart in the second embodiment for calculating and recording risk values for content. Here, it is assumed that the content is an image generated by generative AI based on instructions (prompts, etc.) from the user, and the following processes begin at a timed indicative of image acquisition. However, the processes in S406 and S407 can be performed before the image acquisition in S401. Note that S401 to S408 and S410 are similar to those in the first embodiment, and therefore their description will be omitted.
[0087] In S901, the input unit 802 obtains the input content entered by the user. In this embodiment, the input unit 802 obtains... Figure 7The input content shown is the file path to the image data and a threshold. Note that data regarding the determination of the risk of infringement (such as the threshold) can be included, and other information is not specifically restricted.
[0088] In S902, the condition acquisition unit 804 acquires the generation conditions for the content obtained in S401. The content generation conditions are conditions set to create the content obtained in S401. For example, generation conditions include model name, hints, negative hints, CFG scale, and seed. This embodiment assumes that the model name and hints are acquired, but the generation conditions to be acquired are not limited to these.
[0089] In S903, the risk determination unit 805 compares the risk value of the object calculated in S408 with the threshold obtained in S901 to assess the risk of infringement of the object and presents the result. Note that the method for assessing the risk of infringement is not limited to a specific method. For example, the determination can be based on the relationship between the risk value of the object calculated by a model and the threshold. The risk value of the object can be calculated by multiple models and assessed by the relationship between their average or median and the threshold. Supplementary information from data protected by rights can be considered when performing the assessment.
[0090] Note that the display unit 107 can display one or more evaluation results. All results can be presented, or filtering can be performed under any conditions to limit the results to be presented. In this embodiment ( Figure 7 In step S408, if the risk value of an object calculated in step S408 is greater than a threshold, the risk of infringement is determined to be high, and the object with a high risk of infringement is displayed using highlight 703 and display area 704.
[0091] In S904, the substitute acquisition unit 806 acquires and presents to the user substitute data for the object identified as having a high risk of infringement in S903. Note that upon receiving an instruction to replace the object identified as having a high risk of infringement with a substitute image (such as pressing...), Figure 7 In the case of the "Replace" button, control is executed to generate alternative content in which the object is replaced using a substitute image. The alternative data is not particularly limited, as long as it can replace an object identified as having a high risk of infringement in the content obtained in S401. For example, there is a method where a data set with a low risk of infringement is pre-obtained, recorded in a management server, etc., and retrieved at any given time. The method for retrieving the recorded data set is not particularly limited, and for example, there is a method for retrieving data that has a high degree of similarity to an object identified as having a high risk of infringement.
[0092] Alternatively, data can be selected based on prompts from the generative AI-generated content, or users can select arbitrary data by referring to tag data of recorded data groups. One set of alternative data can be obtained, or multiple sets of alternative data can be obtained for the user to choose from. In this embodiment ( Figure 7 In the image, multiple images that are highly similar to objects identified as having a high risk of infringement are obtained from a pre-prepared group of images with a low risk of infringement and are displayed in the display area 706.
[0093] As described above, according to the second embodiment, information is displayed regarding objects with a high risk of infringement included in the assessment of target content (images), as well as information about related rights-protected data. This can reduce the user's burden when verifying the details of rights-protected data.
[0094] It becomes easier to assess target content (images) by replacing objects with similar alternative images that pose a high risk of infringement, and it becomes easier to obtain images with a low risk of infringement.
[0095] (Modified Example 2-1) In the above description, the threshold is arbitrarily determined and set by the user inputting it into the threshold setting UI 707. However, the method of setting the threshold is not limited to this. For example, there are methods for determining the data group protected by rights. For example, there are methods for determining the maximum distance in the feature representation space of data groups of records of the same category as protected data groups as a threshold for the risk of infringement. Other methods include determining the minimum distance between data groups of records of different categories as protected data groups as a threshold. Here, the category of protected data group is data with the same rights information, and for example, different poses of the same character, different drawing styles, etc. are classified into the same category.
[0096] This allows users to save time and effort in determining thresholds. It enables the adaptive determination of thresholds corresponding to the content of protected data groups to be compared with objects at risk of infringement.
[0097] (Modified Example 2-2) In the above description, a replacement image is obtained from a pre-prepared set of images with low infringement risk. On the other hand, the method is not particularly limited, as long as it is a method for obtaining an image with low infringement risk that can replace an object with high infringement risk. For example, an image can be regenerated based on the generation conditions obtained by the condition acquisition unit 804, and any portion with low infringement risk can be obtained. An object with high infringement risk can be used as the generation condition for generative AI to generate a new image, and an image with low infringement risk can be obtained from the generated image.
[0098] This allows users to save time and effort in preparing image sets with low risk of infringement in advance. It also reduces the recording area for image sets with low risk of infringement.
[0099] Third Embodiment In the third embodiment, an information processing apparatus for evaluating a trained model of generated images will be described. Specifically, the apparatus will describe deriving risk values as described in the first embodiment from multiple images generated using various trained models, recording the risk values in association with generation conditions, and evaluating the form of each trained model.
[0100] summary In this embodiment, for each trained model, the generated images are displayed to show locations (regions) where the risk value tends to be relatively high. The user is then presented with information indicating that the generation conditions at those locations tend to have a relatively high risk value.
[0101] Figure 10 This is a diagram illustrating a usage scenario of the information processing device in the third embodiment. Figure 10 This diagram illustrates a scenario where the information processing device 1001 records in a management table 1002 the location information of objects at risk of infringement in the content (images) of each trained model, along with their associated risk values and generation conditions. The information processing device 1001 displays the content recorded in the management table 1002 and the content calculated from the recording results on a display screen 1003. For example, users can make confirmations when relearning a trained model, when making improvements specific to a trained model, or when considering commercial use.
[0102] Display area 1004 shows the tendency of risk values corresponding to locations in the image generated by the trained model (here, ID=001). For example, areas with a particularly high risk of infringement are displayed as high-risk areas 1005. The risk value of the area and the conditions that influence the generation of the risk value are displayed together. By reviewing these displays, users can understand the locational tendency of the content of each trained model. This allows users to, for example, improve the appropriate trained model based on locational tendency.
[0103] Device configuration Figure 11 This is a diagram illustrating the functional configuration of the information processing device in the third embodiment. Note that the hardware configuration is different from that in the first embodiment ( Figure 1 The hardware configurations of the two are similar, so the description will be omitted.
[0104] Information processing device 1101 includes the first embodiment ( Figure 3 The information processing device 1101 includes a first acquisition unit 302, an identification unit 303, a first feature calculation unit 304, a second acquisition unit 305, a second feature calculation unit 306, and a risk value calculation unit 307, as described in [reference to document 1]. The information processing device 1101 also includes a condition acquisition unit 1102, a condition evaluation unit 1103, a recording unit 1104, and a display unit 1105.
[0105] The condition acquisition unit 1102 acquires the generation conditions (model and hints used) used when generating the content acquired by the first acquisition unit 302. In this embodiment, the acquired generation conditions are recorded in the management table 1002.
[0106] The condition assessment unit 1103 uses the results of the risk determination unit 805 and the risk of infringement associated with the generated condition assessment and location information obtained by the condition acquisition unit 1102. The recording unit 1104 records the content obtained by the condition acquisition unit 1102 and the condition assessment unit 1103 in management tables such as 1002. As a result, management... Figure 10 The information in the management table 1002. The display unit 1105 displays the evaluation results of the trained model based on the contents of the management table 1002 recorded by the recording unit 1104.
[0107] Figure 12 This is a flowchart for evaluating a trained model. Here, it is assumed that the content is an image generated by generative AI based on instructions (prompts, etc.) from the user, and the following processes begin at a time when the image is acquired. However, the processes in S406 and S407 can be performed before the image is acquired in S401.
[0108] In S1201, the condition acquisition unit 1102 acquires the generation conditions for the content acquired in S401. In this embodiment, the trained model ID and a hint are acquired. The trained model ID is an ID that can identify the model file and the learning data used for learning. Note that the generation conditions to be acquired are not limited to these.
[0109] In S1202, the risk value calculation unit 307 confirms whether the risk of infringement has been assessed for all content (images) indicated by the user. If there is still content for which the risk of infringement has not been assessed, the unassessed content is subjected to the processing in S1201 and thereafter, and the risk of infringement is assessed. If the risk of infringement has been assessed for all content, the process proceeds to S1203.
[0110] In S1203, the condition assessment unit 1103 records multiple pieces of information regarding the assessed risk of infringement in the recording unit 1104, including the model used, the content, the location information of each object in the content, the risk value of each object, and the generation conditions. Then, the display unit 1105 displays the model's assessment results from the recorded results.
[0111] Here, the location information of each object in the content is the location information of each object obtained in S403. The risk value is the risk value of the object calculated in S408. The generation conditions obtained in S1201 are related to the features of the object. For example, the generation conditions include hints, negative hints, CFG scale, and seeds. In this embodiment, the trained model ID used to indicate the model used, the content ID used to indicate the content, the region range of the object in the content, the risk value of the object, and the hints used when creating the content are recorded in management table 1002.
[0112] The evaluation of the model is calculated from the location information, risk value, and generation conditions of each content object generated by the model used. For example, by integrating and normalizing the infringement risk value of the objects generated by the model used in association with the location information, it is possible to calculate the trend of the distribution of infringement risk of the model used.
[0113] For example, each word used as a prompt in the generation conditions is integrated in association with the risk value of the object included in the content generated by that word and the location information of the object. By comparing the integration results, it is possible to calculate words (high-risk words) that tend to have a high risk of infringement in the model used, and their location information. In this embodiment, the distribution tendency of the infringement risk of the model used and the prompt words with a high risk of infringement and their location information are calculated. Note that the evaluation method is not limited to the above method, as long as it can assess the infringement risk of the model.
[0114] The presentation of the evaluation results is performed by the display unit 107. The presentation method is not particularly limited, as long as it allows the user to identify the evaluation results. The items to be presented are not particularly limited, as long as they are items whose infringement risk of the used model can be confirmed in association with location information. For example, some or all of the items recorded in management table 1002 can be presented. Some or all of the evaluation results of the used model can be presented.
[0115] The above Figure 10 The display screen 1003 shows the tendency of the risk value corresponding to the location as part of the evaluation results of the model used (ID=001), and highlights the high-risk area 1005. The risk value and high-risk words in the high-risk area 1005 are displayed together.
[0116] As described above, according to the third embodiment, the risk values described in the first embodiment are derived from multiple images generated using various trained models and recorded in association with the generation conditions. Specifically, the risk values and generation conditions are recorded in association with location information in the images. Then, based on the recorded information, an assessment associated with the location in the images generated by each model is calculated and presented to the user. This allows the user to grasp the tendency of infringement risk in each trained model and to perform appropriate improvements to the trained models based on that tendency.
[0117] (Modified Example 3-1) In the above description, the calculation of risk values associated with location information is performed by integrating the risk values of objects in the content generated by the trained model in association with the object's location information. However, this method is not particularly limited, as long as the risk of infringement by the trained model can be assessed. For example, weighting can be performed based on the location information in the image of the object. Specifically, weighting can be performed to reduce the risk of infringement in areas that are unlikely to be used, in cases where the user plans to use only a portion of the content. Alternatively, only objects with risk values exceeding an arbitrary threshold can be integrated and calculated. The risk value of the object and the location information of the object can be associated with each pixel of the content, or with a region in which multiple pixels are connected.
[0118] (Modified Example 3-2) In the above description, as a method for assessing the infringement risk of prompt words, the risk value of the object included in the content generated by each prompt word is integrated in association with the object's location information. Then, the infringement risk of the prompt word is assessed from the integration result. However, the method for assessing the infringement risk of prompt keywords is not limited to the above method. For example, only some prompt keywords that are highly relevant to the object among the multiple prompt words used to generate content can be associated and integrated. The determination of prompt keywords that are highly relevant to the object can be performed by user selection. Alternatively, all prompt words used to generate the object and content are input into, for example, a pre-learned neural network, and feature representations are extracted. Prompt keywords that are highly relevant to the object can be selected by, for example, using feature representations extracted using cosine similarity comparison.
[0119] Other embodiments Embodiments of this disclosure can also be implemented by a computer of a system or apparatus that reads and executes computer-executable instructions (e.g., one or more programs) recorded on a storage medium (which may also be more fully referred to as a 'non-transitory computer-readable storage medium') to perform one or more functions of the above embodiments and / or includes one or more circuits (e.g., application-specific integrated circuits (ASICs)) for performing one or more functions of the above embodiments, and by a method performed by a computer of the system or apparatus by, for example, reading and executing computer-executable instructions from the storage medium to perform one or more functions of the above embodiments and / or controlling one or more circuits to perform one or more functions of the above embodiments. The computer may include one or more processors (e.g., a central processing unit (CPU), a microprocessor unit (MPU)) and may include separate computers or networks of separate processors to read and execute computer-executable instructions. The computer-executable instructions may be provided to the computer, for example, from a network or storage medium. The storage medium may include, for example, a hard disk, random access memory (RAM), read-only memory (ROM), storage devices for distributed computing systems, optical discs (such as CDs, DVDs, or Blu-ray discs). TM One or more of the following: flash memory devices, memory cards, etc.
[0120] Embodiments of the present invention can also be implemented by means of a network or Various storage media provide software (including computer program products of computer programs) that performs the functions of the above embodiments to a system or device, and the computer (central processing unit (CPU), microprocessor unit (MPU)) of the system or device reads and executes the computer program.
[0121] While this disclosure has been described with reference to embodiments, it is to be understood that this disclosure is not limited to the disclosed embodiments. The scope of the appended claims is to be given the broadest interpretation in order to cover all such modifications and equivalent structures and functions.
Claims
1. An information processing apparatus, comprising: An acquisition unit that acquires content generated by generative artificial intelligence (AI); An identification unit that identifies objects that pose a risk of infringement among the objects included in the content; as well as A recording unit records the location information of objects in the content as risk information.
2. The information processing apparatus according to claim 1, wherein... The identification unit identifies objects belonging to specific categories within the content as objects with infringement risks.
3. The information processing apparatus according to claim 1, further comprising: The calculation unit calculates a risk value indicating the degree of likelihood that an object with infringement risk is indeed an infringing object, wherein... The recording unit records the location information and risk value of the object with infringement risk in the risk information in association with each other.
4. The information processing apparatus according to claim 3, wherein The calculation unit calculates the risk value of the object at risk of infringement based on the similarity between the object at risk of infringement and the given protected data.
5. The information processing apparatus according to claim 1, further comprising: The display control unit displays the content and the risk information on the display unit.
6. The information processing apparatus according to claim 5, wherein The display control unit also displays information on the display unit about protected data similar to objects at risk of infringement.
7. The information processing apparatus according to claim 5, further comprising: A substitute obtaining unit obtains a substitute that can replace an object with an infringement risk; A receiving unit that receives an instruction to replace an object in the content that poses a risk of infringement with the substitute; as well as A generation unit generates alternative content based on the instructions, wherein objects with infringement risks are replaced using the alternatives in the alternative content.
8. The information processing apparatus according to claim 1, further comprising: The exporting unit, based on the risk information of multiple pieces of content generated by the generative AI, exports the positions in the content generated by the generative AI that have a relatively high probability of generating objects with infringement risks.
9. The information processing apparatus according to claim 8, further comprising: The condition obtaining unit obtains the generation conditions used when the generative AI generates each of the plurality of content items, wherein... The exporting unit also exports generation conditions in the content generated by the generative AI that indicate a high probability of generating objects with infringement risks.
10. A control method for an information processing device, the control method comprising: Obtain content generated by generative AI; Identify objects that pose a risk of infringement among the objects included in the content; as well as The location information of the objects in the content is recorded as risk information in the storage unit.
11. A computer-readable storage medium storing a program for causing a computer to perform the control method according to claim 10.
12. A computer program product including a program for causing a computer to perform the control method according to claim 10.