Image processing system, image processing method, and program

The image processing system uses AI models for precise feature extraction and correction to ensure critical features are retained in style-converted images, addressing unintended modifications and enhancing user control over the conversion process.

WO2026070346A1PCT designated stage Publication Date: 2026-04-02SONY GROUP CORP

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Existing image generation AI models often result in unintended modifications to features in style-converted images, making it difficult to retain specific characteristics such as identity features of individuals or logos in the original image.

Method used

An image processing system that includes a feature information extraction unit to identify and track specific features before and after style conversion, a determination unit to assess retention, and correction units to modify the image to ensure these features are preserved, using AI models like Grounding DINO and VLM for precise feature extraction and LLM for determination.

Benefits of technology

Ensures that critical features in the input image, like skin color and logos, are accurately retained or corrected in the style-converted output, enhancing user control over the conversion process and maintaining recognizable elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025031759_02042026_PF_FP_ABST
    Figure JP2025031759_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present technology relates to an image processing system, an image processing method, and a program that enable an input image to be subjected to style conversion as intended by a user. An image processing system according to the present technology comprises: a converting unit that subjects an input image to style conversion to generate a converted image; and a determining unit that acquires a first determination result indicating whether a feature to be determined, among features included in the input image, is retained before and after the style conversion. The present technology is applicable to image processing systems that subject input images to style conversion, for example.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing system, image processing method, and program

[0001] The present technology relates to an image processing system, an image processing method, and a program, and particularly relates to an image processing system, an image processing method, and a program that can perform style conversion on an input image as intended by a user.

[0002] In recent years, image generation AI (Artificial Intelligence) models capable of performing various style conversions on images have been developed (see, for example, Patent Document 1).

[0003] Examples of such image generation AI models include Stable Diffusion. In Stable Diffusion, by inputting an image with a desired painting style as a reference image or performing fine-tuning on the model, it is possible to generate images with various painting styles.

[0004] Japanese Patent Application Laid-Open No. 2021-135822

[0005] When performing style conversion on an input image using an image generation AI model, unintended content modifications may occur in the converted image. For example, when converting a real image into an image that evokes a specific anime (animation), i.e., an anime-style image, the features of the people shown in the real image may be overly processed, making it unclear who the original people in the converted image are.

[0006] The present technology has been made in view of such a situation and enables style conversion to be performed on an input image as intended by the user.

[0007] An image processing system according to one aspect of the present technology includes a conversion unit that performs style conversion on an input image to generate a converted image, and a determination unit that obtains a first determination result as to whether a feature to be determined among the features included in the input image is retained before and after the style conversion.

[0008] One aspect of this technology is an image processing method which includes applying a style transformation to an input image to generate a transformed image, and obtaining a determination result of whether or not the feature to be determined among the features contained in the input image is retained before and after the style transformation.

[0009] One aspect of this technology involves a program that causes a computer to perform a process that includes applying a style transformation to an input image to generate a transformed image, and obtaining a determination result of whether or not the feature to be judged among the features contained in the input image is retained before and after the style transformation.

[0010] In one aspect of this technology, a style transformation is applied to an input image to generate a transformed image, and a determination result is obtained as to whether or not the feature to be judged among the features contained in the input image is retained before and after the style transformation.

[0011] This is a diagram illustrating an image processing system to which this technology is applied. This is a diagram illustrating a first example of unintended modifications to the converted image for the user. This is a diagram illustrating a second example of unintended modifications to the converted image for the user. This is a block diagram illustrating an example of the configuration of the image processing system. This is a diagram illustrating an example of a list of features to be preserved. This is a flowchart illustrating the processes performed by the image processing system. This is a diagram illustrating a first example of a screen displayed on the display unit. This is a flowchart illustrating other processes performed by the image processing system. This is a block diagram illustrating a first modified example of the configuration of the image processing system. This is a diagram illustrating a second example of a screen displayed on the display unit. This is a diagram illustrating a third example of a screen displayed on the display unit. This is a diagram illustrating a fourth example of a screen displayed on the display unit. This is a diagram illustrating a fifth example of a screen displayed on the display unit. This is a flowchart illustrating yet another process performed by the image processing system. This is a block diagram illustrating a third modified example of the configuration of the image processing system. This is a block diagram illustrating an example of the configuration of computer hardware.

[0012] The following describes the configurations for implementing this technology. The explanation will proceed in the following order: 1. Configuration and operation of the image processing system 2. Modified examples

[0013] <1. Configuration and Operation of the Image Processing System> Figure 1 is a diagram illustrating an image processing system to which this technology is applied.

[0014] The image processing system 1 of this technology is a system that applies style transformation to an input image using an image generation AI model.

[0015] In the example shown in Figure 1, a live-action image showing two athletes (people) is input to the image processing system 1 as input image P1. The image processing system 1 inputs input image P1 to an image generation AI model, and outputs an image (transformed image) which is the input image P1 with style transformation applied, as output image P2.

[0016] Image processing system 1, for example, applies a style transformation to the input image P1, which transforms the input image into an image reminiscent of a specific manga, a so-called manga-style image, thereby obtaining an output image P2 in which people A and B appear to be immersed in the world of that manga. In this disclosure, style transformation includes changing the drawing style, changing the taste, changing the texture, changing the color tone, deforming the content, adding a mosaic, etc. Hereinafter, the player shown on the left in the input image P1 and the output image P2 will also be referred to as person A, and the player shown on the right will also be referred to as person B.

[0017] By using the image processing system 1 to provide users with images that make it appear as if a person in the input image has entered the world of content such as manga or anime, service providers can enhance fan engagement. Furthermore, by using the image processing system 1 to convert images of a company's products, etc., into images in art styles commonly used in different fields, while retaining the logos and text in the original image, companies can appeal to customer segments that previously had no opportunity to be interested in their products.

[0018] In this case, when a style transformation is applied to an input image using an image generation AI model, the transformed image may undergo modifications that are unintended by the user.

[0019] For example, as shown in Figure 2, the skin color of person A may change before and after style conversion. If characteristics that guarantee a person's identity, such as gender, age, and skin color, change before and after style conversion, it becomes impossible to determine who the original person was that appears in the converted image.

[0020] Furthermore, as shown in Figure 3, for example, the shape of the number and logo printed on the uniform worn by person B may change before and after the style conversion.

[0021] At first glance, it might seem that by using prompts and control networks, users can specify features they want to retain (or should retain), such as features that guarantee a person's identity or features with complex shapes like text and logos, preventing these features from changing before and after style conversion. A control network is a technology that generates images that reproduce complex compositional and pose information that cannot be specified by prompts alone, by inputting conditions such as images and postures in addition to prompts.

[0022] However, if the input image contains more than one person, even if the user specifies detailed characteristics of each person in the prompt, the converted image may not necessarily reflect the user's intended characteristics. Furthermore, even if the user specifies hint information such as outline and depth for logos and text using a control network, it may not be possible to reliably prevent the logo shape or text from changing before and after style conversion.

[0023] This technology was conceived with the above points in mind. It generates a transformed image by applying a style transformation to an input image, and obtains a determination result (first determination result) of whether or not the features that should be preserved among the features contained in the input image are preserved before and after the style transformation. This enables the application of style transformation to the input image according to the user's intentions.

[0024] Figure 4 is a block diagram showing an example configuration of the image processing system 1.

[0025] As shown in Figure 4, the image processing system 1 is configured by connecting a user terminal 11 and a server 12 via wired or wireless communication.

[0026] The user terminal 11 is configured as an information processing device such as a smartphone, tablet, or PC. The user terminal 11 consists of an input unit 31 and a display unit 32.

[0027] The input unit 31 receives input images that are to be style-converted.

[0028] The input image may be a real-life image, a computer graphics (CG) image, or an image created using an AI image generation model. The input image may contain people, landscapes, logos, or text.

[0029] When a real-life image is input, the user can obtain a transformed image that includes image elements such as people and scenery depicted in the real-life image. Furthermore, when a CG image or an image created using an AI image generation model is input, the user can obtain a transformed image that includes image elements such as objects and effects that do not exist in reality.

[0030] The input image may be a still image or a moving image. If the input image is a still image, a still image will be output as the output image; if the input image is a moving image, a moving image will be output as the output image. The input image may also be a frame image that makes up a moving image. If the input image is a still image, examples of input image formats include JPEG (Joint Photographic Experts Group) and PNG (Portable Network Graphics).

[0031] The input unit 31 supplies the input image to the server 12.

[0032] The display unit 32 displays a screen that includes the input image, the output image supplied from the server 12, and the results of the server 12's determination of whether the features to be retained are retained before and after the style conversion.

[0033] Server 12 is configured, for example, as a web server or a server device (information processing device) on the cloud. Server 12 consists of a feature information extraction unit 51, a storage unit 52, an image conversion unit 53, a determination unit 54, a correction unit 55, and an integration unit 56.

[0034] The feature information extraction unit 51 obtains a list of features to be retained before and after style conversion from the storage unit 52. The storage unit 52 stores features to be retained, such as "the person's skin color," "the number on the back of the uniform worn by the person," and "the person's gender."

[0035] The feature information extraction unit 51 extracts feature information for each feature to be retained from the input image. Specifically, for each image element having a feature to be retained, the feature information extraction unit 51 extracts the positional information of the image element within the input image and the feature information for each feature to be retained. In this disclosure, an image element refers to a person or object that appears in an image, such as a person, an object, a landscape, or an effect.

[0036] Feature information is information that indicates the state of the features of each image element, and includes information such as "the skin color of the person on the left side of the image is white," "the skin color of the person on the right side of the image is brown," and "the number on the back of the uniform worn by the person in the center of the image is 6."

[0037] For example, the feature information extraction unit 51 first uses an AI model such as Grounding DINO, which utilizes Zero-shot object detection technology, to obtain a bounding box as positional information that indicates where the image element specified in natural language is located within the input image. For example, the positional information may be obtained in the format of "a 100px square centered at a position 200px from the right and 400px from the top." The feature information extraction unit 51 may also use a segmentation model to obtain a segmentation mask from the input image and use that segmentation mask as positional information for the image element.

[0038] The feature information extraction unit 51 then uses a VLM (Vision Language Model) such as GPT-4o, LLaVA, or BLIP to extract feature information about the features to be retained for each image element of the input image. A VLM is an AI model that takes an image and query text as input and outputs a response text corresponding to the query text.

[0039] For example, when the feature information extraction unit 51 inputs the input image and the query text "What is the skin color of the person in the center of the image?" into the VLM, it can obtain a response such as "a beige color with a hint of light pink." This allows the feature information extraction unit 51 to obtain information that links the positional information and feature information of image elements, such as "Person in the center of the image: Skin color is a beige color with a hint of light pink."

[0040] By using zero-shot object detection technology and VLM, it is possible to handle cases where the specification of image elements with features to be retained, or the specification of the features to be retained themselves, is done in natural language. Since the positional information of each image element is identified, if a feature to be retained is not retained in the converted image, the image processing system 1 can indicate to the user which features of the image elements in the input image are not retained before and after style conversion, for example, by surrounding the image element whose features are not retained.

[0041] Furthermore, the feature information extraction unit 51 can also extract feature information about the features to be retained from the partial image by cutting out a partial image containing the image elements from the input image based on the positional information of the image elements, and inputting the partial image and a query text such as "What is the skin color of the person?" into the VLM. In addition, the feature information extraction unit 51 may extract the color information (color code) of the image elements in the input image as feature information. For example, feature information such as "Skin color of the person on the left side of the image: #efccaa" may be extracted.

[0042] The feature information extraction unit 51 supplies the position information of each image element of the input image and feature information about the features to be retained that each image element possesses to the determination unit 54.

[0043] As described above, a list of features to be retained is stored in the memory unit 52. In the memory unit 52, the features to be retained are stored in a text format such as "skin color". By managing the features to be retained in a text format, it becomes possible to flexibly accept the user's specification of the features to be retained.

[0044] A plurality of tags indicating the features that may be included in the input image may be presented to the user, and the user may be able to select whether each feature should be retained before and after the style conversion. When the user selects the features to be retained using tags, the memory unit 52 can manage the features to be retained in a format other than natural language.

[0045] The memory unit 52 may store, as a template, a list corresponding to, for example, the content of the input image, that is, a combination of the types of image elements in the input image and the type of the input image (whether the input image is a real image or a composite image).

[0046] FIG. 5 is a diagram showing an example of a list of features to be retained.

[0047] As shown in FIG. 5, when the input image is a real image in which a person is shown, the list of features to be retained includes, for example, "skin color", "hair color", "eye color", and "presence or absence of beard".

[0048] Also, when the input image is a real image in which characters are shown, the list of features to be retained includes, for example, "character size", "character font", "character content (spelling)", and "character color".

[0049] Further, when the input image is a composite image in which a logo is shown, the list of features to be retained includes, for example, "logo size", "logo color", and "logo shape".

[0050] When a template of the list of features to be retained is prepared in advance, the user does not need to specify the features to be retained.

[0051] The image conversion unit 53 generates a converted image by applying style conversion to the input image. For example, the image conversion unit 53 generates a converted image using an image generation AI model built on a neural network, which converts the input image to the same style as the image input as a reference image. In this case, the user can specify the style they want to apply to the input image simply by inputting a reference image.

[0052] The image conversion unit 53 may generate the converted image using a diffusion-based image generation AI model such as Stable Diffusion. In this case, the user can specify in detail the style they want to apply to the input image using prompts written in natural language. Alternatively, the image conversion unit 53 may generate the converted image using a diffusion-based image generation AI model that has been fine-tuned to generate images with a specific style. In this case, it is expected that the consistency and quality of the style conversion applied to the input image can be improved.

[0053] The image conversion unit 53 may generate a converted image by applying a filter that converts the color tone or a filter that adds a mosaic effect. When generating a converted image using an image generation AI model, objects that do not exist in the input image may be drawn in the converted image, but when generating a converted image using a filter, such risks can be avoided. The image conversion unit 53 may use multiple image generation AI models in combination, or use an image generation AI model in combination with a filter.

[0054] Furthermore, the user can change the content of the style conversion applied to the input image by the image conversion unit 53 at any time.

[0055] The image conversion unit 53 supplies the input image and the converted image to the determination unit 54.

[0056] The determination unit 54 compares the input image supplied by the image conversion unit 53 with the converted image to determine whether or not the features that should be preserved before and after style conversion are preserved.

[0057] The determination unit 54 determines, for example, whether each feature to be retained is retained before and after style conversion, using VLM. Specifically, first, the determination unit 54 cuts out partial images containing image elements having the features to be retained from the converted image, element by element, based on the position information supplied from the feature information extraction unit 51. The determination unit 54 extracts feature information about the features to be retained from the partial images by inputting the same query text into the VLM as the query text input into the VLM when feature information was extracted from the partial images and the input image.

[0058] The feature information extracted from the output image includes information such as "the skin color of the person on the left side of the output image is white," "the skin color of the person on the right side of the output image is white," and "the number on the back of the uniform worn by the person in the center of the output image is 18."

[0059] For example, the determination unit 54 can input a partial image of the person on the right, extracted from the converted image, and the query text "What is the person's skin color?" into the VLM, thereby obtaining a response text such as "beige" as feature information. The determination unit 54 compares the feature information about the feature to be retained extracted from the input image ("beige with a hint of light pink") with the feature information about the feature to be retained extracted from the converted image ("beige") to determine whether the feature to be retained is retained before and after the style conversion. For example, since "beige with a hint of light pink" and "beige" do not perfectly match, the determination unit 54 determines that the skin color of the person on the right is not retained before and after the style conversion.

[0060] By using VLM to determine whether the features to be retained are preserved before and after style conversion, it is possible to perform the determination of whether the features to be retained are preserved using the same workflow, even if the features to be retained are measured on different scales such as color, shape, and size.

[0061] If the feature information is extracted as numerical values ​​such as color information, the determination unit 54 compares the color information as feature information extracted from the input image with the color information as feature information extracted from the converted image to determine whether or not the features to be retained are retained.

[0062] Furthermore, when the determination unit 54 compares the feature information extracted from the input image with the feature information extracted from the converted image, it is preferable to allow for a certain degree of leeway in the comparison. For example, in the above example, since the skin color of the person on the right is beige in both the input image and the output image, the determination unit 54 may determine that the skin color of the person on the right is preserved before and after the style conversion.

[0063] When comparing feature information with a certain degree of leeway, the determination unit 54 uses LLM (Large Language Models) to determine, for example, whether the feature information extracted from the input image and the feature information extracted from the converted image are similar. When using LLM to determine whether the feature information is similar or not, the user may be prompted to input examples that can be considered similar or dissimilar. This allows the user to adjust the strictness of the determination made by the LLM.

[0064] When feature information is extracted as numerical values, the determination unit 54 calculates, for example, the similarity (e.g., cosine similarity) between the feature information extracted from the input image and the feature information extracted from the converted image. If the similarity is greater than or equal to a threshold, it determines that the features to be retained are retained before and after the style conversion. Because the determination of whether the features to be retained are retained is made based on numerical values, the reason for the determination can be clearly presented to the user. The user can adjust the strictness of the determination by adjusting the threshold.

[0065] If the feature to be retained is "the content of the characters," the determination unit 54 may compare the string extracted by character recognition processing on the input image with the string extracted by character recognition processing on the output image to determine whether the feature to be retained is retained before and after style conversion. By extracting the string using character recognition processing, it is possible to accurately determine whether the content of the characters has changed before and after style conversion.

[0066] If it is determined that certain features that should be preserved are not preserved before and after style conversion, the reason for this determination may be stored in the server 12. This reason for determination can be used, for example, in a later process that modifies the converted image.

[0067] The determination unit 54 compares the input image with the modified converted image supplied by the integration unit 56 to determine whether or not the features that should be retained in the modified converted image are retained.

[0068] The determination unit 54 supplies the converted image and the determination result of whether the features to be retained in the converted image are retained to the modification unit 55. The determination unit 54 also supplies the converted image supplied from the image conversion unit 53 and the modified converted image supplied from the integration unit 56 as output images to the display unit 32, and supplies the determination result of whether the features to be retained are retained before and after style conversion (whether the features to be retained are retained in the converted image before and after modification) to the display unit 32.

[0069] The modification unit 55 and the integration unit 56 modify the converted image so that features that should be retained in the converted image are retained. For example, the modification unit 55 and the integration unit 56 modify the regions within the converted image that contain image elements having features that should be retained but were determined not to be retained before and after the style conversion.

[0070] Specifically, first, the modification unit 55 extracts a partial image from the input image that includes image elements that have been determined not to retain the features to be retained. Since the positional information of the image elements that have the features to be retained within the input image is obtained by the feature information extraction unit 51, the partial image is extracted from the input image based on that positional information.

[0071] Next, the modification unit 55 applies a style transformation to the partial image to generate a partially modified image, which is an image for partially modifying the transformed image.

[0072] The content of the style transformation applied to the partial image may be adjusted based on the reason why the determination unit 54 determines that certain features that should be retained are not retained. For example, prompts or control networks can be used to adjust the content of the style transformation applied to the partial image. For example, as explained with reference to Figure 2, if an input image showing two players is input and the skin color of the player on the left becomes whiter before and after the style transformation, the correction unit 55 generates a partially corrected image by applying a style transformation to a partial image extracted from the input image P1 containing the area of ​​the player on the left. When applying a style transformation to the partial image, the correction unit 55 adjusts the content of the style transformation applied to the partial image by inputting a prompt containing the word "brown" to the image generation AI model that performs the style transformation, based on the reason that the skin color of the player on the left became whiter before and after the style transformation.

[0073] Prompts for adjusting the style transformation applied to a partial image are generated, for example, using LLM. By adjusting the style transformation based on the reason why the determination unit 54 determined that certain features to be retained are not retained, it is expected that a partially modified image suitable for correcting the transformed image can be generated without repeating the style transformation. Since prompts for adjusting the style transformation are generated using LLM, the server 12 can generate prompts automatically.

[0074] If the shape of characters or logos changes before and after style conversion, the correction unit 55 may generate a partially corrected image by changing the color tone of a partial image extracted from the input image using a filter.

[0075] In this way, style transformations are reapplied only to image elements that do not retain the features that should be retained, thus preventing further style transformations from being applied to image elements that do retain the features that should be retained.

[0076] Instead of generating a partially modified image based on the input image, the region in the output image containing image elements that are determined to lack features to be retained may be directly modified by a fill-in process called Inpaint or similar. In this case, the fill-in process is performed on the output image, taking into account the reason why the determination unit 54 determined that the features to be retained are not retained. The region in the output image to which the fill-in process is applied may be specified by the user.

[0077] The modification unit 55 supplies the conversion image, the partially modified image, and information indicating the cropping position of the partial image that will serve as the basis for the partially modified image to the integration unit 56.

[0078] The integration unit 56 integrates the partially modified image supplied from the modification unit 55 with the converted image based on the cropping position of the partial image that will serve as the basis for the partially modified image. For example, the integration unit 56 integrates the partially modified image and the converted image by pasting the partially modified image directly onto the converted image.

[0079] The integration unit 56 can also apply a mask or fill-in process such as Inpaint to areas in the converted image that contain image elements where it has been determined that the features to be retained are not retained, and then paste the partially modified image. In order to make the partially modified image blend in with the converted image, a predetermined transparency may be set for the edges of the partially modified image.

[0080] The integration unit 56 supplies the converted image, which is the result of integrating the partially modified images, to the determination unit 54 as the corrected converted image. If all the features to be retained are not retained in a single modification, the determination unit 54 makes the determination, the modification unit 55 makes the modification, and the integration unit 56 makes the integration, and this process is repeated, thereby optimizing the converted image step by step.

[0081] Next, referring to the flowchart in Figure 6, the processing performed by the image processing system 1 having the above configuration will be explained. In the processing shown in Figure 6, the converted image, which has undergone style conversion by the image conversion unit 53, is output without modification. The processing shown in Figure 6 is started, for example, when an input image is input by the user and a button to instruct the user to start style conversion is clicked by the user.

[0082] In step S1, the image conversion unit 53 generates a converted image by applying a style conversion to the input image.

[0083] The processes in steps S2 and S3 are carried out in parallel with the process in step S1. In step S2, the feature information extraction unit 51 determines the features to be retained before and after style conversion. Specifically, the feature information extraction unit 51 obtains a list of features to be retained corresponding to the content of the input image from the storage unit 52.

[0084] In step S3, the feature information extraction unit 51 extracts feature information about the features to be retained from the input image.

[0085] After the processing in steps S1 and S3 is performed, in step S4, the determination unit 54 determines whether or not the features that should be retained before and after the style conversion are retained.

[0086] In step S5, the determination unit 54 outputs the converted image converted by the image conversion unit 53 as the output image, and the display unit 32 displays the output image together with the determination result from the determination unit 54.

[0087] Figure 7 shows a first example of the screen displayed on the display unit 32.

[0088] In Figure 7, the input image is displayed on the left side of the screen, the output image is displayed in the center, and the results of determining whether each feature to be retained is actually retained before and after style conversion are displayed on the right side. Below the results of the determination, a conversion start button B1 is displayed to instruct the user to start the style conversion.

[0089] When a user inputs an input image into the image processing system 1, the input image and a list of features to be preserved before and after style conversion are displayed on the screen. The user inputs the input image by, for example, dragging and dropping the input image onto the screen or by specifying the location where the input image file is stored.

[0090] With the input image and a list of features to be retained displayed on the screen, when the user clicks the conversion start button B1, a style conversion is performed on the input image, and the output image and the result of determining whether each of the features to be retained was actually retained before and after the style conversion (OK or NG) are displayed on the screen. For items that were determined not to be retained before and after the style conversion, a warning icon is displayed to indicate that the feature was not retained before and after the style conversion.

[0091] In the example in Figure 7, the features to be preserved are listed as "skin color," "hair color," "presence or absence of a beard," and "jersey number." The results of the judgment on whether these features are actually preserved before and after style conversion are displayed for the person shown on the left side of the input and output images.

[0092] For example, the skin color of the person on the left becomes whiter before and after the style conversion, so the result indicates that "skin color" is not preserved. The "hair color" and "presence or absence of beard" of the person on the left do not change before and after the style conversion, so the result indicates that hair color and presence or absence of beard are preserved. In the input image, the number printed on the uniform worn by the person on the left is hidden by the person's arm, making it difficult to determine the number based on the input image, so it is not determined whether the "number" is preserved before and after the style conversion.

[0093] As described above, the image processing system 1 of this technology can show the user whether or not each feature that should be preserved is actually preserved before and after style conversion. The user can check the features that are not preserved before and after style conversion, and if the processing on those features is not acceptable, they can redo the style conversion.

[0094] Next, other processes performed by the image processing system 1 will be explained with reference to the flowchart in Figure 8. In the process shown in Figure 8, the output image is output after it is confirmed that the features to be preserved before and after style conversion are preserved. The process shown in Figure 8 is started, for example, when an input image is input by the user and the conversion start button B1 is clicked by the user.

[0095] The processing in steps S21 to S23 is the same as the processing in steps S1 to S3 in Figure 6.

[0096] After the processing in steps S21 and S23 is performed, in step S24, the determination unit 54 compares the input image and the converted image to determine whether or not the features that should be retained before and after style conversion are retained.

[0097] If it is determined in step S24 that the features to be retained are not retained, the correction unit 55 generates a partially corrected image in step S25.

[0098] In step S36, the integration unit 56 modifies the converted image by integrating the converted image and the partially modified image.

[0099] After the processing in step 36 is performed, the process returns to step S24, where the input image and the modified converted image are compared, and it is determined again whether or not the features that should be preserved before and after style conversion are retained.

[0100] If it is determined in step S24 that the features to be retained are retained, in step S27 the determination unit 54 outputs the converted image as the output image, and the display unit 32 displays the output image together with the determination result from the determination unit 54.

[0101] As described above, by repeatedly checking whether the features to be preserved are retained before and after style conversion, and by correcting the converted image, it becomes possible to ultimately output an image in which all the features to be preserved are confirmed to be retained.

[0102] <2. Modifications> - Example where the user specifies the features to be retained. Figure 9 is a block diagram showing a first modification of the configuration of the image processing system 1. In Figure 9, components identical to those in Figure 4 are denoted by the same reference numerals. Repetitive explanations are omitted as appropriate.

[0103] The image processing system 1 in Figure 9 differs from the image processing system 1 in Figure 4 in that specified information is supplied from the user terminal 11 to the server 12.

[0104] The input unit 31 of the user terminal 11 accepts input images and also accepts user specifications for features to be retained before and after style conversion. The input unit 31 supplies the server 12 with specification information, which is information indicating the features to be retained as specified by the user.

[0105] The storage unit 52 of the server 12 adds the features indicated by the specified information supplied from the user terminal 11 to a list of features to be retained before and after style conversion and stores them.

[0106] Figure 10 shows a second example of the screen displayed on the display unit 32.

[0107] In the example screen shown in Figure 10, the item "+ Add" is displayed at the end of the list of features to be retained.

[0108] When a user inputs an image into the image processing system 1, the input image and a list of features to be retained before and after style conversion are displayed on the screen. Next, the user can specify additional features to be retained by clicking the item indicated by "+ Add" in the list of features to be retained.

[0109] The features to be retained can be specified by the user by entering natural language text such as "person's skin color" or "uniform number," by entering information in CSV format such as "person's skin color, uniform number, apparent gender," or by selecting them using a tagging system.

[0110] When features to be retained are specified by inputting natural language text, users can flexibly specify the features to be retained. When features to be retained are specified by selecting tags, users can specify the features to be retained simply by toggling tags on or off.

[0111] Next, when the user clicks the start conversion button B1, a style conversion is performed on the input image, and the output image and the result of determining whether each of the features to be retained is actually retained before and after the style conversion are displayed on the screen.

[0112] As described above, users can specify the features they want to retain (or should retain) before and after style conversion.

[0113] - An example of highlighting image elements that have features not preserved before and after style conversion. Figure 11 shows a third example of the screen displayed on the display unit 32.

[0114] In the example screen shown in Figure 11, between the result of determining whether each feature to be preserved is actually preserved before and after the style conversion and the conversion start button B1, a modification button B2 is displayed to instruct the user to start modifying the converted image.

[0115] With the output image (converted image) and the results of the style conversion determining whether each feature to be retained is actually preserved before and after the style conversion displayed on the screen, the user selects the desired items from the list of features to be retained. The user selects items by, for example, clicking on the row displaying information about each item in the list, or by hovering the mouse over the row displaying information about each item.

[0116] When an item is selected, as shown in Figure 11, the image elements of the input image that possess the features selected by the user are displayed, for example, surrounded by a frame or bordered. In the example in Figure 11, the user selects the "skin color" item, which is determined to be not preserved before and after style conversion, and person A, who is shown on the left side of the input image, and person B, who is shown on the right side, are displayed surrounded by a dashed frame.

[0117] If an item is selected that is determined not to be retained before and after style conversion, the image elements in the input image whose selected features are not retained before and after style conversion will be highlighted. In the example in Figure 11, person A is highlighted, as indicated by the hatching.

[0118] When the user clicks the edit button B2 with the desired items selected from the list of features to be retained, the user's selected features are modified on the converted image, and the modified converted image is displayed on the screen as the output image.

[0119] In this way, users can check which image elements in the input image lack certain features, making it easier to determine whether or not those features need to be modified.

[0120] - Example of displaying the output history of output images: Figure 12 shows a fourth example of the screen displayed on the display unit 32.

[0121] In the example screen shown in Figure 12, the history of output images generated for a single input image is displayed at the bottom of the screen in chronological order.

[0122] When a user clicks on a desired output image from the output image history, the selected (clicked) output image is displayed in the center of the screen, and the result of the determination of whether the features that should be preserved for the selected output image are displayed on the right side of the screen.

[0123] If repeated modifications to the converted image cause the output image to deviate significantly from the image the user desired, the user can easily revert the output image to its original state.

[0124] - Example Figure 13, in which the user specifies the modifications to the converted image, is a diagram showing a fifth example of the screen displayed on the display unit 32.

[0125] In the example screen shown in Figure 13, the results of determining whether each feature to be retained is retained are displayed, along with suggested modifications for each feature to be retained.

[0126] The user can select and specify corrections from a list of options for features that were determined to be lost before and after style conversion. When the user clicks the correction button B2 with the corrections specified by the user, the corrections specified by the user are applied to the converted image, and the corrected converted image is displayed on the screen as the output image.

[0127] For example, if the user specifies a desired color for "skin color," the correction unit 55 performs segmentation to identify the skin area in the converted image and replaces the color information of that area with the color information of the color specified by the user. Also, for example, if the user specifies that a beard should be added, the correction unit 55 generates a partially corrected image to correct the face of the person in the output image using a prompt that includes the phrase "Please add a beard."

[0128] Suggestions for modifications to the features to be retained may be displayed for each image element in both the input and output images. Alternatively, the user-specified modifications may be applied to all image elements in the output image.

[0129] - An example of modifying the converted image before determining whether the features to be preserved before and after style conversion is maintained. Referring to the flowchart in Figure 14, further processing performed by the image processing system 1 will be explained. In the processing in Figure 14, the converted image is modified before determining whether the features to be preserved before and after style conversion are maintained.

[0130] The processing in steps S41 to S43 is the same as the processing in steps S1 to S3 in Figure 6.

[0131] After the processing in steps S41 and S43 is performed, in step S44, the correction unit 55 generates a partially corrected image. Here, the correction unit 55 extracts partial images from the input image that include image elements having features to be preserved, and generates a partially corrected image based on these partial images.

[0132] In step S45, the integration unit 56 modifies the converted image by integrating the converted image and the partially modified image.

[0133] In step S46, the determination unit 54 compares the input image with the modified converted image to determine whether or not the features that should be preserved before and after style conversion are preserved.

[0134] If it is determined in step S46 that the features that should be retained are not retained, the process returns to step S44, and the conversion image is corrected again.

[0135] On the other hand, if it is determined in step S46 that the features to be retained are retained, in step S47 the determination unit 54 outputs the converted image as an output image, and the display unit 32 displays the output image together with the determination result from the determination unit 54.

[0136] By modifying the converted image before determining whether the features to be retained are preserved before and after style conversion, the probability of determining that the features to be retained are not preserved during the initial determination can be reduced, thereby reducing the processing time.

[0137] - An example of using image processing system 1 for the purpose of increasing image anonymity: Instead of determining whether features that should be preserved are preserved before and after style conversion, the system may be configured to determine whether features that should not be preserved are preserved. This allows, for example, a user to obtain an image that has been style converted in such a way that it is impossible to tell who the people in the input image are.

[0138] For example, server 12 determines whether features (such as skin color, hair color, eye color, apparent gender, and apparent age) necessary to guarantee the identity of the person depicted in the input image are retained before and after style conversion. If a predetermined number of features are not retained, it outputs the converted image as the output image.

[0139] The process may involve a combination of determining whether features that should be preserved are retained before and after the style transformation, and determining whether features that should not be preserved are retained. The image processing system 1 can then perform a style transformation on the input image that reflects user requests, such as wanting to preserve skin color before and after the style transformation but wanting to change hair color.

[0140] As described above, in the image processing system 1, features to be retained may be the target of determination as to whether they are retained before and after style conversion, or features that should not be retained may be the target of determination.

[0141] Figure 15, an example in which the user determines whether or not the features to be preserved are preserved before and after style conversion, is a block diagram showing a third modified configuration of the image processing system 1. In Figure 15, the same components as those in Figure 4 are denoted by the same reference numerals. Repetitive explanations are omitted as appropriate.

[0142] The image processing system 1 in Figure 15 differs from the image processing system 1 in Figure 4 in that the server 12 is provided with a feedback acquisition unit 101 instead of a determination unit 54.

[0143] The feedback acquisition unit 101, like the determination unit 54, extracts feature information about features to be retained from the converted image supplied by the image conversion unit 53. The feedback acquisition unit 101 presents the feature information about features to be retained extracted from the input image and the feature information about features to be retained extracted from the converted image to the user by displaying them on the display unit 32 of the user terminal 11.

[0144] The user compares the feature information about the features to be retained extracted from the input image, displayed on the display unit 32, with the feature information about the features to be retained extracted from the converted image, to determine whether the features to be retained are retained before and after the style conversion.

[0145] The input unit 31 receives user feedback on the feature information about the features to be retained extracted from the input image and the feature information about the features to be retained extracted from the converted image. In other words, the input unit 31 receives input from the user indicating whether or not the features to be retained are retained before and after style conversion. The input unit 31 can also receive input on the reason why it was determined that the features to be retained are not retained.

[0146] The input unit 31 supplies the server 12 with the result of determining whether or not the features to be retained are retained before and after style conversion.

[0147] The feedback acquisition unit 101 acquires the result of the user's determination of whether the features to be retained are retained before and after style conversion (whether the features to be retained are retained in the converted image before and after correction). The feedback acquisition unit 101 supplies the converted image and the result of the user's determination of whether the features to be retained are retained in the converted image to the correction unit 55. If the user determines that the features to be retained are retained before and after style conversion, the determination unit 54 supplies the converted image supplied from the image conversion unit 53 or the corrected converted image supplied from the integration unit 56 to the display unit 32 as an output image.

[0148] The user who determines whether the features to be preserved are retained before and after style conversion, and the user who inputs the input image and obtains the output image, may be different people or the same person.

[0149] As described above, the user may be required to determine whether or not the features that should be preserved are retained before and after the style conversion.

[0150] • Example of computer configuration: The series of processes described above can be executed by hardware or by software. When the series of processes are executed by software, the programs that make up the software are installed from a program storage medium onto a computer that is built into dedicated hardware, or a general-purpose personal computer.

[0151] Figure 16 is a block diagram showing an example of the hardware configuration of a computer that executes the series of processes described above by a program.

[0152] The processing circuit 1001, ROM (Read Only Memory) 1002, and RAM (Random Access Memory) 1003 are interconnected by a bus 1004.

[0153] An input / output interface 1005 is further connected to the bus 1004. An input unit 1006 consisting of a keyboard, mouse, etc., and an output unit 1007 consisting of a display, speaker, etc., are connected to the input / output interface 1005. The input unit 1006 corresponds to the input unit 31 of the image processing system 1. The output unit 1007 corresponds to the display unit 32 of the image processing system 1. In addition, a storage unit 1008 consisting of a hard disk or non-volatile memory, a communication unit 1009 consisting of a network interface, etc., and a drive 1010 that drives removable media 1011 are connected to the input / output interface 1005. The storage unit 1008 corresponds to the storage unit 52 of the image processing system 1.

[0154] For example, when a computer functions as an image processing system 1 according to an embodiment of this technology, the processing circuit 1001 functions as a feature information extraction unit 51, an image conversion unit 53, a determination unit 54, an integration unit 56, a modification unit 55, and a feedback acquisition unit 101 by loading a program stored in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.

[0155] The program executed by the computer (processing circuit 1001) is, for example, recorded on removable media 1011, or provided via a wired or wireless transmission medium such as a local area network, the internet, or digital broadcasting, and installed in the storage unit 1008. The storage unit 1008 is not limited to being inside the computer, but may be located outside the computer. The processing circuit 1001 is an example of an integrated circuit, and CPUs (Central Processing Units), MPUs (Micro Processing Units), GPUs (Graphical Processing Units), APUs (Accelerated Processing Units), ASICs (Application Specific Integrated Circuits), and FPGAs (Field Programmable Gate Arrays) can all be considered integrated circuits.

[0156] The programs executed by the computer may be programs that are processed chronologically in the order described herein, or they may be programs that are processed in parallel or at necessary times, such as when a call is made.

[0157] In this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device containing multiple modules in one enclosure, are both considered systems.

[0158] The effects described herein are illustrative and not limited to those described herein, and other effects may also occur.

[0159] The embodiments of this technology are not limited to those described above, and various modifications are possible without departing from the spirit of this technology.

[0160] For example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.

[0161] Furthermore, each step described in the flowchart above can be performed by a single device, or it can be divided and performed by multiple devices.

[0162] Furthermore, if a single step includes multiple processes, those processes can be executed by a single device or shared among multiple devices.

[0163] Furthermore, among the processes described in the embodiments of this disclosure described above, all or part of the processes described as being performed automatically may be performed manually, or all or part of the processes described as being performed manually may be performed automatically by known methods. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above document and drawings may be changed at will unless otherwise specified. For example, the various information shown in each figure is not limited to the information shown.

[0164] Furthermore, each component of the illustrated device is a functional concept and does not necessarily have to be physically configured as shown. In other words, the specific forms of distribution and integration of each device are not limited to those shown, and all or part of them can be functionally or physically distributed and integrated in any unit according to various loads and usage conditions.

[0165] Furthermore, the embodiments of this disclosure described above can be combined as appropriate in areas that do not contradict the processing content. Also, the steps shown in the sequence diagram or flowchart of this embodiment can be changed in order as appropriate. For example, each step may be processed chronologically, repeatedly, or partially in parallel.

[0166] • Example configurations: This technology can also take the following configurations.

[0167] (1) An image processing system comprising: a conversion unit that generates a converted image by applying a style conversion to an input image; and a determination unit that obtains a first determination result of whether or not a feature to be determined among the features contained in the input image is retained before and after the style conversion. (2) The image processing system according to (1), wherein the conversion unit generates the converted image using an image generation AI model that applies the style conversion to the input image. (3) The image processing system according to (1) or (2), further comprising: a feature information extraction unit that extracts feature information about the feature to be determined from the input image and the converted image for each image element of the input image and the converted image, wherein the determination unit obtains the first determination result based on the feature information extracted from the input image and the converted image. (4) The image processing system according to (3), wherein the feature information extraction unit extracts the feature information from the input image and the converted image using a VLM that takes the input image or the converted image and query text as input and outputs a response text to the query text. (5) The image processing system according to (3) or (4), wherein the determination unit compares the feature information extracted from the input image with the feature information extracted from the converted image to determine whether the feature to be determined is retained before and after the style conversion. (6) The image processing system according to (5), wherein the determination unit determines whether the feature to be determined is retained before and after the style conversion based on the determination result of whether the feature information extracted from the input image and the feature information extracted from the converted image are similar, using LLM. (7) The image processing system according to (5), wherein the determination unit calculates the similarity between the feature information extracted from the input image and the feature information extracted from the converted image, and determines whether the feature to be determined is retained before and after the style conversion based on the similarity. (8) The image processing system according to (3), wherein the determination unit presents the feature information extracted from the input image and the feature information extracted from the converted image to the user and obtains the first determination result by the user.(9) The image processing system according to any one of (1) to (8) above, wherein the feature to be determined is at least one of the features to be retained before and after the style conversion and the features not to be retained. (10) The image processing system according to (9) above, further comprising a modification unit for modifying the converted image with respect to the feature to be determined. (11) The image processing system according to (10) above, wherein the modification unit modifies a region in the converted image that includes an image element having a feature to be retained which has been determined not to be retained before and after the style conversion, or modifies a region in the converted image that includes an image element having a feature not to be retained which has been determined to be retained before and after the style conversion. (12) The image processing system according to (11), wherein the modification unit cuts out a partial image from the input image that includes the image elements having features to be retained that are determined not to be retained before and after the style conversion, and integrates a partially modified image generated by applying the style conversion to the partial image cut out from the input image into the converted image, or cuts out a partial image from the input image that includes the image elements having features to be retained that are determined to be retained before and after the style conversion, and integrates a partially modified image generated by applying the style conversion to the partial image cut out from the input image into the converted image. (13) The image processing system according to (12), wherein the modification unit adjusts the content of the style conversion applied to the partial image based on the first determination result. (14) The image processing system according to (10), wherein the modification unit performs modifications to the converted image with content specified by the user. (15) The image processing system according to any one of (10) to (14), wherein the determination unit obtains a determination result of whether the feature to be determined is retained in the converted image corrected by the correction unit. (16) The image processing system according to any one of (9) to (15), wherein the determination unit outputs the converted image in at least one of the cases in which it is determined that the feature to be retained is retained before and after the style conversion, and in which it is determined that the feature that should not be retained is not retained before and after the style conversion.(17) The image processing system according to any one of (1) to (15), wherein the determination unit presents the first determination result and the converted image to the user. (18) The image processing system according to any one of (1) to (17), wherein the features to be determined include features specified by the user. (19) An image processing method comprising: generating a converted image by applying a style transformation to an input image; and obtaining a determination result of whether or not the features to be determined among the features included in the input image are retained before and after the style transformation. (20) A program for causing a computer to perform a process comprising: generating a converted image by applying a style transformation to an input image; and obtaining a determination result of whether or not the features to be determined among the features included in the input image are retained before and after the style transformation.

[0168] 1 Image processing system, 11 User terminal, 12 Server, 31 Input unit, 32 Display unit, 51 Feature information extraction unit, 52 Storage unit, 53 Image conversion unit, 54 Judgment unit, 55 Correction unit, 56 Integration unit, 101 Feedback acquisition unit

Claims

1. An image processing system comprising: a transformation unit that generates a transformed image by applying a style transformation to an input image; and a determination unit that acquires a first determination result of whether or not a feature to be determined among the features included in the input image is retained before and after the style transformation.

2. The image processing system according to claim 1, wherein the conversion unit generates the converted image using an image generation AI model that applies the style conversion to the input image.

3. The image processing system according to claim 1, further comprising a feature information extraction unit that extracts feature information about the feature to be determined from the input image and the converted image for each image element of the input image and the converted image, wherein the determination unit obtains the first determination result based on the feature information extracted from the input image and the converted image.

4. The image processing system according to claim 3, wherein the feature information extraction unit extracts the feature information from the input image and the converted image using a VLM that takes the input image or the converted image and query text as input and outputs a response text to the query text.

5. The image processing system according to claim 3, wherein the determination unit compares the feature information extracted from the input image with the feature information extracted from the converted image to determine whether or not the feature to be determined is retained before and after the style conversion.

6. The image processing system according to claim 5, wherein the determination unit determines whether the feature to be determined is retained before and after the style conversion, based on the determination result of whether the feature information extracted from the input image and the feature information extracted from the converted image are similar, using LLM.

7. The image processing system according to claim 5, wherein the determination unit calculates the similarity between the feature information extracted from the input image and the feature information extracted from the converted image, and determines whether the feature to be determined is retained before and after the style conversion based on the similarity.

8. The image processing system according to claim 3, wherein the determination unit presents the feature information extracted from the input image and the feature information extracted from the converted image to the user, and obtains the first determination result by the user.

9. The image processing system according to claim 1, wherein the feature to be determined is at least one of the features to be retained and the features not to be retained before and after the style conversion.

10. The image processing system according to claim 9, further comprising a correction unit for correcting the converted image with respect to the characteristics of the object to be determined.

11. The image processing system according to claim 10, wherein the modification unit modifies a region in the converted image that includes an image element having a feature to be retained that was determined not to be retained before and after the style conversion, or modifies a region in the converted image that includes an image element having a feature to be retained that was determined to be retained before and after the style conversion.

12. The image processing system according to claim 11, wherein the modification unit cuts out a partial image from the input image that includes the image elements having features to be retained that are determined not to be retained before and after the style conversion, and integrates the partially modified image generated by applying the style conversion to the partial image cut out from the input image into the converted image, or cuts out a partial image from the input image that includes the image elements having features not to be retained that are determined to be retained before and after the style conversion, and integrates the partially modified image generated by applying the style conversion to the partial image cut out from the input image into the converted image.

13. The image processing system according to claim 12, wherein the modification unit adjusts the content of the style conversion applied to the partial image based on the first determination result.

14. The image processing system according to claim 10, wherein the modification unit modifies the converted image with content specified by the user.

15. The image processing system according to claim 10, wherein the determination unit obtains a determination result of whether the features to be determined are retained in the converted image corrected by the correction unit.

16. The image processing system according to claim 8, wherein the determination unit determines that the features to be retained are retained before and after the style conversion, and determines that the features that should not be retained are not retained before and after the style conversion, in at least one of these cases.

17. The image processing system according to claim 1, wherein the determination unit presents the first determination result and the converted image to the user.

18. The image processing system according to claim 1, wherein the features to be determined include features specified by the user.

19. An image processing method comprising: applying a style transformation to an input image to generate a transformed image; and obtaining a determination result of whether or not a feature to be determined among the features contained in the input image is retained before and after the style transformation.

20. A program for causing a computer to perform a process that includes applying a style transformation to an input image to generate a transformed image, and obtaining a determination result of whether or not a feature to be judged among the features contained in the input image is retained before and after the style transformation.

Citation Information

Patent Citations

  • Training method of generative adversarial network, and image style conversion method and device

    CN112967180A

  • Computer program and image processing apparatus

    JP2024007789A

Cited By

  • Image development parameter adjustment system and method using style vector generation with a visual language model

    JP7903149B1