Information processing systems, information processing methods, and programs
The information processing system addresses the issue of partial learning in generative models by integrating LPIPS and CL loss functions for enhanced disruptive effects on semantic features while maintaining visual quality, effectively preventing model theft and copyright infringement.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2026-03-03
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional methods for adversarial degradation of images in generative models fail to effectively disrupt high-performance pre-trained feature extractors, allowing the generation model to partially learn the original image's meaning, and do not adequately maintain visual quality.
An information processing system that integrates multiple loss functions, including LPIPS and CL, to enhance disruptive effects on semantic features while maintaining structural similarity, using multi-objective optimization and perturbation techniques to maximize feature distances and visual quality.
The system effectively disrupts the learning of original image meaning in generative models while preserving visual quality, enhancing the disruptive effect on semantic features and maintaining structural similarity.
Smart Images

Figure 0007843100000001_ABST
Abstract
Description
[Technical Field]
[0001] This invention relates to an information processing system, an information processing method, and a program. [Background technology]
[0002] With the aim of hindering the additional training of diffusion models such as LoRA, PGD (Pr Applying the (jected gradient descent) attack to pixel-level MSE (Mean Squared Error) Methods have been proposed that primarily use small-amplitude perturbation constraints as the loss (see Non-Patent Document 1). [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] Zheng, B., Liang, C., Wu, X., Liu, Y.: "Understanding and Improving Adversarial Attacks on Latent Diffusion Model", arXiv:2310.04687 (2023). [Overview of the project] [Problems that the invention aims to solve]
[0004] However, in methods such as those proposed in Non-Patent Document 1, the generative model is still It can learn from the original image.
[0005] This invention was made in view of the above background, and makes it easier to perform additional training on the learning model. The objective is to provide technology that can take these factors into consideration. [Means for solving the problem]
[0006] The main invention of the present invention for solving the above problems is an information processing system, which includes a first image acquisition unit that acquires an image and a second image generated by applying a perturbation to the first image, and a feature amount acquisition unit that acquires a first feature amount from the first image and a second feature amount from the second image, and an evaluation value determination unit that determines an evaluation value of the second image such that the greater the distance between the first and second feature amounts, the higher the evaluation value. It is characterized by comprising the above.
[0007] Regarding other problems disclosed in the present application and their solutions, they will be made clearer in the column of the embodiments of the invention and the drawings. More clearly described.
Advantages of the Invention
[0008] According to the present invention, it is possible to consider the difficulty of additional learning of the learning model.
Brief Description of the Drawings
Modes for Carrying Out the Invention
[0010] <Summary of the Invention> Image generation by diffusion models is used in a variety of fields such as content production and design support. On the other hand, new risks have emerged, such as secondary use of generated images, copyright infringement, and illegal learning (model theft) that exploits generated content. Therefore, adversarial degradation techniques that pre-process images that "should not be learned" for generative AI have attracted attention. As a representative method, there is the Mist system, which has been shown to significantly interfere with personalized learning using LoRA (Low-Rank Adaptation), DreamBooth, etc. by adding minute perturbations to the original image. However, conventional Mist mainly focuses on maintaining a high SSIM (Structural Similarity Index Measure) and not degrading human visual quality. Therefore, it cannot significantly disrupt recent high-performance pre-trained feature extractors such as CLIP (Contrastive Language-Image Pre-training), ViT (Vision Transformer), Swin Transformer, ResNet (Residual Network), etc., and there remains a problem that the generation model still partially learns the meaning of the original image. In addition, in this embodiment, the "generation model" refers to a deep learning model that receives at least one of an image or a latent variable (noise vector, text prompt, etc.) as input and generates / modifies a still image or a moving image. Specific examples include diffusion models (Stable Diffusi on Model).
[0011] As a representative method, there is the Mist system, which can significantly interfere with personalized learning using LoRA (Low-Rank Adaptation), DreamBooth, etc. by adding minute perturbations to the original image. It has been shown. However, conventional Mist mainly focuses on maintaining a high SSIM (Structural Similarity Index Measure) and not degrading human visual quality. Therefore, it cannot significantly disrupt recent high-performance pre-trained feature extractors such as CLIP (Contrastive Language-Image Pre-training), ViT (Vision Transformer), Swin Transformer, ResNet (Residual Network), etc., and there remains a problem that the generation model still partially learns the meaning of the original image. In addition, in this embodiment, the "generation model" refers to a deep learning model that receives at least one of an image or a latent variable (noise vector, text prompt, etc.) as input and generates / modifies a still image or a moving image. Specific examples include diffusion models (Stable Diffusi on Model), etc. That is, there remains a problem that the generation model still partially learns the meaning of the original image.
[0012] In addition, in this embodiment, the "generation model" refers to a deep learning model that receives at least one of an image or a latent variable (noise vector, text prompt, etc.) as input and generates / modifies a still image or a moving image. Specifically, it refers to a deep learning model that receives at least one of an image or a latent variable (noise vector, text prompt, etc.) as input and generates / modifies a still image or a moving image. Specific examples include diffusion models (Stable Diffusi Examples include GANs (such as o-on), GANs (such as StyleGAN), and VAEs (such as β-VAE).
[0013] In the information processing system of this embodiment, the loss function of Mist is considered from the perspective of multi-objective optimization. The entire system has been redesigned, and LPIPS (Learned Perceptual Image Patch Similarity) and CL have been implemented. While simultaneously maximizing the diverse feature distances of IP, ViT, Swin, and ResNet, SS The IM hinge loss is used to maintain structural similarity. This is the information processing system of this embodiment. According to the information, in addition to the conventional MSE (Mean Squared Error) function, the LPIPS function and CLI Loss functions for multiple feature distance losses and regularization terms using P, ViT, Swin, and ResNet By combining these, it is possible to enhance the disruptive effect on the semantic features that the generative model learns. According to the information processing system of this embodiment, it is possible to maintain a high level of structural similarity. Therefore, by introducing hinge loss, it is possible to achieve both the maintenance of visual quality and the effectiveness of the attack.
[0014] <First Embodiment> The following describes an information processing system according to one embodiment of the present invention. The information processing system uses a computer to process still images or moving images (hereinafter collectively referred to as "images") In generating or modifying the ) multiple deep feature distances are integrated with weight normalization. We will implement a generative parameter optimization method that uses a loss function.
[0015] Figure 1 is a diagram showing an example of the overall configuration of an information processing system according to one embodiment of the present invention. The information processing system of this embodiment is configured to include a management server 2. The management server 2 is a user The terminal 1 is connected to the communication network 3 so that it can communicate. For example, the internet, public telephone networks, mobile phone networks, wireless communication channels, etc. Built by Sanet (registered trademark), etc.
[0016] User terminal 1 is a computer operated by the user. User terminal 1 is, for example, This can refer to smartphones, tablet computers, personal computers, etc. ru.
[0017] Management Server 2 is a general-purpose computer such as a workstation or personal computer. It can be done as a computer, or logically implemented through cloud computing. It may be revealed.
[0018] <Management Server 2> Figure 2 shows an example of the hardware configuration of management server 2. Note that the illustrated configuration is This is just one example, and other configurations are possible. Management Server 2 has CPU 201, memory Re 202, storage device 203, communication interface 204, input device 205, output device 20 It includes 6. The storage device 203 stores various data and programs, for example, hard disk These include disk drives, solid-state drives, and flash memory. Face 204 is an interface for connecting to a communication network, for example, Adapters for connecting to the Internet Network (registered trademark), and for connecting to the public telephone network. DM, wireless communication device for wireless communication, USB (Universal Serial) for serial communication These include AL Bus connectors and RS232C connectors. Input device 205 inputs data. These include, for example, keyboards, mice, touch panels, buttons, and microphones. The output device 206 outputs data to, for example, a display, printer, speaker, etc. Yes. Furthermore, each functional unit of the management server 2, described later, is stored in the storage device 203 by the CPU 201. This is achieved by reading the program into memory 202 and executing it, and the management server Each storage unit in B2 is implemented as part of the storage area provided by memory 202 and storage device 203. It will be done.
[0019] Figure 3 shows an example of the software configuration of the management server 2. The management server 2 is image capture The acquisition unit 211, the image generation unit 212, the feature acquisition unit 213, the similarity calculation unit 214, and evaluation It comprises a value determination unit 215 and a storage unit 231.
[0020] The storage unit 231 is the storage device of the management server 2. The storage unit 231 is accessed from the user terminal 1. The first image received, the second image generated, the extracted features, the calculated similarity, and the resolution. It stores predetermined evaluation values, etc. The memory unit 231 also stores model data of multiple feature extractors. To remember.
[0021] The image acquisition unit 211 acquires the first image and the second image generated by perturbing the first image. The image acquisition unit 211 receives the image from the user terminal 1 via the communication interface 204. The first image is received. The image acquisition unit 211 records the received first image in the storage unit 231. To remember.
[0022] The image generation unit 212 generates a second image by perturbing the first image. Image 212 can generate a second image by injecting stochastic noise. The image generation unit 212 applies Gaussian noise and Uniphony to each pixel value of the first image. Perturbation is introduced by adding probabilistic noise, such as frame noise.
[0023] The image generation unit 212 injects diffuse noise into the latent spatial representation of the generation model, and the latent spatial representation It is also possible to decode and generate a second image. In this case, the image generation unit 212 first Image 1 is converted into a latent space representation by the encoder of the generation model, and this latent space representation is expanded Scatter noise is injected. Subsequently, the image generation unit 212 generates a latent space representation into which noise has been injected. The decoder decodes the image and generates a second image.
[0024] The feature acquisition unit 213 acquires a first feature from the first image and a second feature from the second image. The feature acquisition unit 213 acquires features using multiple different feature extractors. Specifically, the feature acquisition unit 213 uses CLIP-ViT / 32 and ViT-Base. Select two or more from Swin-Tiny and ResNet-50 to extract features. ru.
[0025] The feature acquisition unit 213 performs a weighted average of the acquired features and normalizes them. For example, feature acquisition When section 213 uses CLIP-ViT / 32 and ResNet-50, each feature extractor The obtained feature vectors are multiplied by pre-set weights to calculate a weighted average. Afterward, the feature acquisition unit 213 normalizes the obtained feature vectors using L2 normalization.
[0026] The similarity calculation unit 214 calculates the similarity between the first and second images. The similarity is calculated using SSIM (Structural Similarity Index Measure). SIM evaluates the similarity between images by considering three elements: brightness, contrast, and structure. This is an indicator. The similarity calculation unit 214 divides the first image and the second image into small regions. Calculate the SSIM value for each domain, and use the average value across all domains as the final similarity score. .
[0027] The evaluation value determination unit 215 determines the value such that the value increases as the distance between the first and second feature quantities increases. The evaluation value of image 2 is determined. The evaluation value determination unit 215 determines the evaluation value of the first feature and the second feature. Calculate the Euclidean or cosine distance between the points, and determine the evaluation value based on this distance. The greater the distance, the greater the difference in the feature space between the first and second images. This means that the evaluation value determination unit 215 determines a higher evaluation value.
[0028] The evaluation value determination unit 215 determines the evaluation value such that it becomes higher as the similarity and distance increase. It can also be determined. In this case, the evaluation value determination unit 215 is calculated by the similarity calculation unit 214. The evaluation value is calculated by combining the calculated similarity and the distance between features. Specifically, the evaluation The value determination unit 215 calculates the product of similarity and distance as the evaluation value, or the product of similarity and distance. Each element is weighted, and the weighted sum is calculated as the evaluation value.
[0029] The evaluation value determination unit 215 adds a penalty to the evaluation value if the similarity is below a predetermined value. The predetermined value is set to a value such as 0.5 or 0.7. This means that the visual difference between the first and second images is too large, thus determining the evaluation value. Section 215 subtracts a predetermined value from the evaluation value or multiplies the evaluation value by a predetermined coefficient to apply a penalty. Add it.
[0030] Figure 4 is a diagram illustrating the processing flow in an information processing system.
[0031] The image acquisition unit 211 receives the first image from the user terminal 1 (S101). Image generation Unit 212 generates a second image by perturbing the first image (S102). Feature extraction The extraction unit 213 extracts a first feature quantity from the first image and a second feature quantity from the second image. Output (S103). The similarity calculation unit 214 calculates the similarity between the first image and the second image using SSI. The value is calculated by M (S104). The evaluation value determination unit 215 determines the value between the first feature and the second feature. The distance is calculated (S105). The evaluation value determination unit 215 determines the similarity based on the calculated distance. Then the evaluation value of the second image is determined (S106).
[0032] As described above, according to the information processing system of this embodiment, perturbation is applied to the original image. Generated images, and the distance between the features extracted using multiple feature extractors and SSIM By combining similarity measures, it is possible to appropriately evaluate the features of an image.
[0033] The above embodiments have been described, but the above embodiments are intended to facilitate understanding of the present invention. This is intended to be a limitation of the present invention. The present invention can be modified and improved without being removed, and its equivalents are also included.
[0034] <Second Embodiment> The second embodiment will be described below. In the second embodiment, the original image (first image) is used In generating the generated image with added motion (the second image), the features of the original image and the generated image are used. The distance between the two images should be as large as possible, and the similarity between the original image and the generated image should be as large as possible. To impose a perturbation.
[0035] (1) Overall structure Figure 5 is a block diagram showing the overall flow of the image generation system 10 according to the second embodiment. The system 10 includes a generation parameter setting unit 11, a loss function calculation unit 12, and an optimization unit 13. The main components are an image storage unit 14 and an image output unit 15. Each part is composed of a CPU, GPU, etc. This has been done and can be implemented using Python / PyTorch.
[0036] (2) Generation parameter setting unit 11 The generation parameter setting unit 11 receives the original image 31, noise vector 32, and text prompt. The generative model 16 receives at least one type of input such as 33, and as shown by arrow a in Figure 5. Determines the initial parameters to pass to Stable Diffusion (e.g., StyleGAN2). If no input is provided... Random initial values may be used for this.
[0037] (3) Loss function calculation unit 12 3-1 Visual Similarity Term Figure 6 is a block diagram showing the structure of the loss function. The visual similarity term 5 is calculated by comparing image x and the generated image x. Implement the structural similarity index SSIM (Equation 1) or LPIPS for image x'. Coefficient w vis teeth It is for weight adjustment. TIFF0007843100000002.tif18165...(Formula 1)
[0038] 3-2 Feature Distance Term The feature distance term 6 in the same figure is processed by the two-stage weight normalization block 8 shown in Figure 7. First, C LIP-ViT / 32, ViT-Base, Swin-Tiny, ResNet-50 Using two or more feature extractors 21a to 21d selected from four types, the distance d k Calculate each d. k w is the base weighting coefficient kMultiply and add, normalize by the sum, and then obtain the dynamic coefficient l k and scaling Multiply by the scaling coefficient α to obtain the two-stage weighted value shown in (Equation 2). TIFF0007843100000003.tif1874···(Equation 2)
[0039] 3-3 Conditional term Conditional term 7 is the hinge SSIM function that adds a penalty only when the SSIM value is less than the threshold μ (Equation 3), and plays a role in guaranteeing the lower limit of visual quality. TIFF0007,843,100,000,004.tif7115···(Equation 3)
[0040] <Total loss function> Using the above visual similarity term 5, feature distance term 6, and conditional term 7, in this embodiment, (Equation 4 ) defines the total loss function L. Here, w LPIPS is the real number coefficient multiplied by the LPIPS distance number, and λ is the real number coefficient multiplied by the regularization term R(δ). TIFF0007,{843,100,000,005}.tif8132···(Equation 4) L feat : Two-stage weighted feature distance term defined in (Equation 2). d LPIPS : Perceptual patch similarity based on VGG16, etc. Lcond: SSIM hinge loss (Equation 3). R(δ): Regularization term representing perturbation spectrum smoothing and potential stabilization. In this embodiment uses the L2 norm, but other regularizations may also be used.
[0041] The loss function (Equation 4) also becomes a configuration that does not use LPIPS if w LIPS is set to 0 even when LPIPS is included .
[0042] (4) Optimization unit 13 4-1 Multi-stage optimization As schematically shown in Figure 8, the optimization unit 13 divides the epoch into multiple stages: Phase 1, 2, ... The process is divided, and the coefficients or learning rates of each loss term (a)-(c) are switched at each stage. Phase 1 In Phase 1, we prioritize SSIM distance; in Phase 2, we prioritize CLIP distance; and in Phase 3, we prioritize multiple feature distances. It is possible to shift the weight of factors, such as prioritizing certain aspects.
[0043] 4-2 PGD iteration + projection clip Within each phase, the PGD algorithm is used to calculate the gradient ∇δ L Update perturbation δ based on this. After the update, δ is restricted to within ∥δ∥∞≦ε by projection clipping process 9.
[0044] 4-3 Noise injection Frequency noise: High-frequency components are extracted using FFT at even steps. f Superimpose Gaussian noise Then, the reverse transform is used to return to image space. Latent noise: σ is added to the VAE latent z at every odd step. z Add spreading noise and re-decode Continue updating.
[0045] (5) Image storage unit 14 The optimization unit 13 sends the generated images to the image storage unit 14 with each iteration and stores them in list format. The saving interval is every step or at a predetermined period. The saved images are used in subsequent sorting processes.
[0046] (6) Image output unit 15 The image output unit 15 selects the image with the largest (or smallest) SSIM or LPIPS from among the saved images. Select the desired image and provide it externally as the final output 20. Output formats are PNG and JPEG. G and others can be used.
[0047] (7) Hardware implementation example The processing unit is equipped with an NVIDIA® RTX-A6000 (48GB VRAM). It can be implemented on a workstation and requires PyTorch 2.1 and CUDA 12.1. It can be used to process 512x512px images.
[0048] <Example 1: PGD-based image degradation> 1. Structure Input: Original image x • Constraint: Perturbation δ is applied within a given l∞ norm range. Loss function: (Equation 4) 2. Optimization Procedure (1) ∇δ with PGD iteration L Based on this, δ is updated and kept within constraints with a projection clip. (2) Each iteration • Frequency noise injection: A small amount of complex Gaussian noise is added to the high-frequency components of the FFT. • Latent noise injection: Diffuse noise is added to the VAE latent representation, and then it is re-decoded. (3) Save the generated images and output the one with the highest visual similarity.
[0049] <Example 2: Training a GAN Generator> 1. Structure • Generating network: Any GAN generator • Discriminant: Conventional adversarial loss • Replace generator loss: adversarial loss with (Equation 4) 2. Weight adjustment • At the end of each epoch, the weight coefficients in (Equation 4) are Bayesian optimized and applied to the next epoch. 3. The aim of the effect • Obtain generated images with expanded deep feature distance while maintaining visual quality.
[0050] <Example 3: Latent Optimization of Diffusion Models> 1. Structure • Base: Text-conditional diffusion model • Initial generation: Obtain latent Z0 via Text-to-Image • Loss function: Use (Equation 4). 2. Optimization Procedure (1) Latent z is used as the learning variable and updated 30 steps by backpropagation. (2) The weights are linearly shifted so that the first half prioritizes SSIM and the second half prioritizes CLIP distance. 3. Output • Decode the updated latent and expand the semantic distance while maintaining prompt fidelity. Obtain the resulting image.
[0051] <Example 4: Application to Super-Resolution Models (SR)> 1. Structure • Generative model: A general super-resolution network (e.g., EDSR, SwinIR, etc.) Loss function: (Equation 4) • Visual similarity term: LPIPS • Feature distance term: CLIP and Swin system distances • Condition: SSIM hinge (threshold μ = 0.92) 2. Learning Procedure • A high-resolution-low-resolution pair is used as input, and the generator loss is replaced with (Equation 4). • Still learning lol CLIP The semantic distance is gradually increased, emphasizing it step by step. 3. Advantages • It becomes possible to generate super-resolution images that are difficult to use for plagiarism learning while maintaining visual quality.
[0052] <Example 5: Application to Style Transfer (CycleGAN-based)> 1. Structure • Conversion Network: Two Generators for Domain A→B and B→A ·Generator loss: adversarial loss + cycle - consistency loss + (Formula 4) • Feature distance term: Uses the CLIP+ResNet distance after weighting and normalization. 2. Learning Procedure (1) Conventional learning is performed using only loss for a predetermined number of epochs. (2) Then, (Equation 4) is added in fine-tuning and short-term additional training is performed. 3. Utilization • Achieves style transformation that is natural to humans but difficult for AI to learn from the original domain.
[0053] <Example 6: Robustification of a low-light image enhancement model> 1. Structure • Basenet: Retinex-type dual-branch emphasized model • Attacker: This algorithm generates degraded images while maintaining visual quality. • Defender: Adversarial training is performed on the collaborative network. 2. Procedure (1) Regularly learn the emphasis network →(2) Generate adversarial images using this method →(3) The model is made more robust by feeding adversarial images into the training process. 3. Advantages • Nighttime surveillance footage can be used to improve attack resistance while maintaining visual quality.
[0054] <Example 7: Application to Audio Spectrogram Generation> 1. Structure • Replace the image generation model with the "MelSpectrogram Image Generation Network". This loss applies. • Visual similarity term: Log-STFT SSIM. • Feature distance term: CLIP-audio embedding distance + CNN feature distance. 2. Procedure • Voice commands are converted into Mel spectrogram images and optimized using this method. • The optimized spectrogram is decoded using Griffin-Lim. Obtain hostile voices. 3.Significance • It can generate speech that disrupts speech recognition while remaining natural to the ear.
[0055] <Example 8: Application to reinforcement learning observation frames> 1. Environment and Agents A reinforcement learning agent that takes an observed frame (image) as input and outputs an action. (Examples: DQN, PPO, etc.) • The final hidden layer features of the agent's policy network are used as a feature extractor. 2. How to apply the loss function • Observation frame x for each time step t t The total loss L is calculated using (Equation 4). ·Visual similarity term: SSIM, conditional term: SSIM hinge. • Feature distance term: Multiple feature distances within the policy net are integrated and normalized. 3. Perturbation Optimization Procedure • Gradient ∇w LPIPS δ L perturbation δ t Update and restrict ||δt|| with projection clipping. Repeat the above steps while minimizing visual changes and maintaining the agent's behavior policy. Generates an observation sequence that changes [something].
[0056] <Disclosure Items> Furthermore, this disclosure also includes the following configurations. [Item 1] Image for obtaining the first image and the second image generated by perturbing the first image. Acquisition section, A first feature is obtained from the first image, and a second feature is obtained from the second image. A feature acquisition unit, The second image is evaluated such that the value increases as the distance between the first and second features increases. An evaluation value determination unit that determines the value, An information processing system equipped with the following features. [Item 2] The system includes a similarity calculation unit that calculates the similarity between the first and second images, The evaluation value determination unit determines the evaluation value such that the value increases as the similarity is high and the distance is large. Determine the evaluation value. The information processing system described in item 1. [Item 3] The evaluation value determination unit imposes a penalty on the evaluation value if the similarity is below a predetermined value. The information processing system described in item 2 is added. [Item 4] The feature acquisition unit acquires the feature quantities using multiple different feature extractors, and before acquiring them The information processing system described in item 1, which performs a weighted average of the listed features and normalizes them. [Item 5] The image acquisition unit acquires the first image, The system includes an image generation unit that generates a second image by perturbing the first image, The information processing system described in item 1. [Item 6] The image generation unit is an information processing system according to item 5, which injects probabilistic noise. [Item 7] The image generation unit injects diffuse noise into the latent space representation of the generation model, and the latent space representation An information processing system as described in item 5, which decodes the present and generates the second image. [Item 8] The process involves obtaining the first image and the second image generated by perturbing the first image. Top and, A first feature is obtained from the first image, and a second feature is obtained from the second image. The steps, The second image is evaluated such that the value increases as the distance between the first and second features increases. The step of determining the value, A method of information processing performed by a computer. [Item 9] The process involves obtaining the first image and the second image generated by perturbing the first image. Top and, A first feature is obtained from the first image, and a second feature is obtained from the second image. The steps, The second image is evaluated such that the value increases as the distance between the first and second features increases. The step of determining the value, A program that causes a computer to execute something.
[0057] <Disclosure Item 2> Furthermore, this disclosure also includes the following configurations. [Item 1] The system receives at least one generation condition for inputs such as the original image, noise vector, and text prompt. The process involves initializing the parameters for the input and pre-set settings, and calculating the loss function. The process involves a function calculation step and an optimization step where parameters are updated one or more times to generate and save an image. The system includes an output step that outputs at least one of the saved generated images, and each step is performed sequentially. In an image generation method that uses an image generation model that executes and outputs a still image, In the loss function calculation process, A visual similarity term is a loss term based on a predetermined visual similarity index, or The distances between features obtained by multiple different feature extractors are weighted and averaged, and then normalized. The feature distance term is a loss term, or If the value of the visual similarity index falls below a predetermined threshold, the entire loss function is increased. Among the conditional terms, terms that contain two or more It has a function to perform calculations, In the optimization process, The generation parameters are adjusted once or multiple times so that the total loss function increases according to a predetermined criterion. The process of updating and saving the generated image obtained during or as a result of the update process. Image generation method characterized by having [Item 2] In the image generation method described in item 1, The aforementioned feature distance term is weighted by assigning coefficients to each of two or more distinct feature distances. The calculation includes at least one step of performing an average and normalizing the weighted average by the sum of the coefficients. An image generation method characterized by the output of an image. [Item 3] In the image generation method described in item 1 or item 2, The image generation method is characterized in that the feature distance term includes one or more of the following: (i) Includes a loss based on the structural similarity index as a visual similarity term. (ii) The visual similarity term includes perceptual patch similarity based on a trained network. (iii) If the value of the visual similarity index is less than a predetermined threshold, the hinge function And impose a penalty. (iv) In the feature distance term, after the weighted average using the real-valued base weight coefficient w_k, dynamic This involves a weighting operation with two or more stages, where the weight coefficient l_k is multiplied, and then the scaling coefficient is multiplied again. . [Item 4] In the image generation method described in any one of items 1 to 3, The feature extractor comprises two or more systems, at least (i) Natural language—image-trained model, (ii) Patch-based visual transformer, (iii) Convolutional neural networks, It is characterized by including at least one feature extractor belonging to two or more of the three systems. A method for generating images. [Item 5] In the image generation method described in any one of items 1 to 4, The optimization process described above is (a) Divided into multiple stages, the configuration of the total loss function or the weight coefficient or learning rate for each stage Switch and update the generation parameters. or (b) Performed by projection-based iterative optimization, where the perturbation amount is measured in each iteration. Clip within the specified range. An image generation method characterized by including either one or both of the following. [Item 6] In the image generation method described in any one of items 1 to 5, The optimization process is characterized by including at least one of the following in at least a part of it: A method for generating images. (i) Apply the Fast Fourier Transform and Discrete Cosine Transform to the generated image or its intermediate representation. Inject stochastic noise via conversion or wavelet transform. (ii) After injecting stochastic noise or spreading noise into the latent space representation of the generative model, The latent spatial representation is decoded and the generated image is updated. (iii) The generated image from among the saved generated images that has the maximum visual similarity term Select and output. (iv) The visual similarity term includes both SSIM and LPIPS, and the condition term is Threshold determination is performed by referring only to the SSIM value. (v) The generation parameters include a noise vector or a diffusion time. [Item 7] A processing unit that executes the image generation method described in any one of items 1 to 6, and the processing unit that actually An image generation device comprising a storage unit that stores a program to be executed. [Item 8] The computer is instructed to perform one of the image generation methods described in item 1 to 6. Program. [Explanation of Symbols]
[0058] 1 User terminal 2 Management Server
Claims
1. An image acquisition unit that acquires a first image, An image generation unit that generates a second image by perturbing the first image, the image generation unit that injects diffuse noise into the latent spatial representation of the generation model, decodes the latent spatial representation, and generates the second image, A feature acquisition unit that acquires a first feature from the first image and a second feature from the second image, An evaluation value determination unit that determines the evaluation value of the second image such that it increases as the distance between the first and second feature quantities increases, An information processing system equipped with the following features.
2. The system includes a similarity calculation unit that calculates the similarity between the first and second images, The evaluation value determination unit determines the evaluation value such that it increases as the similarity and distance increase. The information processing system according to claim 1.
3. The information processing system according to claim 2, wherein the evaluation value determination unit applies a penalty to the evaluation value when the similarity is less than or equal to a predetermined value.
4. The information processing system according to claim 1, wherein the feature acquisition unit acquires the feature quantities using a plurality of different feature extractors, and normalizes the acquired feature quantities by weighting and averaging them.
5. The information processing system according to claim 1, wherein the image generation unit injects probabilistic noise.
6. The first step is to obtain the first image, A step of generating a second image by perturbing the first image, comprising injecting diffuse noise into the latent spatial representation of the generation model, decoding the latent spatial representation to generate the second image, The steps include obtaining a first feature from the first image and obtaining a second feature from the second image, A step of determining the evaluation value of the second image such that it increases as the distance between the first and second feature quantities increases, A method of information processing performed by a computer.
7. The first step is to obtain the first image, A step of generating a second image by perturbing the first image, comprising injecting diffuse noise into the latent spatial representation of the generation model, decoding the latent spatial representation to generate the second image, The steps include obtaining a first feature from the first image and obtaining a second feature from the second image, A step of determining the evaluation value of the second image such that it increases as the distance between the first and second feature quantities increases, A program that causes a computer to execute something.
Citation Information
Patent Citations
Information processing device, server, information processing method, and program
JP2025120878A
Learning device, facial recognition system, learning method, and recording medium
WO2021144857A1
Evaluation device, evaluation method, and evaluation program
WO2024157404A1