A measurement method and apparatus for the memory problem based on a reverse text-image diffusion model.

By constructing an inverse text-generated image diffusion model and optimizing the objective function using noise distribution and prompt text distribution, the randomness detection of the diffusion model memory problem is solved, enabling accurate measurement and risk auditing of arbitrary images and improving the security of image generation.

CN119478094BActive Publication Date: 2025-10-28ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411566591.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-10-28
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

Existing technologies struggle to comprehensively and accurately audit the memory problem in diffusion models, especially during image generation, where randomness has a significant impact, making it impossible to continuously and accurately quantify the memory risk of any single image.

Method used

By constructing a memory problem measurement method based on an inverse text-to-image diffusion model, this method utilizes noise distribution and cue text distribution to define sensitivity and memory problem measurement index functions, optimizes the weighted objective function, and outputs the optimal cue text and noise distribution, thereby achieving memory problem measurement for any image.

Benefits of technology

It achieves continuous and accurate measurement of memory risk in diffusion models, reduces the impact of randomness, provides quantitative results of memory risk in the image generation process, and comprehensively audits the security of image generation models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119478094B_ABST
    Figure CN119478094B_ABST
Patent Text Reader

Abstract

This invention provides a method for measuring memory problems based on a reverse-engineered image diffusion model. Compared to existing methods, this method constructs a weighted objective function based on a pre-trained image diffusion model and an arbitrary image, and inversely outputs a prompt text distribution and a noise distribution. This solves the randomness problem inherent in previous methods that assessed memory problems by detecting similar content in randomly generated images. By defining the sensitivity of the noise distribution, a metric function for measuring memory problems is constructed, enabling continuous and accurate measurement of the memory problem of the image generation model. It provides a quantitative result of the memory risk for any image, thus auditing the memory risk of the image generation model. This invention also provides a device for measuring memory problems based on a reverse-engineered image diffusion model, realizing the method for measuring image memory problems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method and apparatus for measuring the memory problem of a reverse text-based image diffusion model. Background Technology

[0002] Diffusion models have demonstrated exceptional capabilities in image, video, and 3D model generation, leading to a plethora of AIGC (Artificial Intelligence Generated Content) commercial systems and image generation tools. For example, tools like Stable Diffusion, Midjourney, DALLE, and Imagen have attracted millions of users. Diffusion models are a data generation theory and method that, combined with deep learning techniques, can be used to fit real-world data, enabling the generation of various data types, including images, videos, and music. Diffusion model generation can also incorporate conditions for conditional generation, such as text-to-image generation, where an image is generated based on a given text description. Compared to other image synthesis techniques, diffusion models offer stable training, strong scalability, and higher quality generated content.

[0003] However, under certain conditions, diffusion models frequently generate certain images, figuratively described as the image generation model "memorizing" these images, also known as the memory problem of image generation models. This memory problem contradicts the expectations of image generation models; developers do not want models to rigidly memorize certain data, but rather expect them to generate diverse data. Furthermore, if the memory problem occurs with data involving copyright, privacy, or other issues, it will pose potential security risks. Therefore, it is necessary to audit diffusion models for the memory problem to comprehensively and accurately analyze their memory risks. For example, invention application CN117633899 A discloses a method, system, and apparatus for privacy protection of a mask-based face image generation model. This method obtains the attention level of facial features through a face image generation model, and trains a deepfake face generation model based on the attention level of facial features to protect the privacy of face datasets and reduce face data leakage.

[0004] Text-generated image diffusion models also carry the risk of training data memory issues. Specifically, when a certain prompt text is input, the model may generate an image that is very similar to a training image. For example, Reference 1 (Carlini N, Hayes J, Nasr M, et al., Extracting training data from diffusion models, 32nd USENIX Security Symposium (USENIX Security 23), 2023, 5253-5270.) extracts training data based on a text-generated image diffusion model. The input prompt text is used to detect the data memory problem and reduce privacy leaks during image generation. Specifically, a large amount of image description text is collected from the training set of the text-generated image diffusion model. Then, each description text is input into the diffusion model, and different random noise is used to generate a large number of generated images. For each descriptive text, several images (500 images) are generated and detected to determine if these generated images are identical. If more than 10 identical images are generated, the descriptive text is considered to pose a memory risk. Finally, it is determined whether these identical images are the images corresponding to the descriptive text in the training set. The technique proposed in Reference 2 (Somepalli G, Singlea V, Goldblum M, et al., Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models, Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2023, 6048-6058.) is similar to that in Reference 1. The difference is that Reference 1 uses a pre-trained image retrieval model SSCD to detect whether there are images in the generated images that are very similar to the training images.

[0005] These works can address the memory risk of image generation models to some extent and reduce privacy leaks during image generation. However, they are all based on textual analysis techniques, which generate images randomly and then detect whether there is content similar to the training images, thus being affected by the randomness of the generation process. In addition, existing technologies cannot be used to continuously and accurately quantify the memory risk of any single image. Therefore, a comprehensive and accurate audit of the memory risk of diffusion models is an urgent problem to be solved to protect the security of image generation data. Summary of the Invention

[0006] The purpose of this invention is to provide a method and device for measuring memory problems based on the diffusion model of text-generated images. This method and device, developed based on the distribution of prompt text and noise, are used to analyze any image and determine the memory problems generated by the image. It also provides a measurement index for memory problems to accurately quantify the memory risk of image generation, enabling a comprehensive and accurate audit of the memory risk of the diffusion model and reducing the influence of randomness in the image generation process.

[0007] To achieve the above-mentioned objectives, the embodiments provide a method and apparatus for measuring memory problems based on a reverse text-to-text diffusion model, including a method for measuring memory problems based on a reverse text-to-text diffusion model and an apparatus for measuring memory problems.

[0008] In one embodiment, the method for measuring the memory problem based on the inverse text-based graph diffusion model includes the following steps:

[0009] A noise distribution is constructed based on the noise addition process of the image to be evaluated in the text-based image diffusion model, and a prompt text distribution is constructed based on the text encoder of the text-based image diffusion model.

[0010] The sensitivity of the noise distribution is defined as the probability of generating an image to be evaluated from random noise from the noise distribution and the prompt text sampled from the prompt text distribution. A memory problem metric function is constructed based on the sensitivity.

[0011] A weighted objective function is constructed by using a denoising error based on the distribution of prompt text and noise, and a normality regularization based on the noise distribution.

[0012] The optimal prompt text distribution and noise distribution are obtained by optimizing the weighted objective function through an adaptive optimization algorithm. Based on the optimal noise distribution, the memory problem metric function is calculated to obtain the memory problem metric.

[0013] In one embodiment, the text encoder based on the text-generated graph diffusion model constructs the cue text distribution, including:

[0014] (1) The word segmenter in the text encoder encodes the string-formatted text into a prompt text w of length L, wherein the prompt text w is encoded by the word sequence w = [w1, w2, ..., w...]. L ]express;

[0015] (2) Based on the category distribution of each word, obtain the category distribution probability [π] of each word. i,1 , π i,2 , ..., π i,V ],∑ j =1π i,j =1, j∈[0,V],π i,jLet V represent the probability of the j-th category for the i-th word, and let V represent the category. The category probabilities are parameterized as logπ. i,j =φ i,j , where φ i,j This represents the distribution parameter of the j-th category of the i-th word element in the cue text.

[0016] (3) Smooth the category distribution probability of each word, and construct smoothed word based on the smoothed samples and category embedding vectors. word vectors A sequence of word vectors composed of all word vectors The input is fed into the encoding network to obtain text features c = f(e(w)). The cue text w from the text features and the image to be evaluated x0 are used to construct a cue text distribution q with parameter φ. φ (w|x0).

[0017] In one embodiment, the category distribution of each lexical unit is smoothed and sampled, and smoothed lexical units are constructed based on the smoothed samples and category embedding vectors. word vectors include:

[0018] Smooth sampling of the class distribution is performed using the Gumbel-Softmax reparameterization method:

[0019]

[0020] Among them, g i,j and g i,k Random sampling is performed from a Gumbel(0,1) distribution, and each g i,j Independent of each other, π i,k Let τ represent the probability of the k-th category for the i-th word, where τ is a constant called the temperature factor; V represents the category. This indicates that it comes from the category distribution π i A smooth sample, This represents the weight of the j-th category corresponding to the i-th sample.

[0021] Smoothed lexical units are constructed based on smoothed samples and class embedding vector e(j). word vectors

[0022]

[0023] In one embodiment, the sensitivity-based memory problem metric function is:

[0024]

[0025] in, The sensitivity to noise distribution is represented by the parameter. noise distribution The probability of generating the image x0 to be evaluated is generated by randomly sampling Gaussian noise ∈ and combining it with the prompt text w sampled from the prompt text distribution qφ(w|x0) with parameter φ. Indicated based on the image x0 to be evaluated and noise distribution parameters The constructed normality regularization, memory problem metric function indicates that there is at least one cue text w, and the sensitivity Given a value of 1, the minimum value of normality regularity is used as a metric for memory problems, mem(x0).

[0026] In one embodiment, the normality regularity is achieved by using a noise distribution measured by KL divergence. The distance D between the standard normal distribution and the standard normal distribution KL This indicates that when the noise distribution... The fit is a mean of μ and a variance of σ. 2 Normal distribution N(μ, σ) 2 )hour, Normality regularization The calculation method is as follows:

[0027]

[0028] In one embodiment, the denoising error constructed based on the distribution of the prompt text and the noise distribution is represented as l. de :

[0029]

[0030] Where ∈ represents the noise distribution The random sampled Gaussian noise is denoted as x, where t represents the time of random sampling from the uniform time distribution U(1,T), and T represents the maximum time. t Let x0 be the noisy image at time t during the forward diffusion noise addition process in the Wensheng image diffusion model, ∈ θ (x t ,t,f(e(w))) represents based on x t t,f(e(w)) is the predicted noise calculated by a neural network with parameter θ, where f(e(w)) is the text feature, α t This represents a constant related to time t, with a value ranging from 0 to 1.

[0031] In one embodiment, the weighted objective function is expressed as:

[0032]

[0033] Where L represents the noise reduction error Normality regularization The weighted objective function is constructed using the weights λ.

[0034] In one embodiment, the step of optimizing the weighted objective function using an adaptive optimization algorithm to obtain the optimal prompt text distribution and noise distribution includes:

[0035] (1) Initialize internal variables, specifically, the prompt text distribution parameter φ is set to 0, and the noise distribution parameter is set to 0. Following a normal distribution N(0, I), with a weight λ of 1, the previous denoising error... The value tends to positive infinity, the patience value P is the tolerance period ρ, and the number of optimization steps i is 1;

[0036] (2) Calculate the weighted objective function based on the initialized internal variables.

[0037] (3) Calculate the weighted objective function L relative to the prompt text distribution parameter φ and the noise distribution parameter. gradient and

[0038] (4) Based on the gradient descent method, the distribution parameters φ of the prompt text and the noise distribution parameters are analyzed. Perform an update to obtain the updated prompt text distribution parameters φ′ and noise distribution parameters. Specifically Where γ is the learning rate, used to control the step size of parameter updates;

[0039] (5) Check whether the observation period C has been reached. Specifically, observe every C steps. If the number of optimization steps i > 1 and the remainder of the number of optimization steps i with respect to the observation period C is 0, it means that C steps have been passed. Then proceed to step (6). If the remainder of the number of optimization steps i with respect to the observation period C is not 0, then proceed to step (10).

[0040] (6) Verify the noise reduction error Whether it decreases, specifically, if This indicates that the optimization is not going smoothly, proceed to step (7); if This indicates that the optimization was successful, and we proceed to step (8).

[0041] (7) If Then the weight λ decays by a multiple of λ / 2, the patience value P decays by P-1, and then jumps to step (9);

[0042] (8) If Then the weight λ increases with the weight increment δ, and the patience value is restored to the tolerance period ρ;

[0043] (9) Regarding the previous denoising error Update Proceed to step (11);

[0044] (10) Increase the weight δ appropriately;

[0045] (11) If the patience value P = 0, then jump to step (13); otherwise, proceed to step (12) to update the optimization step number i.

[0046] (12) Update the number of optimization steps i. Specifically, the number of optimization steps i is optimized by an increment of 1. If the number of optimization steps i is greater than the total number of iterations S, then proceed to step (13); otherwise, return to step (2) for iteration.

[0047] (13) Validation of the optimized generated results, specifically, based on the distribution q of the prompt text. φ (w|x0) and noise distribution The system generates images from mid-sampled prompt text and random noise. These generated images are then observed manually or compared using a third-party image matching model. If the generated image and the image to be evaluated are very similar, the system outputs the prompt text distribution qφ(w|x0) and the noise distribution. and normality regularity If an image different from the image to be evaluated exists, output the prompt text distribution qφ(w|x0) and noise distribution. and +∞.

[0048] In one embodiment, the process of minimizing the weighted objective function by obtaining the optimal prompt text distribution and noise distribution is expressed as follows:

[0049]

[0050] Where φ * This represents the optimal prompt text distribution parameters. Let λ be the optimal noise distribution parameter. * This is represented as the optimal weight.

[0051] To clearly demonstrate the measurement method for memory problems, a measurement device for memory problems based on a reverse text-image diffusion model is provided, including a memory and a processor. The memory is used to store a computer program, and the processor is used to implement the measurement method for memory problems based on the reverse text-image diffusion model when the computer program is executed.

[0052] Compared with the prior art, the beneficial effects of the present invention include at least the following:

[0053] The present invention provides a method for measuring memory problems based on a reverse-engineered image-generated diffusion model. Compared with existing methods, this method constructs a weighted objective function based on a pre-trained image-generated diffusion model and an arbitrary image, and outputs a prompt text distribution and a noise distribution in reverse. This solves the randomness effect of previous methods that assessed memory problems by detecting similar content in randomly generated images. By defining the sensitivity of the noise distribution, a measurement index function for memory problems is constructed, which continuously and accurately measures the memory problem of the image generation model, provides a quantitative result of the memory risk of any image, and audits the memory risk of the image generation model.

[0054] The present invention also provides a measurement device for the memory problem based on the inverse text-based image diffusion model, and realizes a measurement method for the image memory problem. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0056] Figure 1 A flowchart for measuring memory problems;

[0057] Figure 2 A schematic diagram of the structure of a measurement method for memory problems;

[0058] Figure 3 The flowchart is for the adaptive optimization algorithm. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of this invention.

[0060] To continuously and accurately quantify the memory risk of any image and achieve a comprehensive and accurate audit of the memory risk of the diffusion model, this embodiment provides a memory problem measurement method based on an inverse text-based image diffusion model. This prediction method constructs a weighted objective function given a pre-trained text-based image diffusion model and an arbitrary image, and inversely outputs a prompt text distribution and a noise distribution. Based on the weighted objective function, the distribution parameters are optimized to create a measurement index for image memory problems, which is used to continuously and accurately measure image memory problems.

[0061] In this embodiment, the diffusion model is trained using the following noise addition and denoising methods. For a training image... A noise sample is randomly taken from a standard Gaussian distribution N(0, I) with the same dimension as x0. A time t is randomly sampled from a uniform distribution U(1,T), and noise is superimposed onto the training image in the following manner:

[0062]

[0063] Where α t It is a constant related to time t, α at all times t t∈[1,T] are all pre-defined manually. The larger t is, the greater α is. t The closer it is to 0.

[0064] Add noise to the training image x t Time t and text features c are input into a pre-trained denoising network to obtain predicted noise ∈ θ (x t ,t,c), prediction noise∈ θ The error between the true value and the noise ∈ 0 is expressed by the L2 norm between them. This means that, with the goal of minimizing this error, the parameters θ of the pre-trained denoising network are continuously optimized to complete the training of the pre-trained denoising network.

[0065] The textural diffusion model is essentially a denoiser. Using a trained denoising network, it progressively denoises a random noisy image derived from a standard Gaussian distribution to generate a realistic image. It is typically implemented using a deep neural network, and the function implemented by this denoising network can be expressed as:

[0066]

[0067] Where θ is its pre-trained parameter; It is a noisy image, where H×W×C1 is the height, width and number of channels of the image, respectively; t∈[1,T] represents a time moment, and T represents the maximum time moment, which is a positive integer; This represents the conditions for generating an image from the guiding text, corresponding to the input prompt text. The input prompt text consists of L tokens, each of which is a C2-dimensional real vector.

[0068] The actual text-based image diffusion model uses another pre-trained text encoder to encode the input string-formatted text into continuous features, i.e., c mentioned above. The text encoder consists of a tokenizer, an embedding layer e, and an encoder network f. The tokenizer encodes the input string-formatted text into a sequence of tokens w = [w1, w2, ..., w...] of length L. L ], w i ∈[1, V], where V represents the category, i.e., the number of all possible tokens, w iThis represents the i-th word. Next, the embedding layer e assigns each word w... i Mapped to C2-dimensional word vectors e(w) i Finally, all word vectors e(w) are processed. i The inputs are sequentially fed into the encoding network f to obtain semantically rich prompt text features c = f(e(w)), where e(w) = [e(w1), e(w2), ..., e(w3)]. L )).

[0069] like Figure 1 As shown in the embodiment, the method for measuring the memory problem based on the inverse text-based graph diffusion model includes the following steps:

[0070] S1. Construct a noise distribution based on the noise addition process of the image to be evaluated in the text-based image diffusion model, and simultaneously construct a prompt text distribution based on the text encoder of the text-based image diffusion model.

[0071] The text distribution is modeled in the text encoder embedding layer e: the tokenizer encodes the string-formatted text into prompt text w = [w1, w2, ..., w] of length L. L Assume each word w i It follows a V-dimensional categorical distribution, where the probability of each word's category is [π]. i,1 , π i,2 , ..., π i,V ],∑ j=1 π i,j =1, j∈[0,V],π i,j Let represent the probability of the i-th word belonging to the j-th category, V represent the category, and L words follow L independent category distributions. The probability of each category distribution can be parameterized as log π. i,j =φ i,j , φ i,j This represents the distribution parameter of the j-th category of the i-th word element in the cue text.

[0072] Since the distribution of the prompt text is not differentiable on the sampling, the category distribution probability of each word is reparameterized using the Gumbel-Softmax reparameterization method to smooth the category distribution sampling.

[0073]

[0074] Among them, g i,j and g i,k Random sampling is performed from a Gumbel(0,1) distribution, and each g i,j Independent of each other, π i,kLet τ represent the probability of the k-th category for the i-th word, where τ is a constant called the temperature factor; V represents the category. This indicates that it comes from the category distribution π i A smooth sample, This represents the weight of the j-th category corresponding to the i-th sample.

[0075] Smoothed lexical units are constructed based on smoothed samples and class embedding vector e(j). word vectors

[0076]

[0077] A sequence of word vectors composed of all word vectors The input is fed into the encoding network to obtain text features c = f(e(w)). The cue text w from the text features and the image to be evaluated x0 are used to construct a cue text distribution q with parameter φ. φ (w|x0).

[0078] S2. Define the sensitivity of the noise distribution as the probability of generating the image to be evaluated from random noise from the noise distribution and the prompt text sampled from the prompt text distribution. Construct a memory problem metric function based on the sensitivity.

[0079] In one example, the memory problem metric function is constructed based on sensitivity as follows:

[0080]

[0081] in, The sensitivity to noise distribution is represented by the parameter. noise distribution Randomly sampled Gaussian noise ∈ and jointly derived from the prompt text distribution q with parameter φ φ The probability of generating the image x0 to be evaluated from the prompt text w sampled in (w|x0). Indicated based on the image x0 to be evaluated and noise distribution parameters The constructed normality regularization, memory problem metric function indicates that there is at least one cue text w, and the sensitivity Under the condition of 1, the minimum value of normality regularity is used as the memory problem metric mem(x0); the metric mem(x0) is a non-negative real number that represents the degree to which the image is remembered by the given model. The closer the metric mem(x0) is to 0, the greater the degree of memory, and the closer it is to positive infinity, the smaller the degree of memory.

[0082] S3. Construct a weighted objective function by using the denoising error based on the distribution of prompt text and noise distribution, and the normality regularization based on the noise distribution.

[0083] In the embodiments, such as Figure 2 As shown, based on an arbitrary image, a pre-trained text-based image diffusion model is used to output a cue text distribution and a noise distribution in reverse, and the degree to which it is remembered is measured based on the noise distribution.

[0084] In the embodiment, the noise distribution is measured using the KL divergence metric. The distance D between the standard normal distribution and the standard normal distribution KL This indicates that when the noise distribution... The fit is a mean of μ and a variance of σ. 2 Normal distribution N(μ, σ) 2 )hour, Normality regularization The calculation method is as follows:

[0085]

[0086] in, Indicated based on the image x0 to be evaluated and noise distribution parameters Constructing a normality regularization.

[0087] The denoising error, constructed based on the distribution of the prompt text and the noise distribution, is denoted as l. de :

[0088]

[0089] Where ∈ represents the noise distribution The random sampled Gaussian noise is denoted as x, where t represents the time of random sampling from the uniform time distribution U(1,T), and T represents the maximum time. t Let x0 be the noisy image at time t during the forward diffusion noise addition process in the Wensheng image diffusion model, ∈ θ (x t ,t,f(e(w))) represents based on x t t,f(e(w)) is the predicted noise calculated by a neural network with parameter θ, where f(e(w)) is the text feature, α t This represents a constant related to time t, with a value ranging from 0 to 1.

[0090] Based on denoising error l de and normality regularization nr The constructed weighted objective function is:

[0091]

[0092] Where λ is the introduced weight, and L represents the weight based on the denoising error. Normality regularization The weighted objective function is constructed using the weights λ.

[0093] S4. The optimal prompt text distribution and noise distribution are obtained by optimizing the weighted objective function through an adaptive optimization algorithm. The memory problem metric function is calculated based on the optimal noise distribution to obtain the memory problem metric.

[0094] To optimize the weighted objective function and obtain the optimal prompt text distribution and noise distribution parameters, this invention proposes an adaptive optimization algorithm to minimize the weighted objective function. During the optimization process, the parameters θ of the pre-trained denoising network remain constant. The algorithm input includes: the pre-trained denoising network ∈ θ The text encoder, the image to be evaluated x0, the total number of iterations S, the observation period C, the weight increment δ, the denoising error threshold ξ, the learning rate γ, and the tolerance period ρ;

[0095] like Figure 3 As shown, the steps of the adaptive optimization algorithm are as follows:

[0096] (1) Initialize internal variables, specifically, the prompt text distribution parameter φ is set to 0, and the noise distribution parameter is set to 0. Following a normal distribution N(0, I), with a weight λ of 1, the previous denoising error... The value tends to positive infinity, the patience value P is the tolerance period ρ, and the number of optimization steps i is 1;

[0097] (2) Calculate the weighted objective function based on the initialized internal variables.

[0098] (3) Calculate the weighted objective function L relative to the prompt text distribution parameter φ and the noise distribution parameter. gradient and

[0099] (4) Based on the gradient descent method, the distribution parameters φ of the prompt text and the noise distribution parameters are analyzed. Perform an update to obtain the updated prompt text distribution parameters φ′ and noise distribution parameters. Specifically Where γ is the learning rate, used to control the step size of parameter updates;

[0100] (5) Check whether the observation period C has been reached. Specifically, observe every C steps. If the number of optimization steps i > 1 and the remainder of the number of optimization steps i with respect to the observation period C is 0, it means that C steps have been passed. Then proceed to step (6). If the remainder of the number of optimization steps i with respect to the observation period C is not 0, then proceed to step (10).

[0101] (6) Verify the noise reduction error Whether it decreases, specifically, if This indicates that the optimization is not going smoothly, proceed to step (7); if This indicates that the optimization was successful, and we proceed to step (8).

[0102] (7) If Then the weight λ decays by a multiple of λ / 2, the patience value P decays by P-1, and then jumps to step (9);

[0103] (8) If Then the weight λ increases with the weight increment δ, and the patience value is restored to the tolerance period ρ;

[0104] (9) Regarding the previous denoising error Update Proceed to step (11);

[0105] (10) Increase the weight δ appropriately;

[0106] (11) If the patience value P = 0, then jump to step (13); otherwise, proceed to step (12) to update the optimization step number i.

[0107] (12) Update the number of optimization steps i. Specifically, the number of optimization steps i is optimized by an increment of 1. If the number of optimization steps i is greater than the total number of iterations S, then proceed to step (13); otherwise, return to step (2) for iteration.

[0108] (13) Validation of the optimized generated results, specifically, based on the distribution q of the prompt text. φ (w|x0) and noise distribution The system generates images from mid-sampled prompt text and random noise. These generated images are then observed manually or compared using a third-party image matching model. If the generated image and the image to be evaluated are very similar, the system outputs the prompt text distribution qφ(w|x0) and the noise distribution. and normality regularity If an image different from the image to be evaluated exists, output the prompt text distribution q. φ (w|x0), noise distribution and +∞.

[0109] The optimal prompt text distribution and noise distribution are obtained based on an adaptive optimization algorithm, thereby minimizing the weighted objective function, which is expressed as:

[0110]

[0111] Where φ * This represents the optimal prompt text distribution parameters. Let λ be the optimal noise distribution parameter. * This is represented as the optimal weight.

[0112] The memory problem metric function is solved based on the optimal prompt text distribution and noise distribution to obtain the memory problem metric and realize the measurement of image memory problem.

[0113] To address these issues, the present invention also provides a measurement device for the memory problem based on a reverse text-image diffusion model, comprising a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, implements a measurement method for the memory problem based on a reverse text-image diffusion model.

[0114] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for measuring the memory problem in a reverse text-based graph diffusion model, characterized in that, The following steps are involved: A noise distribution is constructed based on the noise addition process of the image to be evaluated in the text-based image diffusion model. Simultaneously, a cue text distribution is constructed based on the text encoder of the text-based image diffusion model. The construction of the cue text distribution includes: (1) The word segmenter in the text encoder encodes the string-formatted text into a prompt text w of length L, wherein the prompt text w is encoded by the word sequence w = [w1, w2, ..., w...]. L ]express; (2) Based on the category distribution of each word, obtain the category distribution probability [π] of each word. i,1 , π i,2 , ..., π i,V ],∑ j =1π i,j =1, j∈[0,V],π i,j Let V represent the probability of the j-th category for the i-th word, and let V represent the category. The category probabilities are parameterized as logπ. i,j =φ i,j , where φ i,j This represents the distribution parameter of the j-th category of the i-th word element in the cue text. (3) Smooth the category distribution probability of each word, and construct smoothed word based on the smoothed samples and category embedding vectors. word vectors A sequence of word vectors composed of all word vectors The input is fed into the encoding network to obtain text features c = f(e(w)). The cue text w from the text features and the image to be evaluated x0 are used to construct a cue text distribution q with parameter φ. φ (w|x0); The sensitivity of the noise distribution is defined as the probability of generating the image to be evaluated from random noise from the noise distribution and the prompt text sampled from the prompt text distribution. Based on the sensitivity, a memory problem metric function is constructed, which is expressed as: in, The sensitivity to noise distribution is represented by the parameter. noise distribution The probability of generating the image x0 to be evaluated is generated by randomly sampling Gaussian noise ∈ and combining it with the prompt text w sampled from the prompt text distribution qφ(w|x0) with parameter φ. Indicated based on the image x0 to be evaluated and noise distribution parameters The constructed normality regularization, memory problem metric function indicates that there is at least one cue text w, and the sensitivity Given a value of 1, the minimum value of normality regularity is used as a metric for memory problems, mem(x0). A weighted objective function is constructed by using a denoising error based on the distribution of the prompt text and noise distribution, and a normality regularization based on the noise distribution. The normality regularization is achieved by using the KL divergence measure for the noise distribution. The distance D between the standard normal distribution and the standard normal distribution KL This indicates that when the noise distribution... The fit is a mean of μ and a variance of σ. 2 Normal distribution N(μ, σ) 2 )hour, Normality regularization The calculation method is as follows: The optimal prompt text distribution and noise distribution are obtained by optimizing the weighted objective function using an adaptive optimization algorithm. Based on the optimal noise distribution, a memory problem metric function is calculated to obtain the memory problem metric. The optimization of the weighted objective function using the adaptive optimization algorithm to obtain the optimal prompt text distribution and noise distribution includes: (1) Initialize internal variables, specifically, the prompt text distribution parameter φ is set to 0, and the noise distribution parameter is set to 0. Following a normal distribution N(0, I), with a weight λ of 1, the previous denoising error... The value tends to positive infinity, the patience value P is the tolerance period ρ, and the number of optimization steps i is 1; (2) Calculate the weighted objective function based on the initialized internal variables. (3) Calculate the weighted objective function L relative to the prompt text distribution parameter φ and the noise distribution parameter. gradient and (4) Based on the gradient descent method, the distribution parameters φ of the prompt text and the noise distribution parameters are analyzed. Perform an update to obtain the updated prompt text distribution parameters φ′ and noise distribution parameters. Specifically Where γ is the learning rate, used to control the step size of parameter updates; (5) Check whether the observation period C has been reached. Specifically, observe every C steps. If the number of optimization steps i > 1 and the remainder of the number of optimization steps i with respect to the observation period C is 0, it means that C steps have been passed. Then proceed to step (6). If the remainder of the number of optimization steps i with respect to the observation period C is not 0, then proceed to step (10). (6) Verify the noise reduction error Whether it decreases, specifically, if This indicates that the optimization is not going smoothly, proceed to step (7); if This indicates that the optimization was successful, and we proceed to step (8). (7) If Then the weight λ decays by a multiple of λ / 2, the patience value P decays by P-1, and then jumps to step (9); (8) If Then the weight λ increases with the weight increment δ, and the patience value is restored to the tolerance period ρ; (9) Regarding the previous denoising error Update Proceed to step (11); (10) Increase the weight δ appropriately; (11) If the patience value P = 0, then jump to step (13); otherwise, proceed to step (12) to update the optimization step number i. (12) Update the number of optimization steps i. Specifically, the number of optimization steps i is optimized by an increment of 1. If the number of optimization steps i is greater than the total number of iterations S, then proceed to step (13); otherwise, return to step (2) for iteration. (13) Validation of the optimized generated results, specifically, based on the distribution q of the prompt text. φ (w|x0) and noise distribution The system generates images from mid-sampled prompt text and random noise. These generated images are then compared either manually or using a third-party image matching model. If the generated image and the image to be evaluated are very similar, the prompt text distribution q is output. φ (w|x0), noise distribution and normality regularity If an image different from the image to be evaluated exists, output the prompt text distribution q. φ (w|x0), noise distribution and +∞.

2. The method for measuring memory problems according to claim 1, characterized in that, The method described above involves smoothing the category distribution of each lexical unit and constructing smoothed lexical units based on the smoothed samples and category embedding vectors. word vectors include: Smooth sampling of the class distribution is performed using the Gumbel-Soffmax reparameterization method: Among them, g i,j and g i,k Random sampling is performed from a Gumbel(0,1) distribution, and each g i,j Independent of each other, π i,k Let τ represent the probability of the k-th category for the i-th word, where τ is a constant called the temperature factor; V represents the category. This represents a smoothed sample from the class distribution. This represents the weight of the j-th class corresponding to the i-th sample; Smoothed lexical units are constructed based on smoothed samples and class embedding vector e(j). word vectors 3. The method for measuring memory problems according to claim 1, characterized in that, The denoising error constructed based on the distribution of prompt text and noise distribution is denoted as l. de : Where ∈ represents the noise distribution The random sampled Gaussian noise is denoted as x, where t represents the time of random sampling from the uniform time distribution U(1,T), and T represents the maximum time. t Let x0 be the noisy image at time t during the forward diffusion noise addition process in the Wensheng image diffusion model, ∈ θ (x t ,t,f(e(w))) represents based on x t t,f(e(w)) is the predicted noise calculated by a neural network with parameter θ, where f(e(w)) is the text feature, α t This represents a constant related to time t, with a value ranging from 0 to 1.

4. The method for measuring memory problems according to claim 1, characterized in that, The weighted objective function is expressed as: Where L represents the noise reduction error Normality regularization The weighted objective function is constructed using the weights λ.

5. The method for measuring memory problems according to claim 1, characterized in that, The method for minimizing the weighted objective function by obtaining the optimal prompt text distribution and noise distribution is expressed as follows: Where φ * This represents the optimal prompt text distribution parameters. Let λ be the optimal noise distribution parameter. * This is represented as the optimal weight.

6. A measurement device for the memory problem based on a reverse text-image diffusion model, comprising a memory and a processor, wherein the memory is used to store computer programs, characterized in that, The processor is configured to, when executing the computer program, implement the method for measuring the memory problem based on the reverse text-based graph diffusion model as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Mask-based face image generation model privacy protection method, system and device

    CN117633899A