Generative model analysis method and device based on hyper-parameter fingerprints
By learning prior knowledge and generating hyperparameter fingerprints, the problem of low accuracy of hyperparameter fingerprints of generative models in the prior art is solved, and the refined analysis and accuracy of generative models are achieved.
Patent Information
- Application Number
- CN202510135573.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the hyperparameter fingerprint of the generative model is low and lacks the ability to finely distinguish multiple hyperparameter traces, resulting in the accuracy of the analysis results of the generative model.
By learning prior knowledge independent of multiple generative models and perceived prior knowledge of the target image, a hyperparameter fingerprint is generated, and the hyperparameter prediction results of the source model are output by using the target analysis network to determine the generative model of the target image.
It realizes the refined distinction between hyperparameter traces of generative models, improves the accuracy of generative model analysis, provides more accurate information guidance, and continuously improves practical application capabilities through end-to-end training.
Smart Images

Figure CN120107386A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision and machine learning technology, and in particular to a generative model analysis method and device based on hyperparameter fingerprint. Background Art
[0002] With the rapid advancement of deep learning technology, various generative models that can generate highly realistic visual content have emerged. In view of the risk of large-scale dissemination of false information, the academic community has proposed a series of visual analysis tasks for generative models and tried to solve them. One of the most challenging key tasks is model parsing, which aims to infer model hyperparameters from generated images. This task helps to analyze the increasing number of unknown new models and try to reverse engineer these unknown models and restore their structures.
[0003] In the related technology, it mainly relies on a fingerprint estimation network with fixed parameters to extract a single and rough model fingerprint from the image, so as to trace the generative model of the image through the model fingerprint.
[0004] However, in related technologies, the model fingerprint has low accuracy and lacks the ability to finely distinguish traces of various hyperparameters. The analysis results of the generative model are also not accurate, which needs to be solved urgently. Summary of the invention
[0005] The present application provides a generative model analysis method and device based on hyperparameter fingerprints to solve the problems in related technologies, such as low accuracy of model fingerprints, lack of ability to finely distinguish multiple hyperparameter traces, and low accuracy of analysis results for generative models.
[0006] The first aspect of the present application provides a generative model parsing method based on a hyperparameter fingerprint, comprising the following steps: extracting at least one hidden fingerprint of a target image, and using at least one target hypernetwork to learn prior knowledge that is independent of multiple generative models, and using the at least one hidden fingerprint and another target hypernetwork other than the at least one target hypernetwork to learn perceptual prior knowledge of the target image; generating a hyperparameter fingerprint of the target image based on the at least one hidden fingerprint, the prior knowledge that is independent of multiple generative models, and the perceptual prior knowledge; inputting the hyperparameter fingerprint into a target parsing network to output a source model hyperparameter prediction result of the target image, and determining the generative model of the target image through the source model hyperparameter prediction result.
[0007] Optionally, in one embodiment of the present application, the utilizing at least one target hypernetwork to learn prior knowledge that is independent of multiple generative models comprises: based on image instances of multiple generative models, maintaining at least one set of baseline prior vectors and hyperparameter vectors that are independent of the multiple generative models to learn the prior knowledge that is independent of the multiple generative models.
[0008] Optionally, in one embodiment of the present application, the using the at least one hidden layer fingerprint and another target super network other than the at least one target super network to learn the perceptual prior knowledge of the target image includes: encoding the at least one hidden layer fingerprint into an instance prior vector; and determining the perceptual prior knowledge based on the instance prior vector.
[0009] Optionally, in one embodiment of the present application, generating a hyperparameter fingerprint of the target image based on the at least one hidden fingerprint, the prior knowledge independent of multiple generative models and the perceptual prior knowledge includes: fusing the prior knowledge independent of multiple generative models and the perceptual prior knowledge to obtain complete prior knowledge, and using the complete prior knowledge to generate transformation weights of the at least one hidden fingerprint; performing a linear transformation on the at least one hidden fingerprint based on the transformation weights to generate a hidden hyperparameter fingerprint; and inputting the hidden hyperparameter fingerprint into a target fingerprint estimation network to generate a hyperparameter fingerprint of the target image.
[0010] Optionally, in one embodiment of the present application, the generation expression of the hidden layer hyperparameter fingerprint is:
[0011]
[0012] Among them, c b is the reference prior vector, is the hyperparameter prior vector, c it is the instance perception vector, θ represents the parameters of the weight generation hypernetwork h, n hp represents the number of hyperparameters, c com represents the complete prior vector, represents the complete prior vector component corresponding to a single hyperparameter j, W j represents the transformation coefficient component corresponding to a single hyperparameter j, z 2 is the hidden fingerprint obtained after the initial hidden fingerprint is convolved again, z hp represents the hidden layer hyperparameter fingerprint, W represents z 2 Transformation weights for linear transformations.
[0013] The second aspect of the present application provides a generative model parsing device based on a hyperparameter fingerprint, including: a learning module, used to extract at least one hidden fingerprint of a target image, and use at least one target supernetwork to learn prior knowledge independent of multiple generative models, and use the at least one hidden fingerprint and another target supernetwork other than the at least one target supernetwork to learn perceptual prior knowledge of the target image; a generation module, used to generate a hyperparameter fingerprint of the target image based on the at least one hidden fingerprint, the prior knowledge independent of multiple generative models and the perceptual prior knowledge; a parsing module, used to input the hyperparameter fingerprint into a target parsing network to output a source model hyperparameter prediction result of the target image, and determine the generative model of the target image through the source model hyperparameter prediction result.
[0014] Optionally, in one embodiment of the present application, the learning module includes: a learning unit, used to maintain at least one set of baseline prior vectors and hyperparameter vectors independent of the multiple generative models based on image instances of multiple generative models, so as to learn the prior knowledge independent of the multiple generative models.
[0015] Optionally, in one embodiment of the present application, the learning module includes: an encoding unit, used to encode the at least one hidden fingerprint into an instance prior vector; and a determination unit, used to determine the perceptual prior knowledge based on the instance prior vector.
[0016] Optionally, in one embodiment of the present application, the generation module includes: a fusion unit, used to fuse the prior knowledge independent of multiple generative models and the perceptual prior knowledge to obtain complete prior knowledge, so as to use the complete prior knowledge to generate the transformation weights of the at least one hidden layer fingerprint; a transformation unit, used to perform a linear transformation on the at least one hidden layer fingerprint based on the transformation weights to generate a hidden layer hyperparameter fingerprint; and a generation unit, used to input the hidden layer hyperparameter fingerprint into a target fingerprint estimation network to generate a hyperparameter fingerprint of the target image.
[0017] Optionally, in one embodiment of the present application, the generation expression of the hidden layer hyperparameter fingerprint is:
[0018]
[0019] Among them, c b is the reference prior vector, is the hyperparameter prior vector, c it is the instance perception vector, θ represents the parameters of the weight generation hypernetwork h, n hp represents the number of hyperparameters, c com represents the complete prior vector, represents the complete prior vector component corresponding to a single hyperparameter j, w j represents the transformation coefficient component corresponding to a single hyperparameter j, z 2 is the hidden fingerprint obtained after the initial hidden fingerprint is convolved again, z hp represents the hidden layer hyperparameter fingerprint, W represents z 2 Transformation weights for linear transformations.
[0020] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the generative model parsing method based on hyperparameter fingerprints as described in the above embodiment.
[0021] The fourth aspect of the present application provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above-mentioned generative model parsing method based on hyperparameter fingerprints.
[0022] The fifth aspect of the present application provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned generative model parsing method based on hyperparameter fingerprints.
[0023] The embodiment of the present application can learn prior knowledge that is independent of multiple models through a hypernetwork shared between each model, and learn the perceptual prior knowledge of the target image through another hypernetwork, and then fuse the two prior knowledge to generate a hyperparameter fingerprint, so as to realize the analysis of the generative model corresponding to the target image. Thus, it is realized that by mining the prior knowledge that various hyperparameters in different generative models are independent of the generative model, the traces of the hyperparameters of various generative models can be finely distinguished; through the target image instance, the prior knowledge encoding the target image instance is learned, so as to realize the dynamic response to the change of the input instance level; further, the prior knowledge that is independent of multiple generative models and the perceptual prior knowledge of the target image are fused and integrated into the feature representation of the fingerprint estimation network, and the hyperparameter fingerprint corresponding to various hyperparameters is generated, which can provide refined information guidance for more accurate analysis of the generative model, and each prediction process can train the entire network structure end-to-end, which can continuously improve the practical application ability of the present application. Thus, the problems in the related technology that the model fingerprint has low accuracy, lacks the ability to finely distinguish multiple hyperparameter traces, and the accuracy of the analysis results of the generative model is not high are solved.
[0024] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0026] Figure 1 A flowchart of a generative model parsing method based on hyperparameter fingerprints provided according to an embodiment of the present application;
[0027] Figure 2 A schematic diagram showing the comparison of model fingerprints and hyperparameter fingerprints of different face generation models according to an embodiment of the present application;
[0028] Figure 3 This is a schematic diagram of the architecture of a generative model parsing system based on hyperparameter fingerprints according to an embodiment of the present application;
[0029] Figure 4 This is a flow chart of a generative model parsing method based on hyperparameter fingerprints according to an embodiment of the present application;
[0030] Figure 5 A schematic diagram of the structure of a generative model analysis device based on hyperparameter fingerprints provided according to an embodiment of the present application;
[0031] Figure 6 It is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application.
[0032] Reference numerals:
[0033] 10-Generative model parsing device based on hyperparameter fingerprint: 100-learning module, 200-generation module and 300-parsing module; 601-memory, 602-processor and 603-communication interface. DETAILED DESCRIPTION
[0034] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0035] The following describes the generative model parsing method and device based on hyperparameter fingerprints of an embodiment of the present application with reference to the accompanying drawings. In view of the problem that the model fingerprint in the related technologies mentioned in the above background technology has low accuracy and lacks the ability to finely distinguish multiple hyperparameter traces, and the parsing results of the generative model are not accurate, the present application provides a generative model parsing method based on hyperparameter fingerprints, in which a hypernetwork shared by each model can be used to learn prior knowledge that is independent of multiple models, and another hypernetwork can be used to learn the perceptual prior knowledge of the target image, and then the two prior knowledge are fused to generate a hyperparameter fingerprint, thereby realizing the parsing of the generative model corresponding to the target image. Thus, it is achieved that by mining the prior knowledge that various hyperparameters in different generative models are independent of the generative models, the traces of the hyperparameters of various generative models can be finely distinguished; through the target image instance, the prior knowledge encoding the target image instance is learned, so as to achieve dynamic response to changes in the input instance level; further, the prior knowledge that is independent of multiple generative models and the perceptual prior knowledge of the target image are integrated and integrated into the feature representation of the fingerprint estimation network to generate hyperparameter fingerprints corresponding to various hyperparameters, which can provide refined information guidance for more accurate analysis of the generative model, and each prediction process can train the entire network structure end-to-end, which can continuously improve the practical application ability of this application. Thus, the problems in the related technology such as low accuracy of model fingerprints, lack of ability to finely distinguish various hyperparameter traces, and low accuracy of analysis results for generative models are solved.
[0036] Specifically, Figure 1 A flowchart of a generative model parsing method based on hyperparameter fingerprint provided in an embodiment of the present application.
[0037] like Figure 1 As shown, the generative model parsing method based on hyperparameter fingerprint includes the following steps:
[0038] In step S101, at least one hidden fingerprint of a target image is extracted, and at least one target super network is used to learn prior knowledge that is independent of multiple generative models, and at least one hidden fingerprint and another target super network other than the at least one target super network are used to learn perceptual prior knowledge of the target image.
[0039] It is understandable to those skilled in the art that using hyperparameter fingerprints to complete the task of generative model parsing can better identify the traces left by different hyperparameters, thereby predicting the hyperparameters of each generative model in a targeted manner. Figure 2 Schematic diagram of comparison of model fingerprints and hyperparameter fingerprints of different face generation models according to an embodiment of the present application. Figure 2As shown in the figure, the two model fingerprints in the first column from the left are highly similar and difficult to distinguish by human eyes because they come from two face generation models of the same series, stargan1 and stargan2. In contrast, if the hyperparameter fingerprints correspond to different categories of hyperparameters, such as Figure 2 In the six columns of fingerprint images on the right, there are obvious differences between the upper and lower images in the same column (corresponding to stargan1 and stargan2 respectively). In addition, by comparing the hyperparameter fingerprints of different types, such as the minimum square loss (second column) and the cross entropy loss (third column), they can also be easily distinguished from each other.
[0040] As a possible implementation method, in the process of obtaining the hyperparameter fingerprint, the embodiment of the present application can first extract the hidden fingerprint from the target image. The target image here can be understood as an image generated by a generative model, and the generative model that generates the target image can be analyzed. The hidden fingerprint here can be understood as a data pattern or feature vector contained in the hidden layer of the target image generation model that can characterize the unique characteristics or source information of the image.
[0041] Furthermore, in order to fully explore the prior knowledge contained in images generated by a variety of generative models, the embodiments of the present application can learn the information shared between models through at least one hypernetwork shared between the models (such as a model-independent hypernetwork), that is, the prior knowledge that is independent of multiple generative models, thereby having an abstract understanding of the hyperparameter features that are independent of the generative model.
[0042] Secondly, for the current target image instance, the embodiment of the present application can learn the characteristics of the current target image instance, that is, the perceptual prior knowledge of the target image, from the hidden fingerprint through another hypernetwork (such as an instance-aware hypernetwork), so that the hypernetwork that learns the current target image instance can dynamically respond to the current input instance.
[0043] Prior knowledge here refers to the knowledge, information or assumptions about the generative model or target image that are already available before the generative model is parsed. This knowledge can be obtained based on past experience, data statistics, theoretical research, etc. For example, the shape and color features of objects in the target image.
[0044] In addition, the two target hypernetworks in the embodiments of the present application can be understood as related hypernetworks that can learn prior knowledge that is independent of multiple generative models and perceptual prior knowledge of the target image, for example, hypernetworks based on complex network theory, hypernetworks based on neural network architecture, etc., which can be specifically selected or adjusted by professional and technical personnel in this field according to actual conditions, and this application does not impose any specific restrictions.
[0045] Optionally, in one embodiment of the present application, at least one target hypernetwork is used to learn prior knowledge that is independent of multiple generative models, including: based on image instances of multiple generative models, maintaining at least one set of baseline prior vectors and hyperparameter vectors that are independent of multiple generative models to learn prior knowledge that is independent of multiple generative models.
[0046] In the actual implementation process, when using at least one target hypernetwork to learn generative model-independent prior knowledge, the present application can, but is not limited to, maintain at least one set of baseline prior vectors and hyperparameter vectors that are independent of multiple generative models based on image instances of multiple generative models, wherein both vectors can be learned to learn prior knowledge that is independent of multiple generative models.
[0047] For example, for input image instances from multiple generative models, the present application may, but is not limited to, maintain a set of model-independent reference prior vectors c through a model-independent super-network in a model-independent encoder. b and the hyperparameter prior vector c hp , as shown below:
[0048]
[0049] Among them, c b represents the benchmark prior vector, c hp represents the hyperparameter prior vector, and Respectively represent D b and D hp dimensional real number field vector space.
[0050] The reference prior vector can be understood as a vector representation of a specific prior distribution of multiple generative models that generate the target image. It is usually determined based on past experience, theoretical knowledge, or some default assumptions, and provides a starting point for subsequent statistical inference. For example, when estimating the parameters of a multinomial distribution, a uniform reference prior vector may be set, which represents an unbiased initial assumption about the probability of occurrence of each category in the absence of any additional information.
[0051] Hyperparameter prior vector c hpHere, it can also be understood as a vector form of a priori settings for the hyperparameters of multiple generative models that generate the target image. For example, the hyperparameters in the generative model include but are not limited to the number of layers of various convolutional and pooling layers in its neural network architecture, the type of loss function, etc., and the hyperparameter prior vector is the vector representation of these hyperparameters. In general, the hyperparameter prior vector here refers to the feature representation of the prior knowledge of hyperparameters of various generative models in the feature space, which will change continuously according to a large number of input images and contains the hyperparameter prior knowledge learned from all past input images.
[0052] Based on the benchmark prior vector c b and the hyperparameter prior vector c hp , the embodiment of the present application can determine the prior knowledge that is independent of multiple models. By maintaining these vectors, the embodiment of the present application can provide a unified evaluation standard for the performance of the hypernetwork, which helps to promote the learning of feature representation and multi-task learning of the hypernetwork, facilitates the debugging and optimization of the hypernetwork, and thus improves the learning accuracy of the prior knowledge that is independent of multiple generative models in the embodiment of the present application during each maintenance process.
[0053] The embodiments of the present application can learn the prior knowledge that various hyperparameters in different generative models are independent of the generative models, thereby achieving a refined distinction of the traces of hyperparameters of various generative models.
[0054] Optionally, in one embodiment of the present application, at least one hidden layer fingerprint and another target super network other than at least one target super network are used to learn perceptual prior knowledge of a target image, including: encoding at least one hidden layer fingerprint into an instance prior vector; and determining perceptual prior knowledge based on the instance prior vector.
[0055] In other embodiments, when the present application learns the perceptual prior knowledge of a target image instance, for the current target image instance, the hidden fingerprint may be encoded into an instance prior vector, but is not limited to, to thereby determine the perceptual prior knowledge of the target image.
[0056] For example, after obtaining the hidden fingerprint, the embodiment of the present application may, but is not limited to, input the hidden fingerprint into the instance-aware encoder e in the instance-aware hypernetwork. it , to utilize the instance-aware encoder e it Encode the hidden fingerprint as the instance prior vector c it , the formula can be expressed as but not limited to:
[0057]
[0058] Among them, c it represents the instance prior vector, e it represents the instance-aware hypernetwork, z1 represents the hidden fingerprint, θ and denote the parameters of the instance-aware hypernetwork and D it dimensional real number field vector space.
[0059] It should be noted that in the process of learning the perceptual prior knowledge of the target image, the embodiment of the present application will simultaneously maintain the instance-aware encoder e in the instance-aware hypernetwork. it The existing parameters are dynamically updated in real time when a new target image is received, which can effectively improve the learning ability and response speed of the instance-aware hypernetwork in the embodiment of the present application when perceiving prior knowledge of other images.
[0060] Step S102, generating a hyperparameter fingerprint of the target image based on prior knowledge and perceptual prior knowledge that are independent of the plurality of generative models.
[0061] In certain embodiments, after obtaining prior knowledge and perceptual prior knowledge that are independent of multiple generative models, the present application can integrate the prior knowledge that is independent of multiple generative models and the perceptual prior knowledge of the target image instance, and then generate a hyperparameter fingerprint through a certain fingerprint estimation network (such as a broadcast fingerprint estimation network) to utilize the hyperparameter fingerprint prediction to achieve generative model analysis of the target image.
[0062] Next, the process of generating a hyperparameter fingerprint of a target image using prior knowledge and perceptual prior knowledge that are independent of multiple generative models in an embodiment of the present application is further explained.
[0063] Optionally, in one embodiment of the present application, based on at least one hidden fingerprint, prior knowledge and perceptual prior knowledge unrelated to multiple generative models, a hyperparameter fingerprint of a target image is generated, including: fusing prior knowledge and perceptual prior knowledge unrelated to multiple generative models to obtain complete prior knowledge, so as to generate transformation weights of at least one hidden fingerprint using the complete prior knowledge; performing a linear transformation on at least one hidden fingerprint based on the transformation weights to generate a hidden hyperparameter fingerprint; inputting the hidden hyperparameter fingerprint into a target fingerprint estimation network to generate a hyperparameter fingerprint of the target image. The generation expression of the hidden hyperparameter fingerprint can be, but is not limited to, expressed as:
[0064]
[0065] Among them, c b is the reference prior vector, is the hyperparameter prior vector, c it is the instance perception vector, θ represents the parameters of the weight generation hypernetwork h, n ho represents the number of hyperparameters, c com represents the complete prior vector, represents the complete prior vector component corresponding to a single hyperparameter j, W j represents the transformation coefficient component corresponding to a single hyperparameter j, z 2 is the hidden fingerprint obtained after the initial hidden fingerprint is convolved again, z hp represents the hidden layer hyperparameter fingerprint, W represents z 2 Transformation weights for linear transformations.
[0066] In the actual implementation process, when the present application generates the hyperparameter fingerprint of the target image through prior knowledge and perceptual prior knowledge that are independent of multiple generative models, the two are mainly fused, and a certain transformation weight is generated using the fused complete prior knowledge; based on the transformation weight, a linear transformation is performed on the hidden layer fingerprint to generate a hidden layer hyperparameter fingerprint; finally, the hidden layer hyperparameter fingerprint can be integrated into the target fingerprint estimation network to generate the hyperparameter fingerprint of the target image.
[0067] Among them, the target fingerprint estimation network here can be understood as a fingerprint estimation network used to generate hyperparameter fingerprints, which can complete the hyperparameter fingerprint generation task. The specific network architecture and other relevant design information can be selected or adjusted by professional and technical personnel in this field according to actual conditions, and this application does not make specific restrictions.
[0068] For example, the present application can first broadcast and fuse the learned prior knowledge that is independent of multiple generative models and the instance-aware prior knowledge, that is, connect each hyperparameter prior vector with the baseline prior vector and the instance prior vector to obtain the integrated complete prior vector (complete prior knowledge), and input it into a two-layer fully connected hypernetwork that generates weights to obtain a set of transformation weights W.
[0069] Next, the embodiment of the present application can use a certain fingerprint estimation network (such as a two-stage fingerprint estimation network) to transform the hidden fingerprint z 1 Continue convolution and output the hidden fingerprint z 2 Then, z is transformed into 2 Adjust, that is, perform a sublinear transformation, and the transformation coefficient is the weight W generated above, from which a set of hidden layer hyperparameter fingerprints z can be obtained hp , the formula can be expressed as follows but is not limited to:
[0070]
[0071] Among them, z 2 is the hidden fingerprint obtained after the initial hidden fingerprint is convolved again, z hp Represents the hidden layer hyperparameter fingerprint.
[0072] Next, the target fingerprint estimation network, such as the broadcast fingerprint estimation network, is used to estimate the hidden layer hyperparameter fingerprint z again. hp By performing convolution, a set of hyperparameter fingerprints can be generated.
[0073] Also, broadcast fusion here can be understood as integrating prior knowledge that is unrelated to multiple generative models and instance-aware prior knowledge in a way that can be acquired and utilized by all related modules or processing flows. If different prior knowledge is imagined as different information sources, broadcast fusion is to uniformly disseminate and merge the contents of these information sources so that all places that need to use this knowledge can receive the complete fused knowledge. For example, in a multimodal data processing task, the prior knowledge from the image modality (such as the visual features of the object) and the prior knowledge from the text modality (such as the semantic description of the object) are broadcast fused through specific algorithms or mechanisms, so that the modules that process images and texts can use the fused knowledge.
[0074] Through broadcast fusion, these scattered information can be integrated together to form a more comprehensive and complete knowledge system. The fused prior knowledge can provide richer and more accurate guidance information for hyperparameter fingerprint generation, which helps to learn and understand the patterns and laws in the data, making the broadcast fingerprint estimation network more adaptable when facing different prior knowledge changes.
[0075] It should be noted that the method of integrating prior knowledge and instance-aware prior knowledge that are independent of multiple generative models and the various hypernetworks used in the embodiments of the present application can be adjusted by professional and technical personnel in the field according to actual conditions. This is only an illustrative description without specific limitation.
[0076] The embodiments of the present application can align and fuse different prior knowledge and instance-aware prior knowledge that are independent of multiple generative models, thereby improving the representativeness and flexibility of the complete prior knowledge; and the embodiments of the present application can generate hyperparameter fingerprints corresponding to various hyperparameters based on the fused complete prior knowledge, which helps to provide refined information guidance for more accurate analysis of the generative model.
[0077] Step S103, inputting the hyperparameter fingerprint into the target parsing network to output the source model hyperparameter prediction result of the target image, and determining the generative model of the target image through the source model hyperparameter prediction result.
[0078] In some embodiments, after generating the hyperparameter fingerprint, the present application can input the hyperparameter fingerprint into the target parsing network, and output the source hyperparameter prediction result of the target image through the target parsing network, that is, the prediction result of various parameter information contained in the generative model that generates the target image. For example, architecture category, loss category, special layer and other hyperparameters. Based on this information, the embodiment of the present application can further analyze what model the generative model corresponding to the target image may be.
[0079] Among them, the target parsing network here can be understood as a network for parsing the source model hyperparameters of the target image. The specific network architecture is not limited in the embodiment of the present application and can be set or adjusted by professional and technical personnel in this field according to actual conditions.
[0080] For example, given that different architectures have different characteristics and applicable scenarios, the embodiments of the present application can determine the basic architecture of the generative model corresponding to the target image, such as GAN, VAE, etc., based on the architecture information.
[0081] Next, the embodiment of the present application can clarify the loss function used in training the generative model of the target image source according to the loss category in the hyperparameters, because the loss function will affect the learning process and generation effect of the model.
[0082] Furthermore, the embodiments of the present application can add a special layer in the basic architecture according to the special layer information so as to understand the specific functions of the generative model corresponding to the target image.
[0083] Finally, by combining the architecture, loss function, special layers and other hyperparameter information, we can build a complete generative model corresponding to the target image.
[0084] It should be noted that in the embodiment of the present application, each time the source hyperparameters of a target image are predicted, it is equivalent to performing an end-to-end training on multiple encoders, hypernetworks, fingerprint estimation networks, target parsing networks, etc. in the embodiment of the present application, that is, each time multiple encoders, hypernetworks, fingerprint estimation networks, target parsing networks, etc. will be optimized, which can continuously improve the actual application capabilities.
[0085] The present application is described in detail below with reference to a specific embodiment.
[0086] Figure 3 This is a schematic diagram of the architecture of a generative model parsing system based on hyperparameter fingerprints according to an embodiment of the present application. Figure 4 This is a flowchart of a generative model parsing method based on hyperparameter fingerprints according to an embodiment of the present application. Figure 3 and Figure 4 As shown:
[0087] First, the embodiment of the present application can input the image generated by the generative model into the shared one-stage fingerprint estimation network in the instance-aware hypernetwork or feature extractor to extract the one-stage hidden fingerprint of the input image. The one-stage fingerprint estimation network can be composed of six layers of convolution. The first layer can extract the RGB image into a 64-channel feature map. The number of channels does not change in the subsequent five layers to obtain the one-stage hidden fingerprint z 1 .
[0088] Next, the shared one-stage fingerprint estimation network inputs the one-stage hidden fingerprint into the instance-aware encoder in the instance-aware hypernetwork, so as to use the instance-aware hypernetwork to encode the input hidden fingerprint into a 256-dimensional instance prior vector c it , to complete the learning of the perceptual prior knowledge of the input image. The instance-aware encoder may be, but is not limited to, composed of three layers of convolution and three layers of full connection.
[0089] At the same time, the embodiment of the present application can use the model-independent hypernetwork (i.e., the model-independent encoder) to learn the shared information between multiple generative models, that is, the prior knowledge that is independent of the multiple generative models. Specifically, the embodiment of the present application can maintain a set of learnable reference prior vectors c that are independent of the generative model. b and the hyperparameter prior vector c hp , to achieve the learning of prior knowledge that is independent of multiple generative models. Among them, the benchmark prior vector c b Can be a separate 256-dimensional vector, the hyperparameter prior vector c hp It can be composed of 25 256-dimensional vectors corresponding to 25 hyperparameters.
[0090] Then, each hyperparameter prior vector is connected with the benchmark prior vector and the instance prior vector, i.e., broadcast fusion, to obtain the integrated complete prior vector, which contains 25 768-dimensional vectors corresponding to 25 hyperparameters. This is input into a two-layer fully connected hypernetwork that can generate weights, and a set of 25 transformation weights W can be obtained.
[0091] And, the shared one-stage fingerprint estimation network will send the hidden fingerprint of the first stage to the two-stage fingerprint estimation network in the feature extractor, and the two-stage fingerprint estimation network will calculate the hidden fingerprint z of the first stage. 1 Perform convolution again to generate the second-stage hidden fingerprint z 2 Among them, the second-stage fingerprint estimation consists of five layers of convolution, which is the same as the last five layers of the first stage. The input hidden fingerprint z 1 Continue convolution and the output is z 2 .
[0092] Next, z 2Input adjustment layer, perform 25 linear transformations in the channel dimension, the transformation coefficient is the weight W generated above, and get a set of 25 hidden layer hyperparameter fingerprints z hp .
[0093] z hp Input the three-stage fingerprint estimation network, that is, the broadcast fingerprint estimation network in the previous embodiment, to generate a set of 25 hyperparameter fingerprints F hp The third stage is similar to the first stage, with six layers of convolution, but the first layer of convolution is moved to the end.
[0094] It should be noted that in the embodiment of the present application, the process of extracting the latent fingerprint of the input image undergoes three stages of fingerprint estimation networks. Among them, the feature extractor here can be understood to include a fingerprint estimation network of a one-stage fingerprint estimation network and a two-stage fingerprint estimation network, and Figure 1 The shared fingerprint estimation network in the instance-aware supernetwork is also a one-stage fingerprint estimation network, that is, the feature extractor and the instance-aware supernetwork share the same one-stage fingerprint estimation network. Therefore, when the image is actually input, it is input into the one-stage fingerprint estimation network, and then the one-stage fingerprint estimation network outputs the one-stage hidden layer fingerprint to the instance-aware encoder in the instance-aware supernetwork and the two-stage fingerprint estimation network in the feature extractor.
[0095] Finally, the hyperparameter fingerprint F hp Input the target parsing network to predict the values of 25 hyperparameters. The parsing network is divided into two parts. The first part is a shared encoder with five convolutional layers and three fully connected layers, which performs the same operation on the 25 hyperparameter fingerprints. The second part is a branched fully connected, which separates the 25 features output by the first part, inputs them into different layers of fully connected layers, and outputs the predicted values of 25 hyperparameters. Among them, hyperparameters include but are not limited to the architecture category of the generative model corresponding to the input image, loss category, special layers and other hyperparameters.
[0096] According to the generative model parsing method based on hyperparameter fingerprints proposed in the embodiment of the present application, prior knowledge unrelated to multiple models can be learned through a hypernetwork shared between each model, and the perceptual prior knowledge of the target image can be learned through another hypernetwork, and then the two prior knowledge are fused to generate hyperparameter fingerprints, thereby realizing the parsing of the generative model corresponding to the target image. Thus, it is realized that by mining the prior knowledge that various hyperparameters in different generative models are unrelated to the generative model, the traces of the hyperparameters of various generative models can be finely distinguished; through the target image instance, the prior knowledge encoding the target image instance is learned, so as to realize the dynamic response to the changes at the input instance level; further, the prior knowledge unrelated to multiple generative models and the perceptual prior knowledge of the target image are fused and integrated into the feature representation of the fingerprint estimation network to generate hyperparameter fingerprints corresponding to various hyperparameters, which can provide refined information guidance for more accurate parsing of the generative model, and each prediction process can train the entire network structure end-to-end, which can continuously improve the practical application ability of the present application. This solves the problems in related technologies, such as low accuracy of model fingerprints, lack of ability to finely distinguish traces of multiple hyperparameters, and low accuracy of parsing results for generative models.
[0097] Next, a generative model analysis device based on hyperparameter fingerprints proposed according to an embodiment of the present application is described with reference to the accompanying drawings.
[0098] Figure 5 It is a structural diagram of a generative model analysis device based on hyperparameter fingerprints according to an embodiment of the present application.
[0099] like Figure 5 As shown, the generative model parsing device 10 based on hyperparameter fingerprint includes: a learning module 100, a generation module 200 and a parsing module 300.
[0100] Among them, the learning module 100 is used to extract at least one hidden fingerprint of the target image, and use at least one target super network to learn prior knowledge that is independent of multiple generative models, and use at least one hidden fingerprint and another target super network other than the at least one target super network to learn perceptual prior knowledge of the target image.
[0101] The generation module 200 is used to generate a hyperparameter fingerprint of a target image based on at least one hidden layer fingerprint, prior knowledge unrelated to a plurality of generative models, and perceptual prior knowledge.
[0102] The parsing module 300 is used to input the hyperparameter fingerprint into the target parsing network to output the source model hyperparameter prediction result of the target image, and determine the generative model of the target image through the source model hyperparameter prediction result.
[0103] Optionally, in one embodiment of the present application, the learning module 100 includes:
[0104] The learning unit is used to maintain at least one set of baseline prior vectors and hyperparameter vectors independent of the multiple generative models based on image instances of the multiple generative models, so as to learn prior knowledge independent of the multiple generative models.
[0105] Optionally, in one embodiment of the present application, the learning module 100 includes: an encoding unit and a determining unit.
[0106] Wherein, the encoding unit is used to encode at least one hidden layer fingerprint into an instance prior vector;
[0107] The determination unit is used to determine the perceptual prior knowledge according to the instance prior vector.
[0108] Optionally, in one embodiment of the present application, the generation module 200 includes: a fusion unit, a transformation unit and a generation unit.
[0109] The fusion unit is used to fuse prior knowledge and perceptual prior knowledge that are unrelated to multiple generative models to obtain complete prior knowledge, so as to use the complete prior knowledge to generate transformation weights of hidden fingerprints.
[0110] The transformation unit is used to perform a linear transformation on at least one hidden layer fingerprint based on the transformation weight to generate a hidden layer hyperparameter fingerprint.
[0111] The generation unit is used to input the hidden layer hyperparameter fingerprint into the target fingerprint estimation network to generate the hyperparameter fingerprint of the target image.
[0112] Optionally, in one embodiment of the present application, the generation expression of the hidden layer hyperparameter fingerprint can be, but is not limited to, expressed as:
[0113]
[0114] Among them, c b is the reference prior vector, is the hyperparameter prior vector, c it is the instance perception vector, θ represents the parameters of the weight generation hypernetwork h, n hp represents the number of hyperparameters, c com represents the complete prior vector, represents the complete prior vector component corresponding to a single hyperparameter j, W j represents the transformation coefficient component corresponding to a single hyperparameter j, z 2 is the hidden fingerprint obtained after the initial hidden fingerprint is convolved again, z hp represents the hidden layer hyperparameter fingerprint, W represents z 2 Transformation weights for linear transformations.
[0115] It should be noted that the aforementioned explanation of the embodiment of the generative model analysis method based on hyperparameter fingerprint is also applicable to the generative model analysis device based on hyperparameter fingerprint of this embodiment, and will not be repeated here.
[0116] According to the generative model parsing device based on hyperparameter fingerprints proposed in the embodiment of the present application, the prior knowledge unrelated to multiple models can be learned through a hypernetwork shared between each model, and the perceptual prior knowledge of the target image can be learned through another hypernetwork, and then the two prior knowledge are fused to generate a hyperparameter fingerprint, so as to realize the parsing of the generative model corresponding to the target image. Thus, it is realized that by mining the prior knowledge that various hyperparameters in different generative models are unrelated to the generative model, the traces of the hyperparameters of various generative models can be finely distinguished; through the target image instance, the prior knowledge encoding the target image instance is learned, so as to realize the dynamic response to the changes at the input instance level; further, the prior knowledge unrelated to multiple generative models and the perceptual prior knowledge of the target image are fused and integrated into the feature representation of the fingerprint estimation network, and the hyperparameter fingerprints corresponding to various hyperparameters are generated, which can provide refined information guidance for more accurate parsing of the generative model, and each prediction process can train the entire network structure end-to-end, which can continuously improve the practical application ability of the present application. This solves the problems in related technologies, such as low accuracy of model fingerprints, lack of ability to finely distinguish traces of multiple hyperparameters, and low accuracy of parsing results for generative models.
[0117] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include:
[0118] A memory 601 , a processor 602 , and a computer program stored in the memory 601 and executable on the processor 602 .
[0119] When the processor 602 executes the program, the generative model parsing method based on hyperparameter fingerprint provided in the above embodiment is implemented.
[0120] Furthermore, the electronic device further comprises:
[0121] The communication interface 603 is used for communication between the memory 601 and the processor 602 .
[0122] The memory 601 is used to store computer programs that can be executed on the processor 602 .
[0123] The memory 601 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0124] If the memory 601, the processor 602 and the communication interface 603 are implemented independently, the communication interface 603, the memory 601 and the processor 602 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0125] Optionally, in a specific implementation, if the memory 601, the processor 602 and the communication interface 603 are integrated on a chip, the memory 601, the processor 602 and the communication interface 603 can communicate with each other through an internal interface.
[0126] The processor 602 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0127] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned generative model parsing method based on hyperparameter fingerprints.
[0128] An embodiment of the present application also provides a computer program product, including a computer program, which can run computer instructions. When the computer instructions are executed by a processor, the generative model analysis method based on hyperparameter fingerprint provided in the embodiment of the present application is implemented.
[0129] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0130] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0131] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0132] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0133] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one or a combination of multiple of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0134] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0135] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0136] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A generative model analysis method based on hyperparameter fingerprint, characterized in that: The following steps are involved: Extracting at least one hidden fingerprint of a target image and learning prior knowledge unrelated to a plurality of generative models using at least one target super network, and learning perceptual prior knowledge of the target image using the at least one hidden fingerprint and another target super network other than the at least one target super network; generating a hyperparameter fingerprint of the target image based on the at least one hidden fingerprint, the prior knowledge independent of the plurality of generative models, and the perceptual prior knowledge; The hyperparameter fingerprint is input into a target parsing network to output a source model hyperparameter prediction result of the target image, and a generative model of the target image is determined through the source model hyperparameter prediction result.
2. The method according to claim 1, characterized in that The method of using at least one target hypernetwork to learn prior knowledge that is independent of multiple generative models includes: Based on image instances of multiple generative models, at least one set of baseline prior vectors and hyperparameter vectors that are independent of the multiple generative models is maintained to learn the prior knowledge that is independent of the multiple generative models.
3. The method according to claim 1, characterized in that The method of learning the perceptual prior knowledge of the target image by using the at least one hidden fingerprint and another target super network other than the at least one target super network comprises: encoding the at least one hidden fingerprint as an instance prior vector; The perceptual prior knowledge is determined based on the instance prior vector.
4. The method according to claim 1, characterized in that: The step of generating a hyperparameter fingerprint of the target image based on the at least one hidden fingerprint, the prior knowledge independent of the plurality of generative models, and the perceptual prior knowledge comprises: fusing the prior knowledge unrelated to the plurality of generative models and the perceptual prior knowledge to obtain complete prior knowledge, so as to generate a transformation weight of the at least one hidden fingerprint using the complete prior knowledge; Performing a linear transformation on the at least one hidden layer fingerprint based on the transformation weight to generate a hidden layer hyperparameter fingerprint; The hidden layer hyperparameter fingerprint is input into a target fingerprint estimation network to generate a hyperparameter fingerprint of the target image.
5. The method according to claim 4, characterized in that The generation expression of the hidden layer hyperparameter fingerprint is: Among them, c b is the reference prior vector, is the hyperparameter prior vector, c it is the instance perception vector, θ represents the parameters of the weight generation hypernetwork h, n hp represents the number of hyperparameters, c com represents the complete prior vector, represents the complete prior vector component corresponding to a single hyperparameter j, w j represents the transformation coefficient component corresponding to a single hyperparameter j, z2 is the hidden fingerprint obtained after the initial hidden fingerprint is convolved again, and z hp Represents the hidden layer hyperparameter fingerprint, and W represents the transformation weight of the z2 linear transformation.
6. A generative model analysis device based on hyperparameter fingerprint, characterized in that: The following steps are involved: a learning module, configured to extract at least one hidden fingerprint of a target image, and learn prior knowledge unrelated to a plurality of generative models using at least one target supernetwork, and learn perceptual prior knowledge of the target image using the at least one hidden fingerprint and another target supernetwork other than the at least one target supernetwork; A generating module, configured to generate a hyperparameter fingerprint of the target image based on the at least one hidden layer fingerprint, the prior knowledge independent of the plurality of generative models, and the perceptual prior knowledge; The parsing module is used to input the hyperparameter fingerprint into the target parsing network to output the source model hyperparameter prediction result of the target image, and determine the generative model of the target image through the source model hyperparameter prediction result.
7. The device according to claim 6, characterized in that The learning module includes: A learning unit is used to maintain at least one set of reference prior vectors and hyperparameter vectors independent of the multiple generative models based on image instances of the multiple generative models, so as to learn the prior knowledge independent of the multiple generative models.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the generative model parsing method based on hyperparameter fingerprints as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the generative model parsing method based on hyperparameter fingerprint as described in any one of claims 1 to 5.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed, it is used to implement the generative model parsing method based on hyperparameter fingerprint as described in any one of claims 1-5.