Model generation model output characteristic control method, device and equipment and medium
By performing feature map analysis and weight adjustment on the UNet network of the Stable Diffusion model, the problem of uncontrollable output features of the diffusion model is solved, and precise control and quality improvement of the generated images are achieved.
Patent Information
- Application Number
- CN202510108210.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
There are uncontrollable factors when outputting features in the existing diffusion model, which makes it difficult to accurately control the diversity, feature intensity and direction of the generated results, affecting the effect of the model in practical applications.
The set picture is generated through the Stable Diffusion model, and the feature map output of the key level of the UNet network is obtained. The statistical value of each layer of feature map is calculated based on the feature extraction mechanism of PyTorch, the weight is set, and the feature map is adjusted through multiplication operations, the model is rerun to determine the control information, and finally the optimal weight configuration is obtained for controlling the model output.
Accurate control of the output quality and direction of the diffusion model is achieved, ensuring the unified image style and improving the quality and controllability of the generated results.
Smart Images

Figure CN120070616A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model output feature control, and in particular to a method, device, equipment and medium for controlling model output features used in model generation. Background Art
[0002] There are often some uncontrollable factors when existing diffusion models output features. For example, the diversity of output features may lead to output results that do not meet expectations, or the intensity and direction of features cannot be accurately controlled, which affects the effect of the model in practical applications.
[0003] When using diffusion models such as Stable Diffusion to generate model images, there are often uncontrollable factors in the model output features. The main manifestations are:
[0004] 1. The diversity of generated results leads to inconsistent styles;
[0005] 2. The feature intensity and direction are difficult to control accurately;
[0006] 3. The correlation and influence of the output features of each layer of the model are difficult to quantify;
[0007] 4. Lack of precise control mechanism for features at each level of UNet. Summary of the invention
[0008] The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for controlling the output characteristics of a model generated by a model, so as to achieve precise control of the output quality and direction of all common diffusion models.
[0009] In a first aspect, the present invention provides a method for controlling output features of a model generated by a model, comprising the following steps:
[0010] Step 1: Generate the set image through the Stable Diffusion model and obtain the feature map of the key layer output of the UNet network of the Stable Diffusion model;
[0011] Step 2: Extract the feature map of each layer based on the feature extraction mechanism of PyTorch, calculate the mean, standard deviation, maximum value and minimum value of the feature map of each layer in the channel dimension, and record them in the feature distribution dictionary;
[0012] Step 3: Set the weights through the feature distribution dictionary, adjust the feature map of the set layer through multiplication operation, obtain the feature adjustment map, put the feature adjustment map back into the corresponding layer of the UNet network, and then re-run the StableDiffusion model to obtain the generated result, compare the generated result with the image, and obtain the control information of the layer;
[0013] Step 4: Repeat Step 3 to determine the control information for each layer in the key hierarchy, set the weight parameters for each layer in the key hierarchy according to the control information, and obtain the optimal weight configuration, which is used to generate images after the Stable Diffusion model is loaded.
[0014] In a second aspect, the present invention provides a device for controlling the output features of a model for generating models, including:
[0015] A feature map acquisition module that generates a set of pictures through the Stable Diffusion model and acquires the feature maps output by the key hierarchy of the UNet network of the Stable Diffusion model;
[0016] A feature calculation module that extracts the feature maps of each layer based on the feature extraction mechanism of PyTorch, calculates the mean, standard deviation, maximum value, and minimum value of each layer of feature maps in the channel dimension, and records them in the feature distribution dictionary;
[0017] A layer control information acquisition module that sets weights through the feature distribution dictionary, adjusts the feature maps of the set layer through multiplication operations to obtain feature-adjusted maps, returns the feature-adjusted maps to the corresponding hierarchy of the UNet network, then runs the Stable Diffusion model again to obtain the generation result, compares the generation result with the picture, and obtains the control information of this layer;
[0018] A model control information module that repeats the layer control information acquisition module to determine the control information for each layer in the key hierarchy, sets the weight parameters for each layer in the key hierarchy according to the control information, and obtains the optimal weight configuration, which is used to generate images after the Stable Diffusion model is loaded.
[0019] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method described in the first aspect.
[0020] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method described in the first aspect.
[0021] One or more technical solutions provided by the present invention have at least the following technical effects or advantages:
[0022] The present invention can accurately control the output quality and direction of the diffusion model, enabling users to ensure the unity of the image style and the quality of the generated images meeting user requirements when using the diffusion model to generate images.
[0023] By directly loading this configuration file, the present invention quickly applies the verified optimal control parameters. The function automatically parses the weight range and control parameters in the configuration file and applies them to the corresponding levels of the UNet network, thereby achieving precise regulation of the generation process. This export and loading mechanism of the configuration file makes the feature control scheme have good reusability and practicability, and can significantly improve the quality and controllability of the model generation results.
[0024] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically gives the specific implementation manners of the present invention. Brief Description of the Drawings
[0025] The present invention will be further described below with reference to the drawings in conjunction with the embodiments.
[0026] Figure 1 It is a flowchart of the method in Embodiment 1 of the present invention;
[0027] Figure 2 It is a structural schematic diagram of the device in Embodiment 2 of the present invention. Detailed Description of the Invention
[0028] The overall idea of the technical solution in the embodiments of the present application is as follows:
[0029] 1. Realize the output analysis of the Unet layer of the diffusion model;
[0030] 2. Realize the hierarchical control output of the Unet layer of the diffusion model;
[0031] First, use the register_hooks() function to register hook functions at the key levels of the UNet network in Stable Diffusion, including: the input layer, the first downsampling block (128 layers), the middle layer (256 layers), the last upsampling block (64 layers), and the output layer. The feature map outputs of these layers are collected through the hook functions, and the feature statistical data of each layer is recorded for subsequent analysis.
[0032] Extract the feature maps of each layer of the UNet based on the feature extraction mechanism of PyTorch, and use torch.mean(), torch.std(), torch.max(), and torch.min() to calculate the mean, standard deviation, maximum value, and minimum value of each layer of feature maps in the channel dimension, and record these statistical values in the feature distribution dictionary.
[0033] Design feature debugging experiments, and the weight adjustment is realized through multiplication operations. The specific steps are as follows:
[0034] 1. Traverse the feature maps of each layer of the UNet. For each layer:
[0035] - Use a for loop to traverse the preset list of weight values [0.1, 0.5, 1.0, 1.5, 2.0]
[0036] - For each weight value w, multiply the feature map by the weight: feature_map = feature_map * w;
[0037] - Among them, 1.0 means keeping the original features unchanged, <1.0 suppresses the feature influence through multiplication, and >1.0 enhances the feature influence through multiplication;
[0038] 2. After each weight adjustment, record and evaluate the changes in the generated results:
[0039] - Facial details: Evaluate the clarity of facial features, skin texture, etc.;
[0040] - Pose performance: Evaluate the naturalness of the standing pose, limb coordination, etc.;
[0041] - Background effect: Evaluate the background texture details, spatial hierarchy, etc.;
[0042] Through this systematic weight traversal and feature adjustment experiment, the influence degree of different weight values on the generation effect can be quantitatively analyzed. If modifying a certain weight has a huge impact on the generated result, it indicates that this layer is the UNet layer that controls the corresponding feature.
[0043] After obtaining the information of the specific feature control layer, for example, the downsampling block (128 layers) mainly controls the facial features and skin texture, the middle layer (256 layers) mainly controls the model pose and overall composition, and the last upsampling block (64 layers) mainly controls the background and overall color tone. By separately controlling the features of these three levels, precise adjustment of different aspects of the generated image can be achieved. For example:
[0044] - 128-layer downsampling block:
[0045] - Layers 1 - 8 control the clarity of facial features. Among them, layers 1 - 3 mainly control the details of the eyes, layers 4 - 6 control the nose contour, and layers 7 - 8 control the mouth features;
[0046] - Layers 9 - 16 control the fineness of skin texture. Among them, layers 9 - 12 are responsible for skin texture, and layers 13 - 16 are responsible for skin color uniformity;
[0047] - Layers 17 - 24 control the naturalness of facial expressions. Among them, layers 17 - 20 control the expressions of eyebrows and eyes, and layers 21 - 24 control the curvature of the mouth corners
[0048] -......
[0049] -256 intermediate layers:
[0050] - Layers 1-20 control the naturalness of the model's posture;
[0051] - Layers 21-40 control the coordination of limb movements;
[0052] - Layers 41-60 control the overall aesthetic of the composition
[0053] -…….
[0054] -64-layer upsampling block:
[0055] - Layers 1-10 control the degree of reality of the background;
[0056] - Layers 11-20 control the color atmosphere of the overall picture;
[0057] - Layers 21-30 control the spatial layering and depth of field effects;
[0058] -…….
[0059] It is necessary to determine the optimal weight control range for each layer. In order to avoid extreme values in the feature map caused by weight adjustment and damage to the generated results, we need to analyze the distribution characteristics of each layer of features. According to the feature distribution dictionary, we can get the normal distribution range of each layer of features. When the feature value exceeds the mean ±3 times the standard deviation, it means that the weight adjustment is too large, resulting in feature abnormality. In this way, the safe adjustment range of each layer of features can be determined to avoid generation crashes.
[0060] For example, if the mean of a layer of features is 0.5 and the standard deviation is 0.1, the range after weight adjustment should be limited to [0.7, 1.3] to ensure that the adjusted feature value is still within the normal distribution range. Specifically, when the feature mean is 0.5, if the weight is set to 0.7, the feature value will be scaled to about 0.35; if the weight is set to 1.3, the feature value will be amplified to about 0.65. This range corresponds exactly to the interval of mean ± 3 times standard deviation [0.2, 0.8]. In this way, we can effectively regulate the feature strength while maintaining the stability of the feature distribution. For features at other levels, a similar method can be used to determine the appropriate weight adjustment range based on their mean and standard deviation, thereby achieving precise control of the model generation process as a whole.
[0061] Finally, the verified optimal weight configuration is exported. This function saves the control parameters of each layer as a configuration file in JSON format, containing key information such as:
[0062] -128-layer downsampling block weight range:
[0063] - Facial feature control layer (layers 1 - 8): [0.7, 1.2];
[0064] - Skin texture control layer (layers 9 - 16): [0.8, 1.3];
[0065] - Expression control layer (layers 17 - 24): [0.7, 1.3];
[0066] -......
[0067] - Weight range of the 256 - layer intermediate layer:
[0068] - Standing pose control layer (layers 1 - 20): [0.8, 1.2];
[0069] - Action control layer (layers 21 - 40): [0.85, 1.15];
[0070] - Composition control layer (layers 41 - 60): [0.9, 1.1];
[0071] -......
[0072] - Weight range of the 64 - layer upsampling block:
[0073] - Background control layer (layers 1 - 10): [0.9, 1.1];
[0074] - Color control layer (layers 11 - 20): [0.85, 1.15];
[0075] - Space control layer (layers 21 - 30): [0.95, 1.05];
[0076] -......
[0077] For a new generation task, the configuration file can be directly loaded through the load_control_config() function to quickly apply the verified optimal control parameters. The function will automatically parse the weight range and control parameters in the configuration file and apply them to the corresponding levels of the UNet network, thereby achieving precise regulation of the generation process. This mechanism for exporting and loading the configuration file makes the feature control scheme have good reusability and practicability, and can significantly improve the quality and controllability of the model generation results.
[0078] Example 1
[0079] As Figure 1 shown, this example provides a method for controlling the output features of a model for generating models, including the following steps:
[0080] Step 1: Generate a set picture through the Stable Diffusion model, and obtain the feature map output by the key levels of the UNet network of the Stable Diffusion model;
[0081] Step 2: Extract the feature map of each layer based on the feature extraction mechanism of PyTorch, calculate the mean, standard deviation, maximum value and minimum value of each layer of feature map in the channel dimension, and record them in the feature distribution dictionary;
[0082] Step 3: Set weights through the feature distribution dictionary, adjust the feature map of the set layer through multiplication operation to obtain a feature adjustment map, put the feature adjustment map back into the corresponding level of the UNet network, then re-run the StableDiffusion model to obtain the generation result, compare the generation result with the picture, and obtain the control information of this layer;
[0083] Step 4: Repeat Step 3 to determine the control information of each layer in the key level, set the weight parameters of each layer in the key level according to the control information to obtain the optimal weight configuration, and the optimal weight configuration is used to generate images after the Stable Diffusion model is loaded.
[0084] In this embodiment, preferably, Step 1 is specifically: Generate a set picture through the Stable Diffusion model, use the register_hooks function to register a hook function at the key levels of the UNet network of the Stable Diffusion model, and the key levels include the input layer, the first downsampling block, the middle layer, the upsampling block and the output layer; Collect the feature maps output by the key levels through the register_hooks function and record the feature maps of each layer.
[0085] In this embodiment, preferably, Step 2 is specifically: Extract the feature map of each layer based on the feature extraction mechanism of PyTorch, and use the torch.mean function, torch.std function, torch.max function and torch.min function to calculate the mean, standard deviation, maximum value and minimum value of each layer of feature map in the channel dimension, and record them in the feature distribution dictionary.
[0086] In this embodiment, preferably, Step 3 is specifically:
[0087] Set the weight list [0.1, 0.5, 1.5, 2.0] through the feature distribution dictionary. Multiply each weight value in the weight list with the feature map of the set layer through multiplication operation to obtain the feature adjustment maps respectively. Put the feature adjustment maps back into the corresponding levels of the UNet network. Then run the Stable Diffusion model again to obtain the generation result. Compare the generation result with the said picture to obtain the change description, where the change description includes face details, pose performance or background effect, so as to obtain the control information of the set layer;
[0088] The multiplication operation is specifically as follows: for each weight value w, multiply the feature map by the weight value: feature_map[i + 1] = feature_map[i] × w; where, when w is less than 1.0, it means suppressing the feature influence; when w is greater than 1.0, it means enhancing the feature influence.
[0089] Based on the same inventive concept, the present application also provides a device corresponding to the method in Embodiment 1, details of which are shown in Embodiment 2.
[0090] Embodiment 2
[0091] As Figure 2 shown, in this embodiment, a device for controlling the output features of a model for generating models is provided, including:
[0092] A feature map acquisition module, which generates a set picture through the Stable Diffusion model and acquires the feature maps output by the key levels of the UNet network of the Stable Diffusion model;
[0093] A feature calculation module, which extracts the feature maps of each layer based on the feature extraction mechanism of PyTorch, calculates the mean, standard deviation, maximum value and minimum value of each layer of feature maps in the channel dimension, and records them in the feature distribution dictionary;
[0094] A layer control information acquisition module, which sets weights through the feature distribution dictionary, adjusts the feature maps of the set layer through multiplication operation to obtain the feature adjustment maps, puts the feature adjustment maps back into the corresponding levels of the UNet network, then runs the Stable Diffusion model again to obtain the generation result, and compares the generation result with the said picture to obtain the control information of this layer;
[0095] A model control information module, which repeats the layer control information acquisition module to determine the control information of each layer in the key levels, sets the weight parameters of each layer in the key levels according to the control information to obtain the optimal weight configuration, and the optimal weight configuration is used to generate images after the Stable Diffusion model is loaded.
[0096] In this embodiment, preferably, the feature map acquisition module is specifically: generating a set of pictures through the Stable Diffusion model, using the register_hooks function to register hook functions at key levels of the UNet network of the Stable Diffusion model, where the key levels include the input layer, the first downsampling block, the middle layer, the upsampling block, and the output layer; collecting the feature maps output by the key levels through the register_hooks function and recording the feature maps of each layer.
[0097] In this embodiment, preferably, the feature calculation module is specifically: extracting the feature maps of each layer based on the feature extraction mechanism of PyTorch, using the torch.mean function, torch.std function, torch.max function, and torch.min function to calculate the mean, standard deviation, maximum value, and minimum value of each layer of feature maps in the channel dimension, and recording them in the feature distribution dictionary.
[0098] In this embodiment, preferably, the layer control information acquisition module is specifically:
[0099] Setting a weight list [0.1, 0.5, 1.5, 2.0] through the feature distribution dictionary, adjusting each weight value in the weight list with the feature map of the set layer through multiplication operations to obtain feature adjustment maps respectively, putting the feature adjustment maps back into the corresponding levels of the UNet network, then running the Stable Diffusion model again to obtain the generation result, comparing the generation result with the picture to obtain a change description, where the change description includes face details, pose performance, or background effect, to obtain the control information of the set layer;
[0100] The multiplication operation is specifically: for each weight value w, multiplying the feature map by the weight value: feature_map[i + 1] = feature_map[i] × w; where when w is less than 1.0, it means suppressing the feature influence; when w is greater than 1.0, it means enhancing the feature influence.
[0101] Since the device introduced in the second embodiment of the present invention is the device used to implement the method of the first embodiment of the present invention, based on the method introduced in the first embodiment of the present invention, those skilled in the art can understand the specific structure and variations of the device, so it will not be elaborated here. Any device used in the method of the first embodiment of the present invention belongs to the scope protected by the present invention.
[0102] Based on the same inventive concept, this application provides an electronic device embodiment corresponding to the first embodiment, as detailed in the third embodiment.
[0103] Embodiment Three
[0104] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, any implementation manner in Embodiment 1 can be implemented.
[0105] Since the electronic device introduced in this embodiment is the device used to implement the method in Embodiment 1 of this application, based on the method introduced in Embodiment 1 of this application, those skilled in the art can understand the specific implementation manner and various variation forms of the electronic device in this embodiment. Therefore, the specific implementation of how this electronic device implements the method in the embodiments of this application will not be introduced in detail here. As long as the device used by those skilled in the art to implement the method in the embodiments of this application belongs to the scope protected by this application.
[0106] Based on the same inventive concept, this application provides a storage medium corresponding to Embodiment 1, as detailed in Embodiment 4.
[0107] Embodiment 4
[0108] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any implementation manner in Embodiment 1 can be implemented.
[0109] The technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0110] Through the technical solution of this embodiment, the control information of each layer in the UNet network can be obtained, and then corresponding weight parameters are set for each layer, so that the Stable Diffusion model can control the output according to the weight parameters when generating pictures, greatly meeting the needs of users.
[0111] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0114] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or means for implementing the functions specified in multiple blocks.
[0115] Although the specific embodiments of the present invention have been described above, those skilled in the art of this technology should understand that the specific embodiments we described are illustrative rather than used to limit the scope of the present invention. Equivalent modifications and variations made by those skilled in the art in accordance with the spirit of the present invention should all be covered by the scope protected by the claims of the present invention.
Claims
1. A method for controlling output features of a model for model generation, characterized in that: The steps include: Step 1: Generate the set image through the Stable Diffusion model and obtain the feature map of the key layer output of the UNet network of the Stable Diffusion model; Step 2: Extract the feature map of each layer based on the feature extraction mechanism of PyTorch, calculate the mean, standard deviation, maximum value and minimum value of the feature map of each layer in the channel dimension, and record them in the feature distribution dictionary; Step 3: Set the weights through the feature distribution dictionary, adjust the feature map of the set layer through multiplication operation, obtain the feature adjustment map, put the feature adjustment map back to the corresponding layer of the UNet network, and then re-run the Stable Diffusion model to obtain the generated result, compare the generated result with the image, and obtain the control information of the layer; Step 4: Repeat step 3 to determine the control information of each layer in the key hierarchy, set the weight parameters of each layer in the key hierarchy according to the control information, and obtain the optimal weight configuration, which is used to generate an image after the Stable Diffusion model is loaded.
2. A method for controlling output characteristics of a model generation model according to claim 1, characterized in that: The step 1 is specifically as follows: a set image is generated through a Stable Diffusion model, and a hook function is registered at a key level of a UNet network of the Stable Diffusion model using a register_hooks function, where the key levels include an input layer, a first downsampling block, an intermediate layer, an upsampling block, and an output layer; the feature maps output by the key levels are collected through the register_hooks function, and the feature maps of each layer are recorded.
3. The method for controlling output characteristics of a model for model generation according to claim 1, characterized in that: The step 2 is specifically as follows: extracting the feature map of each layer based on the feature extraction mechanism of PyTorch, using the torch.mean function, torch.std function, torch.max function and torch.min function to calculate the mean, standard deviation, maximum value and minimum value of the feature map of each layer in the channel dimension, and recording them in the feature distribution dictionary.
4. The method for controlling output characteristics of a model generation model according to claim 1, characterized in that: The step 3 is specifically as follows: A weight list [0.1, 0.5, 1.5, 2.0] is set through a feature distribution dictionary, and each weight value in the weight list is adjusted by a multiplication operation with the feature map of the set layer to obtain feature adjustment maps respectively, and the feature adjustment maps are put back into the corresponding layer of the UNet network, and then the Stable Diffusion model is re-run to obtain a generated result, and the generated result is compared with the image to obtain a change description, which includes facial details, posture performance or background effects, and obtains control information of the set layer; The multiplication operation is specifically as follows: for each weight value w, multiply the feature map by the weight value: feature_map[i+1]=feature_map[i]×w; wherein, when w is less than 1.0, it indicates suppressing the feature influence; and when w is greater than 1.0, it indicates enhancing the feature influence.
5. A device for controlling output characteristics of a model for model generation, characterized in that: include: Get the feature map module, generate the set image through the Stable Diffusion model, and obtain the feature map of the key layer output of the UNet network of the Stable Diffusion model; The feature calculation module extracts the feature map of each layer based on the feature extraction mechanism of PyTorch, calculates the mean, standard deviation, maximum value and minimum value of each layer of feature map in the channel dimension, and records them in the feature distribution dictionary; Get the layer control information module, set the weight through the feature distribution dictionary, adjust the feature map of the set layer through multiplication operation to obtain the feature adjustment map, put the feature adjustment map back to the corresponding layer of the UNet network, and then re-run the Stable Diffusion model to obtain the generated result, compare the generated result with the image, and obtain the control information of the layer; The model control information module repeatedly obtains the layer control information module, determines the control information of each layer in the key layer, sets the weight parameters of each layer in the key layer according to the control information, and obtains the optimal weight configuration, which is used to generate an image after the Stable Diffusion model is loaded.
6. The device for controlling output characteristics of a model for model generation according to claim 5, characterized in that: The feature map acquisition module is specifically as follows: a set image is generated through a Stable Diffusion model, and a hook function is registered at a key level of the UNet network of the Stable Diffusion model using the register_hooks function, where the key levels include an input layer, a first downsampling block, an intermediate layer, an upsampling block, and an output layer; the feature maps output by the key levels are collected through the register_hooks function, and the feature maps of each layer are recorded.
7. The device for controlling output characteristics of a model for model generation according to claim 5, characterized in that: The feature calculation module is specifically as follows: based on the feature extraction mechanism of PyTorch, the feature map of each layer is extracted, and the mean, standard deviation, maximum value and minimum value of the feature map of each layer in the channel dimension are calculated using the torch.mean function, torch.std function, torch.max function and torch.min function, and recorded in the feature distribution dictionary.
8. The device for controlling output characteristics of a model for model generation according to claim 5, characterized in that: The acquisition layer control information module is specifically: A weight list [0.1, 0.5, 1.5, 2.0] is set through a feature distribution dictionary, and each weight value in the weight list is adjusted by a multiplication operation with the feature map of the set layer to obtain feature adjustment maps respectively, and the feature adjustment maps are put back into the corresponding layer of the UNet network, and then the Stable Diffusion model is re-run to obtain a generated result, and the generated result is compared with the image to obtain a change description, which includes facial details, posture performance or background effects, and obtains control information of the set layer; The multiplication operation is specifically as follows: for each weight value w, multiply the feature map by the weight value: feature_map[i+1]=feature_map[i]×w; wherein, when w is less than 1.0, it indicates suppressing the feature influence; and when w is greater than 1.0, it indicates enhancing the feature influence.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 4 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.