Image generation method and device, electronic equipment and storage medium

By acquiring text description content and image generation model, and using area masks and control parameters to generate multi-subject frame images, the problem of difficulty in controlling multiple subject positions and local effects in the prior art is solved, and refined control and efficient multi-subject frame image generation are achieved.

CN120107390APending Publication Date: 2025-06-06NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510252013.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the field of image generation, it is difficult to control the positional relationship of multiple subjects on the image based on text description statements, and the part can only be simply redrawn through text description statements, making it difficult to achieve precise local effect control.

Method used

By obtaining the input text description content and the pre-trained image generation model, the area mask and type description information of multiple subjects are obtained, the area control parameters are obtained according to the area mask, the global display information of the subject in the target image area is controlled, and the target image containing multiple subjects is generated.

Benefits of technology

The image generation of multiple subjects in the same frame is realized, ensuring that the positional relationship of multiple subjects is consistent with the definition of the area mask, and the subject effect of the target image area is refined to meet user needs, and improving the same frame effect of multiple subjects in the same frame image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107390A_ABST
    Figure CN120107390A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, electronic equipment and a storage medium, and relates to the technical field of image generation. The method comprises the steps that input text description content and a pre-trained image generation model are obtained, and the text description content comprises type description information of multiple subjects; obtaining region masks of the plurality of main bodies, wherein the region masks are used for indicating image regions where the plurality of main bodies are located in the to-be-generated image; obtaining an area control parameter of a target image area in the plurality of image areas, wherein the area control parameter is used for controlling global display information of a corresponding main body in the target image area; and according to the type description information and the area control parameters of the plurality of subjects, generating a target image containing the plurality of subjects by adopting an image generation model. According to the invention, multi-main-body same-frame image generation can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image generation technology, and in particular to an image generation method, device, electronic device and storage medium. Background Art

[0002] With the continuous development of artificial intelligence (AI) technology, in the field of image generation, AI's text-based image technology, which automatically generates images based on some text description sentences, has also been widely used in various computer service scenarios.

[0003] In the process of generating images containing multiple subjects through text description sentences in the field of image generation, it is difficult to control the positional relationship of multiple subjects on the image based on the text description sentences, and only simple redrawing of the local part can be performed through the text description sentences, which makes it difficult to achieve precise local effect control. Summary of the invention

[0004] The purpose of the present application is to provide an image generation method, device, electronic device and storage medium to address the deficiencies in the above-mentioned prior art, so as to realize the image generation of multiple subjects in the same frame.

[0005] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows:

[0006] In a first aspect, an embodiment of the present application provides an image generation method, the method comprising:

[0007] Acquire input text description content and a pre-trained image generation model, wherein the text description content includes: type description information of multiple subjects;

[0008] Acquire region masks of the multiple subjects, where the region masks are used to indicate multiple image regions where the multiple subjects are located in the image to be generated;

[0009] Acquire, according to the region masks of the multiple subjects, a region control parameter of a target image region in the multiple image regions, wherein the region control parameter is used to control global display information of a corresponding subject in the target image region, and the global display information is used to indicate an overall display parameter of the corresponding subject in the target image region;

[0010] According to the type description information of the multiple subjects and the region control parameters, the image generation model is used to generate a target image containing the multiple subjects.

[0011] In a second aspect, an embodiment of the present application further provides an image generating device, the device comprising:

[0012] A model and text acquisition module, used to acquire input text description content and a pre-trained image generation model, wherein the text description content includes: type description information of multiple subjects;

[0013] A mask acquisition module, used for acquiring region masks of the plurality of subjects, wherein the region masks are used for indicating a plurality of image regions where the plurality of subjects are located in the image to be generated;

[0014] A parameter acquisition module, used for acquiring, according to the area masks of the multiple subjects, area control parameters of a target image area in the multiple image areas, the area control parameters being used for controlling global display information of the corresponding subject in the target image area, the global display information being used for indicating the overall display parameters of the corresponding subject in the target image area;

[0015] An image generation module is used to generate a target image containing the multiple subjects using the image generation model according to the type description information of the multiple subjects and the area control parameters.

[0016] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores program instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the program instructions to perform the steps of the image generation method as described in any one of the first aspects.

[0017] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image generation method as described in any one of the first aspects are executed.

[0018] The beneficial effects of this application are:

[0019] The image generation method, device, electronic device and storage medium provided by the present application determine the positions of multiple subjects in an image based on a region mask, and control the global display information of the subjects in the target image region based on the region control parameters, so that in the target image generated by the image generation model according to the type description information of the multiple subjects and the region control parameters, the positional relationship of the multiple subjects is consistent with the positional relationship defined by the region mask, and the effect of the subjects in the target image region is finely controlled, which meets user needs and improves the same-frame effect of images with multiple subjects in the same frame. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0021] Figure 1 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 1 ;

[0022] Figure 2 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 2 ;

[0023] Figure 3 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 3 ;

[0024] Figure 4 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 4 ;

[0025] Figure 5 A schematic diagram of multiple entities in the same frame provided by an embodiment of the present application;

[0026] Figure 6 A schematic diagram of multiple subjects in a same frame with multiple styles provided in an embodiment of the present application;

[0027] Figure 7 A schematic diagram of the structure of an image generating device provided in an embodiment of the present application;

[0028] Figure 8 A schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0030] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.

[0031] In addition, the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] It should be noted that, in the absence of conflict, the features in the embodiments of the present application may be combined with each other.

[0033] Figure 1 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 1 ,like Figure 1 As shown, the method may include:

[0034] S101, obtaining input text description content and a pre-trained image generation model, wherein the text description content includes: type description information of multiple subjects.

[0035] In this embodiment, the text description content is content input by the user, which is used to describe the information contained in the image that the user expects to generate. In the image generation of multiple subjects in the same frame, the text description content at least includes type description information of multiple subjects, such as "a cat and a dog are playing on the grass", where "cat" and "dog" are two subjects. Furthermore, the text description content can also include: interaction description information between multiple subjects, such as "the cat is chasing the dog, and the dog is running away", and can also include scene description information where multiple subjects are located, such as "a group of people are having a party on the beach", where the scene description information includes beach and party.

[0036] The image generation model is a pre-trained artificial intelligence model. The image generation model can be determined according to the style of the image that the user expects to generate. If the user expects to generate a realistic-style image, an image generation model that is good at realistic style can be selected. If the user expects to generate a cartoon-style image, an image generation model that is good at cartoon style can be selected. Among them, the model loading technology "CreateHookModelAsLora" can be used to load an image generation model from the model library.

[0037] S102 . Obtain region masks of multiple subjects, where the region masks are used to indicate multiple image regions where the multiple subjects are located in the image to be generated.

[0038] In this embodiment, in order to specify the position of the subject in the generated image, a regional mask mask can be set for the subject. If the position of the entire subject needs to be specified, a regional mask mask can be set for the entire subject. If only the position of part of the subject needs to be specified, a regional mask mask can be set for part of the subject. This embodiment does not impose any restrictions on this.

[0039] In some embodiments, a regional mask can be set for each of the multiple subjects whose positions need to be specified, that is, each subject corresponds to a regional mask, and each regional mask is used to indicate the image area where the corresponding subject is located. For example, if the multiple subjects are the sky and the ground, two regional masks can be created, one covering the sky area and the other covering the ground area.

[0040] In other embodiments, a region mask can be set for multiple subjects whose positions need to be specified, and multiple subjects corresponding to the multiple regions of the region mask can be marked respectively, wherein different color values ​​or gray values ​​can be used to distinguish the multiple regions in the region mask. For example, white can be used to represent the sky area, gray to represent the ground area, and black to represent the area where no operation is performed; different colors can be used to represent different characters, such as red for character A, green for character B, and blue for character C.

[0041] In some implementations, the region mask may be generated by manually drawing the region mask, automatically generating the region mask, or using a predefined region mask.

[0042] Manually drawn area mask means that the user uses the drawing tool to manually draw on the canvas to specify the shape and position of the area mask; automatically generated area mask means that the system automatically generates it according to preset rules or algorithms. For example, if the text description content describes "cat" or "dog", an area mask with the corresponding shape of "cat" or "dog" is automatically generated; predefined area mask means an area mask with a fixed shape predefined by the system, such as a circle, rectangle, etc. The user can directly select and use it, and determine the image area corresponding to the selected area mask in the canvas.

[0043] S103. Acquire region control parameters of a target image region among the multiple image regions according to region masks of the multiple subjects, wherein the region control parameters are used to control global display information of the corresponding subject in the target image region, and the global display information is used to indicate overall display parameters of the corresponding subject in the target image region.

[0044] In this embodiment, in order to achieve refined control of local effects, a target image area that needs to be refined controlled can be selected, and regional control parameters for refined control can be set for the target image area. The regional control parameters define the type of operation performed on the target image area, as well as the intensity of the operation. Based on the regional control parameters, the global display information of the corresponding subject in the target image area can be controlled.

[0045] The overall display parameter may be parameters of multiple display attribute types or parameters of one display attribute type, for example, the overall display parameter may include brightness, color, texture, etc. of the corresponding subject in the target image area.

[0046] In some embodiments, the region mask and region control parameters of the target image region are transferred to the region controller RegionController to perform region registration on the target control region, and the registered region controller RegionController is stored.

[0047] S104: Generate a target image containing multiple subjects using an image generation model according to the type description information and region control parameters of the multiple subjects.

[0048] In this embodiment, the Clip (Contrastive Language–Image Pretraining) text encoding model is used to perform text encoding on the type description information, interaction description information, scene description information, etc. of multiple subjects to generate a feature vector, and the feature vector, region mask, and region control parameter are input into the image generation model. The image generation model generates a target image containing multiple subjects according to the feature vector and the region mask, and determines the global display information of the subject corresponding to the target image region according to the region control parameters to obtain a target image containing multiple subjects.

[0049] The image generation method provided by the above embodiment determines the positions of multiple subjects in the image based on the region mask, and controls the global display information of the subjects in the target image region based on the region control parameters, so that in the target image generated by the image generation model according to the type description information of the multiple subjects and the region control parameters, the positional relationship of the multiple subjects is consistent with the positional relationship defined by the region mask, and the effect of the subjects in the target image region is finely controlled, which meets user needs and improves the same-frame effect of the multi-subject same-frame image.

[0050] In one possible implementation, Figure 2 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 2 ,like Figure 2As shown, the process of obtaining the region control parameter of the target image region among the multiple image regions according to the region masks of the multiple subjects in S103 may include:

[0051] S201: Obtain at least one attribute parameter value corresponding to a preset display attribute type of a target image area.

[0052] S202: Generate a region control parameter of the target image region according to a region mask, a display attribute type, and at least one attribute parameter value of the target image region.

[0053] In this embodiment, the image to be generated has multiple display attribute types, each of which is used to control the display content of the image in a display dimension. The user can customize the display attribute type of the target image area. The display attribute type (Operation Type) is used to indicate the type of operation performed on the target image area, such as style transfer, color adjustment, content replacement, etc., and according to the selected display attribute type, at least one attribute parameter value (Parameter Values) can be set for the display attribute type. The attribute parameter value is used to indicate the operation parameter value performed on the target image area. For example, if the display attribute type is style transfer, the attribute parameter value is style information. If the display attribute type is color adjustment, the attribute parameter value is hue offset, saturation scaling ratio, brightness adjustment value, etc. If the display attribute type is content replacement, the attribute parameter value is replacement of image content.

[0054] The region mask, the display attribute type and at least one attribute parameter value of the target image region are associated to obtain the region control parameter of the target image region, and the region controller of the target image region is obtained by registering the region.

[0055] In some embodiments, the regional control parameters may further include: Blend Strength, which is used to indicate the degree of integration of the control effect of the regional control parameters with the image. A value of 1 indicates that the control effect of the regional control parameters is fully applied, and a value of 0 indicates that the control effect of the regional control parameters is not applied.

[0056] In other embodiments, the region control parameters may further include: a feather radius, where the feather radius is used to indicate the region edge transition range of the target image region controlled by the region control parameters, wherein the larger the feather radius, the softer and more natural the transition of the region edge, and the smoother the transition between different regions.

[0057] In some other embodiments, the region control parameters may further include: a control model, and the control model may process the generation of the target image region separately.

[0058] The image generation method provided in the above embodiment can finely adjust the attribute parameter values ​​of various display attribute types of the image region through the region control parameters, thereby improving the refinement effect of the generated target image.

[0059] In another possible implementation, Figure 3 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 3 ,like Figure 3 As shown, the method may also include:

[0060] S201: Obtain at least one attribute parameter value corresponding to a preset display attribute type of a target image area.

[0061] S203: Acquire at least one condition control information of the target image area, where the condition control information is used to control the detailed display information of the subject in the target image area.

[0062] S204: Generate a region control parameter of the target image region according to the region mask of the target image region, the display attribute type, at least one attribute parameter value, and conditional control information of the target image region.

[0063] In this embodiment, in addition to defining the display attribute type and attribute parameter value of the target image area, attribute information that cannot be defined using attribute parameter values ​​can be represented by conditional control information.

[0064] The conditional control information (Conditioning) can control the detailed display information of the subject in the target image area. The detailed display information can be the local display information of the subject in the target image area, or the refined display information of the subject in the target image area in the target display attribute type.

[0065] Based on the conditional control information, conditional control of the target image area can be achieved, wherein the attribute parameter value is the overall parameter value of the display attribute type of the target image area, and the detailed display information can be the local display information of the sub-area in the target image area. For example, the attribute parameter value specifies the color of the cat in the target image area, and the conditional control information can define the texture fineness of the cat's hair. The attribute parameter value specifies the overall color of the person in the target image area, and the conditional control information can define the person's clothing color, hat color, etc.; or, the detailed display information can be the refined display information of the target display attribute type for the main body of the target image area, for example, the attribute parameter value specifies the brightness value of the target image area, and the conditional control information can describe "making the texture of the target image area finer" or "making the color of the target image area closer to the color of the reference image".

[0066] The region mask, display attribute type, at least one attribute parameter value and conditional control information of the target image region are associated to obtain the region control parameter of the target image region, and the region controller of the target image region is obtained by registering the region.

[0067] In some embodiments, the process of obtaining the condition control information of the target image area in S203 may include:

[0068] According to the input conditional text information, conditional control information of the target image area is determined.

[0069] In this embodiment, the conditional control information of the target image area can be determined by inputting conditional text information, wherein the description information of the target image area in the text description content can be extracted, or the conditional text information for the target image area can be directly input to determine the conditional control information of the target image area. For example, the conditional text information can be "make the sky bluer", "make the character's expression happier", "make the texture of the target image area more delicate".

[0070] In some other embodiments, the process of obtaining the condition control information of the target image area in S203 may include:

[0071] By extracting the target features of the reference image, the conditional control information of the target image area is determined.

[0072] In this embodiment, if the display effect of the target image area in the specified display dimension is expected to be close to the reference image, the reference image can be input, and the conditional control information of the target image area can be determined by extracting the target features of the reference image in the specified display dimension. For example, if the color of the target image area is expected to be close to the reference image, the color features of the reference image can be extracted as the conditional control information, and if the style of the target image area is expected to be close to the reference image, the style features of the reference image can be extracted as the conditional control information.

[0073] The image generation method provided in the above embodiment further controls the detail display information of the target image area through the condition control information, so that the main body of the target image area in the generated image has more specific details, thereby improving the image generation effect.

[0074] In a possible implementation, before the above S104 uses the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, the method may further include:

[0075] The region weights of the plurality of image regions are determined according to region attribute information of the plurality of image regions, wherein the region attribute information includes at least one of the following information: region shape, region size, region edge smoothness, region overlap, and region content.

[0076] The process of generating a target image containing multiple subjects using an image generation model according to the type description information and the region control parameters of the multiple subjects in S104 may include:

[0077] According to the type description information of multiple subjects, the region control parameters and the region weights of multiple image regions, an image generation model is used to generate a target image.

[0078] In this embodiment, after different image areas are independently controlled using regional control parameters, multiple image areas need to be fused to achieve a smooth transition and a natural fusion effect between different image areas and avoid obvious stitching traces, wherein the regional weight of each image area can be determined based on the contribution of each image area to the generated target image.

[0079] Specifically, the region weight of each image region is determined according to the region attribute information of each image region.

[0080] Among them, the region shape is the shape defined by the region mask. The more regular the region shape and the clearer the boundary, the higher the region weight. The more irregular the region shape and the more blurred the boundary, the lower the region weight.

[0081] The region size is the size of the area covered by the region mask. The larger the region size, the more important the corresponding subject is in the image. The larger the region weight, the smaller the region size, and the smaller the region weight.

[0082] The regional edge smoothness (feathering) is the degree of feathering of the regional mask edge. When calculating the regional weight, the weight of the edge region will gradually decrease, thereby achieving a smooth transition of the edges of different regions. The higher the regional edge smoothness, the faster the regional weight of the edge region decreases. The regional weight can be calculated based on the distance from the pixel to the edge of the regional mask. The farther from the edge of the regional mask, the higher the regional weight, and the closer to the edge of the regional mask, the lower the regional weight.

[0083] The region overlap is used to indicate that when there is overlap between regions, in order to avoid excessive amplification of the features of the overlapping regions, the weights of the regions in the overlapping parts need to be normalized to ensure that the sum of the weights of the regions in the overlapping parts is 1.

[0084] The regional content is the information contained in the region. The regional weight can be determined according to the importance of the information contained in the region in the image, or according to the difference between the information contained in the region and the information contained in other regions. The higher the importance of the information contained in the region in the image, the greater the regional weight, and the greater the difference between the information contained in the region and the information contained in other regions, the smaller the regional weight. The attention mechanism can be used to learn the importance or difference of the information contained in different regions, and dynamically calculate the regional weight of each image region.

[0085] In some embodiments, the region attribute information may also include a user-defined weight or a user-defined region preference.

[0086] According to the one or more regional attribute information, a weight value is calculated for each pixel point in each image region, and the weight value determines the degree to which the feature of each pixel point shares the feature of the final image.

[0087] The image generation model uses the regional weights of multiple image regions to perform weighted fusion based on the regional features of the generated multiple image regions to generate the final target image. The mathematical formula of weighted fusion can be expressed as:

[0088]

[0089] Among them, B(x,y) is the pixel value of the fused image at the pixel point (x,y), W i (x, y) is the weight of the i-th image region at the pixel point (x, y), R i (x,y) is the eigenvalue of the i-th image region at the pixel point (x,y).

[0090] It can be seen that the larger the regional weight, the greater the influence of the regional features on the generated image.

[0091] For example, assuming that there are two overlapping image areas A and image area B, for a pixel point P, it is located at the center of the regional mask mask of image area A and far away from the center of the regional mask mask of image area B. The regional weight of the pixel point P in image area A is close to 1, and the weight in image area B is close to 0. The final pixel value of the pixel point P is mainly determined by the features of image area A.

[0092] The image generation method provided in the above embodiment calculates the regional weight of the image area according to the regional attribute information to fuse multiple image areas that are independently controlled, thereby achieving a smooth transition and natural fusion effect between different areas in the generated target image, thereby improving the image generation effect.

[0093] In a possible implementation, the text description content may further include: style description information of multiple subjects. Before the above S104 generates a target image containing multiple subjects using an image generation model according to the type description information of the multiple subjects and the region control parameters, the method may further include:

[0094] According to the style description information of the multiple subjects, a style control model corresponding to the style description information is obtained, and a mapping relationship between the style control model and the image regions where the multiple subjects are located is determined. The style control model is used to control the style of the corresponding subject.

[0095] In this embodiment, the text description content includes style description information of the subject to indicate the display style of the subject. Since the image generation model is a basic model, it can only generate multi-subject same-frame images with unified subject styles, and cannot generate a separate style for each subject. In order to generate subjects with different styles in the same frame, it is necessary to introduce a style control model.

[0096] Specifically, according to the style description information of multiple subjects contained in the text description content, a style control model consistent with the style description information is loaded, and the style control model is associated with the region mask of the corresponding image area to determine the mapping relationship between the style control model and the corresponding image area. The style control model only acts on the corresponding image area.

[0097] In some embodiments, the style control module is a LoRA (Low-Rank Adaptation) model. If the style description information of the subject is a cartoon style, a cartoon-style LoRA model is loaded; if the style description information of the subject is an oil painting style, a oil painting-style LoRA model is loaded; if the style description information of multiple subjects is of different styles, LoRA models of different styles are loaded for multiple subjects; if the text description content only describes one style description information, a LoRA model of one style is loaded; if the style description information described in the text description content is for a single subject, the LoRA model of this style is associated with the area mask of the subject; if the style description information described in the text description content is for the entire image, the LoRAzepam model of this style is associated with the entire image.

[0098] For example, if the text description is "generate a Van Gogh-style cafe under the starry sky", the image generation model can be a basic model for generating landscape paintings to better draw the basic forms of the starry sky and the cafe, and then load a Van Gogh-style LoRA model to add Van Gogh's painting style to the image.

[0099] If the text description is "generate a picture containing an ink-style Monkey King and a cyberpunk-style Iron Man", then an ink-style LoRA model is loaded for Monkey King, and a cyberpunk-style LoRA model is loaded for Iron Man.

[0100] Furthermore, the style control model set for each image region can be used as a part of the region control parameters and registered together with other region control parameters as a region controller.

[0101] The process of generating a target image containing multiple subjects using an image generation model according to the type description information and the region control parameters of the multiple subjects in S104 may include:

[0102] According to the type description information and region control parameters of multiple subjects, an image generation model and a style control model are used to generate a target image.

[0103] In this embodiment, the style control model is a lightweight model that can generate images of a specific style without modifying the parameters of the image generation model.

[0104] Specifically, a micro-network is connected in parallel to the feature extraction layer of the image generation model. The micro-network consists of a dimensionality reduction layer (Down-projection Layer) and a dimensionality increase layer (Up-projection Layer). The image features of the corresponding image area are reduced in dimensionality through the dimensionality reduction layer to reduce the number of parameters. The dimensions of the reduced features are then restored to the original dimensions through the dimensionality increase layer. The original image features and the features output by the micro-network are weightedly summed through the LoRA model to obtain the final image features, wherein the weights of the weighted summation are the parameters of the LoRA model obtained through learning.

[0105] Among them, the image features of the image area are the features of the basic image generated by the image generation model according to the text description content.

[0106] In some embodiments, the style control model set for each image region may be used as a part of the region control parameters and registered together with other region control parameters as a region controller.

[0107] The image generation model generates a basic image based on the text description content, and then fine-tunes the display content of the corresponding image area in the basic image based on the display attribute type, attribute parameter value, conditional control information, and style control model of the image area to output a multi-subject same-frame image with multi-style fusion.

[0108] The image generation method provided in the above embodiment loads a style control model for the subject based on the style description information of the subject, so that when the target image is generated based on the image generation model, the style of the subject displayed in the corresponding image area can be adjusted based on the style control model, so that subjects of different styles can be framed in the same frame, providing users with flexible and diverse image generation methods and improving image generation effects.

[0109] In one possible implementation, Figure 4 Schematic diagram of the process of the image generation method provided in the embodiment of the present application Figure 4 ,like Figure 4 As shown, before the above S104 uses the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, the method may further include:

[0110] S301. Obtain model adjustment information of at least one adjustment type, where the model adjustment information of each adjustment type includes: type parameter, action target, parameter value, and adjustment parameter, where the adjustment parameter includes at least: parameter basic strength and associated area mask.

[0111] In this embodiment, in order to generate images more finely, the model parameters of the image generation model can be modified. Specifically, for the parameter type to be modified, the adjustment type and the model adjustment information corresponding to the adjustment type are determined. The adjustment type indicates the parameter type to be modified. The type parameter (Type) contained in the model adjustment information indicates the function of the model adjustment information, that is, the adjustment type, for example, it is used to adjust color, change texture, enhance details, etc. Model adjustment information of different adjustment types acts on different modules or levels within the image generation.

[0112] The target (Target) indicates that the model adjustment information of each adjustment type acts on the module or layer in the image generation model, and determines the specific impact range of the model adjustment information in the image generation process.

[0113] Parameters are the parameter types and values ​​adjusted by the model adjustment information. For example, the Strength parameter is used to control the effect strength of the model adjustment information. The model adjustment information for enhancing contrast will have a "Strength" parameter. The higher the value, the more obvious the contrast enhancement effect. The Range parameter is used to control the scope of the model adjustment information. The model adjustment information for adjusting color will have a "Hue Range" parameter, indicating that it is effective within a specific color range. The Factor parameter is used to control the value of the multiplication or scaling operation of the model adjustment information. The model adjustment information for adjusting brightness will have a "Brightness Factor" parameter. A value greater than 1 makes the image brighter, and a value less than 1 makes the image darker.

[0114] The parameter base strength (S_base) is the adjustment range of the parameter value. For example, if the parameter value for adjusting the saturation is 1.5 (indicating a 50% increase in saturation) and the parameter base strength is 0.8, then the final saturation adjustment range is 1.5*0.8=1.2, i.e., a 20% increase in saturation. The parameter base strength can be used to adjust the actual parameter value without modifying the parameter value.

[0115] The associated region mask is used to indicate the image region on which the model conditioning information is applied.

[0116] In some embodiments, a model parameter adjuster (Hook) is created based on the model adjustment information of each adjustment type, and the created Hook is registered and stored to record the information of all created Hooks. The Hooks are grouped and managed to facilitate search and calling. For example, all Hooks related to control color can be placed in the same group, and the Hooks are scheduled and managed to determine the execution order and timing of the Hooks to ensure that the Hooks can work according to the expected process. The dependencies between the Hooks are managed to ensure that each Hook can work normally when the effectiveness of each Hook depends on other Hooks, and conflict handling rules between Hooks are defined to enable adjustments when conflicts occur.

[0117] S302: Calculate the adjustment weight corresponding to each adjustment type according to the adjustment parameters of each adjustment type.

[0118] In this embodiment, weighting is performed according to the parameter base strength (S_base) and the mask value (A_mask) corresponding to the associated region mask to obtain the adjustment weight of each adjustment type in the corresponding image region.

[0119] In some embodiments, the adjustment parameters further include: a timing factor, and the timing factor is used to control the change of the adjustment weight over time.

[0120] In this embodiment, during the image generation process, the image generation model will gradually iterate to generate images, and the timing factor (T_factor) will dynamically change according to the current generation time step (timestep) to control the strength of the model adjustment information at different generation stages. For example, in order to show that the clarity of the image is gradually improved during the generation process, the timing factor can be used to adjust the model adjustment information of the corresponding clarity to gradually increase.

[0121] For example, the formula for calculating the adjustment weight can be expressed as: W_hook = ɑ*S_base+β*T_factor+γ*A_mask, where ɑ, β and γ are adjustment coefficients respectively.

[0122] S303: Adjust the model parameters of the image generation model corresponding to the target according to the model adjustment information and adjustment weight corresponding to each adjustment type.

[0123] In this embodiment, the target in the image generation model is determined according to the target corresponding to the model adjustment information of each adjustment type, and the model parameters of the corresponding type of the target are adjusted according to the adjustment weight and the adjustment parameter type. Based on the target and parameter value after the adjustment of the model parameters, the parameter value of the corresponding image area is adjusted.

[0124] The image generation method provided in the above embodiment calculates the adjustment weight based on the model adjustment information to modify the model parameters of the target in the image generation model, thereby achieving refined image generation and improving the image generation effect.

[0125] It should be noted that the image parameters targeted by the region control information and the model adjustment information may be the same image parameters, but the image parameters set by the region control information are input into the image generation model as the features of the image generation model, while the model adjustment information directly modifies the features of the specified modules or levels of the model. The two have different dimensions of action.

[0126] In a possible implementation, after acquiring the model adjustment information of at least one adjustment type in S301, the method may further include:

[0127] A conflict handling rule for model adjustment information of at least one adjustment type is determined, where the conflict handling rule is used to indicate a parameter adjustment method when a conflict occurs between different model adjustment information.

[0128] In this embodiment, after a Hook is generated according to the model adjustment information of each adjustment type, in order to avoid conflicts when multiple Hooks act on the image generation model, it is necessary to set conflict handling rules for the Hooks in advance.

[0129] The conflict handling rules include one or more of the following: priority rules, weighted average rules, conditional judgment rules, sequential execution rules, and mutual exclusion rules.

[0130] Specifically, the conflict types are divided into direct conflict, indirect conflict and resource conflict. Direct conflict means that two Hooks act on the same image area and are of the same adjustment type. For example, one Hook wants to increase the red value of a pixel, and the other Hook wants to decrease the red value of the same pixel.

[0131] Indirect conflict means that two hooks are of different adjustment types, but their effects will affect each other, resulting in the final result not meeting expectations. For example, if one hook enhances the contrast of an image and another hook enhances the saturation of the image, the superposition of the two may cause the image to be too glaring.

[0132] Resource conflict occurs when two Hooks attempt to use the same resource (such as the same LoRA model), but their usage methods are incompatible.

[0133] Conflicts between Hooks can be resolved by using conflict handling rules.

[0134] Specifically, the priority rule indicates setting priorities for different Hooks, and when a conflict occurs, the effect of the Hook with a higher priority takes precedence. For example, the Hook manually set by the user has a higher priority, overwriting the effect of the Hook automatically generated by the system.

[0135] The weighted average rule is that if two Hooks try to modify the same property, they can be weighted averaged according to their respective weights. For example, if the weight of one Hook is 0.7 and the weight of the other Hook is 0.3, the final result will be more biased towards the effect of the first Hook.

[0136] The conditional judgment rule indicates the target Hook to be used based on the preset conditions. For example, if the brightness of the current pixel is lower than a certain threshold, a Hook that increases the brightness is used, otherwise another Hook that keeps the brightness unchanged is used.

[0137] The order of execution rule indicates that for indirect conflicts, the order of execution of the Hooks is adjusted to resolve them. For example, executing a Hook that enhances contrast first and then executing a Hook that enhances saturation may have a better effect than executing them in reverse.

[0138] Mutually exclusive rules indicate that you can set Hooks to be mutually exclusive, that is, they cannot be effective at the same time. For example, you can set a "black and white filter" Hook and all color adjustment Hooks to be mutually exclusive, ensuring that no color adjustments will be made after applying the black and white filter.

[0139] Furthermore, conflict detection and prompts can be set up: if the system cannot automatically resolve the conflict, a warning can be issued to the user, prompting the user to manually adjust the parameters or priority of the Hook.

[0140] In a possible implementation, after acquiring the model adjustment information of at least one adjustment type in S301, the method may further include:

[0141] The legitimacy of the model adjustment information of at least one adjustment type is verified, and the verification content includes at least one of the following: parameter verification, type verification, target verification, dependency verification, conflict verification, and security verification.

[0142] In this embodiment, parameter verification is to check whether the parameters of the Hook are within the allowed range and whether the type is correct. For example, if the hue parameter of a color control Hook requires an integer between 0 and 360, then if a negative number or string is passed in, it will be judged as illegal.

[0143] Type checking is to confirm whether the type of the Hook is a known type in the system to prevent the creation of undefined or incorrect Hook types.

[0144] Target verification is to check whether the target specified by the Hook exists and is valid. For example, if a Hook is to act on a specific layer of the model, it is necessary to confirm that the layer exists in the current model.

[0145] Dependency checking is to check whether other Hooks or resources that a Hook depends on are available. For example, if a Hook requires another Hook to execute first, it is necessary to confirm that the dependent Hook has been registered and is in an activated state.

[0146] Common types of dependencies include:

[0147] Execution order dependency: Hook A must be executed before Hook B to work properly. For example, a Hook responsible for adjusting brightness may need to be executed after a Hook responsible for color correction to get the correct input.

[0148] Resource dependency: The operation of Hook A depends on a specific resource, such as a specific model or dataset.

[0149] State dependency: The operation of Hook A depends on the state of another Hook B, for example, Hook B must be in the activated state.

[0150] The management system needs to track and manage these dependencies to ensure that all dependency conditions are met when executing the Hook. If a Hook's dependency condition is not met, the management system may delay the execution of the Hook or issue a warning message. Reasonable management of dependencies can avoid Hook execution errors and ensure the stability and correctness of the image generation process.

[0151] Conflict checking is to check whether the newly registered Hook will conflict with the existing Hook. For example, if two Hooks try to modify the same model parameter, conflict detection and processing are required.

[0152] Security verification is required for some Hooks that involve code execution or external resource access to prevent malicious code or illegal access.

[0153] In a possible implementation, before the above S104 uses the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, the method may further include:

[0154] Image features of multiple style types are obtained from multiple style transfer images; and fusion features are calculated according to the image features of multiple style types and type weights of multiple style types.

[0155] The process of generating a target image containing multiple subjects using an image generation model according to the type description information and the region control parameters of the multiple subjects in S104 may include:

[0156] According to the type description information, regional control parameters and fusion features of multiple subjects, an image generation model is used to generate a target image.

[0157] In some embodiments, in order to transfer image features in other images to the image to be generated, image features may be obtained from the style transfer image, and fusion features may be calculated based on the type weight of the feature type.

[0158] In other embodiments, if it is necessary to migrate image features of different style types in multiple style transfer images to the target to be generated, the corresponding types of image features can be extracted from the multiple style transfer images respectively, and the fusion features can be calculated according to the type weights of the respective feature types.

[0159] In the image generation process, the fused features are applied to the target image so that the effect of the corresponding image area in the target image on the feature type is consistent with the style transfer image.

[0160] Among them, the type weight can be calculated according to the adjustment parameters in the corresponding model adjustment information, that is, the adjustment weight.

[0161] In the process of feature fusion, feature preprocessing is required. The purpose of feature preprocessing is to enable features from different sources to be better fused. Features extracted from different images or models may have different scales, ranges or representations. The preprocessing step will perform some standardization or conversion operations, such as:

[0162] Normalization: Scale the feature values ​​to a uniform range (for example, between 0 and 1) to eliminate the effects of different scales.

[0163] Standardization: Transform the feature values ​​into a distribution with mean 0 and standard deviation 1.

[0164] Dimensionality reduction: For high-dimensional features, dimensionality reduction techniques (such as PCA) can be used to reduce the dimension of the features and improve fusion efficiency.

[0165] Alignment: For features from different models, it may be necessary to align them to ensure that they correspond semantically.

[0166] The specific method of preprocessing depends on the type of features and the requirements for fusion, and this embodiment does not limit this.

[0167] Calculate fusion weights (calculate_weights) to control the contribution of different features to the final fusion result. The higher the weight, the greater the impact of the feature on the final result. The calculation of fusion weights can be based on a variety of parameters:

[0168] User-specified weights: Users can directly set the weight of each feature, for example, to make a feature of a certain model contribute more.

[0169] Attention mechanism: The system can learn the importance of different features and automatically assign weights.

[0170] Mask: Adjust the feature weights according to the Mask value to use different fusion strategies in different regions. For example, in one region, the feature weight of model A is higher, while in another region, the feature weight of model B is higher.

[0171] Timing information: Dynamically adjust the weight of features according to the generated time steps to achieve a fusion effect that changes over time.

[0172] If the adjustment weight is calculated based on the adjustment parameters in the model adjustment information, the adjustment weight of the feature acting on each region can be adjusted by adjusting the adjustment coefficient of the region mask and the adjustment coefficient of the timing factor.

[0173] If multiple Hooks are set for an image area, it is necessary to perform Hook combination (apply_fusion) on multiple Hooks. The execution of Hook combination is to fuse the preprocessed features according to the calculated weights, and multiply different features by their corresponding weights through weighted averaging and then sum them up.

[0174] For example, when the Hook combination adopts the weighted average mechanism, the image features can be understood as the numerical representation of the image display attributes, and different features represent different display attributes of the image. For example, the color feature can use a set of numerical values ​​to represent the color distribution of the image, such as the average value or histogram of the red, green, and blue channels; the texture feature can use a set of numerical values ​​to represent the image's roughness, directionality and other texture information; the edge feature can use a set of numerical values ​​to represent the strength and direction of the edge in the image; the style feature can use a set of numerical values ​​to represent the image's artistic style, such as brushstrokes, color matching, etc.

[0175] When performing weighted averaging, each image feature is multiplied by a weight. The size of the weight determines the importance of the image feature in the final fusion result. The larger the weight, the greater the impact of the image feature on the final fused image.

[0176] For example, suppose we have two Hooks, one is responsible for extracting the color features of the style transfer image (F1), and the other is responsible for extracting the texture features of the style transfer image (F2). If we want the final generated image to be more biased towards the color of the first image, then when performing weighted averaging, we can set a higher weight (W1) for F1 and a lower weight (W2) for F2. The final fused feature F\final = W1*F1+W2*F2, this fused feature will retain more color information of the first image, and the image generation model will generate the final image based on F\final, making it closer to the first image in color.

[0177] You can also use methods such as splicing different features together in a certain dimension to generate fused features, using an attention mechanism to dynamically select and combine different features, or a fusion method based on a neural network, which is not limited in this embodiment.

[0178] After feature fusion, post-processing can be used to further optimize the fused features to improve the quality and visual effect of the final generated image. Post-processing operations can include: denoising, sharpening, color correction, contrast adjustment, etc.

[0179] The image generation method provided in the above embodiment calculates fusion features based on different types of image features and type weights in the style transfer image to transfer the image features in the style transfer image to the generated target image, thereby realizing image style transfer. If multiple style transfer images are selected, multi-style fusion effects can also be guaranteed.

[0180] For example, Figure 5 A schematic diagram of multiple agents in the same frame provided in the embodiment of the present application, such as Figure 5 As shown, by using the image generation method provided in the embodiment of the present application, a target image with multiple subjects in the same frame can be generated, and the positions of the multiple subjects in the image can be flexibly set, and each subject also has a refined display effect.

[0181] Figure 6 A schematic diagram of multiple styles and multiple subjects in the same frame provided by the embodiment of the present application, such as Figure 6 As shown, by using the image generation method provided in the embodiment of the present application, target images of multiple subjects in the same frame in various styles can be generated, and each subject also has a refined display effect.

[0182] Based on the above method embodiment, the embodiment of the present application also provides an image generating device. Figure 7A schematic diagram of the structure of an image generating device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the device may include:

[0183] The model and text acquisition module 501 is used to acquire input text description content and a pre-trained image generation model, wherein the text description content includes: type description information of multiple subjects;

[0184] A mask acquisition module 502 is used to acquire region masks of multiple subjects, where the region masks are used to indicate multiple image regions where the multiple subjects are located in the image to be generated;

[0185] A parameter acquisition module 503, configured to acquire a region control parameter of a target image region in the plurality of image regions according to the region masks of the plurality of subjects, wherein the region control parameter is used to control global display information of the corresponding subject in the target image region, and the global display information is used to indicate an overall display parameter of the corresponding subject in the target image region;

[0186] The image generation module 504 is used to generate a target image containing multiple subjects by using an image generation model according to the type description information and the region control parameters of the multiple subjects.

[0187] Optionally, the parameter acquisition module 503 is specifically used to acquire at least one attribute parameter value corresponding to a preset display attribute type of the target image area;

[0188] A region control parameter of the target image region is generated according to the region mask, the display attribute type, and at least one attribute parameter value of the target image region.

[0189] Optionally, the parameter acquisition module 503 is also used to acquire at least one conditional control information of the target image area, and the conditional control information is used to control the detail display information of the subject in the target image area; and generate the area control parameters of the target image area according to the area mask, display attribute type, at least one attribute parameter value and the conditional control information of the target image area.

[0190] Optionally, the parameter acquisition module 503 is further used to determine the conditional control information of the target image area according to the input conditional text information; and / or determine the conditional control information of the target image area by extracting the target features of the reference image.

[0191] Optionally, the device may further include:

[0192] A region weight determination module is used to determine the region weights of the plurality of image regions according to the region attribute information of the plurality of image regions, wherein the region attribute information includes at least one of the following information: region shape, region size, region edge smoothness, region overlap, and region content;

[0193] The image generation module 504 is specifically configured to generate a target image using an image generation model according to type description information of multiple subjects, region control parameters and region weights of multiple image regions.

[0194] Optionally, the text description content also includes: style description information of multiple subjects; the model and text acquisition module 501 is further used to acquire a style control model corresponding to the style description information according to the style description information of the multiple subjects, and determine a mapping relationship between the style control model and the image area where the multiple subjects are located, and the style control model is used to control the style of the corresponding subject;

[0195] The image generation module 504 is further used to generate a target image according to the type description information and the region control parameters of the multiple subjects by using the image generation model and the style control model.

[0196] Optionally, the device may further include:

[0197] An adjustment information acquisition module, used to acquire model adjustment information of at least one adjustment type, wherein the model adjustment information of each adjustment type includes: a type parameter, an action target, a parameter value, and an adjustment parameter, wherein the adjustment parameter includes at least: a parameter base strength and an associated region mask;

[0198] An adjustment weight calculation module is used to calculate the adjustment weight corresponding to each adjustment type according to the adjustment parameters of each adjustment type;

[0199] The model parameter adjustment module is used to adjust the model parameters of the corresponding action targets in the image generation model according to the model adjustment information and adjustment weights corresponding to each adjustment type.

[0200] Optionally, the adjustment parameters further include: a timing factor, and the timing factor is used to control the change of the adjustment weight over time.

[0201] Optionally, the adjustment information acquisition module is further used to determine a conflict handling rule for model adjustment information of at least one adjustment type, where the conflict handling rule is used to indicate a parameter adjustment method when a conflict occurs between different model adjustment information.

[0202] Optionally, the conflict handling rules include one or more of the following: a priority rule, a weighted average rule, a conditional judgment rule, a sequential execution rule, and a mutual exclusion rule.

[0203] Optionally, the adjustment information acquisition module is also used to verify the legitimacy of model adjustment information of at least one adjustment type, and the verification content includes at least one of the following: parameter verification, type verification, target verification, dependency verification, conflict verification, and security verification.

[0204] Optionally, the device may further include:

[0205] A feature extraction module is used to obtain image features of multiple style types from multiple style migration images; and calculate fusion features according to the image features of multiple style types and type weights of multiple style types;

[0206] The image generation module 504 is further used to generate a target image using an image generation model according to the type description information, region control parameters and fusion features of multiple subjects.

[0207] The image generating device provided by the above-mentioned embodiment determines the positions of multiple subjects in the image based on the region mask, and controls the global display information of the subjects in the target image region based on the region control parameters, so that in the target image generated by the image generation model according to the type description information of the multiple subjects and the region control parameters, the positional relationship of the multiple subjects is consistent with the positional relationship defined by the region mask, and the effect of the subjects in the target image region is finely controlled, which meets the user needs and improves the same-frame effect of the multi-subject same-frame image.

[0208] The above-mentioned device is used to execute the method provided by the aforementioned embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.

[0209] The above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASICs), or one or more microprocessors, or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).

[0210] Figure 8 A schematic diagram of an electronic device provided in an embodiment of the present application, such as Figure 8 As shown, the electronic device 600 may include: a processor 601, a storage medium 602 and a bus, wherein the storage medium 602 stores program instructions executable by the processor 601. When the electronic device 600 is running, the processor 601 communicates with the storage medium 602 via the bus, and the processor 601 executes the program instructions to execute the above method embodiment.

[0211] Specifically, the processor executes the steps of the above-mentioned image generation method, which may include:

[0212] The input text description content and the pre-trained image generation model are obtained, wherein the text description content includes: type description information of multiple subjects; region masks of the multiple subjects are obtained, the region masks are used to indicate multiple image regions where the multiple subjects are located in the image to be generated; region control parameters of a target image region in the multiple image regions are obtained according to the region masks of the multiple subjects, the region control parameters are used to control global display information of the corresponding subject in the target image region, and the global display information is used to indicate the overall display parameters of the corresponding subject in the target image region; according to the type description information and the region control parameters of the multiple subjects, an image generation model is used to generate a target image containing multiple subjects.

[0213] Optionally, the processor executes the step of acquiring the region control parameter of the target image region among the multiple image regions according to the region masks of the multiple subjects, which may include:

[0214] At least one attribute parameter value corresponding to a preset display attribute type of the target image area is obtained; and a region control parameter of the target image area is generated according to a region mask, a display attribute type, and at least one attribute parameter value of the target image area.

[0215] Optionally, the processor executes the steps of the above-mentioned image generation method, and may further include: acquiring at least one condition control information of the target image area, the condition control information being used to control the detailed display information of the subject in the target image area;

[0216] The processor executes the above step of acquiring the region control parameter of the target image region in the multiple image regions according to the region masks of the multiple subjects, which may include:

[0217] The region control parameter of the target image region is generated according to the region mask of the target image region, the display attribute type, at least one attribute parameter value and the conditional control information of the target image region.

[0218] Optionally, the processor executes the above step of acquiring the condition control information of the target image area, which may include:

[0219] Determine the conditional control information of the target image area according to the input conditional text information; and / or determine the conditional control information of the target image area by extracting the target features of the reference image.

[0220] Optionally, before the processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, the processor executes the steps of the image generation method, which may also include:

[0221] Determining region weights of the plurality of image regions according to region attribute information of the plurality of image regions, wherein the region attribute information includes at least one of the following information: region shape, region size, region edge smoothness, region overlap, and region content;

[0222] The processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, which may include:

[0223] According to the type description information of multiple subjects, the region control parameters and the region weights of multiple image regions, an image generation model is used to generate a target image.

[0224] Optionally, the text description content further includes: style description information of the multiple subjects; before the processor executes the above-mentioned steps of using the image generation model to generate a target image containing multiple subjects according to the type description information of the multiple subjects and the area control parameters, the processor executes the steps of the image generation method, which may also include:

[0225] According to the style description information of the multiple subjects, a style control model corresponding to the style description information is obtained, and a mapping relationship between the style control model and the image regions where the multiple subjects are located is determined, wherein the style control model is used to control the style of the corresponding subject;

[0226] The processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, which may include:

[0227] According to the type description information and region control parameters of multiple subjects, an image generation model and a style control model are used to generate a target image.

[0228] Optionally, before the processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, the processor executes the steps of the image generation method, which may also include:

[0229] Obtain model adjustment information of at least one adjustment type, wherein the model adjustment information of each adjustment type includes: type parameter, target, parameter value, and adjustment parameter, wherein the adjustment parameter includes at least: parameter base strength and associated area mask; calculate the adjustment weight corresponding to each adjustment type according to the adjustment parameter of each adjustment type; adjust the model parameter of the corresponding target in the image generation model according to the model adjustment information and adjustment weight corresponding to each adjustment type.

[0230] Optionally, the adjustment parameters further include: a timing factor, and the timing factor is used to control the change of the adjustment weight over time.

[0231] Optionally, after the processor executes the above-mentioned acquisition of model adjustment information of at least one adjustment type, the processor executes the steps of the image generation method, which may also include:

[0232] A conflict handling rule for model adjustment information of at least one adjustment type is determined, where the conflict handling rule is used to indicate a parameter adjustment method when a conflict occurs between different model adjustment information.

[0233] Optionally, the conflict handling rules include one or more of the following: a priority rule, a weighted average rule, a conditional judgment rule, a sequential execution rule, and a mutual exclusion rule.

[0234] Optionally, after the processor executes the above-mentioned acquisition of model adjustment information of at least one adjustment type, the processor executes the steps of the image generation method, which may also include:

[0235] The legitimacy of the model adjustment information of at least one adjustment type is verified, and the verification content includes at least one of the following: parameter verification, type verification, target verification, dependency verification, conflict verification, and security verification.

[0236] Optionally, before the processor executes the above-mentioned steps of generating an image containing multiple subjects using an image generation model according to the text description content and the area control parameters, the processor executes the steps of the image generation method, which may also include:

[0237] Obtain image features of multiple style types from multiple style transfer images; calculate fusion features based on the image features of multiple style types and type weights of multiple style types;

[0238] The processor executes the above-mentioned step of using the image generation model to generate an image containing multiple subjects according to the text description content and the area control parameters, which may include:

[0239] According to the type description information, regional control parameters and fusion features of multiple subjects, an image generation model is used to generate a target image.

[0240] The image generation method executed by the processor of the above-mentioned embodiment determines the positions of multiple subjects in the image based on the region mask, and controls the global display information of the subjects in the target image region based on the region control parameters, so that in the target image generated by the image generation model according to the type description information of the multiple subjects and the region control parameters, the positional relationship of the multiple subjects is consistent with the positional relationship defined by the region mask, and the effect of the subjects in the target image region is finely controlled, which meets user needs and improves the same-frame effect of the multi-subject same-frame image.

[0241] Optionally, the present application further provides a computer-readable storage medium having a computer program stored thereon, and the computer program executes the above method embodiment when executed by a processor.

[0242] Specifically, the processor executes the steps of the above-mentioned image generation method, which may include:

[0243] The input text description content and the pre-trained image generation model are obtained, wherein the text description content includes: type description information of multiple subjects; region masks of the multiple subjects are obtained, the region masks are used to indicate multiple image regions where the multiple subjects are located in the image to be generated; according to the region masks of the multiple subjects, region control parameters of the target image region in the multiple image regions are obtained, the region control parameters are used to control the global display information of the corresponding subject in the target image region, and the global display information is used to indicate the overall display parameters of the corresponding subject in the target image region; according to the type description information and the region control parameters of the multiple subjects, the image generation model is used to generate a target image containing multiple subjects.

[0244] Optionally, the processor executes the step of acquiring the region control parameter of the target image region among the multiple image regions according to the region masks of the multiple subjects, which may include:

[0245] At least one attribute parameter value corresponding to a preset display attribute type of the target image area is obtained; and a region control parameter of the target image area is generated according to a region mask, a display attribute type, and at least one attribute parameter value of the target image area.

[0246] Optionally, the processor executes the steps of the above-mentioned image generation method, and may further include: acquiring at least one condition control information of the target image area, the condition control information being used to control the detailed display information of the subject in the target image area;

[0247] The processor executes the above step of acquiring the region control parameter of the target image region in the multiple image regions according to the region masks of the multiple subjects, which may include:

[0248] The region control parameters of the target image region are generated according to the region mask of the target image region, the display attribute type, at least one attribute parameter value and the conditional control information of the target image region.

[0249] Optionally, the processor executes the above step of acquiring the condition control information of the target image area, which may include:

[0250] Determine the conditional control information of the target image area according to the input conditional text information; and / or determine the conditional control information of the target image area by extracting the target features of the reference image.

[0251] Optionally, before the processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, the processor executes the steps of the image generation method, which may also include:

[0252] Determining region weights of the plurality of image regions according to region attribute information of the plurality of image regions, wherein the region attribute information includes at least one of the following information: region shape, region size, region edge smoothness, region overlap, and region content;

[0253] The processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, which may include:

[0254] According to the type description information of multiple subjects, the region control parameters and the region weights of multiple image regions, an image generation model is used to generate a target image.

[0255] Optionally, the text description content further includes: style description information of the multiple subjects; before the processor executes the above-mentioned steps of using the image generation model to generate a target image containing multiple subjects according to the type description information of the multiple subjects and the area control parameters, the processor executes the steps of the image generation method, which may also include:

[0256] According to the style description information of the multiple subjects, a style control model corresponding to the style description information is obtained, and a mapping relationship between the style control model and the image regions where the multiple subjects are located is determined, wherein the style control model is used to control the style of the corresponding subject;

[0257] The processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, which may include:

[0258] According to the type description information and region control parameters of multiple subjects, an image generation model and a style control model are used to generate a target image.

[0259] Optionally, before the processor executes the above-mentioned step of using the image generation model to generate a target image containing multiple subjects according to the type description information and the region control parameters of the multiple subjects, the processor executes the steps of the image generation method, which may also include:

[0260] Obtain model adjustment information of at least one adjustment type, wherein the model adjustment information of each adjustment type includes: type parameter, target, parameter value, and adjustment parameter, wherein the adjustment parameter includes at least: parameter base strength and associated area mask; calculate the adjustment weight corresponding to each adjustment type according to the adjustment parameter of each adjustment type; adjust the model parameter of the corresponding target in the image generation model according to the model adjustment information and adjustment weight corresponding to each adjustment type.

[0261] Optionally, the adjustment parameters further include: a timing factor, and the timing factor is used to control the change of the adjustment weight over time.

[0262] Optionally, after the processor executes the above-mentioned acquisition of model adjustment information of at least one adjustment type, the processor executes the steps of the image generation method, which may also include:

[0263] A conflict handling rule for model adjustment information of at least one adjustment type is determined, where the conflict handling rule is used to indicate a parameter adjustment method when a conflict occurs between different model adjustment information.

[0264] Optionally, the conflict handling rules include one or more of the following: a priority rule, a weighted average rule, a conditional judgment rule, a sequential execution rule, and a mutual exclusion rule.

[0265] Optionally, after the processor executes the above-mentioned acquisition of model adjustment information of at least one adjustment type, the processor executes the steps of the image generation method, which may also include:

[0266] The legitimacy of the model adjustment information of at least one adjustment type is verified, and the verification content includes at least one of the following: parameter verification, type verification, target verification, dependency verification, conflict verification, and security verification.

[0267] Optionally, before the processor executes the above-mentioned steps of generating an image containing multiple subjects using an image generation model according to the text description content and the area control parameters, the processor executes the steps of the image generation method, which may also include:

[0268] Obtain image features of multiple style types from multiple style transfer images; calculate fusion features based on the image features of multiple style types and type weights of multiple style types;

[0269] The processor executes the above-mentioned step of using the image generation model to generate an image containing multiple subjects according to the text description content and the area control parameters, which may include:

[0270] According to the type description information, regional control parameters and fusion features of multiple subjects, an image generation model is used to generate a target image.

[0271] The image generation method executed by the processor of the above-mentioned embodiment determines the positions of multiple subjects in the image based on the region mask, and controls the global display information of the subjects in the target image region based on the region control parameters, so that in the target image generated by the image generation model according to the type description information of the multiple subjects and the region control parameters, the positional relationship of the multiple subjects is consistent with the positional relationship defined by the region mask, and the effect of the subjects in the target image region is finely controlled, which meets user needs and improves the same-frame effect of the multi-subject same-frame image.

[0272] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0273] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0274] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0275] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to perform some steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (English: Read-Only Memory, abbreviated: ROM), random access memory (English: Random Access Memory, abbreviated: RAM), disk or optical disk and other media that can store program codes.

[0276] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. An image generation method, characterized in that: The method comprises: Acquire input text description content and a pre-trained image generation model, wherein the text description content includes: type description information of multiple subjects; Acquire region masks of the multiple subjects, where the region masks are used to indicate multiple image regions where the multiple subjects are located in the image to be generated; Acquire, according to the region masks of the multiple subjects, a region control parameter of a target image region in the multiple image regions, wherein the region control parameter is used to control global display information of a corresponding subject in the target image region, and the global display information is used to indicate an overall display parameter of the corresponding subject in the target image region; According to the type description information of the multiple subjects and the region control parameters, the image generation model is used to generate a target image containing the multiple subjects.

2. The method according to claim 1, characterized in that The acquiring, according to the region masks of the multiple subjects, a region control parameter of a target image region among the multiple image regions comprises: Acquire at least one attribute parameter value corresponding to a preset display attribute type of the target image area; The region control parameter of the target image region is generated according to the region mask of the target image region, the display attribute type, and the at least one attribute parameter value.

3. The method according to claim 2, characterized in that The method further comprises: Acquire at least one condition control information of the target image area, where the condition control information is used to control the detail display information of the subject in the target image area; The acquiring, according to the region masks of the multiple subjects, a region control parameter of a target image region among the multiple image regions comprises: The region control parameter of the target image region is generated according to the region mask of the target image region, the display attribute type, the at least one attribute parameter value and the conditional control information of the target image region.

4. The method according to claim 3, characterized in that The acquiring of the condition control information of the target image area includes: Determining conditional control information of the target image area according to the input conditional text information; and / or, By extracting the target features of the reference image, conditional control information of the target image area is determined.

5. The method according to claim 1, characterized in that Before generating a target image containing the multiple subjects using the image generation model according to the type description information of the multiple subjects and the region control parameters, the method further includes: Determining the region weights of the plurality of image regions according to region attribute information of the plurality of image regions, the region attribute information including at least one of the following information: region shape, region size, region edge smoothness, region overlap, and region content; The step of generating a target image containing the multiple subjects by using the image generation model according to the type description information of the multiple subjects and the region control parameters includes: The target image is generated using the image generation model according to the type description information of the multiple subjects, the region control parameters of the multiple image regions and the region weights.

6. The method according to claim 1, characterized in that The text description content also includes: style description information of the multiple subjects; before generating a target image containing the multiple subjects using the image generation model according to the type description information of the multiple subjects and the region control parameters, the method also includes: According to the style description information of the multiple subjects, a style control model corresponding to the style description information is acquired, and a mapping relationship between the style control model and the image area where the multiple subjects are located is determined, wherein the style control model is used to control the style of the corresponding subject; The step of generating a target image containing the multiple subjects by using the image generation model according to the type description information of the multiple subjects and the region control parameters includes: The target image is generated according to the type description information of the multiple subjects and the region control parameters by using the image generation model and the style control model.

7. The method according to claim 1, characterized in that Before generating a target image containing the multiple subjects using the image generation model according to the type description information of the multiple subjects and the region control parameters, the method further includes: Acquire model adjustment information of at least one adjustment type, where the model adjustment information of each adjustment type includes: a type parameter, an action target, a parameter value, and an adjustment parameter, where the adjustment parameter includes at least: a parameter base strength and an associated region mask; Calculating the adjustment weights corresponding to the various adjustment types according to the adjustment parameters of the various adjustment types; According to the model adjustment information and adjustment weights corresponding to the various adjustment types, the model parameters of the corresponding action targets in the image generation model are adjusted.

8. The method according to claim 7, characterized in that The adjustment parameters also include: a timing factor, and the timing factor is used to control the change of the adjustment weight over time.

9. The method according to claim 7, characterized in that After obtaining the model adjustment information of at least one adjustment type, the method further includes: A conflict handling rule for the model adjustment information of the at least one adjustment type is determined, where the conflict handling rule is used to indicate a parameter adjustment method when a conflict occurs between different model adjustment information.

10. The method according to claim 9, characterized in that The conflict handling rules include one or more of the following: priority rules, weighted average rules, conditional judgment rules, sequential execution rules, and mutual exclusion rules.

11. The method according to claim 7, characterized in that After obtaining the model adjustment information of at least one adjustment type, the method further includes: The legitimacy of the model adjustment information of the at least one adjustment type is verified, and the verification content includes at least one of the following: parameter verification, type verification, target verification, dependency verification, conflict verification, and security verification.

12. The method according to claim 1, characterized in that Before generating the image containing the multiple subjects by using the image generation model according to the text description content and the region control parameter, the method further includes: Obtain image features of multiple style types from multiple style transfer images; Calculating a fusion feature according to the image features of the multiple style types and the type weights of the multiple style types; The step of generating an image including the multiple subjects by using the image generation model according to the text description content and the region control parameter includes: The target image is generated using the image generation model according to the type description information of the multiple subjects, the region control parameters and the fusion features.

13. An image generating device, characterized in that: The device comprises: A model and text acquisition module, used to acquire input text description content and a pre-trained image generation model, wherein the text description content includes: type description information of multiple subjects; A mask acquisition module, used for acquiring region masks of the plurality of subjects, wherein the region masks are used for indicating a plurality of image regions where the plurality of subjects are located in the image to be generated; A parameter acquisition module, used for acquiring, according to the area masks of the multiple subjects, area control parameters of a target image area in the multiple image areas, the area control parameters being used for controlling global display information of the corresponding subject in the target image area, the global display information being used for indicating the overall display parameters of the corresponding subject in the target image area; An image generation module is used to generate a target image containing the multiple subjects using the image generation model according to the type description information of the multiple subjects and the area control parameters.

14. An electronic device, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores program instructions executable by the processor, and when the electronic device is running, the processor and the storage medium communicate through the bus, and the processor executes the program instructions to perform the steps of the image generation method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the image generation method according to any one of claims 1 to 12 are executed.

Citation Information

Cited By

  • Display method, intelligent terminal and storage medium

    CN120832113A