Image generation model training method and image generation method
By building a training set that includes the corresponding relationship between images and structural features, and training an image generation model, the problem that image generation models in the prior art is difficult to generate accurate images, achieving higher quality image generation.
Patent Information
- Application Number
- CN202510151237.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The prior art is difficult to generate more accurate images through other types of content descriptions, limiting the scope of application of image generation models.
A method of training image generation model is proposed. By constructing an image training set, including an image training subset and a structural feature training subset, the two models to be trained to obtain the target image generation model based on the one-to-one correspondence relationship of the structural feature training subset.
It realizes the integration of structural features with image features during image generation, which improves the quality and accuracy of image generation.
Smart Images

Figure CN120147446A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and particularly to an image generation model training and an image generation method. Background Art
[0002] Image generation describes the information of the image to be generated through text content, and uses an image generation model to generate the image corresponding to the text content. However, in the process of generating an image using text content, other types of content cannot be used to describe the information of the image to be generated, and thus a more accurate image cannot be generated. Summary of the Invention
[0003] The present disclosure provides an image generation model training and an image generation method.
[0004] The first aspect embodiment of the present disclosure provides an image generation model training method, including:
[0005] Construct an image training set, where the image training set includes an image training subset and a structural feature training subset, and the image training subset and the structural feature training subset correspond one by one;
[0006] Based on the first feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, train a first model to be trained to obtain a first target image generation model;
[0007] Based on the second feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, train a second model to be trained to obtain a second target image generation model, where the second feature points are the target feature points in the first feature points.
[0008] In the embodiments of the present disclosure, the constructing the image training set includes:
[0009] Obtain an initial image set;
[0010] For any initial image in the initial image set, extract the segmentation mask of the initial image;
[0011] Based on the initial image and the segmentation mask, extract the image feature map of the initial image;
[0012] Based on the relationship between the segmentation mask and the image feature map, determine whether the initial image meets a preset standard;
[0013] Screen the grayscale images corresponding to the initial images that meet the preset criteria in the initial image set to construct the image training subset; and use the image feature maps corresponding to the initial images in the image training subset to construct the structural feature training subset, where the grayscale image corresponding to the initial image is obtained by setting the pixel values of the non-image features in the initial image to grayscale values, and the non-image features are obtained by subtracting the image feature map from the initial image.
[0014] In an embodiment of the present disclosure, the extracting the image feature map of the initial image based on the initial image and the segmentation mask includes:
[0015] Extract the initial structural features of the initial image using the segmentation mask;
[0016] Map the pixel values of the initial structural features to the pixel points at the corresponding positions of the preset image to obtain the image feature map.
[0017] In an embodiment of the present disclosure, the determining whether the initial image meets the preset criteria based on the relationship between the segmentation mask and the image feature map includes:
[0018] Calculate the intersection over union metric based on the segmentation mask and the image feature map;
[0019] If the intersection over union metric is greater than the preset threshold, the initial image meets the preset criteria;
[0020] If the intersection over union metric is less than or equal to the preset threshold, the initial image does not meet the preset criteria.
[0021] In an embodiment of the present disclosure, the training the first model to be trained based on any one of the structural features in the structural feature training subset and the image features of the image corresponding to the structural feature to obtain the first target image generation model includes:
[0022] For any one of the structural features in the structural feature training subset, use the first model to be trained to extract the first structural point features of the first feature points of the structural feature;
[0023] Calculate the first loss value of the first model to be trained based on the first structural point features and the image features of the corresponding image in the image training subset;
[0024] Optimize the first model to be trained based on the relationship between the first loss value and the first preset loss criterion, and output the first target image generation model.
[0025] In an embodiment of the present disclosure, training a second model to be trained based on the second feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature to obtain a second target image generation model includes:
[0026] For any structural feature in the structural feature training subset, use the second model to be trained to extract the second structural point features of the second feature points of the structural feature;
[0027] Fuse the second structural point features and the image features of the corresponding images in the image training subset to obtain the semantic features of the second structural point features;
[0028] Calculate the second loss value of the second model to be trained based on the semantic features and the second feature points;
[0029] Based on the relationship between the second loss value and the second preset loss criterion, optimize the second model to be trained and output the second target image generation model.
[0030] An embodiment of the second aspect of the present disclosure provides an image generation method, the method including:
[0031] Obtain the structural features of the image to be generated;
[0032] Use the first target image generation model to obtain an initial target image based on the structural features and a preset generated image;
[0033] Use the second target image generation model to optimize the initial target generated image based on the structural features to generate a target image, where the first target image generation model and the second target image generation model are obtained according to the method in the first aspect and any optional implementation manner of the first aspect.
[0034] In an embodiment of the present disclosure, the preset generated image is an image generated based on the text information corresponding to the image to be generated.
[0035] In an embodiment of the present disclosure, the step of using the first target image generation model to obtain an initial target image based on the structural features and a preset generated image includes:
[0036] Use the first target image generation model to generate first structural point features based on the first feature points of the structural features, where the first feature points are all the feature points of the structural features;
[0037] Calculate the pixel loss value between the first structural point features and the preset generated image;
[0038] Optimize the preset generated image based on the relationship between the pixel loss value and the first preset pixel standard to obtain an initial target generated image.
[0039] In an embodiment of the present disclosure, the method of using the second target image generation model to optimize the initial target generated image based on the structural feature to generate a target image includes:
[0040] Use the second target image generation model to obtain a second structural point feature based on the second feature point of the structural feature, where the second feature point is the target feature point of the structural feature;
[0041] Perform feature fusion on the second structural point feature and the image feature of the initial target generated image to obtain a fusion feature;
[0042] Calculate the structural loss value of the second model to be trained based on the fusion feature and the second structural point feature;
[0043] Optimize the fusion feature based on the relationship between the structural loss value and the second preset structural standard, and output the target image.
[0044] An embodiment of the third aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor runs the computer program to implement the methods described in the first aspect or any optional implementation manner of the first aspect, the second aspect, and any optional implementation manner of the second aspect.
[0045] An embodiment of the fourth aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. The program is executed by a processor to implement the methods described in the first aspect and any optional implementation manner of the first aspect, the second aspect, and any optional implementation manner of the second aspect.
[0046] The technical solutions provided in the embodiments of the present disclosure at least have the following technical effects or advantages:
[0047] Construct an image training set, where the image training set includes an image training subset and a structural feature training subset, and the image training subset and the structural feature training subset correspond one by one; train a first model to be trained based on the first feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, and obtain a first target image generation model, realizing the fusion of structural features and image features in the image generation process; further, train a second model to be trained based on the second feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, and obtain a second target image generation model; since the second feature points are the target feature points in the first feature points, the quality of subsequent image generation is improved on the basis of the first training model.
[0048] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present disclosure. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0050] In the drawings:
[0051] Figure 1 shows a flowchart of a method for training an image generation model provided by an embodiment of the present disclosure;
[0052] Figure 2 shows a schematic diagram of the training of the first model to be trained in a method for training an image generation model provided by an embodiment of the present disclosure;
[0053] Figure 3 shows a schematic diagram of the training of the second model to be trained in a method for training an image generation model provided by an embodiment of the present disclosure;
[0054] Figure 4 shows a flowchart of an image generation method provided by an embodiment of the present disclosure;
[0055] Figure 5 shows a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure;
[0056] Figure 6 shows a schematic diagram of a storage medium provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0058] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present disclosure should have the ordinary meaning understood by those skilled in the art to which the present disclosure pertains.
[0059] An embodiment of the present disclosure provides an image generation model training method, as Figure 1 shown in an embodiment of the present disclosure, an image generation model training method includes the following steps:
[0060] In step S11, an image training set is constructed.
[0061] Among them, the image training set includes an image training subset and a structural feature training subset, and the image training subset and the structural feature training subset correspond one by one.
[0062] Exemplarily, the image training subset and the structural feature training subset in the image training set correspond. According to the application scenario of the embodiment of the present disclosure, the image training set can be constructed according to the type of the application scenario. Specifically, if the image generation model is used for generating images of the human body structure type after training, then the image training subset in the image training set includes human body images, and the structural feature training subset is the structural features corresponding to the human body images. If the image generation model is used for generating images of the building type after training, then the image training subset in the image training set includes building images, and the structural feature training subset is the structural features corresponding to the building images.
[0063] It should be noted that for the scenarios corresponding to the same application type, the images should also be of the same type. For example, for the scenarios of the human body structure, the images are all images including the human body structure. The embodiment of the present disclosure takes the application scenario of generating images of the human body structure type as an example for introduction.
[0064] In some embodiments, when obtaining the image training set, it is necessary to correspond the image training subset and the structural feature training subset. Therefore, the corresponding structural features in the structural feature training subset can be generated according to the images in the image training subset. Specifically, the image training set can be determined by the following method.
[0065] Obtain an initial image set; for any initial image in the initial image set, extract the segmentation mask of the initial image; based on the initial image and the segmentation mask, extract the image feature map of the initial image; based on the relationship between the segmentation mask and the image feature map, determine whether the initial image meets a preset standard; screen the grayscale images corresponding to the initial images that meet the preset standard in the initial image set to construct an image training subset; and use the image feature maps corresponding to the initial images in the image training subset to construct a structural feature training subset.
[0066] Among them, the grayscale image corresponding to the initial image is obtained by setting the pixel values of the non-image features in the initial image to grayscale values, and the non-image features are obtained by subtracting the image feature map from the initial image.
[0067] Exemplarily, the images in the initial training image set are all images including the human body. For the initial images in the initial training image set, the Mask R-CNN network can be used to extract the segmentation masks of the human bodies in the initial images; among them, the image feature map can be obtained by extracting the human feature map in the image or by encoding the image, and calculate the intersection over union index between the segmentation mask and the image feature map; screen the grayscale images corresponding to the initial images that meet the preset standard in the initial image set to construct an image training subset, where the preset standard can be that the intersection over union index greater than 0.5 is considered to meet the preset standard.
[0068] Specifically, based on the segmentation mask and the image feature map, calculate the intersection over union index; if the intersection over union index is greater than the preset threshold, the initial image meets the preset standard; if the intersection over union index is less than or equal to the preset threshold, the initial image does not meet the preset standard.
[0069] In some embodiments, the image feature map can also be obtained in the following way: use the segmentation mask to extract the initial structural features of the initial image; map the pixel values of the initial structural features to the pixel points at the corresponding positions of the preset image to obtain the image feature map.
[0070] Exemplarily, the initial image can be a grayscale image or a color image.
[0071] The segmentation mask and the initial image are two-dimensional arrays of the same size, and the pixel values therein represent the category or region belonging of each pixel. Use the segmentation mask to extract the pixels of a specific region from the initial image. The specific operation is to multiply the mask pixel by pixel with the image pixel by pixel (the region with a mask value of 1 is retained, and the region with a value of 0 is ignored).
[0072] Extract the structural features from the region extracted by the mask. The preset image can be a blank image of the same size as the initial image, or other images. Map the pixel values of the extracted structural features (such as the structure diagram) to the corresponding positions of the preset image. This is usually done by assigning values pixel by pixel.
[0073] In step S12, based on the first feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, train the first model to be trained to obtain the first target image generation model.
[0074] Exemplarily, in the embodiments of the present disclosure, a Vision Transformer (ViT) network can be used to extract the first feature points in any structural feature in the structural feature training subset, where the first feature points can be the feature points that retain all human body structures, for example, 27,533 points. In addition, a Contrastive Language–Image Pre-training (CLIP) image encoder can be used to encode the images in the image training subset to obtain image features.
[0075] Locate the key first feature points from the structural features. For example, a corner detection algorithm (such as Harris corner detection) or a feature point extraction algorithm (such as SIFT, SURF) can be used to locate the feature points.
[0076] The input of the first model to be trained includes the first feature points of the structural features. The output of the first model to be trained is the predicted image features, that is, the global image features predicted from the feature points of the input structural features.
[0077] The first model to be trained can select a suitable model architecture. For example: a fully connected neural network: if the input is a feature vector, a fully connected neural network can be used. A convolutional neural network (CNN): if the input is a feature map, a CNN can be used. A Transformer architecture: if sequential feature points need to be processed, a Transformer can be used.
[0078] Use the data pair of the structural point features and image features of the corresponding structural features, that is, the image features of the corresponding images in the image training subset as the ground truth, and the predicted image features obtained by using the first model to be trained according to the structural point features of the structural features as the training values; select a suitable loss function, such as mean squared error (MSE) or cross-entropy loss, to calculate the difference between the ground truth and the training values. Use an optimization algorithm (such as Adam or SGD) to adjust the model parameters to minimize the loss function, and then obtain the first target image generation model.
[0079] Furthermore, the first target image generation model can be further optimized and evaluated, and the model architecture or training parameters can be adjusted according to the evaluation results, such as increasing the number of layers, adjusting the learning rate, or using data augmentation. Repeat the training and evaluation process until the first target image generation model reaches satisfactory performance.
[0080] In some embodiments, the above step S12 can also be implemented in the following manner: for any structural feature in the structural feature training subset, use the first model to be trained to extract the first structural point feature of the first feature point of the structural feature; based on the first structural point feature and the image feature of the corresponding image in the image training subset, calculate the first loss value of the first model to be trained; based on the relationship between the first loss value and the first preset loss criterion, optimize the first model to be trained, and output the first target image generation model.
[0081] Exemplarily, as Figure 2 shown, in the embodiments of the present disclosure, the first model to be trained can be a ViT model. Use the ViT model to be trained to extract the first structural point feature of the first feature point of the structural feature; the cosine similarity contrast loss function can be used to calculate the first loss value between the first structural point feature and the image feature. Optimization algorithms (such as SGD, Adam) can be used to update the VIT model parameters according to the gradient of the first loss value. Dynamically adjust the learning rate according to the training progress, for example, use the learning rate decay strategy. If the loss value no longer decreases within a certain number of iterations, stop the training to prevent overfitting.
[0082] In step S13, based on the second feature point of any structural feature in the structural feature training subset and the image feature of the image corresponding to the structural feature, train the second model to be trained to obtain the second target image generation model, where the second feature point is the target feature point among the first feature points.
[0083] Exemplarily, the difference in the way of training the second model to be trained from the above step S12 is that after predicting the second structural point feature through the second feature point of any structural feature, the second structural point feature needs to be fused with the image feature, and then the loss value is calculated with the second feature point, and then the parameters of the second model to be trained are adjusted according to the loss value for optimization.
[0084] In some embodiments, the above step S13 includes: for any structural feature in the structural feature training subset, use the second model to be trained to extract the second structural point feature of the second feature point of the structural feature; fuse the second structural point feature and the image feature of the corresponding image in the image training subset to obtain the semantic feature of the second structural point feature; calculate the second loss value of the second model to be trained based on the semantic feature and the second feature point; based on the relationship between the second loss value and the second preset loss criterion, optimize the second model to be trained, and output the second target image generation model.
[0085] Exemplarily, the second structural point feature corresponding to the second feature point can be generated by the first target image generation model obtained in step S12, or can be generated by other models or networks, which is not limited here.
[0086] As shown Figure 3 in the figure, the human body image passes through the CLIP image encoder to obtain the image feature K vector and V vector. The second structural point feature is used as the Q vector as the input of the second model to be trained, and feature fusion is performed to obtain the semantic feature. Further, based on the semantic feature and the second structural point feature, the second loss value is calculated, and the parameters of the second model to be trained are further adjusted until the second loss value meets the requirements, and the second target image generation model is obtained.
[0087] Through the image generation model training method of the embodiments of the present application, an image training set is constructed, where the image training set includes an image training subset and a structural feature training subset, and the image training subset and the structural feature training subset correspond one by one; based on the first feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, the first model to be trained is trained to obtain the first target image generation model, realizing the fusion of structural features and image features in the image generation process; further, based on the second feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, the second model to be trained is trained to obtain the second target image generation model; since the second feature points are the target feature points in the first feature points, the quality of subsequent image generation is improved on the basis of the first training model.
[0088] Corresponding to the above embodiments, the embodiments of the present disclosure further provide an image generation method, which is an application method of the image reconstruction model obtained by the image generation model training method as shown Figure 1 in the figure, as shown Figure 4 in the figure, the method includes:
[0089] In step S41, the structural features of the image to be generated are obtained.
[0090] Exemplarily, when generating an image, according to the structural features, generation guidance in terms of structural features can be provided on the basis of relevant image generation, improving the accuracy of the image. For example, in the embodiments of the present disclosure, for an image used to generate a human body structure, the structural features may be the human body structure information corresponding to the image to be generated. Specifically, it may be the positions corresponding to each key point in the human body structure, as well as the mutual position relationships between the points, etc.
[0091] In step S42, using the first target image generation model, based on the structural features and the preset generated image, an initial target image is obtained.
[0092] Exemplarily, the preset generated image is a blank image with the same height and width as the image to be generated. The blank image refers to an image in which the pixel values are default values. In some embodiments, the preset generated image may also be a preliminary image pre-generated according to the text content.
[0093] Further, input the structural features into the first target image generation model trained in the above embodiments to generate an initial target image. Specifically, to further obtain a more accurate initial target image, in some embodiments, the above step S42 can also be implemented in the following manner: Use the first target image generation model to generate first structural point features based on the first feature points of the structural features, where the first feature points are all the feature points of the structural features; calculate the pixel loss value between the first structural point features and a preset generated image; and optimize the preset generated image based on the relationship between the pixel loss value and a first preset pixel standard to obtain an initial target generated image.
[0094] Exemplarily, use the first target image generation model to process the first target image and extract all the feature points (first feature points) of its structural features. According to the extracted first feature points, generate first structural point features. This may involve converting the feature points into feature vectors or feature maps for subsequent processing.
[0095] Design a loss function to calculate the pixel loss value between the first structural point features and the preset generated image. Commonly used loss functions include mean squared error (MSE), mean absolute error (MAE), or more complex perceptual loss.
[0096] Compare the first structural point features with the preset generated image, and use the defined loss function to calculate the pixel loss value. Select an optimization algorithm to minimize the pixel loss value. Commonly used optimization algorithms include gradient descent (such as SGD, Adam, etc.) or advanced optimization methods based on genetic algorithms, particle swarm optimization, etc. Initialize the parameters of the optimization algorithm (such as learning rate, number of iterations, etc.). In each iteration, update the pixel values of the preset generated image using the optimization algorithm according to the current pixel loss value of the preset generated image. Repeat the iteration process until a preset stop condition is reached (such as the loss value converges, the maximum number of iterations is reached, etc.). When the optimization process is completed, obtain the optimized preset generated image, which is the initial target generated image. Verify whether the initial target generated image meets the expected structural feature requirements.
[0097] In step S43, use the second target image generation model to optimize the initial target generated image based on the structural features to generate a target image, where the first target image generation model and the second target image generation model are trained according to the above embodiments.
[0098] Exemplarily, the accuracy of the image corresponding to the initial target generated image obtained in the above step S42 in practical applications may not meet the application requirements. Therefore, it is also necessary to further use the second target image generation model to optimize the initial target generated image according to the second feature points in the structural features to obtain a target image.
[0099] Different from the above step S42, after predicting the second structural point feature through the second feature point of any structural feature, the second structural point feature needs to be fused with the initial target generated image feature, and then the loss value is calculated with the second feature point. Furthermore, the parameters of the second model to be trained are adjusted based on the loss value for optimization.
[0100] Specifically, the above step S43 can also be implemented in the following way: Using the second target image generation model, based on the second feature point of the structural feature, obtain the second structural point feature, where the second feature point is the target feature point of the structural feature; perform feature fusion on the second structural point feature and the image feature of the initial target generated image to obtain a fused feature; calculate the structural loss value of the second model to be trained based on the fused feature and the second structural point feature; optimize the fused feature based on the relationship between the structural loss value and the second preset structural standard, and output the target image.
[0101] Specifically, this optimization method is basically the same as the training method of the second target image generation model in the above embodiment. The difference is that the parameter adjustment of the model becomes the adjustment of the initial target generated image, and then the target image is obtained after reaching the iteration stop condition.
[0102] Through the image generation method of the embodiments of the present application, obtain the structural feature of the image to be generated; use the first target image generation model to obtain the initial target image based on the structural feature and the preset generated image; use the second target image generation model to optimize the initial target generated image based on the structural feature and generate the target image. Through the two-stage generation process, the initial target image is generated in the first stage, and the image details are further optimized in the second stage. This phased method can gradually improve the quality of the generated image and reduce the errors that may be brought about by one-time generation. The second target image generation model can focus on optimizing the details, textures, and structures of the image on the basis of the first stage, so as to generate more realistic and higher-quality images. It can better capture the core information of the input features and avoid overfitting the noise of the training data during the one-time generation process.
[0103] Corresponding to the implementation manners of the above methods, the embodiments of the present disclosure also provide an image generation model training device for executing the image generation model training method of any of the above Figure 1 illustrated embodiments. The device for training the image generation model includes:
[0104] A dataset construction module for constructing an image training set, where the image training set includes an image training subset and a structural feature training subset, and the image training subset and the structural feature training subset correspond one by one;
[0105] The first training module is used to train the first feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, and train the first model to be trained to obtain the first target image generation model;
[0106] The second training module is used to train the second feature points of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, and train the second model to be trained to obtain the second target image generation model, where the second feature points are the target feature points among the first feature points.
[0107] The image generation model training device provided in the above embodiments of the present disclosure and the image generation model training method provided in the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0108] Corresponding to the implementation manners of the above methods, the embodiments of the present disclosure also provide an image reconstruction device for performing the image reconstruction method of any one of the above Figure 4 schematic embodiments. The image reconstruction device includes:
[0109] The feature acquisition module is used to acquire the structural features of the image to be generated;
[0110] The first generation module is used to use the first target image generation model to obtain an initial target image based on the structural feature and a preset generated image;
[0111] The second generation module is used to use the second target image generation model to optimize the initial target generated image based on the structural feature and generate a target image, where the first target image generation model and the second target image generation model are obtained according to the above embodiments.
[0112] The image generation device provided in the above embodiments of the present disclosure and the image generation method provided in the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0113] The embodiments of the present disclosure also provide an electronic device for performing the above method. Please refer to Figure 5 , which shows a schematic diagram of an electronic device provided in some embodiments of the present disclosure. As Figure 5 shown, the electronic device includes: a processor 500, a memory 501, a bus 502, and a communication interface 503. The processor 500, the communication interface 503, and the memory 501 are connected through the bus 502; a computer program that can run on the processor 500 is stored in the memory 501, and when the processor 500 runs the computer program, it executes the foregoing Figure 1 orFigure 4 The method provided by any of the schematic embodiments.
[0114] Among them, the memory 501 may include a high-speed random access memory (Random Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between this system network element and at least one other network element is realized through at least one communication interface 503 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.
[0115] The bus 502 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory 501 is used to store a program. After receiving an execution instruction, the processor 500 executes the program. The foregoing Figure 1 or Figure 5 The method disclosed by any of the schematic embodiments can be applied to the processor 500 or implemented by the processor 500.
[0116] The processor 500 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 500 or by instructions in software form. The above-mentioned processor 500 may be a general-purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory or electrically erasable programmable memory, register, etc. This storage medium is located in the memory 501, and the processor 500 reads the information in the memory 501 and combines its hardware to complete the steps of the above method.
[0117] The electronic device provided by the embodiments of the present disclosure and the method provided by the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the method adopted, run or implemented by it.
[0118] Embodiments of the present disclosure also provide a computer-readable storage medium corresponding to the method provided in the foregoing embodiments. Please refer to Figure 6 , which shows that the computer-readable storage medium is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the method provided in any of the foregoing embodiments.
[0119] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0120] The computer-readable storage medium provided in the above embodiments of the present disclosure and the method provided in the embodiments of the present disclosure are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.
[0121] It should be noted that:
[0122] In the specification provided here, a large number of specific details are described. However, it can be understood that the embodiments of the present disclosure can be practiced without these specific details. In some instances, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this specification.
[0123] Similarly, it should be understood that, in order to streamline the present disclosure and help understand one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present disclosure, the various features of the present disclosure are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the following schematic diagram: that is, the claimed present disclosure requires more features than those explicitly recited in each embodiment. The inventive aspect lies in less than all the features of the single embodiments disclosed previously. Therefore, the implementation manners following the specific implementation manners are hereby clearly incorporated into the specific implementation manners, where each implementation manner itself is a separate embodiment of the present disclosure.
[0124] In addition, those skilled in the art can understand that, although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present disclosure and forms different embodiments.
[0125] As described above, this is only a preferred specific embodiment of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed by the present disclosure should be covered within the protection scope of the present disclosure.
Claims
1. A method for training an image generation model, characterized in that: The method comprises: Constructing an image training set, wherein the image training set includes an image training subset and a structural feature training subset, and the image training subset corresponds to the structural feature training subset one by one; Based on the first feature point of any structural feature in the structural feature training subset and the image feature of the image corresponding to the structural feature, training a first model to be trained to obtain a first target image generation model; Based on the second feature point of any structural feature in the structural feature training subset and the image features of the image corresponding to the structural feature, a second model to be trained is trained to obtain a second target image generation model, wherein the second feature point is a target feature point among the first feature points.
2. The method according to claim 1, characterized in that The step of constructing an image training set comprises: Get an initial set of images; For any initial image in the initial image set, extract a segmentation mask of the initial image; Extracting an image feature map of the initial image based on the initial image and the segmentation mask; Determining whether the initial image meets a preset standard based on a relationship between the segmentation mask and the image feature map; The image training subset is constructed by screening the grayscale images corresponding to the initial images that meet the preset criteria in the initial image set; and the structural feature training subset is constructed by using the image feature map corresponding to the initial images in the image training subset, wherein the grayscale image corresponding to the initial image is obtained by setting the pixel values of the non-image features in the initial image to grayscale values, and the non-image features are obtained by subtracting the image feature map from the initial image.
3. The method according to claim 2, characterized in that The step of extracting an image feature map of the initial image based on the initial image and the segmentation mask comprises: extracting initial structural features of the initial image using the segmentation mask; The pixel values of the initial structural features are mapped to pixel points at corresponding positions of a preset image to obtain the image feature map.
4. The method according to claim 2, characterized in that: The determining whether the initial image meets a preset standard based on the relationship between the segmentation mask and the image feature map includes: Calculating an intersection-over-union (IoU) index based on the segmentation mask and the image feature map; If the intersection-over-union ratio index is greater than a preset threshold, the initial image meets the preset standard; If the intersection-over-union ratio index is less than or equal to the preset threshold, the initial image does not meet the preset standard.
5. The method according to claim 1, characterized in that The step of training a first model to be trained based on any structural feature in the structural feature training subset and the image feature of the image corresponding to the structural feature to obtain a first target image generation model includes: For any structural feature in the structural feature training subset, extracting a first structural point feature of a first feature point of the structural feature using a first model to be trained; Calculating a first loss value of the first to-be-trained model based on the first structure point feature and image features of corresponding images in the image training subset; Based on the relationship between the first loss value and the first preset loss standard, the first model to be trained is optimized and a first target image generation model is output.
6. The method according to claim 1, characterized in that The step of training a second model to be trained based on a second feature point of any structural feature in the structural feature training subset and an image feature of an image corresponding to the structural feature to obtain a second target image generation model comprises: For any structural feature in the structural feature training subset, extracting a second structural point feature of a second feature point of the structural feature using a second model to be trained; Performing feature fusion on the second structure point feature and the image feature of the corresponding image in the image training subset to obtain a semantic feature of the second structure point feature; Calculate a second loss value of the second to-be-trained model based on the semantic feature and the second feature point; Based on the relationship between the second loss value and the second preset loss standard, the second model to be trained is optimized and a second target image generation model is output.
7. An image generation method, characterized in that: The method comprises: Obtaining structural features of the image to be generated; Using the first target image generation model, based on the structural features and the preset generation image, an initial target image is obtained; The target image is generated by optimizing the initial target image based on the structural features using a second target image generation model, wherein the first target image generation model and the second target image generation model are obtained according to the method as described in any one of claims 1 to 6.
8. The method according to claim 7, characterized in that The preset generated image is an image generated based on text information corresponding to the image to be generated.
9. The method according to claim 7, characterized in that: The method of using the first target image generation model to generate an image based on the structural features and the preset to obtain an initial target image includes: Generate a first structural point feature based on the first feature point of the structural feature by using the first target image generation model, where the first feature point is all the feature points of the structural feature; Calculating a pixel loss value of the first structure point feature and a preset generated image; Based on the relationship between the pixel loss value and the first preset pixel standard, the preset generated image is optimized to obtain an initial target generated image.
10. The method according to claim 7, characterized in that The step of using the second target image generation model to optimize the initial target generation image based on the structural features to generate the target image includes: Using the second target image to generate a model, based on a second feature point of the structural feature, a second structural point feature is obtained, where the second feature point is a target feature point of the structural feature; Performing feature fusion on the second structure point feature and the image feature of the initial target generated image to obtain a fused feature; Calculate the structural loss value of the second to-be-trained model based on the fusion feature and the second structural point feature; Based on the relationship between the structural loss value and a second preset structural standard, the fusion feature is optimized and the target image is output.
Citation Information
Patent Citations
Restoration body image generation method and device, equipment and storage medium
CN113888615A
Training-free microscopic image key area identification method and device
CN118710662A
System and method for processing colon image data
US20220122263A1
Cited By
Automatic canning control method and system for fruit and vegetable cans
CN120972644A