Image generation method and device, equipment and storage medium
By obtaining and comparing the semantic differences between the first generated image and the target image output by the image generation model, adjusting the parameters of the image generation model, and generating the second generated image, the problem of semantic inconsistency of image data in the existing technology is solved, and the semantic processing accuracy of the image generation model and the quality of the generated image are improved.
Patent Information
- Application Number
- CN202510679596.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-19
AI Technical Summary
In the existing technology, the pre-generated image data does not meet the semantic requirements of image data processing tasks, which causes the image processing model to capture incorrect semantic features during training, affecting the task processing effect.
By obtaining the first semantics of the first generated image output by the image generation model and the second semantics of the target image, the semantic difference between the two is determined, and the parameters of the image generation model are adjusted based on the semantic difference to generate a second generated image to improve image quality and semantic consistency.
The accuracy of the semantic processing function of the output images of the image generation model is improved, the semantic consistency between the generated images and the target images is improved, and the semantic capture ability of the data processing model during training is enhanced.
Smart Images

Figure CN120672883A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image generation technology, and in particular to an image generation method, apparatus, device and storage medium. Background Art
[0002] In practical applications, with the continuous development of deep learning model applications, image data processing tasks, including image generation, data augmentation, and semantic segmentation, need to be implemented using image processing models. These models, in turn, rely on pre-generated image data for training. However, in related technologies, pre-generated image data often fails to meet the semantic requirements of image data processing tasks. Therefore, generating image data that meets these semantic requirements has become a pressing technical challenge. Summary of the Invention
[0003] Based on the above technical problems, the embodiments of the present application provide an image generation method, apparatus, device and storage medium.
[0004] The technical solution provided by the embodiments of this application is as follows:
[0005] The present invention first provides an image generation method, which includes:
[0006] Obtain a first semantic meaning of a first generated image output by an image generation model; obtain a second semantic meaning of a target image; determine a semantic difference between the first semantic meaning and the second semantic meaning; adjust parameters of the image generation model based at least on the semantic difference to obtain an image generation model with adjusted parameters; generate a second generated image using the generation model with adjusted parameters; wherein a first degree of difference between the first generated image and the target image is greater than a second degree of difference between the second generated image and the target image.
[0007] An embodiment of the present application also provides an image generation device, which includes: an acquisition module for acquiring a first semantic meaning of a first generated image output by an image generation model; acquiring a second semantic meaning of a target image; a processing module for determining a semantic difference between the first semantic meaning and the second semantic meaning; adjusting the parameters of the image generation model based at least on the semantic difference to obtain an image generation model with adjusted parameters; a generation module for generating a second generated image using the generation model with adjusted parameters; wherein the first degree of difference between the first generated image and the target image is greater than the second degree of difference between the second generated image and the target image.
[0008] An embodiment of the present application further provides an electronic device, comprising a processor and a memory; a computer program is stored in the memory; and when the computer program is executed by the processor, it can implement any of the above-described image generation methods.
[0009] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored; when the computer program is executed by a processor of an electronic device, it can implement any of the above-described image generation methods.
[0010] An embodiment of the present application also provides a computer program product; the program product includes a computer program; when the computer program is executed by a processor of an electronic device, it can implement any of the above-mentioned image generation methods.
[0011] The technical solution provided in the embodiments of the present application has the following beneficial effects:
[0012] The image generation method provided in the embodiment of the present application achieves tracking and acquisition of the semantics of the first generated image output by the image generation model and the target image by acquiring the first semantics of the first generated image and the second semantics of the target image; and, by determining the semantic difference between the first semantics and the second semantics, achieves accurate tracking and characterization of the difference between the first semantics and the second semantics; on this basis, the parameters of the image generation model are adjusted at least based on the semantic difference to obtain the image generation model after parameter adjustment, which can achieve targeted adjustment of the parameters in the image generation model, thereby improving the accuracy of the semantic processing function of the parameter-adjusted image generation model in the image generation process; at the same time, a second generated image is generated by the parameter-adjusted image generation model, and the first difference between the first generated image and the target image is greater than the second difference between the second generated image and the target image. In this way, the image quality of the second generated image is improved relative to the first generated image, and in particular, the consistency between the semantics of the second generated image and the semantics of the target image can be improved; on the other hand, when the second generated image is applied to the training of the data processing model, the ability of the data processing model to accurately capture the semantics it carries can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A flowchart of an image generation method provided in an embodiment of the present application;
[0014] Figure 2 A schematic diagram of a network structure for correcting image data provided in an embodiment of the present application;
[0015] Figure 3 A schematic diagram of the structure of an image generating device provided in an embodiment of the present application;
[0016] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0018] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0019] In the following description, the terms "first\second" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0020] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0021] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0022] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in examples, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0023] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0024] 1) Texture: Texture refers to the repetitive or regular patterns and structures that appear visually in local areas of an image. These patterns are typically not directly determined by the shape or edges of an object, but rather by the arrangement, intensity, and color distribution of local pixels. Texture is represented by the grayscale distribution of a pixel and its surrounding spatial neighborhood. The varying degrees of repetitiveness of local texture information contribute to global texture information.
[0025] 2) Semantics: This refers to the meaning of image content. Image semantics can be expressed through language, including natural language and symbolic language (mathematical language). However, the expression of image semantics is not limited to natural language; its extension corresponds to all ways in which the human visual system understands images.
[0026] 3) Image generation model: refers to a machine learning model that can generate new images, usually based on deep learning techniques such as Convolutional Neural Networks (CNN) and Generative Adversarial Networks (GAN).
[0027] In order to better understand the image generation method provided in the embodiments of the present application, the relevant technology is first described below.
[0028] In practical applications, with the continuous development of deep learning model applications, image data processing tasks including image generation, data enhancement, and semantic segmentation need to be implemented through image processing models; and image processing models need to rely on pre-generated image data for training.
[0029] However, in practical applications, there are often some unreasonable aspects in pre-generated image data. For example, the pre-generated image data does not match the semantics of the image data processing tasks, which causes the model corresponding to the image data processing tasks to capture incorrect semantic features during the training process, and then causes the task processing effect of the above data processing tasks to be unable to meet the actual image processing requirements.
[0030] Based on the above technical problems, the embodiments of the present application provide an image generation method, apparatus, device and storage medium.
[0031] The embodiment of the present application first provides an image generation method; Figure 1 A flow chart of the image generation method provided in the embodiment of the present application is shown as follows: Figure 1 As shown, the method may include the following steps:
[0032] S101. Obtain a first semantic meaning of a first generated image output by an image generation model.
[0033] In one embodiment, the image generation model may include a neural network model having an image generation function; illustratively, the image generation model may include Stable Difffusion or GAN.
[0034] Among them, GAN can include a generator and a discriminator. The generator and the discriminator compete with each other during the data processing process and improve together. The ultimate goal is to enable the generator to generate image data that is difficult to distinguish from real image data.
[0035] Stable Difffusion can include U-Net, Contrastive Language–Image Pre-training (CLIP), and Variational Auto-Encoders (VAE). U-Net is one of the core components of the Stable Diffusion model, primarily used for denoising. Stable Diffusion uses CLIP to process text input data, converting it into a format the model can understand. CLIP's role in Stable Diffusion is to convert text cues into a latent space representation, thereby guiding image generation. VAE uses the encoder to convert the input image into the latent space, and the decoder to restore the high-resolution image from the latent space.
[0036] In one embodiment, the first generated image may include K images generated by an image generation model processing input data; exemplarily, the input data may include text data, original images, and voice data, etc., wherein the text data and voice data may include a set of features that can be included in the image generated by the image generation model; K may be an integer greater than or equal to 1.
[0037] In one embodiment, the first semantics may include content contained in the first generated image that can be understood and interpreted by humans; exemplarily, the first semantics may specifically include elements such as objects, object type identifiers, scenes, and events in the first generated image, and may also include the relationships between the above elements; exemplarily, objects may have life signs, for example, objects may include humans or animals, and objects may also not have life signs, for example, objects may include streets, vehicles, traffic lights, flowing water, and electronic equipment, etc.; exemplarily, the type identifier may include labels such as the name or number of the type to which the object belongs; exemplarily, the scene may include the environment in which the object is located, for example, the environment may include the natural environment and the social environment, and may also include home and office, etc.; exemplarily, the event may include a general description of the behavior and / or status of the object in the above scene, which is not limited in the embodiments of the present application.
[0038] In one embodiment, the first semantics can be obtained by:
[0039] The first generated image is subjected to semantic extraction by a neural network model having a semantic extraction function, thereby obtaining a first semantic meaning; wherein the neural network model may include a CNN. The CNN generally includes an input layer, a convolutional layer, a pooling layer, and an output layer; the input layer is used to input the first generated image, the convolutional layer is used to extract image features of the first generated image through the convolution kernels contained therein, the pooling layer is used to reduce the dimensionality of the image features to extract feature invariance and reduce the risk of overfitting, and the output layer is used to classify or probabilistically output the results of the dimensionality reduction processing of the pooling layer, and determine the first semantic meaning based on the classification or probabilistic output results.
[0040] S102: Acquire the second semantics of the target image.
[0041] In one embodiment, the target image can be actually acquired by an image acquisition device, and the target image can include K images corresponding to the first generated image, that is, the target image can be a real image corresponding to the objective world.
[0042] In one embodiment, the second semantics may include the content contained in the target image that can be understood and interpreted by humans; illustratively, the second semantics may specifically include elements such as objects, object type identifiers, scenes, and events in the target image, and may also include the relationships between the above elements.
[0043] Exemplarily, the method for acquiring the second semantics can be the same as the method for acquiring the first semantics, and the neural network models used to acquire the first semantics and the second semantics can be independent of each other, but the two can have the same parameters and structure. That is, two completely identical neural network models can be used to process the first generated image to obtain the first semantics, and to process the target image to obtain the second semantics, so as to improve the consistency between the first semantics and the second semantics determination process.
[0044] S103: Determine the semantic difference between the first semantics and the second semantics.
[0045] In one embodiment, the semantic difference may include the degree of matching between the first semantics and the second semantics; illustratively, the degree of matching may be reflected in the form of a matching score, and the degree of matching may increase as the matching score increases. Accordingly, the semantic difference may decrease as the degree of matching increases.
[0046] In one embodiment, semantic differences may be determined by:
[0047] The first semantics and the second semantics are compared and judged by a network model having a semantic recognition function, thereby obtaining a semantic difference; illustratively, the network model may include an artificial intelligence (AI) model.
[0048] Among them, the above-mentioned AI model may include Transformer, the core idea of which is to parse the first semantics and the second semantics of the input through the attention mechanism, so as to better understand the contextual relationship contained in the first semantics and the second semantics. Exemplarily, Transformer may include an encoder, a decoder, a self-attention mechanism, and positional encoding. Among them, the encoder is used to process the input sequence represented by the first semantics and the second semantics, encode the input sequence through the self-attention mechanism, and generate a series of vector representations; the decoder is used to generate the output sequence, and also uses the self-attention mechanism to process the output of the encoder, and interacts with the output of the encoder through the attention mechanism to generate the final output sequence; the self-attention mechanism is the core of Transformer, which calculates the correlation between each element in the input sequence and other elements and assigns different weights to each element, so as to better understand the contextual information of the sequence; positional encoding is used to add position information to each element to help the model understand the sequential relationship of elements.
[0049] S104: Adjust parameters of the image generation model based at least on the semantic difference to obtain an image generation model with adjusted parameters.
[0050] In one embodiment, the parameter-adjusted image generation model can be obtained by any of the following methods:
[0051] A first parameter is determined based on the semantic difference, and the first parameter of the image generation model is adjusted to obtain the image generation model after parameter adjustment; wherein the first parameter may include a parameter in the image generation model corresponding to the above-mentioned semantic dimension.
[0052] A parameter adjustment step is determined based on the degree of semantic difference, and a first parameter of the image generation model is adjusted based on the parameter adjustment step to obtain an image generation model after parameter adjustment.
[0053] It should be noted that the above-mentioned process of adjusting the parameters of the image generation model can be a cyclic recursive adjustment process. For example, if the qth semantic difference between the semantics of the qth first generated image and the semantics of the target image is greater than the semantic threshold, the parameters of the image generation model are adjusted for the qth time to obtain the image generation model after the qth parameter adjustment. If the q+1th semantic difference between the q+1th first generated image generated by the image generation model after the qth parameter adjustment and the target image is greater than the semantic threshold, then through the above-mentioned method, at least based on the q+1th semantic difference between the semantics of the q+1th first generated image and the semantics of the target image, the image generation model after the qth parameter adjustment is adjusted to obtain the image generation model after the q+1th parameter adjustment, until the difference between the semantics of the first generated image and the semantics of the target image is less than or equal to the semantic threshold, at this time, the image generation model after the parameter adjustment can be obtained; wherein, q can be an integer greater than or equal to 1.
[0054] S105 . Generate a second generated image using the image generation model after parameter adjustment.
[0055] A first degree of difference between the first generated image and the target image is greater than a second degree of difference between the second generated image and the target image.
[0056] In one embodiment, the first degree of difference may include the degree of difference between the first generated image and the target image in thematic dimension, semantic dimension, and picture quality dimension; accordingly, the second degree of difference may include the degree of difference between the second generated image and the target image in thematic dimension, semantic dimension, and picture quality dimension; exemplarily, the picture quality dimension may include image integrity, authenticity, and clarity, etc.
[0057] In one embodiment, the second generated image may be generated by any of the following methods:
[0058] The first generated image is corrected by the image generation model after parameter adjustment to obtain a second generated image.
[0059] The data is processed by the image generation model with adjusted parameters to generate a second generated image.
[0060] In one embodiment, the second generated image can be used as sample data, specifically for training a data processing model for a target task.
[0061] From the above, it can be seen that the image generation method provided in the embodiment of the present application achieves tracking and acquisition of the semantics of the first generated image output by the image generation model and the target image by obtaining the first semantics of the first generated image and the second semantics of the target image; and, by determining the semantic difference between the first semantics and the second semantics, it achieves accurate tracking and characterization of the difference between the first semantics and the second semantics; on this basis, the parameters of the image generation model are adjusted at least based on the semantic difference to obtain the image generation model after parameter adjustment, which can achieve targeted adjustment of the parameters in the image generation model, thereby improving the accuracy of the semantic processing function of the parameter-adjusted image generation model in the image generation process; at the same time, the second generated image is generated by the parameter-adjusted image generation model, and the first difference between the first generated image and the target image is greater than the second difference between the second generated image and the target image. In this way, the image quality of the second generated image is improved relative to the first generated image, and in particular, the consistency between the semantics of the second generated image and the semantics of the target image can be improved; on the other hand, when the second generated image is applied to the training of the data processing model, the ability of the data processing model to accurately capture the semantics it carries can be improved.
[0062] Based on the above embodiments, in the image generation method provided in the embodiments of the present application, determining the semantic difference between the first semantics and the second semantics can be achieved by:
[0063] Context data included in the first semantics is identified to obtain a semantic recognition result; if the semantic recognition result indicates that the first semantics satisfies the target rule, a semantic difference between the first semantics and the second semantics is determined.
[0064] Accordingly, if the semantic recognition result indicates that the first semantics does not satisfy the target rule, the operation of determining the semantic difference between the first semantics and the second semantics may not be performed.
[0065] In one embodiment, the target rules may be predetermined or adjusted, and the target rules may be related to the scene represented by the target image. For example, if the target image represents a traffic scene, the target rules may include spatial position relationships that should be satisfied between vehicles, between vehicles and pedestrians, and between vehicles and pedestrians and traffic lights and lanes.
[0066] In one embodiment, the target rules may also include scientific principles and / or scientific logic for describing the objective world. For example, the scientific principles and / or scientific logic may include the relative positional relationship that objects in the first generated image should have, and may also include the state or condition that objects in the first generated image should have in their environment. For example, scientific logic may include that people walking in space should wear spacesuits, and airplanes should fly in the air rather than in the water.
[0067] In one embodiment, the context data may include the relationship between objects included in the first semantics, and may also include the state and / or action of the objects included in the first semantics.
[0068] In one embodiment, the semantic recognition result can be obtained by:
[0069] The states and actions of the objects contained in the first semantics are identified and analyzed to obtain first data, the relative positional relationships between the objects contained in the first semantics are identified to obtain second data, and then the first data and the second data are combined to obtain a semantic recognition result; wherein the state may include the position of the object in the first generated image, the color and brightness of the object, and the size of the pixel area occupied by the object.
[0070] From the above, it can be seen that in the image generation method provided in the embodiment of the present application, the context data contained in the first semantics is identified to obtain a semantic recognition result, thereby realizing tracking and identifying the context data in the first semantics; and, if the semantic recognition result represents that the first semantics satisfies the target rule, the semantic difference between the first semantics and the second semantics is determined, which not only realizes the precise control of the operation of determining the semantic difference, but also realizes the precise judgment of whether the first semantics satisfies the target rule, and also improves the effectiveness of the semantic difference.
[0071] Based on the above embodiments, in the image generation method provided in the embodiments of the present application, determining the semantic difference between the first semantics and the second semantics can also be achieved by the following methods:
[0072] The first semantics and the second semantics are aligned to obtain aligned first semantics and second semantics; the aligned first semantics and the second semantics are processed by the discriminator in the GAN to obtain semantic differences.
[0073] Exemplarily, the discriminator may include an input layer, a convolutional layer, and a fully connected layer; wherein the input layer is used to input the first generated image and the target image, the convolutional layer can perform feature extraction operations on the first generated image and the target image, the output data of the convolutional layer is flattened and input into the fully connected layer, and finally a scalar is output to represent the probability that the first generated image is the target image.
[0074] In one embodiment, the alignment of the first semantics and the second semantics may be achieved in the following manner:
[0075] The first semantics are analyzed to obtain a first object set and a first state set, and the second semantics are analyzed to obtain a second object set and a second state set. Then, according to the first coordinate set corresponding to the first object set and the second coordinate set corresponding to the second object set, the first object set and the second object set are coordinate-aligned to obtain an aligned object set. Then, the state data in the first state set and the second state set are added to the aligned object set to obtain the aligned first semantics and the second semantics. The first object set may include a set of objects included in the first semantics, the first state set may include a set of states of the objects in the first object set, the second object set may include a set of objects included in the second semantics, and the second state set may include a set of states of the objects in the second object set.
[0076] In one embodiment, the discriminator in the GAN judges the consistency between the objects respectively contained in the aligned first semantics and the second semantics to determine whether the objects in the aligned first semantics and the second semantics are the same, and obtains an object discrimination result. At the same time, the states of the corresponding objects in the aligned first semantics and the second semantics are judged to determine whether the objects in the aligned first semantics and the second semantics have the same state, and obtains a state discrimination result. The combination of the object discrimination result and the state discrimination result is determined as a semantic difference.
[0077] From the above, it can be seen that in the image generation method provided in the embodiment of the present application, the first semantics and the second semantics are aligned to obtain the aligned first semantics and the second semantics, thereby improving the corresponding correlation between the data in the aligned first semantics and the second semantics; and, the aligned first semantics and the second semantics are processed by the discriminator in the GAN to obtain semantic differences, which not only can achieve efficient processing of the aligned first semantics and the second semantics, but also can improve the accuracy and pertinence of the semantic differences.
[0078] Based on the above embodiments, the image generation method provided in the embodiments of the present application may further perform the following operations:
[0079] SA1. Obtain a first image feature.
[0080] Among them, the first image feature includes the first texture feature, the first illumination feature and the first spatial feature of the first generated image; the first spatial feature includes at least the spatial positions of N objects contained in the first generated image; N is an integer greater than or equal to 1.
[0081] In one embodiment, the first texture feature may include roughness, contrast, directionality, regularity, etc. of the first generated image, and may also include details and surface features of the first generated image.
[0082] In one embodiment, the first illumination feature may include illumination intensity, illumination intensity variation, brightness distribution, light quality, and light source position in the first generated image.
[0083] In one embodiment, the first spatial feature may include the position, size, perspective, proportion, and shape of the objects included in the first generated image, and may also include the relative position and arrangement of different objects included in the first generated image.
[0084] In one embodiment, the above-mentioned various features in the first image features can be obtained by processing the first generated image through a feature extraction model with texture feature extraction, illumination feature extraction and spatial feature extraction; illustratively, the feature extraction model can include CNN.
[0085] SA2. Obtain a second image feature.
[0086] The second image feature includes a second texture feature, a second illumination feature, and a second spatial feature of the target image; the second spatial feature includes the spatial positions of M objects contained in the target image; and M is an integer greater than or equal to 1.
[0087] In one embodiment, the second texture feature may include roughness, contrast, directionality, regularity, etc. of the target image, and may also include details and surface features of the target image.
[0088] In one embodiment, the second illumination feature may include illumination intensity, illumination intensity variation, brightness distribution, light quality, and light source position in the target image.
[0089] In one embodiment, the second spatial feature may include the position, size, perspective, proportion, and shape of the objects contained in the target image, and may also include the relative position and arrangement of different objects contained in the target image.
[0090] In one embodiment, the above-mentioned various features in the second image feature can be obtained by processing the target image through a feature extraction model with texture feature extraction, illumination feature extraction, and spatial feature extraction.
[0091] It should be noted that the feature extraction model used to process the target image to obtain the second image features can be the same as the feature extraction model used to process the first generated image to obtain the first image features. That is to say, two identical feature extraction models can be used to perform feature extraction on the first generated image and the target image, respectively, so as to obtain the first image features and the second image features respectively.
[0092] Exemplarily, the steps of obtaining the first image feature and the steps of obtaining the second image feature may be adjusted sequentially or performed in parallel, and this embodiment of the present application does not limit this.
[0093] It should be noted that the values of M and N may be the same or different. When the degree of consistency between the first generated image and the target image is greater than the first threshold, the values of M and N may be the same. However, if the degree of consistency between the first generated image and the target image is less than or equal to the first threshold, it may indicate that there is a large difference between the first generated image and the target image. In this case, the values of M and N may be different.
[0094] SA3. Process the first image feature and the second image feature through the discriminator in the GAN to determine the difference in image features.
[0095] The image feature difference includes a texture difference between the first texture feature and the second texture feature, a spatial difference between the first spatial feature and the second spatial feature, and an illumination difference between the first illumination feature and the second illumination feature.
[0096] In one embodiment, the discriminator may include a context perception module, which can judge the consistency between the first texture feature and the second texture feature, the matching degree between the first illumination feature and the second illumination feature, and the matching degree between the first spatial feature and the second spatial feature, thereby achieving a comprehensive evaluation of the image quality of the first generated image to determine whether there is texture distortion, illumination difference and spatial feature difference between the first generated image and the target image. If there is at least one of texture distortion, illumination difference and spatial feature difference, at least one of texture distortion, illumination difference and spatial feature difference can be determined as image feature difference.
[0097] For example, if there is texture distortion, it may mean that the first generated data does not match the real image corresponding to the target image in terms of pixel details and texture; in this case, if the first generated image is used as sample data to train the data processing model, it will cause the data processing model to incorrectly capture the texture details.
[0098] For example, if there is a lighting difference, it can be characterized that the naturalness, smoothness or realism of the lighting in the first generated image is different from the lighting state in the real world. At this time, if the first generated image is used as sample data to train the data processing model, the data processing model will become less sensitive to lighting and lighting changes.
[0099] In one embodiment, the discriminator can determine whether the first generated image is consistent with the target image from a macro perspective while determining the difference in image features, thereby achieving authenticity judgment of the first generated image.
[0100] From the above, it can be seen that the image generation method provided in the embodiment of the present application obtains the first image feature and the second image feature, and then processes the first image feature and the second image feature through the discriminator in the GAN to determine the image feature difference, thereby realizing a rapid judgment of the degree of difference between the first image feature and the second image feature; and, since the first image feature includes the first texture feature, the first illumination feature and the first spatial feature of the first generated image, and the second image feature includes the second texture feature, the second illumination feature and the second spatial feature of the target image, the image feature difference includes the texture difference between the first texture feature and the second texture feature, the spatial difference between the first spatial feature and the second spatial feature, and the illumination difference between the first illumination feature and the second illumination feature. In this way, a comprehensive and refined difference quantification of the texture feature, illumination feature and spatial feature dimensions of the first generated image and the target image is realized, thereby improving the comprehensiveness and accuracy of the data in the image feature difference.
[0101] Based on the aforementioned embodiment, in the image generation method provided in the embodiment of the present application, adjusting the parameters of the image generation model based on at least the semantic difference to obtain the parameter-adjusted image generation model includes:
[0102] Based on the image feature differences and semantic differences, the parameters of the image generation model are adjusted to obtain a generator image generation model with adjusted parameters.
[0103] In one embodiment, the parameter-adjusted image generation model can be obtained by:
[0104] First, a first parameter of an image generation model is adjusted based on semantic differences to obtain an image generation model after the first parameter is adjusted. Then, based on image feature differences, a second parameter included in the image generation model after the first parameter is adjusted is adjusted to obtain an image generation model after the parameter is adjusted. The second parameter may include a set of parameters related to the processing process of texture features, illumination features, and spatial features in the image generation model.
[0105] Specifically, based on the illumination difference corresponding to the first illumination feature and the second illumination feature in the image feature difference, the illumination processing parameters related to the illumination feature processing process in the image generation model can be adjusted, so that the illumination features of the image generated by the above model can be consistent with the illumination features of the target image; based on the texture difference corresponding to the first texture feature and the second texture feature in the image feature difference, the texture processing parameters associated with the texture processing process of the image generation model can be adjusted using texture mapping technology, so that the texture of the generated image generated by the above model is more delicate and consistent with the texture features of the target image; exemplarily, based on the spatial difference corresponding to the first spatial feature and the second spatial feature in the image feature difference, the spatial parameters corresponding to the relative position processing process of the object in the image generation model can be adjusted, so that the objects in the generated image generated by the above model and the objects in the target image are spatially consistent in the relative position relationship dimension.
[0106] Exemplarily, semantic parameters related to semantics in the image generation model can be adjusted based on semantic differences, so that the objects and their category identifiers contained in the image generated by the above model can be consistent with the objects and their category identifiers contained in the target image, respectively.
[0107] In related technologies, if there are texture differences, illumination differences, spatial differences, and semantic differences between the first generated image used to train the data processing model and the target image, it indicates that the image quality of the first generated image is less than or equal to the second threshold, and the rationality of the first generated image is less than or equal to the third threshold. In this case, if the data processing model is trained based on the sample data composed of the first generated image, it will affect the training process of the data processing model, making the trained data processing model have poor feature capture capabilities in terms of texture features, illumination features, spatial features, and semantics, thereby ultimately reducing the data processing performance of the data processing model.
[0108] As can be seen from the above, in the image generation method provided in the embodiment of the present application, the parameters of the image generation model are adjusted based on the image feature differences and semantic differences to obtain the image generation model after the parameters are adjusted. In this way, through the above operation, a comprehensive and diversified targeted adjustment of the parameters in the generator image generation model is achieved, thereby improving the efficiency of parameter adjustment of the image generation model and improving the image generation effect of the image generation model after the parameters are adjusted.
[0109] Based on the aforementioned embodiments, in the image generation method provided in the embodiments of the present application, adjusting the parameters of the image generation model based on at least the semantic difference to obtain the parameter-adjusted image generation model can also be achieved in the following manner:
[0110] Based on the semantic differences and image feature differences, a parameter adjustment strategy is determined; based on the parameter adjustment strategy, the parameters of the image generation model are adjusted to obtain the image generation model after parameter adjustment.
[0111] In one embodiment, the parameter adjustment strategy may include steps, methods, conditions, etc. for adjusting the parameters of the image generation model.
[0112] In one embodiment, the parameter adjustment strategy may be determined by any of the following methods:
[0113] Based on the degree of difference between the semantic difference and the image feature difference, the parameter adjustment priority is determined, and the parameter adjustment priority is used as the parameter adjustment strategy; for example, if the degree of difference corresponding to the semantic difference is smaller than the degree of difference corresponding to the image feature difference, the parameter adjustment priority may include: the priority of adjusting the first parameter is smaller than the priority of adjusting the second parameter.
[0114] If the correlation between the semantic difference and the image feature difference is greater than the third threshold, the parameter adjustment strategy may be determined as: jointly adjusting the first parameter and the second parameter.
[0115] In one embodiment, after the parameter adjustment strategy is determined, the parameters included in the image generation model may be adjusted sequentially or jointly based on the parameter adjustment strategy, thereby obtaining the image generation model after parameter adjustment.
[0116] As can be seen from the above, the image generation method provided in the embodiments of the present application determines a parameter adjustment strategy based on semantic differences and image feature differences, and then adjusts the parameters of the image generation model based on the parameter adjustment strategy to obtain the image generation model after parameter adjustment. In this way, through the above steps, targeted adjustment of the parameters of the image generation model is achieved, thereby improving the efficiency and pertinence of the parameter adjustment of the image generation model.
[0117] Based on the foregoing embodiment, in the image generation method provided in the embodiment of the present application, the image generation model includes a generator in a GAN.
[0118] Exemplarily, the generator is typically a deep neural network that post-processes input data to produce output data, including high-dimensional vector images, text, or speech data. Exemplarily, the input data can be low-dimensional vectors, such as random noise. The generator's specific structure can include an input layer and hidden layers. The input layer is used to receive input data, and the hidden layers can include fully connected layers and convolutional layers, etc., for learning the mapping from input data to output data.
[0119] Accordingly, determining the semantic difference between the first semantics and the second semantics can be achieved in the following manner:
[0120] The first semantics and the second semantics are processed by the discriminator in GAN to obtain the semantic difference.
[0121] Illustratively, the method provided in the aforementioned embodiment can be used to process the first semantics and the second semantics through a discriminator to obtain a semantic difference.
[0122] Accordingly, the above method may further perform the following steps:
[0123] The discriminator determines the image feature difference between the first generated image and the target image, and jointly adjusts the parameters of the generator and the parameters of the generator based on at least the image feature difference and the semantic difference to obtain a GAN with adjusted parameters.
[0124] Illustratively, the method provided in the aforementioned embodiment can be used to process the first image feature corresponding to the first generated image and the second image feature corresponding to the target image through a discriminator to determine the image feature difference.
[0125] In one embodiment, the discriminator may first determine whether the first generated image is consistent with the target image. If the discriminator determines that the first generated image is inconsistent with the target image, the image feature difference between the first generated image and the target image may be determined by the method provided in the aforementioned embodiment.
[0126] In one embodiment, jointly adjusting the parameters of the discriminator and the generator can be achieved by:
[0127] The generator and discriminator are jointly trained through adversarial training, so that the parameters of the discriminator and the generator can be jointly adjusted.
[0128] For example, during adversarial training, the generator and the discriminator compete with each other, wherein the generator attempts to generate first image data that can deceive the discriminator, while the discriminator strives to improve its judgment ability to distinguish whether the target image is consistent with the first image data.
[0129] Generally speaking, before adversarial training, the parameters of the generator and discriminator need to be reasonably initialized first to reduce the probability of their parameters falling into local optimal solutions during training; secondly, during adversarial training, the discriminator can be trained first so that it can better distinguish the difference between the target image and the first generated image, and then the generator can be trained so that the first image data it generates can deceive the discriminator as much as possible.
[0130] In the above training process, the goal of the generator is to maximize the probability of the discriminator misjudging the first image data, while the goal of the discriminator is to maximize the probability of correctly classifying the target image and the first generated image.
[0131] For example, during the adversarial training process, the learning rate can be dynamically adjusted to reduce the instability during the adversarial training process.
[0132] From the above, it can be seen that in the image generation method provided in the embodiment of the present application, the image generation model includes a generator in a GAN, and the first semantics and the second semantics are processed by the discriminator in the GAN to obtain a semantic difference, and the image feature difference between the first generated image and the target image is determined by the discriminator. In this way, the discriminator in the GAN realizes a comprehensive determination of the semantic difference and the image feature difference; at the same time, at least based on the image feature difference and the semantic difference, the parameters of the discriminator and the parameters of the generator are jointly adjusted to obtain the parameter-adjusted GAN. In this way, through the above method, not only the joint training of the generator and the discriminator is realized, but also the efficiency of the joint training can be improved, and the quality of the second generated image generated by the generator in the parameter-adjusted GAN can also be improved.
[0133] Figure 2 A schematic diagram of the network structure for correcting image data provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the network structure for correcting image data may include a first network 201, a second network 202, a generator 203, a discriminator 204 and a post-processing unit 205.
[0134] Exemplarily, the first network 201 can be a neural network model with a semantic extraction function in the aforementioned embodiment, and a feature extraction model with texture feature extraction, illumination feature extraction and spatial feature extraction, which is used to perform feature extraction on the first generated image 206, thereby obtaining a first data set 207, and inputting the first data set 207 into the discriminator 204; wherein, the illumination features, texture features, spatial relationships and semantic information in the first data set 207 can respectively correspond to the first illumination features, first texture features, first spatial features and first semantics in the aforementioned embodiment.
[0135] Exemplarily, the network structure and network parameters of the second network 202 can be exactly the same as those of the first network 201, which is used to extract features from the real image corresponding to the target image 208, obtain a second data set 209, and input the second data set into the discriminator 204; wherein the illumination features, texture features, spatial relationships and semantic information in the second data set can respectively correspond to the second illumination features, second texture features, second spatial features and second semantics in the aforementioned embodiments.
[0136] Exemplarily, the discriminator 204 can judge the first data set and the second data set to obtain semantic differences and image feature differences, and based on the semantic differences and image feature differences, perform adversarial training on the discriminator 204 and the generator 203 to obtain a trained GAN. At this time, the previously generated image can be corrected by the trained GAN to obtain a corrected generated image 210; exemplarily, the corrected generated image can correspond to the second generated image in the aforementioned embodiment.
[0137] Exemplarily, after obtaining the corrected generated image 210, the corrected generated image can be smoothed by a post-processing unit to remove unnatural edges or blemishes in the generated image. The post-processing unit can also perform a noise removal operation on the generated image to improve the clarity of the generated data, thereby further improving the image quality of the corrected generated image.
[0138] In related technologies, the generated image data has poor quality in terms of lighting, texture, and spatial relationships, as well as irrationality due to semantic errors. This makes it easy for the data processing model to be unable to capture the effective features carried in the real image when training the data processing model using the above image data, resulting in poor training effect of the data processing model and weak model generalization ability.
[0139] In the embodiment of the present application, through the above-mentioned network structure, the image context data of the first generated image and the target image, including texture features, lighting features, spatial relationships and semantic information, combined with an automated and intelligent image correction scheme, can improve the image quality of the corrected generated image, so that the corrected generated image can contain the pixel features of the real image, and improve the consistency between the texture, lighting, spatial relationships and semantic dimensions of the corrected generated image and the target data processing task, thereby improving the quality of the corrected generated image; and, through the above-mentioned network structure, the integration of training and application of the generator and the discriminator is realized; at the same time, with the help of the mutual combination between the generator and the discriminator, the efficiency of training the generator and the discriminator can be improved, and the effect of training the generator and the discriminator can be improved; on the other hand, by training the data processing model with sample data composed of the corrected generated image, the generalization ability of the data processing model corresponding to the target task can be improved.
[0140] The embodiment of the present application also provides an image generating device, Figure 3 A schematic diagram of the structure of the image generating device provided in the embodiment of the present application is shown as follows: Figure 3 As shown, the image generating device 3 includes:
[0141] An acquisition module 301 is configured to acquire a first semantic meaning of a first generated image output by an image generation model; and acquire a second semantic meaning of a target image;
[0142] A processing module 302 is configured to determine a semantic difference between the first semantics and the second semantics; and adjust parameters of the image generation model based at least on the semantic difference to obtain an image generation model with adjusted parameters.
[0143] The generating module 303 is configured to generate a second generated image using the parameter-adjusted generating model; wherein a first degree of difference between the first generated image and the target image is greater than a second degree of difference between the second generated image and the target image.
[0144] In some embodiments, the processing module 302 is configured to identify the context data contained in the first semantics and determine a semantic recognition result; if the semantic recognition result indicates that the first semantics satisfies the target rule, determine the semantic difference between the first semantics and the second semantics.
[0145] In some embodiments, the processing module 302 is used to align the first semantics and the second semantics to obtain aligned first semantics and second semantics; and process the aligned first semantics and second semantics through the discriminator in the GAN to obtain semantic differences.
[0146] In some embodiments, the acquisition module 301 is configured to acquire a first image feature; wherein the first image feature includes a first texture feature, a first illumination feature, and a first spatial feature of the first generated image; the first spatial feature includes the spatial positions of N objects contained in the first generated image; N is an integer greater than or equal to 1;
[0147] The acquisition module 301 is further configured to acquire a second image feature; wherein the second image feature includes a second texture feature, a second illumination feature, and a second spatial feature of the target image; the second spatial feature includes the spatial positions of M objects contained in the target image; M is an integer greater than or equal to 1;
[0148] The processing module 302 is used to process the first image feature and the second image feature through the discriminator in the GAN to determine the image feature difference; wherein the image feature difference includes the texture difference between the first texture feature and the second texture feature, the spatial difference between the first spatial feature and the second spatial feature, and the illumination difference between the first illumination feature and the second illumination feature.
[0149] In some embodiments, the processing module 302 is configured to adjust parameters of the image generation model based on image feature differences and semantic differences to obtain an image generation model with adjusted parameters.
[0150] In some embodiments, the processing module 302 is configured to determine a parameter adjustment strategy based on semantic differences and image feature differences; and adjust parameters of the image generation model based on the parameter adjustment strategy to obtain an image generation model after parameter adjustment.
[0151] In some embodiments, the image generation model includes a generator in a GAN;
[0152] The processing module 302 is configured to process the first semantics and the second semantics through a discriminator in the GAN to obtain a semantic difference; determine the image feature difference between the first generated image and the target image through the discriminator; and jointly adjust the parameters of the discriminator and the parameters of the generator based on at least the image feature difference and the semantic difference to obtain a GAN with adjusted parameters.
[0153] The embodiment of the present application also provides an electronic device, Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application, such as Figure 4 As shown, the electronic device 4 includes a processor 401 and a memory 402; a computer program 403 is stored in the memory; when the computer program is executed by the processor, it can implement any of the above-mentioned image generation methods.
[0154] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored; when the computer program is executed by a processor of an electronic device, it can implement any of the above-described image generation methods.
[0155] An embodiment of the present application also provides a computer program product, which includes a computer program; when the computer program is executed by a processor of an electronic device, it can implement any of the above-mentioned image generation methods.
[0156] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0157] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0158] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0159] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0160] It should be noted that the above-mentioned computer-readable storage medium can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface storage, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various electronic devices that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0161] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0162] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0163] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus necessary general hardware nodes, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0164] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0165] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0167] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An image generation method, characterized in that: The method comprises: Obtaining a first semantic meaning of a first generated image output by the image generation model; Acquire the second semantics of the target image; determining a semantic difference between the first semantics and the second semantics; Adjusting parameters of the image generation model based at least on the semantic difference to obtain an image generation model with adjusted parameters; A second generated image is generated by the image generation model after parameter adjustment; wherein a first degree of difference between the first generated image and the target image is greater than a second degree of difference between the second generated image and the target image.
2. The method according to claim 1, characterized in that The determining of the semantic difference between the first semantics and the second semantics includes: Identifying context data contained in the first semantics to obtain a semantics recognition result; If the semantic recognition result indicates that the first semantics satisfies the target rule, the semantic difference between the first semantics and the second semantics is determined.
3. The method according to claim 1, characterized in that The determining of the semantic difference between the first semantics and the second semantics includes: performing alignment processing on the first semantics and the second semantics to obtain aligned first semantics and second semantics; The aligned first semantics and the second semantics are processed by a discriminator in a generative adversarial network (GAN) to obtain the semantic difference.
4. The method according to claim 1, wherein The method further comprises: Obtaining a first image feature; wherein the first image feature includes a first texture feature, a first illumination feature, and a first spatial feature of the first generated image; the first spatial feature includes the spatial positions of N objects contained in the first generated image; N is an integer greater than or equal to 1; Acquire a second image feature; wherein the second image feature includes a second texture feature, a second illumination feature, and a second spatial feature of the target image; the second spatial feature includes the spatial positions of M objects contained in the target image; M is an integer greater than or equal to 1; The first image feature and the second image feature are processed by a discriminator in the GAN to determine image feature differences; wherein the image feature differences include texture differences between the first texture feature and the second texture feature, spatial differences between the first spatial feature and the second spatial feature, and illumination differences between the first illumination feature and the second illumination feature.
5. The method according to claim 4, characterized in that The adjusting the parameters of the image generation model at least based on the semantic difference to obtain the image generation model after parameter adjustment includes: Based on the image feature difference and the semantic difference, the parameters of the image generation model are adjusted to obtain an image generation model with adjusted parameters.
6. The method according to claim 4, characterized in that The adjusting the parameters of the image generation model at least based on the semantic difference to obtain the image generation model after parameter adjustment includes: determining a parameter adjustment strategy based on the semantic difference and the image feature difference; The parameters of the image generation model are adjusted based on the parameter adjustment strategy to obtain an image generation model after parameter adjustment.
7. The method according to claim 1, characterized in that The image generation model includes a generator in a GAN; and determining the semantic difference between the first semantics and the second semantics includes: Processing the first semantics and the second semantics by a discriminator in the GAN to obtain the semantic difference; The method further comprises: determining, by the discriminator, a difference in image features between the first generated image and the target image; Based at least on the image feature difference and the semantic difference, parameters of the discriminator and parameters of the generator are jointly adjusted to obtain a GAN with adjusted parameters.
8. An image generating device, characterized in that: The image generating device comprises: An acquisition module is configured to acquire a first semantic meaning of a first generated image output by an image generation model; and acquire a second semantic meaning of a target image; a processing module, configured to determine a semantic difference between the first semantics and the second semantics; and adjust parameters of the image generation model based at least on the semantic difference to obtain an image generation model with adjusted parameters; A generation module is used to generate a second generated image using the parameter-adjusted generation model; wherein a first degree of difference between the first generated image and the target image is greater than a second degree of difference between the second generated image and the target image.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory; a computer program is stored in the memory; when the computer program is executed by the processor, the image generation method according to any one of claims 1 to 7 can be implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program; when the computer program is executed by a processor of an electronic device, the image generation method according to any one of claims 1 to 7 can be implemented.