Image generation method based on ontology knowledge base and attention generative adversarial network

By building an ontology knowledge base and using attention generation adversarial network, the problem of poor image generation quality under complex conditions in the prior art is solved, and high-quality and efficient image generation is achieved.

CN114386428BActive Publication Date: 2025-05-06SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111537463.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-15
Publication Date
2025-05-06
Estimated Expiration
2041-12-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively generate images that meet the description using complex conditions, especially in image generation tasks, where traditional models generate poor images when processing complex conditions and are unfavorable to industrial use.

Method used

By constructing an ontological knowledge base based on image datasets and generating an adversarial network in combination with attention, structured representation of the conditions of the generation task, thereby generating images that meet the description.

Benefits of technology

It realizes effective structured representation and image generation of complex conditions, improves the quality and efficiency of image generation, and is suitable for industrial use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114386428B_ABST
    Figure CN114386428B_ABST
Patent Text Reader

Abstract

The present invention provides an image generation method based on an ontology knowledge base and an attention generation adversarial network, comprising: constructing an ontology knowledge base based on images in an image data set so that the constructed ontology knowledge base has a corresponding relationship with the image; searching the ontology knowledge base according to the conditions of the generation task, and vectorizing the retrieved information to obtain a semantic word vector; using a generation adversarial network with an attention mechanism to generate an image according to the semantic word vector. By using the network provided by the present invention, the generation task of high-quality images under complex conditions can be completed by constructing different ontology knowledge bases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and knowledge system, and specifically relates to an image generation method based on an ontology knowledge base and an attention generative adversarial network. Background Art

[0002] Image generation technology can process white noise to generate more realistic images by using deep neural networks. It can also generate specific images that meet the conditions based on different conditional information (such as image category information, image style information, image description text information, image scene triples, etc.). It has been widely used in image, video and animation production, data enhancement, image restoration and other fields.

[0003] In the task of conditional image generation, the condition can be simple, such as specifying the category of the object in the generated image (generating a cat image), or complex, such as an object with special features (such as a dancing person). In traditional conditional generation models such as DCGAN [Alec Radford & Luke Metz, Soumith Chintala, Unsupervised representation learning with deep convolutional generative adversarial networks], the entire condition is directly used as the input of the model, which often fails to model complex conditions, resulting in the generated image not meeting the condition or the generated image quality is poor.

[0004] In the patent document "Image Generation Method Based on Conditional Generative Adversarial Convolutional Neural Network" with application number CN202010031013.0, it first collects images according to categories, then uses autoencoders to extract image features, and obtains the average features of a certain type of samples through clustering, and finally uses the conditions and sample average features as input to generate the network model. Although this method effectively alleviates the mode collapse problem compared to directly inputting conditions, the overall process is more cumbersome, and the computational complexity of feature extraction for all images is large, which is not conducive to industrial use; in addition, the conditions that can be generated are relatively simple.

[0005] Some models can generate images that match the description from some more complex conditions (such as a descriptive text sentence, which can be a structured representation, such as a scene with multiple objects and a person riding a horse on the grass), such as AttnGAN [Tao Xu, Pengchuan Zhang, Qiuyuan Huang, Han Zhang, Zhe Gan, Xiaolei Huang, Xiaodong He, "AttnGAN: Fine-Grained Text to Image Generation with Attentional Generative Adversarial Networks"], but the sentence still needs to be segmented. Such a processing process lacks structure and is not suitable for generation under other complex conditions. Summary of the invention

[0006] The purpose of the present invention is to propose an image generation method based on an ontology knowledge base and an attention-generated adversarial network, so as to be suitable for representing descriptions in a structured manner, thereby generating images that conform to the descriptions.

[0007] In order to achieve the above object, the present invention provides an image generation method based on an ontology knowledge base and an attention-generated adversarial network, comprising:

[0008] S1: construct an ontology knowledge base based on the images in the image dataset, so that the constructed ontology knowledge base has a corresponding relationship with the images;

[0009] S2: According to the conditions of the generated task, the ontology knowledge base is searched and the retrieved information is vectorized to obtain semantic word vectors;

[0010] S3: Use a generative adversarial network with an attention mechanism to generate an image based on the semantic word vector of step S2.

[0011] The step S1 comprises: constructing an ontology knowledge base that can be used for image generation tasks from a class of images in the image data set; and the constructed ontology knowledge base comprises: ontologies, relationships between ontologies and attributes of ontologies.

[0012] The relationship between the entities includes at least one of a subordinate relationship, an environmental relationship and an interactive relationship; the attributes of the entity include the semantic information of the entity and / or the position of the entity in the image.

[0013] The condition of the generation task is an ontology of a single category or an ontology with special attributes.

[0014] The step S2 comprises:

[0015] S21: Input the conditions for generating the task, search the conditions for generating the task in the ontology knowledge base, retrieve the first ontology and the attributes of the first ontology that meet the input conditions, and obtain the first ontology set O i and the first ontology attribute set A i ;

[0016] S22: Continue to search the ontology knowledge base for at least one second ontology that is related to the first ontology and all possible attributes of the second ontology in the image where the first ontology is located, and obtain a second ontology set O. j And the second ontology attribute set A i ;

[0017] S23: Set the first entity set O i and the second ontology set O j Together they constitute the total ontology set O, and the first ontology attribute set A i and the second ontology attribute set A j Together they constitute the total attribute set A. The total ontology set O and the total attribute set A together constitute the knowledge set K after matching.

[0018] S24: Use the pre-trained natural language processing word vector model to vectorize the information of the knowledge set K into a semantic word vector V.

[0019] The pre-trained natural language processing word vector model includes a word vector model Bert, a word vector model word2vec and LSTM.

[0020] The generative adversarial network with attention mechanism includes at least one of AttnGAN, WGAN with attention mechanism, PGGAN with attention mechanism and DCGAN with attention mechanism.

[0021] The image generation method based on ontology knowledge base and attention generative adversarial network of the present invention inputs a condition description and constructs an ontology knowledge base related to it to structure the description, and uses a generative adversarial network with an attention mechanism to complete the image generation task by constructing different ontology knowledge bases to generate images that meet certain conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a flowchart of the image generation method based on ontology knowledge base and attention generative adversarial network of the present invention.

[0023] Figure 2 It is a flowchart of an image generation method based on an ontology knowledge base and an attention-generated adversarial network according to embodiment 1 of the present invention.

[0024] Figure 3It is a structural schematic diagram of the ontology knowledge base of the image generation method based on the ontology knowledge base and the attention generative adversarial network according to the first embodiment.

[0025] Figure 4 It is a schematic diagram of the specific structure of a multi-layer attention generation network for an image generation method based on an ontology knowledge base and an attention generation adversarial network.

[0026] Figure 5 This is a schematic diagram of the attention mechanism of the image generation method based on the ontology knowledge base and the attention generative adversarial network of Example 1.

[0027] Figure 6 It is a flowchart of an image generation method based on an ontology knowledge base and an attention-generated adversarial network according to the second embodiment of the present invention.

[0028] Figure 7 It is a structural schematic diagram of the ontology knowledge base of the image generation method based on the ontology knowledge base and the attention generative adversarial network according to the second embodiment of the present invention. DETAILED DESCRIPTION

[0029] The present invention is further described below in conjunction with specific examples. It should be understood that the following examples are only used to illustrate the present invention and are not used to limit the scope of the present invention.

[0030] like Figure 1 As shown, the image generation method based on ontology knowledge base and attention generative adversarial network of the present invention is used for various image generation tasks, and includes the following steps:

[0031] Step S1: constructing an ontology knowledge base based on the images in the image data set, so that the constructed ontology knowledge base has a corresponding relationship with the images;

[0032] In the prior art, an ontology knowledge base is generally composed of a large number of ontologies and the relationships between ontologies (see [Lisa Ehrlinger and Wolfram "Towards a Definition of KnowledgeGraphs"]), each ontology has its semantic categories and attributes. However, most existing ontology knowledge bases use sentence descriptions, that is, starting from the perspective of the part of speech and meaning of a word, containing a large and comprehensive information about a word. However, in a certain scene and image, such an ontology knowledge base with a large and comprehensive information on a specific word cannot describe a specific scene. As the conditions become more complex and more information needs to be generated, it may be difficult to effectively describe using sentence descriptions and the sentences become very long, making it more difficult for the model to capture information and thus more difficult to effectively generate images.

[0033] The ontology knowledge base constructed by the present invention is oriented to image data sets, which describes the ontology in the image and the attributes of each ontology. By describing the ontology information in the image and connecting these ontologies with relationships, the information is structured to achieve structured description. Compared with the description using statements, such structured description makes the control granularity of image generation finer and can achieve effective description under complex conditions.

[0034] Therefore, the step S1 includes: constructing an ontology knowledge base that can be used for image generation tasks from a class of images in the image data set; and the composition of the constructed ontology knowledge base is as follows:

[0035] 1. Ontology: Represents an entity of a certain category in the image, such as: <man>、<left_leg>、 <forest> 、 <wood>wait.

[0036] 2. Relationships between ontologies: The ontology knowledge base composed of ontologies is structured, that is, ontologies are connected by relationships. Such relationships may include but are not limited to: subordinate relationships (<left_hand> and <man>), environmental relations ( <sea>and <surfboard>) and interactions (e.g. <man>and <horse>, the action is ride).

[0037] 3. Attributes of the ontology: The attributes of the ontology may include but are not limited to: semantic information of the ontology (such as color, shape, size, action, purpose, etc.), the position of the ontology in the image, etc.

[0038] It should be noted that in the present invention, the relationship between entities is extracted from the image. And the relationship here is for a specific image dataset. In an image dataset, the extracted relationship type should be as consistent as possible. For example, in a dataset specifically for the interaction between people and objects, information that may be an interactive action should be constructed as an interactive relationship instead of an environmental relationship (such as a person on top of a horse).

[0039] The relationship between entities can be extracted by combining manual extraction and software extraction. For example, the target detection model is first used to extract the objects in the image of the image dataset, and the environmental relationship is preliminarily established according to the coordinate position. After that, other types of relationships are manually extracted and labeled.

[0040] Step S2: According to the conditions of the generated task, the ontology knowledge base is searched, and the retrieved information is vectorized to obtain a semantic word vector;

[0041] Therefore, after the ontology knowledge base is constructed, the ontology knowledge base is searched by inputting the conditions for generating tasks, and the retrieved information is vectorized.

[0042] The input generation task conditions are diverse, that is, it can be a single category of ontology (such as <sunflower>), or it can be an entity with special attributes (such as a (dance) <man>).

[0043] The step S2 specifically includes:

[0044] Step S21: Input the conditions for generating the task, search the conditions for generating the task in the ontology knowledge base, retrieve the first ontology and the attributes of the first ontology that meet the input conditions, and obtain the first ontology set O. i and the first ontology attribute set A i ;

[0045] Step S22: Continue to search the ontology knowledge base for at least one second ontology that is related to the first ontology and all possible attributes of the second ontology in the image where the first ontology is located, and obtain a second ontology set O j And the second ontology attribute set A j ;

[0046] It should be noted that the attributes of the same first entity and the second entity in different images may be different. At the same time, the attributes of the second entity must be the attributes of the second entity in the image where the first entity meets the input conditions. The attributes of the second entity in the image that does not have the first entity do not belong to the second entity attribute set A. j elements.

[0047] Step S23: The first ontology set O i and the second ontology set O j Together they constitute the total ontology set O, and the first ontology attribute set A i and the second ontology attribute set A j Together they constitute the total attribute set A. The total ontology set O and the total attribute set A together constitute the knowledge set K after matching.

[0048] The obtained knowledge set K is K = {(o, a) | o∈O i ∪O j , a∈A i ∪A j }, as the output of retrieving the ontology knowledge base.

[0049] Step S24: Use a pre-trained natural language processing word vector model to vectorize the information of the knowledge set K into a semantic word vector V as the vectorization result for the next step of image generation.

[0050] Pre-trained natural language processing word vector models, including the word vector model BERT (see [Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, "BERT: Pre-training of DeepBidirectional Transformers for Language Understanding", arXiv preprint arXiv:1810.04805, 2018]), the word vector model word2vec (see [Tomas Mikolov, Kai Chen, GregCorrado, Jeffrey Dean, "Efficient Estimation of Word Representations in VectorSpace"]), and LSTM, etc.

[0051] Step S3: Generate an image based on the semantic word vector of step S2 using a generative adversarial network (GAN) with an attention mechanism. Thus, the generated image meets the conditions of the generation task.

[0052] Among them, the generative adversarial network (GAN) with attention mechanism can use a variety of existing models and network structures, such as AttnGAN, WGAN with attention mechanism [Martin Arjovsky, Soumith Chintala, Léon Bottou, "Wasserstein GAN"], PGGAN with attention mechanism [Tero Karras, Timo Aila, Samuli Laine, Jaakko Lehtinen, "Progressive Growing of GANs for Improved Quality, Stability, and Variation"], DCGAN (deep convolutional generative adversarial network) with attention mechanism, etc.

[0053] In summary, in order to solve the problem that conditional information cannot be structured in conditional image generation tasks, the present invention constructs different ontology knowledge bases according to images to make the constructed ontology knowledge bases have a corresponding relationship with the images, and then structuredly utilizes the conditional information to make information processing simpler and more efficient; and by using the attention generative adversarial network, it can be widely applied to the generation of high-quality images under simple and complex conditions.

[0054] Embodiment 1 Image generation method based on ontology knowledge base and attention generation adversarial network

[0055] This embodiment 1 describes an image generation method based on an ontology knowledge base and an attention-generated adversarial network for generating images of specific human actions. The flowchart is as follows: Figure 2 As shown, the specific steps are as follows:

[0056] First, according to step S1, images are collected from a human action dataset, and a human action ontology knowledge base is constructed based on the images. The human action ontology knowledge base can be obtained by collecting information (including the position and action category of a person) from a labeled human action dataset, and then manually extracting other ontologies (such as human limbs) and ontology attributes (actions of various limbs of a person) in the image.

[0057] like Figure 3 As shown in the figure, the composition of the constructed human action ontology knowledge base is as follows:

[0058] 1. Ontology: Human (i.e. <man>) and human body parts (specifically including:<right_foot> ,<right_leg> ,<left_leg> ,<left_foot> , <hip>, <right_hand>, <right_arm>, <left_arm>, <left_hand> (a total of 10 body parts)

[0059] 2. Relationship of the body: There are relationship edges indicating a subordinate relationship between a person and their body parts.

[0060] 3. Attributes of the body: For a person, they have the attribute of actions, such as (dance), (bend), etc. For body parts, they also have the attribute of actions. For example, <left_hand> has the attribute of (hold). <hip>has the attribute (sit_on),<right_foot> There is an attribute of (stand_on). In this embodiment, since the generation task is aimed at the generation of human actions, the attributes of the ontology are only action attributes.

[0061] According to step S2, when a specific human body action (dance) is input, the established ontology knowledge base is searched for the (dance) attribute. <man>Main body (immediately main body <man>:(dance)) as the first entity, and these with (dance) attributes <man>The ontology (i.e., the first ontology) has a representative limb relationship of subordinate relationship as the second ontology. The first ontology and the second ontology together constitute the total ontology set O, and the attributes of the two together constitute the total attribute set A. The total ontology set O and the total attribute set A are matched to form the knowledge set K. The schematic diagram of the ontology library construction and retrieval process is shown in Figure 3 shown.

[0062] Figure 3 In the first entity <man>:(dance) has a subordinate second ontology (i.e., the limb part ontology) as follows:

[0063] <right_foot> :(stand)

[0064] <right_leg> :(no_action)

[0065] <left_leg> :(no_action)

[0066] <left_foot> :(stand)

[0067] <hip>:(no_action)

[0068] :(no_action)

[0069] <right_hand> :(wave)

[0070] <right_arm> :(swing)

[0071] <left_arm> :(swing)

[0072] <left_hand> :(wave)

[0073] It should be noted that the attributes of each second entity (i.e., body part entity) may not be unique and may be different for each image. In this embodiment, the entities (people and body parts) are the same, but the attributes of the same entity in different images corresponding to it may be different.

[0074] The corresponding relationship between the attributes of the (dance) attribute and the corresponding limb part ontology is obtained by extracting information from the ontology knowledge base constructed in step S1. The corresponding relationship between the attributes and the ontology in the final knowledge set K must be a combination relationship that can be found in at least one image corresponding to the ontology knowledge base in step S1.

[0075] Step S2 then uses the pre-trained natural language processing word vector model Bert to convert the ontology and the attributes of the ontology (i.e., action categories) into semantic word vectors V for the retrieved knowledge set K. j (i.e. the limb part itself) and its attributes a j The feature vectors transformed by (i.e. action category) are vectors v(o j ) and the vector v(a j ), and then concatenated to form a semantic word vector V.

[0076] In this embodiment, the first entity and the second entity can be regarded as different levels of information. In the process of image generation by the network, the attention mechanism is used to process the second level of information, and each area of ​​the image uses more or less information from different second entities (different limb parts and attributes). The first entity describes the overall information of the entire image, so it does not need to be sent to G1, G2, and G3. However, in other embodiments, according to other references, the information of the first entity can indeed be input into G1 as guidance information and the information of the second entity.

[0077] Among them, the word vector model Bert is a pre-trained model. As mentioned above, this pre-trained model is an open source tool that can be used directly to convert words into feature vectors.

[0078] Finally, according to step S3, a generative adversarial network (GAN) with an attention mechanism is used to generate images.

[0079] In this embodiment, the generative adversarial network (GAN) with an attention mechanism is a multi-layer DCGAN (deep convolutional generative adversarial network) with an attention mechanism. The deep convolutional generative adversarial network consists of a multi-layer generative network and multiple discriminative networks D, wherein the generative network has an attention mechanism, thereby forming an attention generative network G.

[0080] In this embodiment, the attention generation network G is composed of multiple layers, and the images generated by each layer are from small to large. The discriminant network D is not multi-layered, but multiple. Each discriminant network D discriminates the authenticity of the image generated by each layer of the attention generation network G.

[0081] The following is a detailed description of the multi-layer attention generation network G and the discriminant network D.

[0082] Multi-layer attention generation network G:

[0083] The multi-layer attention generation network G is responsible for generating real images with specific body movements from the feature layout map. Its specific structure is as follows: Figure 4 As shown in Figure 1, the multi-layer attention generation network G is divided into three layers: G1, G2, and G3, which gradually generate high-resolution images. The G1 layer uses a noise z sampled from a normal distribution as the input of the first layer G1 of the attention generation network G, and uses the semantic word vector V (i.e., v(o j ) and v(a j ) of the concatenation vector [v(o j ), v(a j )]) is used as the input of the attention mechanism of the three layers G1, G2, and G3 of the attention generation network G, thereby generating a lower resolution image using several layers of deep neural networks with attention mechanisms. The first layer G1 of the attention generation network G includes four upsampling modules connected in sequence (i.e., the first, second, third, and fourth upsampling modules connected in sequence), each of which is composed of an upsampling layer, a 3×3 convolution layer, and a BatchNorm layer. Figure 5 As shown, in the process of generating an image, the second upsampling module of the first layer G1 of the attention generation network G is set to output a feature map F; its attention mechanism calculates the correlation between the feature map F and the semantic word vector V, and based on this correlation, the semantic word vector V is weighted on the feature map F to obtain a weighted feature map F', which is then spliced ​​onto the feature map F, and then provided to the third upsampling module of the first layer G1 to continue the image generation of the deep neural network.

[0084] The first, second, third, and fourth upsampling modules here are in the same layer of the attention generation network G. The overall order is the first layer G1 (which performs upsampling->upsampling->attention mechanism F and F' concatenation->upsampling->upsampling)->the second layer G2 (which performs upsampling->upsampling->attention mechanism F and F' concatenation->upsampling->upsampling)->the third layer G3 (which performs upsampling->upsampling->attention mechanism F and F' concatenation->upsampling->upsampling). Therefore, the result of the concatenation of F and F' of each layer of the attention generation network G is directly provided to the third upsampling module of the layer of the attention generation network G.

[0085] Among them, when each layer of each attention generation network G (for example, the first layer G1) outputs its feature map F, on the one hand, it directly provides this feature map F to the next G as input for subsequent operations, and on the other hand, it also maps this feature map F to the image space through a 3×3 convolution layer, that is, generates an image to be judged (this 3×3 convolution layer is not in our backbone network, it is only used to generate the image to be judged).

[0086] The second layer G2 of the attention generation network G takes the feature map F output by the last upsampling module (i.e., the fourth upsampling module) of the first layer G1 of the attention generation network G as input, thereby using the attention mechanism to generate a higher resolution image to be identified. The third layer G3 of the attention generation network G is similar to the second layer G2, and takes the feature map F output by the last upsampling module of the second layer G2 as input, thereby using the attention mechanism to generate a higher resolution image to be identified.

[0087] Therefore, the three layers of the attention generation network G generate images to be judged with resolutions from small to large for the three discriminative networks D to judge.

[0088] In other embodiments, the GAN with attention mechanism can also be transformed into other GAN structures, but any GAN must use attention mechanism to generate images according to semantic word vector V. That is to say, in other embodiments, the GAN with attention mechanism can be based on the existing GAN including the generative network and the discriminative network (such as PGGAN, WGAN, etc.), and on this basis, the generative network needs to have attention mechanism to form an attention generative network.

[0089] Discriminant network D:

[0090] After the three layers G1, G2, and G3 of the attention generation network G generate images of three resolutions respectively, the discriminant network takes the real image, the image to be judged generated by the attention generation network, and the information of the semantic word vector obtained in step S2 as input to judge the truth or falsity of each image to be judged: if a real image is used as input, it is judged to be true; if an image generated by the attention generation network is used as input, it is judged to be false.

[0091] There are multiple discriminant networks, which are composed of a first discriminant network D1, a second discriminant network D2 and a third discriminant network D3, which discriminate images of three resolutions generated by G1, G2 and G3 respectively.

[0092] In other embodiments, the discriminant network structure does not include an attention mechanism, which can be implemented by the discriminant network in an existing generative adversarial network (such as PGGAN, WGAN, etc.). It should be noted that the attention generation network G and the discriminant network D are matched with each other, so generally the attention generation network G and the discriminant network D based on the same GAN are used.

[0093] Example 2 Image generation method based on ontology knowledge base and attention generation adversarial network

[0094] Embodiment 2 describes an image generation method based on an ontology knowledge base and an attention-generated adversarial network for generating image generation tasks for specific scenes, and its flowchart is as follows: Figure 6 As shown, the specific steps are as follows:

[0095] First, according to step S1, we first collect images from the scene dataset COCO-Stuff and build a scene ontology knowledge base. Figure 7 As shown in the figure, the composition of the constructed ontology knowledge base is as follows:

[0096] 1. Ontology: Entities in the image scene (e.g. <man> 、 <tree> 、 <sky> 、 <river>wait)

[0097] 2. Ontological relationships: There are environmental relationships between entities in the scene, such as <sky>and <tree>There is an environmental relationship, that is <sky>exist <tree>Above.

[0098] 3. Attributes of the entity: In the scene, each entity has its own set of attributes, such as <sky>It has attributes such as (blue) and (cloudy). <man>There are (is surfing), (wears T-shirt) and so on.

[0099] like Figure 7 As shown, according to step S2, when a specific scene (man is surfing in the sea) is input, the established ontology knowledge base is retrieved for the scene with the attribute (is surfing in the sea). <man>The ontology is used as the first ontology, and then the second ontology representing the scene entity that has an environmental relationship with these first ontologies is retrieved, and the set of the first ontologies and the set of the second ontologies together constitute the total ontology set O, and the set of their attributes together constitute the total attribute set A, and the total ontology set O and the total attribute set A together constitute the knowledge set K.

[0100] Figure 7 Secondary One Body <man>:(is surfing in the sea) has the following scene ontology with environmental relationship:

[0101] <sea>:(blue),(board)

[0102] <sky>:(cloudy)

[0103] <surfboard>:(on the sea),(under foot)

[0104] <wave>:(white)

[0105] <sun>:(red),(in the sky)

[0106] Next, step S2 converts the retrieved information into semantic word vectors using the pre-trained natural language processing word vector model Word2vec, including the ontology and the attributes of the ontology.

[0107] According to step S3, a single-layer DCGAN with an attention mechanism is used to generate images.

[0108] The above is only a preferred embodiment of the present invention, and is not intended to limit the scope of the present invention. The above embodiments of the present invention can also be modified in various ways. All simple, equivalent changes and modifications made according to the claims and the description of the present invention fall within the scope of protection of the claims of the present invention. The contents not described in detail in the present invention are all conventional technical contents.< / sun> < / wave> < / surfboard> < / sky> < / sea> < / man> < / man> < / man> < / sky> < / tree> < / sky> < / tree> < / sky> < / river> < / sky> < / tree> < / man> < / hip> < / man> < / man> < / man> < / man> < / hip> < / hip> < / man> < / man> < / sunflower> < / horse> < / man> < / surfboard> < / sea> < / man> < / wood> < / forest> < / man>

Claims

1. An image generation method based on ontology knowledge base and attention generative adversarial network, characterized in that: include: Step S1: constructing an ontology knowledge base based on the images in the image data set, so that the constructed ontology knowledge base has a corresponding relationship with the images; Step S2: According to the conditions of the generated task, the ontology knowledge base is searched, and the retrieved information is vectorized to obtain a semantic word vector; Step S3: using a generative adversarial network with an attention mechanism to generate an image based on the semantic word vector of step S2; The step S2 comprises: Step S21: Input the conditions for generating the task, search the conditions for generating the task in the ontology knowledge base, retrieve the first ontology and the attributes of the first ontology that meet the input conditions, and obtain the first ontology set and the first ontology attribute set ; Step S22: Continue to search the ontology knowledge base for at least one second ontology that is related to the first ontology and all possible attributes of the second ontology in the image where the first ontology is located, and obtain a second ontology set and the second ontology attribute set ; Step S23: Assemble the first entity and the second ontology set Together they constitute the total ontology set O, and the first ontology attribute set and the second ontology attribute set Together they constitute the total attribute set A. The total ontology set O and the total attribute set A together constitute the knowledge set after matching. ; Step S24: Use the pre-trained natural language processing word vector model to transform the knowledge set The information is vectorized into semantic word vectors .

2. The image generation method based on ontology knowledge base and attention generative adversarial network according to claim 1 is characterized in that: The step S1 comprises: constructing an ontology knowledge base that can be used for image generation tasks from a class of images in the image data set; and the constructed ontology knowledge base comprises: ontologies, relationships between ontologies and attributes of ontologies.

3. The image generation method based on ontology knowledge base and attention generative adversarial network according to claim 2 is characterized in that: The relationship between the entities includes at least one of a subordinate relationship, an environmental relationship and an interactive relationship; the attributes of the entity include the semantic information of the entity and / or the position of the entity in the image.

4. The image generation method based on ontology knowledge base and attention generative adversarial network according to claim 1 is characterized in that: The condition of the generation task is an ontology of a single category or an ontology with special attributes.

5. The image generation method based on ontology knowledge base and attention generative adversarial network according to claim 1, characterized in that: The pre-trained natural language processing word vector model includes a word vector model Bert, a word vector model word2vec and LSTM.

6. The image generation method based on ontology knowledge base and attention generative adversarial network according to claim 1, characterized in that: The generative adversarial network with attention mechanism includes at least one of AttnGAN, WGAN with attention mechanism, PGGAN with attention mechanism and DCGAN with attention mechanism.

Citation Information

Patent Citations

  • Image generation method based on conditional generative adversarial convolutional neural network

    CN111242216A

  • A super-resolution image reconstruction method of a generative adversarial network based on an attention mechanism

    CN109816593A

  • Text sequence image generation method based on generative adversarial network

    CN113239961A