A text-to-image generation algorithm based on an improved text parser
By introducing an improved text parser and style transfer model into the text-to-image generation algorithm, the difficulties of redundant information processing and object interaction understanding in the prior art are solved, and the layout rationality and style diversity of image generation are realized.
Patent Information
- Application Number
- CN202210560027.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-23
AI Technical Summary
The prior art has problems such as difficulty in redundant information processing, lack of understanding of object interaction relationships, complex network systems and lack of exploration of diversified scene layout in text-to-image generation.
A text-to-image generation algorithm based on an improved text parser was designed, and the semantic relationship-to-geometric relationship mapping was implemented using LSTM and MLP, and the style diversity of image generation was enhanced through the style transfer model.
It effectively improves the layout rationality and content style diversity of image generation, solves the problems of redundant information processing and object interaction understanding in the existing technology, and realizes diversified image generation for more complex scenes.
Smart Images

Figure CN115018941B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a text-to-image generation algorithm based on an improved text parser. Background Art
[0002] Background related to image generation: At present, the development of the field of artificial intelligence has attracted much attention. In the field of computer vision, deep learning has shined in the fields of image recognition, image classification, image segmentation, and semantic description of images, and has demonstrated excellent performance. However, to this day, the problem of image generation remains a daunting challenge, especially the task of cross-modal generation from text to image. According to the different generated objects, the task can be specifically divided into two types: single-target object image generation and multi-target complex scene image generation. The former task will focus on generating high-quality, detailed individual objects, while the latter task is aimed at the generation of multiple objects, and different objects have diverse relationships, which is a more complex and challenging task. Therefore, this patent mainly focuses on the generation of multi-target complex scene images, and designs an effective text parser to improve image generation performance.
[0003] Text-to-image related background: Text-to-scene image generation requires a model to extract useful information from text to assist in the generation of scene images. However, most existing methods have the following problems: (1) The existence of redundant information such as prepositions and copulas in text descriptions increases the difficulty of extracting text information; (2) The model lacks understanding of the interaction between objects in the text, which may lead to unreasonable scene layout; (3) The high-quality text feature extraction network system is relatively large and the training process is relatively complex; (4) Existing work mostly focuses on improving image quality and lacks exploration of the diversification of scene layouts for generated images. In summary, how to extract concise semantic information from complex text has become an important challenge facing the direction of text-to-image generation.
[0004] Background of baseline methods: In 2018, Johnson et al. proposed a scene graph to image generation algorithm, which achieved the generation of complex scenes through a structured scene graph that can reflect the semantic relationship between objects. The method also supplemented that the Stanford syntactic analyzer can be used to extract text semantic information more concisely. However, in practical applications, the syntactic analyzer cannot achieve good analysis for complex texts, resulting in errors in semantic structure. In 2019, Wei Sun and Tianfu Wu proposed LostGANs, which achieved image processing optimization by reconfigurable layout and style; in 2016, Justin Johnson, Alexandre Alahi, and Li Fei-Fei proposed Real-Time Style Transfer, which achieved fast and high-resolution style transfer. Based on this, the present invention designs a text parser for complex relationship vocabulary, automatically converts text into scene graphs, and builds an information conversion bridge for the text-to-image generation process.
[0005] Background related to network design: In the text parser involved in this invention, the mapping of semantic class relations to geometric relations is realized based on LSTM (Long Short-Term Memory Network) and MLP (Multi-layer Perceptron). Specifically, the above two networks are both neural networks. Neural networks were originally inspired by biological nervous systems and emerged to simulate biological nervous systems. They are composed of a large number of nodes (or neurons) that are interconnected. The neural network adjusts the weights according to the changes in the input, improves the behavior of the system, and automatically learns a model that can solve the problem.
[0006] LSTM (Long Short Term Memory Network) is a special form of RNN (Recurrent Neural Network), which effectively solves the gradient vanishing and gradient exploding problems of multi-layer neural network training and can handle long-term time-dependent sequences. LSTM network consists of LSTM units, which consist of input gate, output gate and forget gate.
[0007] MLP (Multi-layer Perceptron) is a generalization of PLA (Perceptron). Its main feature is that it has multiple neuron layers, so it is also called DNN (Deep Neural Network). It has an input layer, some intermediate layers and an output layer. Summary of the invention
[0008] The present invention proposes a text-to-image generation algorithm based on an improved text parser, wherein the improved text parser is an improvement based on the Stanford text parser, based on artificial classification data, long short-term memory network (LSTM) and multi-layer perceptron (MLP). In addition, the present invention embeds the style transfer model into the image generation process, achieving style diversity of the generated results.
[0009] The present invention utilizes an improved text parser to achieve the diversity of semantic understanding, maps complex relationships to geometric layout relationships, and realizes the extraction of text information into a number of (subject, predicate, object) triples. Through the triples, the generation model can pay more attention to the relationship between objects, and generate layouts and images based on this. Finally, through the embedding of the style transfer model, the image is stylized. Using the improved text parser and style transfer module, the text-to-image generation algorithm of the present invention can achieve the rationality of scene layout and the diversity of image content and style.
[0010] The technical solution of the present invention is as follows:
[0011] A text-to-image generation algorithm based on an improved text parser. The specific implementation steps are as follows:
[0012] Step S1: extract text information from the COCO dataset and perform statistics and classification to complete information statistics;
[0013] Step S2: construct a relationship mapping dataset based on fine classification, and divide it into a training set, a validation set, and a test set;
[0014] Step S3: construct an automatic relationship classification network and perform pre-training based on the classification data set in step S2 to achieve mapping of complex semantic relationships to geometric spatial relationships;
[0015] Step S4: constructing a text automatic processing module to extract key information from the input text;
[0016] Step S5: Based on the automatic relationship classification network in step S3 and the automatic text processing module in step S4, an improved text parser is constructed, the text description is input, and the parsed structured triples are output, thereby obtaining a scene graph;
[0017] Step S6: construct a layout prediction network based on the scene graph to image generation algorithm sg2im, and input the scene graph into the layout prediction network to obtain the scene layout;
[0018] Step S7: Combining the Real-Time Style Transfer style transfer and the LostGANs image generation model to build a stylized image generation network, and inputting the layout into the stylized image generation network to obtain images with different artistic styles;
[0019] Step S8: Based on the improved text parser in step S5, the layout prediction network in step S6, and the stylized image generation network in step S7, the overall text-to-image generation algorithm is implemented in the order of S5, S6, and S7, and the algorithm is embedded in the web page background to implement network design for user convenience.
[0020] Beneficial effects of the present invention:
[0021] The difference between the present invention and the existing methods is that compared with the existing text-to-image generation algorithm for complex scenes, the improved text parser proposed in the present invention uses scenes Figure 3 The automatic construction of tuples builds a good bridge between text images, allowing the image generation process to better focus on layout relationships. In addition, from the perspective of diversity, on the one hand, the classification network design involved in the present invention realizes a diverse mapping of triple relationships to layouts, thereby bringing semantic diversity to scene layouts. On the other hand, the image generation module design involved in the present invention realizes style diversity in generating scene images in terms of style. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 This is the overall process of the present invention applied to the text-to-image generation algorithm we designed;
[0023] Figure 2 It is the relationship automatic classification network structure in the present invention;
[0024] Figure 3 It is the specific process of extracting text features by the LSTM module in the relational automatic classification network of the present invention;
[0025] Figure 4 It is the specific process of extracting triple features by the LSTM module in the relational automatic classification network of the present invention;
[0026] Figure 5 is the specific details of the LSTM unit in the present invention, where x t It refers to the embedding vector obtained by embeddinglayer;
[0027] Figure 6 This is the specific process of the present invention to realize the generation of multiple images that meet the semantic description from one text;
[0028] Figure 7 It is the specific process of the present invention applied to the algorithm for generating an algorithm from text to image. Specific implementation methods
[0029] The technical solution of the present invention will be further described below in conjunction with specific embodiments and drawings.
[0030] A text-to-image generation algorithm based on an improved text parser (such as Figure 1 ), the steps are as follows:
[0031] Step S1: Extract text information from the COCO dataset and perform statistics and classification to complete information statistics.
[0032] The step S1 is specifically as follows:
[0033] Step S11: Parse the text information in the COCO dataset. First, tag all words in a sentence with parts of speech; second, search and record the nouns and their modifiers in the sentence; then, determine the subject of the verb (including noun subject and preposition object); then, find the relationship between each noun; finally, generate a structured triple in the form of (subject, predicate, object) based on the found nouns and relationships;
[0034] Step S12: extract and integrate all the relation words into a set as the relation set to be learned.
[0035] Step S13: Roughly classify the relation words. That is, roughly classify the relation words with a frequency greater than or equal to 30 into four categories: Geometric (geometric relationship), Possessive (subordinate relationship), Semantic (semantic relationship), and Misc (others), and complete preliminary statistics of the data set information.
[0036] Step S2: construct a relationship mapping dataset based on fine classification and divide it into a training set, a validation set and a test set.
[0037] The step S2 is specifically as follows:
[0038] Step S21: Combined with the analysis of the text in step S11, the relationship words in the text are subdivided and processed, and all the relationships in the relationship set are mapped to 6 geometric relationships (Left of, Right of, Above, Below, Surrounding, Inside);
[0039] Step S22: convert the six geometric relationship categories into a six-dimensional vector, wherein the value of the geometric relationship category manually classified in step S21 is set to 1, and the values of the other categories are set to 0, and the vector is used as the classification label of the original relationship word to complete the data processing;
[0040] Step S23: Based on the input text in step S11, the parsed triples, the relational words and the category labels obtained in step S22, a relational mapping dataset is constructed, and it is further divided into a training set, a test set and a validation set in a ratio of 80%, 10% and 10%.
[0041] Step S3: Constructing a relational automatic classification network (such as Figure 2 ), and pre-training is performed based on the classification data set in step S2 to achieve the mapping of complex semantic information to geometric spatial relationships. Specifically: let the input sentence be t, and the set of triples initially parsed by step S11 be c i , the relative word is ri , the predicted classification result is a 6-dimensional vector representing 6 geometric relationships.
[0042] The step S3 is specifically as follows:
[0043] Step S31: construct an embeddinglayer module, that is, use the pre-trained word2vec model to obtain the word embedding vectors corresponding to the text, triples, and relation words. Specifically: in this module, the text t, each triple c i and the relative word r i Both are input into the word2vec model loaded with pre-trained weights to obtain text embedding vectors respectively Triplet embedding vector Features with word vectors
[0044] Step S32: construct an LSTM network, further process the embedding vectors of the text and triples, and extract the semantic feature vector. That is, in each LSTM unit (such as Figure 5 ) Use the forget gate to control the decision to discard the text feature information in the previous layer, use the input gate to store valid text feature information, and use the output gate to filter the output text information of each layer. Input LSTM network, pass through LSTM unit, output text feature f t (like Figure 3 ), embedding the triplets from the text into vectors Input LSTM network, pass through LSTM unit, output triple feature (like Figure 4 );
[0045] Step S33: Based on the embedding layer module in step S31 and the LSTM module in step S32, the MLP module is integrated to jointly construct a relational automatic classification network. Specifically, the relational word vector Text feature f t , triple feature Splicing together to get the feature f, that is, definition Among them, [;] represents concatenation. Input f into a multi-layer perceptron (MLP) to obtain a 6-dimensional vector, each element of which represents a type of geometric position relationship that can be processed in the COCO dataset.
[0046] Step S34: Use the relationship mapping dataset constructed in step S2 to pre-train the relationship automatic classification network constructed in step S33, and use the Adam optimizer to minimize the loss.
[0047] Step S4: Construct a text automatic processing module (such as Figure 6 ), realizes the extraction of key information in the input text, and specifically improves three types of problems existing in the parsing of complex text.
[0048] The step S4 is specifically as follows:
[0049] Step S41: Improve the problem of poor extraction of parallel relations in texts containing conjunctions before and after "and". First, identify and divide the texts connected by conjunctions such as "and", and then perform part-of-speech tagging to extract the structured information of (subject, predicate, object) triples;
[0050] Step S42: Improve the problem that only one object modified by a quantifier can be extracted. First, use spacy to determine whether the modifier is a quantifier. If so, add the corresponding number of objects and (subject, predicate, object) structured triples according to the number of recognized quantifiers.
[0051] Step S43: Improve the problem of poor extraction of text information containing the be verb. First, perform part-of-speech tagging, and before extracting the (subject, predicate, object) triple, identify and delete the be verb.
[0052] Step S44: Implement the construction of the text automatic processing module. After the text is input, the text is processed in the order of step S41, step S42, and step S43.
[0053] Step S5: Based on the automatic relational classification network in step S3 and the automatic text processing model in step S4, an improved text parser is constructed, and the text description is input to automatically parse out the structured triples that reflect the spatial layout, thereby obtaining a scene graph.
[0054] Step S51: extracting the initial triples of the text based on the Standford grammatical analyzer, and recording the extracted triples;
[0055] Step S52: input the text description into the text automatic processing module in step S4, extract the relationship information, and implement the preprocessing of the text;
[0056] Step S53: inputting the triplet obtained in step S51 and the text processed in step S52 into the relation automatic classification network in S3 to predict the geometric relation category corresponding to each complex relation in the triplet;
[0057] Step S54: Based on the geometric relationship category in S53 and the subject and object of the triples obtained by parsing in step S51, a (subject, predicate, object) triple that can briefly reflect the spatial layout relationship is recombined, and the triples are combined into a scene graph.
[0058] Step S6: construct a layout prediction network based on the scene graph to image generation algorithm sg2im, input the scene graph, and output the scene layout.
[0059] The step S6 is specifically as follows:
[0060] Step S61: Use the graph convolutional network to extract features from the scene graph. That is, firstly, an initial vector is given to each object and relationship in the scene graph; secondly, the initial vectors of the objects and relationships are input into the multi-layer graph convolution; finally, the embedding vector corresponding to each object is output;
[0061] Step S62: Obtain the overall layout based on a multi-layer perceptron (MLP). That is, first, input the embedding vector corresponding to each object into the MLP to predict the coordinates of each object; second, combine the categories and corresponding coordinates of all objects to obtain the scene layout.
[0062] Step S7: Combining Real-Time Style Transfer with the LostGANs image generation model to build a stylized image generation network, and inputting the layout into the stylized image generation network to obtain images with different artistic styles.
[0063] The step S7 is specifically as follows:
[0064] Step S71: input the layout into the existing LostGANs network to generate the original scene image;
[0065] Step S72: Based on the Real-Time Style Transfer algorithm, a style transferor with three styles is constructed and trained. Specifically, the image to be converted is input, and after passing through multiple sets of convolutional layers, residual layers, and convolutional layers, an output image with the same size as the input is obtained. In the high-level image feature space extracted by the VGG-16 network, the distance between the input image and the output image features is used as the loss function, and the style transferor is trained in combination with the content loss and the style loss.
[0066] Step S73: Input the original scene image obtained in step S71 into the style transfer device in S72 to obtain outputs of multiple artistic styles.
[0067] Step S8: Based on the improved text parser in step S5, the layout prediction network in step S6, and the stylized image generation network in step S7, the stylized image generation network is generated in the order of S5, S6, and S7 (e.g. Figure 1 ) to implement the overall text-to-image generation algorithm (such as Figure 7 ), and embed the algorithm into the web page background to realize network design for the convenience of users.
Claims
1. A text-to-image generation algorithm based on an improved text parser, It is characterized in that The method comprises the following steps: Step S1: extract text information from the COCO dataset and perform statistics and classification to complete information statistics; Step S2: construct a relationship mapping dataset based on fine classification, and divide it into a training set, a validation set, and a test set; Step S3: construct an automatic relationship classification network and perform pre-training based on the classification data set in step S2 to achieve mapping of complex semantic relationships to geometric spatial relationships; The step S3 is specifically as follows: Step S31: construct an embedding layer module, that is, use the pre-trained word2vec model to obtain the word embedding vectors corresponding to the text, triples, and relation words. Specifically: in this module, the text t, each triple c i and the relative word r i Both are input into the word2vec model loaded with pre-trained weights to obtain text embedding vectors respectively Triplet embedding vector Features with word vectors Step S32: construct an LSTM network, further process the embedding vectors of the text and triples, and extract the semantic feature vector; that is, use the forget gate control in each LSTM unit to decide to discard the text feature information in the previous layer, use the input gate to store the valid text feature information, and use the output gate to filter the output text information of each layer; embed the text into the vector Input LSTM network, pass through LSTM unit, output text feature f t ; Embed the triplets from the text into vectors Input LSTM network, pass through LSTM unit, output triple feature Step S33: Based on the embedding layer module in step S31 and the LSTM module in step S32, the MLP module is integrated to jointly construct a relationship automatic classification network; specifically, the relationship word vector Text feature f t , triple feature Splicing together to get the feature f, that is, definition Among them, [;] represents concatenation; f is input into a multi-layer perceptron (MLP) to obtain a 6-dimensional vector, each element of which represents a geometric position relationship that can be processed in a class of COCO datasets; Step S34: pre-train the relational automatic classification network constructed in step S33 using the relational mapping dataset constructed in step S2, and use the Adam optimizer to minimize the loss; Step S4: constructing a text automatic processing module to extract key information from the input text; Step S5: Based on the automatic relationship classification network in step S3 and the automatic text processing module in step S4, an improved text parser is constructed, the text description is input, and the parsed structured triples are output, thereby obtaining a scene graph; Step S6: construct a layout prediction network based on the scene graph to image generation algorithm sg2im, and input the scene graph into the layout prediction network to obtain the scene layout; Step S7: Combining the Real-Time Style Transfer style transfer and the LostGANs image generation model to build a stylized image generation network, and inputting the layout into the stylized image generation network to obtain images with different artistic styles; Step S8: Based on the improved text parser in step S5, the layout prediction network in step S6, and the stylized image generation network in step S7, the overall text-to-image generation algorithm is implemented in the order of S5, S6, and S7, and the algorithm is embedded in the web page background to implement network design for user convenience.
2. A text-to-image algorithm based on an improved text parser according to claim 1, It is characterized in that The step S1 is specifically as follows: Step S11: parse the text information in the COCO dataset; first, tag all words in a sentence with parts of speech; second, search and record the nouns and their modifiers in the sentence; then, determine the subject of the verb (including noun subject and preposition object); then, find the relationship between each noun; finally, generate a structured triple in the form of (subject, predicate, object) based on the found nouns and relationships; Step S12: extract and integrate all the relation words into a set as the relation set to be learned; Step S13: Roughly classify the relationship words; that is, roughly classify the relationship words with a frequency greater than or equal to 30 into four categories: Geometric (geometric relationship), Possessive (subordinate relationship), Semantic (semantic relationship), and Misc (others), and complete preliminary statistics on the data set information.
3. A text-to-image algorithm based on an improved text parser according to claim 1 or 2, It is characterized in that The step S2 is specifically as follows: Step S21: Combined with the analysis of the text in step S11, the relationship words in the text are subdivided and processed, and all the relationships in the relationship set are mapped to 6 geometric relationships (Left of, Right of, Above, Below, Surrounding, Inside); Step S22: convert the six geometric relationship categories into a six-dimensional vector, wherein the value of the geometric relationship category manually classified in step S21 is set to 1, and the values of the other categories are set to 0, and the vector is used as the classification label of the original relationship word to complete the data processing; Step S23: Based on the input text in step S11, the parsed triples, the relational words and the category labels obtained in step S22, a relational mapping dataset is constructed, and it is further divided into a training set, a test set and a validation set in a ratio of 80%, 10% and 10%.
4. A text-to-image algorithm based on an improved text parser according to claim 1 or 2, It is characterized in that The step S4 is specifically as follows: Step S41: improving the problem of poor extraction of parallel relations containing conjunctions before and after "and" in the text; first, the text containing conjunctions such as "and" is first identified and divided, and then part-of-speech tagging is performed to extract the structured information of the (subject, predicate, object) triples; Step S42: Improve the problem that only one object modified by a quantifier can be extracted. First, use spacy to determine whether the modifier is a quantifier. If so, add the corresponding number of objects and (subject, predicate, object) structured triples according to the number of recognized quantifiers. Step S43: improving the problem of poor extraction of text information containing the be verb; first, perform part-of-speech tagging, and identify and delete the be verb before extracting the (subject, predicate, object) triple; Step S44: Implement the construction of the text automatic processing module; after the text is input, the text is processed in the order of step S41, step S42, and step S43.
5. A text-to-image algorithm based on an improved text parser according to claim 3, It is characterized in that The step S4 is specifically as follows: Step S41: improving the problem of poor extraction of parallel relations containing conjunctions before and after "and" in the text; first, the text containing conjunctions such as "and" is first identified and divided, and then part-of-speech tagging is performed to extract the structured information of the (subject, predicate, object) triples; Step S42: Improve the problem that only one object modified by a quantifier can be extracted. First, use spacy to determine whether the modifier is a quantifier. If so, add the corresponding number of objects and (subject, predicate, object) structured triples according to the number of recognized quantifiers. Step S43: improving the problem of poor extraction of text information containing the be verb; first, perform part-of-speech tagging, and identify and delete the be verb before extracting the (subject, predicate, object) triple; Step S44: Implement the construction of the text automatic processing module; after the text is input, the text is processed in the order of step S41, step S42, and step S43.
6. A text-to-image algorithm based on an improved text parser according to claim 1, 2 or 5, It is characterized in that The step S5 is specifically as follows: Step S51: extracting the initial triples of the text based on the Standford grammatical analyzer, and recording the extracted triples; Step S52: input the text description into the text automatic processing module in step S4, extract the relationship information, and implement the preprocessing of the text; Step S53: inputting the triplet obtained in step S51 and the text processed in step S52 into the relation automatic classification network in S3 to predict the geometric relation category corresponding to each complex relation in the triplet; Step S54: Based on the geometric relationship category in S53 and the subject and object of the triples parsed in step S51, recombine to obtain (subject, predicate, object) triples that can briefly reflect the spatial layout relationship; and combine the triples into a scene graph.
7. A text-to-image algorithm based on an improved text parser according to claim 3, It is characterized in that The step S5 is specifically as follows: Step S51: extracting the initial triples of the text based on the Standford grammatical analyzer, and recording the extracted triples; Step S52: input the text description into the text automatic processing module in step S4, extract the relationship information, and implement the preprocessing of the text; Step S53: inputting the triplet obtained in step S51 and the text processed in step S52 into the relation automatic classification network in S3 to predict the geometric relation category corresponding to each complex relation in the triplet; Step S54: Based on the geometric relationship category in S53 and the subject and object of the triples parsed in step S51, recombine to obtain (subject, predicate, object) triples that can briefly reflect the spatial layout relationship; and combine the triples into a scene graph.
8. A text-to-image algorithm based on an improved text parser according to claim 4, It is characterized in that The step S5 is specifically as follows: Step S51: extracting the initial triples of the text based on the Standford grammatical analyzer, and recording the extracted triples; Step S52: input the text description into the text automatic processing module in step S4, extract the relationship information, and implement the preprocessing of the text; Step S53: inputting the triplet obtained in step S51 and the text processed in step S52 into the relation automatic classification network in S3 to predict the geometric relation category corresponding to each complex relation in the triplet; Step S54: Based on the geometric relationship category in S53 and the subject and object of the triples parsed in step S51, recombine to obtain (subject, predicate, object) triples that can briefly reflect the spatial layout relationship; and combine the triples into a scene graph.
9. A text-to-image algorithm based on an improved text parser according to claim 1 or 2 or 5 or 7 or 8, It is characterized in that The step S7 is specifically as follows: Step S71: input the layout into the existing LostGANs network to generate the original scene image; Step S72: Based on the Real-Time Style Transfer algorithm, a style transfer device with three styles is constructed and trained; specifically, the image to be converted is input, and after passing through a structure of multiple groups of convolutional layers, residual layers, and convolutional layers, an output image with the same size as the input is obtained; in the high-level image feature space extracted by the VGG-16 network, the distance between the input image and the output image features is used as the loss function, and the style transfer device is trained in combination with the content loss and the style loss; Step S73: Input the original scene image obtained in step S71 into the style transfer device in S72 to obtain outputs of multiple artistic styles.
10. A text-to-image algorithm based on an improved text parser according to claim 3, It is characterized in that The step S7 is specifically as follows: Step S71: input the layout into the existing LostGANs network to generate the original scene image; Step S72: Based on the Real-Time Style Transfer algorithm, a style transfer device with three styles is constructed and trained; specifically, the image to be converted is input, and after passing through a structure of multiple groups of convolutional layers, residual layers, and convolutional layers, an output image with the same size as the input is obtained; in the high-level image feature space extracted by the VGG-16 network, the distance between the input image and the output image features is used as the loss function, and the style transfer device is trained in combination with the content loss and the style loss; Step S73: Input the original scene image obtained in step S71 into the style transfer device in S72 to obtain outputs of multiple artistic styles.
Citation Information
Patent Citations
Image description generation method and system combined with abstract semantic representation and medium
CN111612103A
Scene map generation method, device and equipment
CN111931928A