Image Style Transfer System and Method Based on Multimodal Semantic Matching
Through the multimodal semantic matching image style transfer system, the multimodal pre-trained model and attention mechanism are used to compare the multimodal pre-trained model and attention mechanism in the existing technology, the problem that image style transfer methods in the prior art is difficult to match the content semantic areas, and high-quality, natural and beautiful image stylization results are achieved.
Patent Information
- Application Number
- CN202211578075.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-12-06
AI Technical Summary
When existing image style transfer methods use text to describe style preferences, it is difficult to ensure that the content semantic region of the output image is rendered by the style semantic region matching the style image, resulting in semantic confusion in the stylized result and damaged content structure.
The image style transfer system with multimodal semantic matching is adopted to construct a text image retrieval module through the multimodal pre-trained model of graphic comparison, supporting multiple modal data such as text and images as input, and using attention mechanism and interpolation operations in the image style transfer module to adjust the style image feature distribution to align it with the content image feature distribution.
This realizes image style transfer that supports multiple modal data such as text and images as input, ensuring that the content semantic area of the output image is rendered by the style semantic area matching the style image, improving the flexibility and quality of the image style transfer process.
Smart Images

Figure CN115829830B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image style transfer system and method based on multimodal semantic matching, belonging to the fields of computer and culture. Background Art
[0002] Image style transfer refers to transferring the style of an image to a natural image, so that the natural image retains the original content while having a unique style. With the popularization of mobile devices, the beauty functions in cameras and short videos are widely used by people, and the public's requirements for the effect of image style transfer are getting higher and higher. In addition, image style transfer technology plays a huge role in film and television production, animation rendering, etc. Therefore, further research on image style transfer technology is beneficial to exploring its more potential value and wider application space.
[0003] As an application of deep learning technology in the cultural field, since Gatys et al. first applied the VGG network to the field of image style transfer in 2015, it has led the research trend of combining deep learning with image style transfer algorithms, and a large number of excellent algorithms have emerged.
[0004] Image style transfer methods need to use a style image and a content image as inputs to provide style and content information. However, in many practical situations, users may not have a suitable style image for reference, and it is easier to obtain and adjust the description of style preferences using text rather than using a style image. Therefore, it is very necessary to construct an image style transfer model that supports multimodal data such as text and images as inputs. Currently, most image style transfer methods assume that the image style can be represented by the global statistics of its deep features, such as the Gram matrix or covariance matrix. This global statistics captures the style from the entire image and applies it to the content image, resulting in problems such as semantic style confusion and content structure damage in the final stylized result, where different semantic regions of the content image are rendered by non-matching semantic regions in the style image. Therefore, it is very necessary to use a multimodal semantic matching-based image style transfer method. Summary of the Invention
[0005] The object of the present invention is to support multiple modal data such as text and images as inputs, and ensure that the semantic regions of the output image content are rendered by the style semantic regions matching the style image, so as to improve the flexibility of the image style transfer process and achieve natural, beautiful, and high-quality image stylized results.
[0006] To achieve the above object, a technical solution of the present invention is to provide an image style transfer system based on multimodal semantic matching, which is characterized by including a content image input module, a style information input module, a style image vector library, a text-image retrieval module, an image style transfer module, and a result output module, wherein:
[0007] The content image input module is used to input a content image to the image style transfer module and provide content information for the final output result of the image style transfer module;
[0008] The style information input module is used to input style information to the image style transfer module. The style information is text data for describing the style or a style image for describing the style, so as to support using data in two modalities, namely text or image, as input to provide style information for the final output result of the image style transfer module;
[0009] The style image vector library: A style image vector library is established based on a style image data set. After creating text labels for each style image in the style image data set, a multimodal pre-training model for text-image comparison is used to encode each style image with a text label in the style image data set to obtain a style image vector, and a vector library is established based on all style image vectors;
[0010] The text-image retrieval module: uses a multimodal pre-training model for text-image comparison to encode the text data input through the style information input module into a text vector, then retrieves the style image vector with the highest semantic matching degree with the current text vector in the style image vector library, and outputs the corresponding style image to the image style transfer module;
[0011] The result output module: restores the stylized image features obtained after being processed by the image style transfer module back to an image and outputs it.
[0012] Preferably, the image sizes of the style image and the content image are the same.
[0013] Preferably, the text label includes the name of the creator of the current style image and a text description of the semantic content of the current style image.
[0014] Preferably, the result output module saves the stylized result processed by the image style transfer module to a specified local folder.
[0015] Another technical solution of the present invention is to provide an image style transfer method based on multimodal semantic matching, which is characterized by including the following steps:
[0016] S100. Original image processing:
[0017] Convert the content image input by the user into an image of a set size. If the style information input by the user through the style information input module is a style image, convert the style image into an image of the same size as the content image;
[0018] Obtain a style image dataset and convert the style images in the style image dataset into images of a set size;
[0019] S200. Style image annotation:
[0020] Create a text label for each style image in the style image dataset. The content of the text label includes at least a literal description of the semantic content of the current style image, and finally form a table. Each row in the table records the path of a style image in the style image dataset and its corresponding text label;
[0021] S300. Construct a style image vector library:
[0022] Based on the table obtained in step S200, read the style images at the corresponding paths in the table in index order, and use a multi-modal pre-training model for image-text comparison to encode each style image to obtain a style image vector, thereby constructing a style image vector library;
[0023] S400. Image style transfer, and obtain the final stylized result by selecting different methods according to the input data modality:
[0024] If the user inputs text data for describing the style through the style information input module, encode the input text data into a text vector through the text image retrieval module, then retrieve the style image vector with the highest matching degree with the current text vector from the style image vector library, and restore it to a style image and input it together with the content image input through the content image into the image style transfer module to obtain the final stylized result;
[0025] If the user inputs a style image for providing style information through the style information input module, directly input the style image and the content image input through the content image into the image style transfer module to obtain the final stylized result;
[0026] S500. Analysis result display: The result output module restores the stylized image features processed by the image style transfer module back to an image and outputs it.
[0027] Preferably, in step S200, the text label further includes the name of the creator of the current style image.
[0028] Preferably, step S300 includes the following steps:
[0029] S301. According to the table created in step S200, read the style images under the path in the index order, and extract the image features through the image encoder in the text-image comparison multi-modal pre-trained model to obtain the style image vectors;
[0030] S302. Use the Milvus cloud-native vector database to save the style image vectors.
[0031] Preferably, in the step S400, the specific operations of the text-image retrieval module include the following steps:
[0032] S401. Extract the text features of the input text data through the text encoder in the text-image comparison multi-modal pre-trained model to obtain the text vectors;
[0033] S402. Compare the text vectors with the style image vectors in the style image vector library, calculate the Euclidean distance between the two, and retrieve the style image vector with the highest matching degree with the text vector in the style image vector library;
[0034] S403. According to the index of the style image vector, query the path of the style image corresponding to the index in the table obtained in step S200, and return the style image.
[0035] Preferably, in the step S400, the specific operations of the image style transfer module include the following steps:
[0036] S404. Extract the image features of the content image and the style image through the pre-trained VGG network to obtain the content image feature vector and the style image feature vector;
[0037] S405. Normalize and embed the content image features and the style image features to calculate the attention map;
[0038] S406. Use the attention map obtained in step S405 to rearrange the distribution of the style image features by affine transformation, so that the semantic regions of the content image and the relevant semantic regions of the style image correspond on the feature map, and obtain the style image features with semantic matching;
[0039] S407. Concatenate the content image features and the style image features after adjusting the distribution through the channel connection operation, identify the local differences between the corresponding content image feature and style image feature semantic regions, and then fill in this difference through the interpolation operation to solve the local distortion of the final stylized result;
[0040] S408. Input the stylized image features obtained in the previous step into the next style conversion module, and perform attention operations and interpolation operations according to steps S405, S406 and S407 to achieve further refinement;
[0041] S409. Input the stylized image features that have passed through the three-style conversion module into the result output module to output the final stylized result.
[0042] The structure of the present invention is reasonably designed. The text-image retrieval module is constructed by using the multimodal pre-training model for image-text contrast (CLIP), introducing the text modality to describe the style information for the input of the image style transfer system, so that the overall image style transfer system can support data in two modalities, namely text and style images, as style information. At the same time, the image style transfer module adjusts the feature distribution of the style image through the attention mechanism and interpolation operation, so that the content semantic region of the final transfer result is rendered by the style semantic region that matches it, achieving a high-quality stylized effect on the basis of reducing the loss of the content semantic structure. Brief Description of the Drawings
[0043] Figure 1 It is a flowchart of the image style transfer system based on multimodal semantic matching of the present invention;
[0044] Figure 2 It is an overall framework diagram of the image style transfer system based on multimodal semantic matching of the present invention;
[0045] Figure 3 It is a structural block diagram of the text-image retrieval module of the image style transfer system based on multimodal semantic matching of the present invention;
[0046] Figure 4 It is a structural block diagram of the image style transfer module and the result output module of the image style transfer system based on multimodal semantic matching of the present invention;
[0047] Figure 5 It is a structural block diagram of the style conversion module in the image style transfer module of the image style transfer system based on multimodal semantic matching of the present invention. Detailed Embodiments
[0048] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.
[0049] As Figure 1 shown, an embodiment of the present invention proposes an image style transfer system based on multimodal semantic matching, including a content image input module, a style image input module, a style image vector library, a text-image retrieval module, an image style transfer module, and a result output module, where:
[0050] A content image input module for inputting a content image to an image style transfer module, providing content information for the final output result of the image style transfer module.
[0051] A style information input module for inputting style information to the image style transfer module. The style information includes text data for describing the style or a style image for describing the style, providing style information for the final output result of the image style transfer module.
[0052] In this embodiment, the content image and the style image are obtained through an open-source image dataset. For the content image dataset and the style image dataset, the image sizes are different. Therefore, the image sizes of the content image and the style image dataset are first converted to 512 * 512 pixels to facilitate subsequent processing of the image data.
[0053] Style image vector library: Based on the style image dataset, a style image vector library is established. After creating text labels for each style image in the style image dataset, a multimodal pre-training model for image-text contrast (CLIP) is used to encode each style image with a text label in the style image dataset to obtain style image vectors, and a vector library is established based on all style image vectors. In this embodiment, the content of the text label includes the name of the creator of the current style image and a short description of the semantic content of the current style image. Finally, a table corresponding to the style image dataset is formed, and each row in the table records the path of a style image in the style image dataset and its corresponding text label.
[0054] Text image retrieval module: Using a multimodal pre-training model for image-text contrast (CLIP) to encode the text data input through the style information input module into text vectors, then retrieving the style image vector with the highest semantic matching degree with the current text vector in the style image vector library, and outputting the corresponding style image to the image style transfer module.
[0055] An image style transfer module for transferring the style of the input style image to the input content image while maintaining the integrity of the content structure of the content image.
[0056] Result output module: Restoring the stylized image features obtained after being processed by the image style transfer module back to an image and outputting it. In this embodiment, the result output module will save the stylized result processed by the image style transfer module to a specified local folder for convenient subsequent display and evaluation of the stylized result.
[0057] The following are preferred embodiments of the multimodal semantic matching-based image style transfer system to clearly illustrate the content of the present invention. It should be clear that the content of the present invention is not limited to the following embodiments, and other improvements by conventional technical means of ordinary technicians in the art are also within the scope of the idea of the present invention.
[0058] As Figure 2 shown, an embodiment of the present invention proposes an image style transfer method based on multimodal semantic matching, including the following steps:
[0059] S100. Original image processing: The original content image and style image are of different sizes. To facilitate subsequent processing of image data, the original content image and style image are uniformly converted into images of 512*512 pixels. Specifically, in this embodiment, the third-party library Opencv-python of Python can be used to quickly adjust the size of the image.
[0060] S200. Style image annotation: Create a text label for each style image in the style image dataset. The content of the text label includes the name of the creator of the current style image and a short description of the semantic content of the current style image. Finally, a table corresponding to the style image dataset is formed. Each row in the table records the path of a style image in the style image dataset and its corresponding text label.
[0061] Specifically, the annotated style image needs to extract features through the image encoder in the image-text contrast multimodal pre-trained model (CLIP) according to its index and storage path in the table, and save the extracted style image vector to the style image vector library according to the index.
[0062] S300. Construct a style image vector library:
[0063] According to the table obtained in step S200, use the image-text contrast multimodal pre-trained model (CLIP) to encode each style image in the style image dataset to obtain a style image vector, so as to construct a style image vector library. In this embodiment, the processing of constructing the style image vector library specifically includes the following steps:
[0064] S301. According to the table created in step S200, read the style images in the corresponding paths in index order, and extract their image features through the image encoder in the image-text contrast multimodal pre-trained model (CLIP) to obtain style image vectors;
[0065] S302. Use the Milvus cloud-native vector database to save the style image vectors to form a style image vector library, which has the characteristics of high availability, high performance, and easy expansion, and can perform vector similarity retrieval on a large-scale dataset and achieve real-time recall.
[0066] S400, Image style transfer, and different methods are selected according to the input data modality to obtain the final stylized result:
[0067] If the text data used to describe the style is input through the style information input module, the input text data is encoded into a text vector through the text image retrieval module, and then the style image vector with the highest matching degree with the current text vector is retrieved from the style image vector library, and after restoring it to a style image, it is input into the image style transfer module together with the content image input through the content image to obtain the final stylized result;
[0068] If the style image used to provide style information is input through the style information input module, the style image is directly input into the image style transfer module together with the content image input through the content image to obtain the final stylized result.
[0069] In this embodiment, the specific operations of the text image retrieval module include the following steps:
[0070] S401. Extract text features of the input text data through the text encoder in the multi-modal pre-training model (CLIP) of text-image comparison to obtain a text vector;
[0071] S402. Compare the text vector with the style image vectors in the style image vector library, calculate the Euclidean distance between the two, and retrieve the style image vector with the highest matching degree with the text vector in the style image vector library;
[0072] S403. According to the index of the style image vector, query the path of the style image corresponding to the index in the table obtained in step S200, and return the style image. As Figure 3 shown.
[0073] In this embodiment, the specific operations of the image style transfer module include the following steps:
[0074] S404. Extract image features of the content image and the style image through the pre-trained VGG network to obtain a content image feature vector and a style image feature vector;
[0075] S405. Normalize and embed the content image features and the style image features to calculate an attention map;
[0076] S406. Use the attention map obtained in step S405 to re-arrange the distribution of the style image features through affine transformation, so that the semantic regions of the content image and the relevant semantic regions of the style image correspond on the feature map, and obtain a stylized image feature with semantic matching;
[0077] S407. Concatenate the content image features and the stylized image features after adjusted distribution together through a channel connection operation, identify the local differences between the corresponding content image feature semantic regions and stylized image feature semantic regions, and then fill in such differences through an interpolation operation to solve the local distortion situation of the final stylized result;
[0078] S408. Input the stylized image features obtained in the previous step into the next style conversion module, and perform attention operation and interpolation operation according to steps S405, S406 and S407 to achieve further refinement. The style conversion module is as Figure 5 shown.
[0079] S409. Input the stylized image features that have passed through the three style conversion modules into the result output module to output the final stylized result.
[0080] S500. Analysis result display: The result output module restores the stylized image features processed by the image style transfer module back to an image and then outputs it. Specifically, the result output by the result output module will be saved to a specified folder for subsequent display and evaluation of the stylized result.
[0081] The specific operation process of the present invention on a computer is as follows:
[0082] (1) Enter the image style transfer web page;
[0083] (2) Select the stylization method. You can click with the mouse to select a preset stylized image or click the stylized image upload button to upload a custom stylized image, or you can enter a style description language in the text box to describe the style;
[0084] (3) Click the content image upload button to upload the content image;
[0085] (4) Click the image style transfer button and wait for the system to generate the stylized result;
[0086] (5) After waiting for a period of time, the web page will pop up and display "Image style transfer completed", then display the stylized result on the web page and save the result to the specified folder.
[0087] In summary, the technical solution disclosed in this embodiment has the following advantages compared with the prior art:
[0088] The present invention extracts features of style images in an open-source style image dataset through an image encoder in a text-image contrastive multi-modal pre-training model (CLIP), saves the style image vectors to a Milvus cloud-native vector database, and constructs a text-image retrieval module for retrieving style images with high semantic matching degree through text. In the image style transfer module, the attention mechanism and interpolation operation are used to gradually adjust the alignment of the style image feature distribution and the content image feature distribution, so that the content semantic region and the style semantic region of the final stylized result match each other, and a better stylization effect is obtained while ensuring the integrity of the content structure of the stylized result. The present invention combines the text-image retrieval module and the image style transfer module to construct an image style transfer system based on multi-modal semantic matching, which supports image style transfer that provides style information for both text-driven and image-driven modal data.
Claims
1. An image style transfer system based on multi-modal semantic matching, characterized in that, it includes a content image input module, a style information input module, a style image vector library, a text-image retrieval module, an image style transfer module, and a result output module, where: The content image input module is used to input a content image to the image style transfer module and provide content information for the final output result of the image style transfer module; The style information input module is used to input style information to the image style transfer module. The style information is text data for describing the style or a style image for describing the style, realizing that data in two modalities, text or image, are supported as inputs to provide style information for the final output result of the image style transfer module; Style image vector library: A style image vector library is established based on the style image dataset. After creating text labels for each style image in the style image dataset, a multi-modal pre-training model for text-image comparison is used to encode each style image with a text label in the style image dataset to obtain style image vectors, and a vector library is established based on all style image vectors; Text-image retrieval module: The multi-modal pre-training model for text-image comparison is used to encode the text data input through the style information input module into text vectors, and then retrieve the style image vector with the highest semantic matching degree with the current text vector in the style image vector library, and output the corresponding style image to the image style transfer module; if the style image input by the user through the style information input module is a style image for providing style information, the style image and the content image input through the content image are directly input into the image style transfer module to obtain the final stylized result; Result output module: The stylized image features obtained after being processed by the image style transfer module are restored to an image and output.
2. The image style transfer system based on multi-modal semantic matching according to claim 1, characterized in that, the image sizes of the style image and the content image are the same.
3. The image style transfer system based on multi-modal semantic matching according to claim 1, characterized in that, the text label includes the name of the creator of the current style image and a text description of the semantic content of the current style image.
4. The image style transfer system based on multi-modal semantic matching according to claim 1, characterized in that, the result output module saves the stylized result processed by the image style transfer module to a specified local folder.
5. An image style transfer method based on multi-modal semantic matching, characterized in that, it includes the following steps: S100. Original image processing: Convert the content image input by the user through the content image into an image of a set size. If the style information input by the user through the style information input module is a style image, convert the style image into an image of the same size as the content image; Obtain the style image dataset, and convert the style images in the style image dataset into images of a set size; S200. Style image annotation: Create a text label for each style image in the style image dataset. The content of the text label should at least include a literal description of the semantic content of the current style image, and finally form a table. Each row in the table records the path of a style image in the style image dataset and its corresponding text label; S300. Build a style image vector library: Based on the table obtained in step S200, read the style images at the corresponding paths in the table in index order, and use the text-image contrastive multimodal pre-trained model to encode each style image to obtain a style image vector, thereby constructing a style image vector library; S400. Image style transfer, obtaining the final stylized result by different methods according to the input data modality: If the user inputs text data for describing the style through the style information input module, encode the input text data into a text vector through the text-image retrieval module, then retrieve the style image vector with the highest matching degree with the current text vector from the style image vector library, and restore it to a style image and input it together with the content image input through the content image into the image style transfer module to obtain the final stylized result; If the user inputs a style image for providing style information through the style information input module, directly input the style image and the content image input through the content image into the image style transfer module to obtain the final stylized result; S500. Analysis result display: The result output module restores the stylized image features processed by the image style transfer module back to an image and outputs it.
6. A method for image style transfer based on multimodal semantic matching as claimed in claim 5, wherein, in step S200, the text label further includes the name of the creator of the current style image.
7. A method for image style transfer based on multimodal semantic matching as claimed in claim 5, wherein, step S300 includes the following steps: S301. According to the table created in step S200, read the style images at the paths in index order, and extract image features through the image encoder in the text-image contrastive multimodal pre-trained model to obtain style image vectors; S302. Use the Milvus cloud-native vector database to store the style image vectors.
8. A method for image style transfer based on multimodal semantic matching as claimed in claim 5, wherein, in step S400, the specific operations of the text-image retrieval module include the following steps: S401. Extract text features from the input text data through the text encoder in the text-image contrastive multimodal pre-trained model to obtain a text vector; S402. Compare the text vector with the style image vectors in the style image vector library, calculate the Euclidean distance between the two, and retrieve the style image vector with the highest matching degree with the text vector in the style image vector library; S403. According to the index of the style image vector, query the path of the style image corresponding to the index in the table obtained in step S200, and return the style image.
9. A method for image style transfer based on multimodal semantic matching according to claim 5, characterized in that, in the step S400, the specific operations of the image style transfer module include the following steps: S404. Extract image features of the content image and the style image through a pre-trained VGG network to obtain a content image feature vector and a style image feature vector; S405. Normalize and embed the content image features and the style image features to calculate an attention map; S406. Use the attention map obtained in step S405 to rearrange the distribution of the style image features through affine transformation, so that the semantic region of the content image corresponds to the relevant semantic region of the style image on the feature map, and obtain a semantically matched stylized image feature; S407. Concatenate the content image features and the stylized image features after adjusting the distribution through channel connection operations, identify the local differences between the corresponding content image feature and style image feature semantic regions, and then fill in this difference through interpolation operations to solve the local distortion of the final stylized result; S408. Input the stylized image features obtained in the previous step into the next style conversion module, and perform attention operations and interpolation operations according to steps S405, S406, and S407 to achieve further refinement; S409. Input the stylized image features that have passed through the three style conversion modules into the result output module to output the final stylized result.
Citation Information
Patent Citations
Image processing method and device based on residual network
CN114155542A
Stylized image description generation method based on transfer learning
CN115294427A