Image art style migration model training method and system, and computer program
Through the image art style transfer model training method based on text instructions, the visual and text encoder, discriminator and AdaIN personalized retention layer are used to solve the problem that users' personalized needs are difficult to meet and the migration effect is poor in the prior art, efficient and personalized image style migration is achieved, and user privacy is guaranteed.
Patent Information
- Application Number
- CN202411873442.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-23
AI Technical Summary
The prior art is difficult to meet the user's personalized image art style migration needs, and the migration effect is poor, and there are user privacy and copyright issues.
Using text instructions-based image art style transfer model training method, features are extracted through visual encoder and text encoder, model training is performed using discriminator and reconstruction loss, and parameter aggregation and update through AdaIN personalized retention layer and federated learning.
It realizes the ability to meet the user's personalized style migration needs without leaking user privacy data, and improves the migration effect and generalization capabilities of the model.
Smart Images

Figure CN120031707A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image art style transfer model training method, system and computer program. Background Art
[0002] Generative AI improves the flexibility, visual appeal, and diversity of computer vision content creation by adopting different generative deep learning models, especially in language-to-image generation tasks. The most widely used application in the field of generative AI is image style transfer, which is to merge the content of one image with the style of another image to generate an image with a new style. Style transfer has considerable application value for creative visual designs such as image style filters, artistic style design, and video streaming effects in smart mobile devices. For example, an application with a built-in style transfer network model can quickly render user photos into sketch or ink style. However, such applications usually require users or service providers to prepare target style image datasets in advance. For users, it is time-consuming and laborious to prepare or download the expected target style images in advance, and the limited reference styles provided by the platform make it difficult to provide customized services for different users and cannot meet personalized needs. At the same time, traditional style transfer methods require collecting target images of private photos and artworks from different users, which may involve user privacy and copyright issues of artworks. Therefore, the availability of such applications is seriously affected. Summary of the invention
[0003] In view of this, the embodiments of the present invention provide an image art style transfer model training method, system and computer program to eliminate or improve one or more defects existing in the prior art, and solve the problem that the prior art is limited by user data privacy requirements and is difficult to meet the user's personalized image art style transfer needs and the transfer effect is poor.
[0004] One aspect of the present invention provides a method for training an image art style transfer model based on text instructions, the method being implemented by a plurality of clients, each of which is connected to a server aggregation end, and each of the clients comprising the following steps:
[0005] Based on the set link, obtain a content image to be style-converted, a text instruction describing the target style conversion requirements, and a desired style image as an example;
[0006] Extracting image content features and style features from the content image based on the visual encoder, and extracting instruction semantic features of the text instruction based on the text encoder; taking the image content features and the instruction semantic features as input and outputting a target image after style conversion based on the visual decoder, and taking the image content features and the style features as input and outputting a reconstructed image based on the visual decoder; establishing a reconstruction loss based on the reconstructed image and the original content image, and using a discriminator to establish a discriminant loss for the target image and an expected style image that meets the target style, and minimizing the reconstruction loss and the discriminant loss to perform parameter updates on the visual encoder, the text encoder, the visual decoder, and the discriminator;
[0007] AdaIN is used as the personalized retention layer of each client, and the parameters of the personalized retention layer are uploaded to the personalized parameter library. Each client randomly downloads the reference personalized parameters uploaded by other clients, and uses the reference personalized parameters to convert the local content image according to the target style, and builds a consistency loss with the target image to update the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator;
[0008] Sending the parameters of the local visual encoder, the text encoder, the visual decoder and the discriminator to the server aggregation end for parameter aggregation;
[0009] Receive the updated parameters of the visual encoder, the text encoder, the visual decoder, and the discriminator returned by the server aggregation end;
[0010] After multiple rounds of iterations, the updated visual encoder, the text encoder, and the visual decoder are constructed into an image art style transfer model.
[0011] In some embodiments, the visual encoder and the visual decoder adopt a VGG-19 model, and the text encoder adopts a BERT model.
[0012] In some embodiments, the reconstruction loss is calculated using the L2 distance, and the calculation formula is:
[0013]
[0014] Among them, L r represents the reconstruction loss, represents the original content image, represents the reconstructed image;
[0015] Using a discriminator to establish a discriminative loss for the target image and an expected style image satisfying the target style, including:
[0016] The target image and the desired style image are randomly cropped, and the cropped image fragments and the text instructions are input into the discriminator for recognition. The calculation formula of the discrimination loss is:
[0017]
[0018] Among them, L D represents the discrimination loss, represents the image fragments obtained by randomly cropping the target image, C S represents the image fragments obtained by randomly cropping the desired style image, T represents the text instruction, represents the discriminant result of the image fragment of the target image, D(C S ,T) represents the discrimination result of the image fragments of the desired style image.
[0019] In some embodiments, AdaIN is used as the personalized retention layer of each client, and the expression is:
[0020]
[0021] Among them, x represents the characteristics of the content image input by the corresponding network layer, y represents the characteristics of the text instruction input by the corresponding network layer, μ represents the mean, and σ represents the variance.
[0022] In some embodiments, the server aggregation end aggregates the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator uploaded by each client in a weighted average form.
[0023] In some embodiments, after acquiring the content image to be style-converted, the text instruction describing the target style conversion requirements, and the desired style image as an example based on the set link, the method further includes:
[0024] The content image is preprocessed, and the preprocessing includes format unification and data normalization.
[0025] In some embodiments, the method further includes: sharing the trained image art style transfer model externally through a preset link.
[0026] On the other hand, the present invention also provides an image art style transfer model training system based on text instructions, including multiple clients, each of which is connected to a server aggregation end;
[0027] The client comprises:
[0028] A multi-element acquisition module, for acquiring a content image to be style-converted, a text instruction describing the target style conversion requirements, and an example desired style image based on a set link;
[0029] A multivariate information extraction module is used to extract image content features and style features from the content image based on a visual encoder, and to extract instruction semantic features of the text instruction based on a text encoder; to use the image content features and the instruction semantic features as inputs and output a target image after style conversion based on a visual decoder, and to use the image content features and the style features as inputs and output a reconstructed image based on the visual decoder;
[0030] Generate a model training module, which is used to establish a reconstruction loss based on the reconstructed image and the original content image, and use a discriminator to establish a discriminant loss for the target image and the desired style image that meets the target style, and minimize the reconstruction loss and the discriminant loss to update the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator;
[0031] A consistency alignment module, used to use AdaIN as the personalized retention layer of each client, upload the parameters of the personalized retention layer to the personalized parameter library, each client randomly downloads the reference personalized parameters uploaded by other clients, and uses the reference personalized parameters to convert the local content image according to the target style, and constructs a consistency loss with the target image to update the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator;
[0032] The server aggregation end is used to receive the parameters of the local visual encoder, the text encoder, the visual decoder and the discriminator sent by the client, perform parameter aggregation and send back.
[0033] On the other hand, the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the above method when the computer program / instruction is executed by a processor.
[0034] On the other hand, the present invention also provides a computer program product, comprising a computer program / instruction, which implements the steps of the above method when executed by a processor.
[0035] The beneficial effects of the present invention are at least:
[0036] The image art style transfer model training method, system and computer program of the present invention do not interact with each other in the model training process, and will not cause privacy leakage. The model transfers the style of the original content image based on the text instructions that describe the target style conversion requirements, which can meet the requirements of the user's language expression style transfer needs. Each client performs personalized training locally to ensure the accuracy of the model in the art style transfer ability with reconstruction loss and discrimination loss. The local model is updated by interacting with other clients to personalize the retention layer and verify the style transfer results of local data to build a consistency loss, thereby improving the generalization ability of the model. The model parameters updated by the local training of each client are aggregated and redistributed at the server aggregation end, and multiple rounds of iterations are performed. The server aggregation end can obtain a model that can efficiently adapt to various art style transfer requirements.
[0037] Additional advantages, purposes, and features of the present invention will be described in part in the following description, and will become apparent to those skilled in the art after studying the following, or may be learned from the practice of the present invention. The purposes and other advantages of the present invention may be achieved and obtained by the structures specifically indicated in the specification and the accompanying drawings.
[0038] Those skilled in the art will appreciate that the objectives and advantages that can be achieved with the present invention are not limited to the above specific description, and the above and other objectives that can be achieved by the present invention will be more clearly understood from the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of the present application, and do not constitute a limitation of the present invention. In the drawings:
[0040] Figure 1 The figure is a flowchart of a method for training an image art style transfer model based on text instructions according to an embodiment of the present invention.
[0041] Figure 2 This is a logical schematic diagram of a method for training an image art style transfer model based on text instructions according to an embodiment of the present invention.
[0042] Figure 3 This is a logical schematic diagram of the federated learning model training and application method described in one embodiment of the present invention.
[0043] Figure 4 The figure is a schematic diagram of the model operation logic in the federated learning model training and application method described in one embodiment of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0045] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, only structures and / or processing steps closely related to the solutions according to the present invention are shown in the accompanying drawings, while other details that are not closely related to the present invention are omitted.
[0046] It should be emphasized that the term “include / comprises” when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0047] It should also be noted that, unless otherwise specified, the term “connection” herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0048] In mobile image style transfer applications, the collected and stored images usually contain a large number of user privacy images, and the use of attractive reference art pictures involves intellectual property issues. The traditional centralized training method of image data collection is difficult to meet data privacy and personalization requirements. In the application scenario of image style transfer, there are the following problems:
[0049] First, image-to-image style transfer methods are difficult to meet the personalized needs of users. Requiring users to provide target style images will severely limit the usability of the application, while the reference style images preset by the service provider lack flexibility due to limited quantity or style.
[0050] Second, in addition to the lack of flexibility, the user needs of mobile devices and the original images are different in terms of content, size, and style. The data exists in a non-independent and identically distributed form. Simply applying the federated learning framework directly can easily lead to poor style transfer performance during the collaborative optimization process.
[0051] Taking the above problems into consideration, it is necessary to design a training method and system for an image art style transfer model based on text instructions. The user end can flexibly control the expected style directly through language description to improve the work efficiency of computer vision solutions in mobile device style transfer scenarios.
[0052] Federated learning is a privacy-preserving computing method widely recognized by academia and industry. It aims to collaboratively optimize models without sharing original data. In the scenario of computer vision applications, different devices with the same modeling requirements can be trained collaboratively through a shared cloud server. Therefore, the federated learning method has shown great potential in the application of computer vision style transfer. At present, the research on federated learning for style transfer is very limited, and few of them can simultaneously consider the problems of text instruction control and non-independent and identically distributed training data.
[0053] The present invention provides a method for training an image art style transfer model based on text instructions, wherein the method is implemented by multiple clients, and the clients are all connected to a server aggregation end, such as Figure 1 As shown, each of the client terminals includes the following steps S101 to S106:
[0054] Step S101: based on the set link, a content image to be style-converted, a text instruction describing the target style conversion requirements, and an example desired style image are obtained.
[0055] Step S102: Figure 2 As shown, based on the visual encoder, image content features and style features are extracted from the content image, and based on the text encoder, instruction semantic features of the text instruction are extracted; based on the visual decoder, image content features and instruction semantic features are used as input and output a target image after style conversion, and based on the visual decoder, image content features and style features are used as input and output a reconstructed image; a reconstruction loss is established based on the reconstructed image and the original content image, and a discriminant loss is established for the target image and the expected style image that meets the target style using a discriminant, and the parameters of the visual encoder, text encoder, visual decoder and discriminator are updated by minimizing the reconstruction loss and the discriminant loss.
[0056] Step S103: AdaIN is used as the personalized retention layer of each client, and the parameters of the personalized retention layer are uploaded to the personalized parameter library. Each client randomly downloads the reference personalized parameters uploaded by other clients, and uses the reference personalized parameters to convert the local content image according to the target style, and builds a consistency loss with the target image to update the parameters of the visual encoder, text encoder, visual decoder and discriminator.
[0057] Step S104: Sending parameters of the local visual encoder, text encoder, visual decoder, and discriminator to the server aggregation end for parameter aggregation.
[0058] Step S105: receiving the updated parameters of the visual encoder, text encoder, visual decoder and discriminator aggregated and returned by the server aggregation end.
[0059] Step S106: After multiple rounds of iterations, the updated visual encoder, text encoder, and visual decoder are constructed into an image art style transfer model.
[0060] Steps S101 to S106 are used to implement on the client side, and the federated learning-based solution uses local data to update model parameters and aggregate parameters on the server aggregation side. In step S101, the setting link is the path for the user to provide content images, text instructions, and desired style images. In the specific implementation process, data can be acquired by associating the client album, camera, and microphone. The content image is the object for artistic style transfer, and a diversified content image can be used. In some embodiments, the content image is preprocessed, and the preprocessing includes format unification and data normalization. Format unification includes resizing the image (scaling or cropping the image to a fixed size required by the network), channel unification (ensuring that the image has a consistent number of channels, such as converting a grayscale image to RGB format), and data type conversion (converting image pixel values to the data type required by the network, such as float32). Data normalization includes pixel value normalization (mapping pixel values to a specific range, such as [0,1] or [-1,1], to improve the stability of network training), and color space standardization (for example, converting from BGR to RGB, or performing color adjustment to match the input of the pre-trained model). These steps can ensure the consistency of input data and reduce the impact of data deviation on model performance, thereby improving the effectiveness and efficiency of computer vision tasks.
[0061] The text instructions are target style conversion requirements expressed in natural language. In some cases, the text instructions may be required to be input in a certain format, and expressed one by one according to color description, style description, personal preference, etc.
[0062] The desired style images are used to provide reference for model training. The client can establish a database and provide pre-stored desired style images for users to choose from, or the users can directly upload them.
[0063] In step S102, in some embodiments, the visual encoder and the visual decoder may adopt the VGG-19 model, and the text encoder may adopt the BERT model. In other embodiments, the visual encoder and the visual decoder may also adopt the ResNet, DenseNet, MobileNet or EfficientNet model. The text encoder may also adopt the GPT series, RoBERTa or ALBERT model.
[0064] The client performs feature extraction and image style transfer through the visual encoder, text encoder, and visual decoder, which can achieve multimodal fusion. The text encoder introduces semantic features and directly maps language instructions to the image style transfer task, giving the model the ability to understand natural language instructions, thereby achieving flexible text-based style control. Compared with pure visual style transfer, this multimodal method can handle more diverse task requirements. The model does not need to be trained separately for each style, but dynamically adjusts the style output through text instructions, which is suitable for any style transfer task. Through the reconstruction loss constraint, it ensures that the generated image is as close to the original content image as possible in terms of content preservation, thereby improving content consistency. The discriminative loss can guide the style of the target image to be closer to the expected style, improving the realism and style accuracy of the generated image.
[0065] In some embodiments, the reconstruction loss is calculated using the L2 distance, and the calculation formula is:
[0066]
[0067] Among them, L r represents the reconstruction loss, represents the original content image, Represents the reconstructed image.
[0068] The discriminant loss is established for the target image and the desired style image that meets the target style by using the discriminator, including: randomly cropping the target image and the desired style image, and inputting the cropped image fragments and the text instructions into the discriminator for recognition. The calculation formula of the discriminant loss is:
[0069]
[0070] Among them, L D represents the discrimination loss, represents the image fragments obtained by randomly cropping the target image, C S represents the image fragments obtained by randomly cropping the desired style image, T represents the text instruction, represents the discriminator's discrimination result on the image fragments of the target image, D(C S ,T) represents the discrimination result of the image fragments of the expected style image.
[0071] In step S103, the core idea of step S103 is to introduce a personalized retention layer AdaIN (adaptive instance normalization) in distributed training to achieve personalized style transfer, and at the same time enhance the generalization ability of the model through consistency loss. In the specific implementation, first, each client uses AdaIN as a personalized retention layer. The parameters of this layer (mean μ and standard deviation σ) can capture the unique style statistics of the current client, and upload these parameters to a shared personalized parameter library. Subsequently, each client will randomly download the reference personalized parameters uploaded by other clients to convert the local content image according to the target style. The purpose of this step is to simulate the migration process of different styles through the exchange of personalized parameters across clients, and to improve the adaptability of the model to multi-style conversion.
[0072] In some embodiments, AdaIN is used as the personalized retention layer for each client, and the expression is:
[0073]
[0074] Among them, x represents the characteristics of the content image input by the corresponding network layer, y represents the characteristics of the text instruction input by the corresponding network layer, μ represents the mean, and σ represents the variance.
[0075] During the implementation process, the model constructs a consistency loss between the generated style transfer image and the target image generated by the local client. This consistency loss can constrain the generation result so that the style transfer output with reference to personalized parameters is consistent with the target style image in terms of content and style, thereby promoting the model to learn a wider range of style expression capabilities. Through this mechanism, the visual encoder, text encoder, visual decoder, and discriminator can collaboratively learn in cross-client style information, effectively improving the generalization performance and robustness of the model on different style transfer tasks.
[0076] Step S103 not only realizes the sharing of personalized styles between clients, but also strengthens the model's adaptability to multiple styles through consistency loss while maintaining the consistency of content features. This design can make full use of the diversity of distributed data, avoid over-reliance on a single style data, and ultimately make the model have stronger generalization capabilities, enabling more flexible and high-quality image style transfer.
[0077] In step S104, after completing the local model training, each client sends the updated parameters of the visual encoder, text encoder, visual decoder and discriminator to the server aggregation end. This step is based on the idea of federated learning. The local training of the client is performed on its own data, and the parameter update does not expose the specific image data or text data, thereby ensuring data privacy and security. In actual operation: each client performs several rounds of local training, minimizes the reconstruction loss, discrimination loss and consistency loss, and updates the model parameters. The model parameters (such as weights and biases) after local training are uploaded to the server aggregation end. The uploaded data only contains model parameters, not the original training data, thereby achieving data privacy protection. Through this mechanism, the server side can collect model parameters from multiple clients, realize the collaborative update of distributed training, and lay the foundation for the next step of parameter aggregation. In some embodiments,
[0078] In steps S105 and S106, the server aggregation end aggregates the parameters of the visual encoder, text encoder, visual decoder, and discriminator uploaded by each client in the form of average weighting. The server distributes the aggregated updated model parameters to all clients for the next round of local training.
[0079] In some embodiments, the method further includes: sharing the trained image art style transfer model externally through a preset link.
[0080] On the other hand, the present invention also provides an image art style transfer model training system based on text instructions, including multiple clients, each of which is connected to a server aggregation end;
[0081] The client includes: a multivariate acquisition module, a multivariate information extraction module, a generation model training module and a consistency alignment module.
[0082] The multi-element acquisition module is used to obtain the content image to be style-converted, the text instruction describing the target style conversion requirements, and the desired style image as an example based on the set link.
[0083] The multivariate information extraction model is used to extract image content features and style features from the content image based on the visual encoder, and to extract instruction semantic features from the text instruction based on the text encoder; based on the visual decoder, the image content features and instruction semantic features are used as input and output the target image after style conversion, and based on the visual decoder, the image content features and style features are used as input and output the reconstructed image.
[0084] The generative model training module is used to establish a reconstruction loss based on the reconstructed image and the original content image, and use the discriminator to establish a discriminant loss for the target image and the expected style image that meets the target style. The minimum reconstruction loss and discriminant loss are used to update the parameters of the visual encoder, text encoder, visual decoder and discriminator.
[0085] The consistency alignment module is used to use AdaIN as the personalized retention layer of each client, upload the parameters of the personalized retention layer to the personalized parameter library, and each client randomly downloads the reference personalized parameters uploaded by other clients, and uses the reference personalized parameters to convert the local content image according to the target style, and constructs a consistency loss with the target image to update the parameters of the visual encoder, text encoder, visual decoder and discriminator.
[0086] The server aggregation end is used to receive the parameters of the visual encoder, text encoder, visual decoder and discriminator sent by the client, aggregate the parameters and send them back.
[0087] The description of the text instruction-based image art style transfer model training system may refer to the above steps S101 to S106.
[0088] The present invention is described below in conjunction with a specific embodiment:
[0089] This embodiment proposes a method for training and applying a federated learning model that simultaneously solves text-guided style transfer and complex and diverse training data. This embodiment first addresses the problem of text instruction control transfer style, designs a text encoder and a feature comparison module, extracts features from the text instructions input by the user for style guidance in generating images, and the feature comparison module aligns the content structure and style representation of the original image and the generated image. The loss function encourages the network to generate content by encoding the pixel-level structure and semantics of the image to ensure that the image structure remains unchanged while presenting the expected style. Secondly, this patent designs a content comparison module to address the distribution problem of training data for different devices. The unique style parameters of different clients are retained through a personalized style layer. The optimization direction of different client models is constrained by comparing the content of the generated results of different style parameters. Information sharing can be performed without forgetting the knowledge learned in local training, thereby improving the efficiency of collaborative training.
[0090] This embodiment improves and optimizes the collaborative training method of the art style transfer model in the image style transfer scenario in the IoT mobile device. In a given IoT environment, the main goal is to perform arbitrary style transfer on the images in the user device according to text instructions. Such image data may contain sensitive information such as the user's private photos and life pictures. For example, the mobile device can be a smart phone or tablet computer, which captures images or video data through various image applications. These applications range from photo applications, beauty applications to film and television editing. This embodiment proposes a federal style transfer learning framework for IoT devices, which enables the style transfer model to be collaboratively learned between different user devices, and realizes arbitrary image style transfer tasks based on text instructions while ensuring the privacy and security of sensitive data of the device, so as to enhance the flexibility and practicality of the IoT style transfer method.
[0091] This embodiment processes the image and text information on the user's device, combines the federated learning technology to protect the device's data privacy when training the generated model, and uses generative artificial intelligence and image processing technology to achieve arbitrary image style transfer based on text instructions. The technical roadmap is shown in the attached Figure 3 shown.
[0092] The technical solution mainly includes multi-dimensional acquisition module, multi-dimensional information extraction module, generation model training module, consistency alignment module, federated parameter aggregation module, and style transfer application module.
[0093] The overall process is as follows:
[0094] Step 1. The multi-information acquisition module obtains content images, text instructions, expected style images and other information on the user device side through the camera, microphone, photo album and other applications, and transmits them to the multi-information extraction module.
[0095] Step 2. In the multivariate information extraction module, the visual semantic features of the content image and text instructions are extracted through the visual encoder and text encoder based on the convolutional neural network, and these feature information are input into the visual decoder for image reconstruction and target style image generation.
[0096] Step 3. In the generative model training module, the parameters of the visual encoder, text encoder, and visual decoder are updated based on the designed image reconstruction loss function and discriminant loss function to obtain updated model parameters.
[0097] Step 4. In the consistency alignment module, each client retains local personalized parameters through the personalized parameter layer, uploads the parameters to the cloud service to build a shared personalized parameter library, randomly selects personalized parameters of other clients from the personalized parameter library to construct images, and constructs a consistency loss function by comparing the content and style differences of images generated by different personalized parameters to alleviate the problem of performance degradation of the shared model caused by non-independent and identically distributed data.
[0098] Step 5. In the federated parameter aggregation module, the cloud server performs weighted aggregation of all model parameters uploaded by each client according to the amount of local data, and then delegates the aggregated updated model parameters to all clients. The above steps are repeated until the global model converges. The local client then performs image style transfer application through the image encoder, text encoder, and visual decoder.
[0099] Specifically, the client device obtains the original image, the style text instruction set and the reference style image. The image data is normalized in size through preprocessing, and the text instructions are stored in the current device in the form of a string to prepare for subsequent feature extraction.
[0100] In order to jointly train the style transfer model for all client devices, for each device, the style transfer model mainly adopts a generative fusion algorithm, which mainly includes three parts: image generation module, content comparison consistency module, and federal parameter aggregation module.
[0101] Image generation module:
[0102] (1) First, use the visual encoder (using VGG-19) to extract the original image content features f c, and style features s , use the text encoder (using BERT) to extract the semantic features f of the user input text instruction T t , and f c and f t As the network input of the visual decoder G (using VGG-19), the visual decoder fuses the two types of features and generates the target image It can be expressed as:
[0103]
[0104] In order to keep the image content unchanged during the style transfer process and prevent unexpected deformation and edge distortion, a reconstruction loss function L is proposed. r To ensure content consistency, the visual decoder reconstructs the image based on the original image content and style features to obtain the reconstructed image The reconstruction loss function can be expressed as
[0105]
[0106] Here, L2 distance is used to measure the pixel-level distance between the original image and the reconstructed image.
[0107] (2) To ensure that the style transfer direction of the model is strictly in accordance with the user text instructions, the present invention constructs an image discriminator (using VGG-16) to determine whether the image comes from the generated image or the reference style image. In order to reduce the computational pressure, the generated image and the reference style image S are randomly cropped, and the cropped image fragments are used as the input of the image discriminator network. The discriminator loss function can be expressed as:
[0108]
[0109] in, and C S They represent the cropping results of the generated image and the reference style image respectively, and the performance of the style transfer network is further improved by constraining the style similarity between the generated image and the reference style image.
[0110] Content comparison consistency module:
[0111] In the style transfer network, a regularization layer is usually added after each neural network convolution layer to improve the computational efficiency. Retaining different regularization layers for different styles can make the model more robust. The present invention uses AdaIN as the personalized retention layer for each client, which can be expressed as
[0112]
[0113] Among them, x and y represent the content features and style features of the network layer input, respectively, and μ and σ represent the mean and variance of the input features, respectively (hereinafter collectively referred to as personalized parameters). Since each client has its own personalized parameters μ and σ, these parameters also represent the unique migration style of the client. Therefore, after the client training is completed, the personalized parameters are uploaded to the cloud server, and the personalized parameters of other clients are randomly downloaded from the personalized parameter library to generate images. A consistency loss function is constructed to further ensure the generation consistency of different clients. The consistency loss function uses L2 distance to compare with the locally generated image to alleviate the problem of performance degradation of the shared model caused by non-independent and identically distributed data.
[0114] Federated parameter aggregation: The federated average weighted algorithm is used to weight the parameters of the visual encoder, text encoder, and visual decoder. The above steps are repeated until the three shared models converge. After the style transfer network training is completed, the client can directly input the original image and text instructions to generate the expected style image. The implementation process of any style transfer is as follows: Figure 4 shown.
[0115] Compared with the prior art, the main advantages of this embodiment are:
[0116] In computer vision, existing style transfer technologies usually only transfer by referring to style images, and it is difficult to perform arbitrary style transfer according to the user's text instructions. At the same time, the privacy and security of data are not considered during model training. The present invention proposes a training and application method for an image art style transfer model based on text instructions. By integrating visual images and text instructions, the style transfer model is optimized using data from different clients while protecting personal privacy. At the same time, considering the non-independent and identically distributed problem of data from different devices, language can be used to control the generation of arbitrary style images.
[0117] This embodiment is based on computer vision technology, semantic extraction technology, and image style transfer technology to realize the training and end-to-end application of arbitrary style transfer models based on text instructions, which can effectively improve the availability and flexibility of the image style transfer model while ensuring data privacy, thereby realizing image style transfer applications in different scenarios.
[0118] Corresponding to the above method, the present invention also provides an apparatus / system, which includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, the processor is used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the apparatus / system implements the steps of the method described above.
[0119] The embodiment of the present invention also provides a computer-readable storage medium on which a computer program is stored, and when the computer program is executed by a processor, the steps of the aforementioned edge computing server deployment method are implemented. The computer-readable storage medium can be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0120] In summary, the image art style transfer model training method, system and computer program described in the present invention do not interact with each client's local privacy data during the model training process, and will not cause privacy leakage. The model transfers the style of the original content image based on the text instructions that describe the target style conversion requirements, which can meet the requirements of the user's language expression style transfer needs. Each client performs personalized training locally to ensure the accuracy of the model in the art style transfer ability with reconstruction loss and discrimination loss. The local model is updated by interacting with other clients to personalize the retention layer and verify the style transfer results of local data to build a consistency loss, thereby improving the generalization ability of the model. The model parameters updated by the local training of each client are aggregated and redistributed at the server aggregation end, and multiple rounds of iterations are performed. The server aggregation end can obtain a model that can efficiently adapt to various art style transfer requirements.
[0121] It should be understood by those skilled in the art that the exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0122] It should be clear that the present invention is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present invention.
[0123] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with features of other embodiments or replace features of other embodiments.
[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the embodiments of the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for training an image art style transfer model based on text instructions, characterized in that: The method is implemented by multiple clients, each of which is connected to a server aggregation end, and each of the clients includes the following steps: Based on the set link, obtain a content image to be style-converted, a text instruction describing the target style conversion requirements, and a desired style image as an example; Extracting image content features and style features from the content image based on the visual encoder, and extracting instruction semantic features of the text instruction based on the text encoder; taking the image content features and the instruction semantic features as input and outputting a target image after style conversion based on the visual decoder, and taking the image content features and the style features as input and outputting a reconstructed image based on the visual decoder; establishing a reconstruction loss based on the reconstructed image and the original content image, and using a discriminator to establish a discriminant loss for the target image and an expected style image that meets the target style, and minimizing the reconstruction loss and the discriminant loss to perform parameter updates on the visual encoder, the text encoder, the visual decoder, and the discriminator; AdaIN is used as the personalized retention layer of each client, and the parameters of the personalized retention layer are uploaded to the personalized parameter library. Each client randomly downloads the reference personalized parameters uploaded by other clients, and uses the reference personalized parameters to convert the local content image according to the target style, and builds a consistency loss with the target image to update the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator; Sending the parameters of the local visual encoder, the text encoder, the visual decoder and the discriminator to the server aggregation end for parameter aggregation; Receive the updated parameters of the visual encoder, the text encoder, the visual decoder, and the discriminator returned by the server aggregation end; After multiple rounds of iterations, the updated visual encoder, the text encoder, and the visual decoder are constructed into an image art style transfer model.
2. The image art style transfer model training method based on text instructions according to claim 1, characterized in that: The visual encoder and the visual decoder adopt the VGG-19 model, and the text encoder adopts the BERT model.
3. The image art style transfer model training method based on text instructions according to claim 1, characterized in that: The reconstruction loss is calculated using the L2 distance, and the calculation formula is: Among them, L r represents the reconstruction loss, represents the original content image, represents the reconstructed image; Using a discriminator to establish a discriminative loss for the target image and an expected style image satisfying the target style, including: The target image and the desired style image are randomly cropped, and the cropped image fragments and the text instructions are input into the discriminator for recognition. The calculation formula of the discrimination loss is: L D =log(1-D(C o ,T))+log(D(C S ,T)) Among them, L D represents the discrimination loss, represents the image fragments obtained by randomly cropping the target image, C S represents the image fragments obtained by randomly cropping the desired style image, T represents the text instruction, represents the discriminant result of the image fragment of the target image, D(C S ,T) represents the discrimination result of the image fragments of the desired style image.
4. The image art style transfer model training method based on text instructions according to claim 1, characterized in that: AdaIN is used as the personalized retention layer of each client, and the expression is: Among them, x represents the characteristics of the content image input by the corresponding network layer, y represents the characteristics of the text instruction input by the corresponding network layer, μ represents the mean, and σ represents the variance.
5. The method for training an image art style transfer model based on text instructions according to claim 1, characterized in that: The server aggregation end aggregates the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator uploaded by each client in the form of average weighting.
6. The method for training an image art style transfer model based on text instructions according to claim 1, characterized in that: After obtaining the content image to be style-converted, the text instruction describing the target style conversion requirements, and the desired style image as an example based on the set link, it also includes: The content image is preprocessed, and the preprocessing includes format unification and data normalization.
7. The method for training an image art style transfer model based on text instructions according to claim 1, characterized in that: The method further includes: sharing the trained image art style transfer model externally through a preset link.
8. An image art style transfer model training system based on text instructions, characterized in that: It includes multiple clients, each of which is connected to the server aggregation end; The client comprises: A multi-element acquisition module, for acquiring a content image to be style-converted, a text instruction describing the target style conversion requirements, and an example desired style image based on a set link; A multivariate information extraction module is used to extract image content features and style features from the content image based on a visual encoder, and to extract instruction semantic features of the text instruction based on a text encoder; to use the image content features and the instruction semantic features as inputs and output a target image after style conversion based on a visual decoder, and to use the image content features and the style features as inputs and output a reconstructed image based on the visual decoder; Generate a model training module, which is used to establish a reconstruction loss based on the reconstructed image and the original content image, and use a discriminator to establish a discriminant loss for the target image and the desired style image that meets the target style, and minimize the reconstruction loss and the discriminant loss to update the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator; A consistency alignment module, used to use AdaIN as the personalized retention layer of each client, upload the parameters of the personalized retention layer to the personalized parameter library, each client randomly downloads the reference personalized parameters uploaded by other clients, and uses the reference personalized parameters to convert the local content image according to the target style, and constructs a consistency loss with the target image to update the parameters of the visual encoder, the text encoder, the visual decoder and the discriminator; The server aggregation end is used to receive the parameters of the local visual encoder, the text encoder, the visual decoder and the discriminator sent by the client, perform parameter aggregation and send back.
9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method as claimed in any one of claims 1 to 7 are implemented.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Multi-scene image style migration and edge calculation optimization method
CN120953099A