Model training method, device and equipment and readable storage medium

By using sample images and reference text containing text in the image processing model and combining text loss training model, the problem of text blurring and distortion in the prior art is solved, and higher text fidelity and readability are achieved.

CN120125965APending Publication Date: 2025-06-10ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192728.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing image stylization processing algorithms cannot distinguish different elements in the image when processing images, resulting in blurring and distorting text in the image, reducing the readability and clarity of text in the processed image.

Method used

By obtaining a sample image containing text and a reference text for describing the reference image style, the sample image and the reference text are input to the image processing model to be trained, a generated image conforming to the reference image style is obtained, and the text loss is determined based on the similarity between the first text contained in the sample image and the second text contained in the generated image, and the image processing model is trained based on the text loss.

Benefits of technology

It improves the text fidelity of the image stylization algorithm, ensures that the text information in the image remains consistent before and after the style transformation, and improves the readability and clarity of the text in the processed image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125965A_ABST
    Figure CN120125965A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, device and equipment and a readable storage medium, and the method comprises the steps: obtaining a sample image containing a text and a reference text used for describing the style of a reference image, inputting the sample image and the reference text into a to-be-trained image processing model, obtaining a generated image according with the style of the reference image, and carrying out the training of the generated image. And determining text loss at least according to the similarity between the first text contained in the sample image and the second text contained in the generated image, and training an image processing model according to the text loss. Visibly, the text information of the sample image is compared with the text information of the generated image after the image style transformation, and the image processing model is trained based on the text information, so that the character fidelity of the image stylization algorithm is improved, and the character information in the image before and after the style transformation is kept consistent, and the security of privacy data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and particularly to a model training method, apparatus, device, and readable storage medium. Background Art

[0002] With the increasing attention to privacy data, image processing technology has received extensive attention. In the current era of the rapid development of visual art and artificial intelligence technology, image stylization, as an important means of visual art transformation, has been widely applied in many fields such as art creation, advertising design, and social media.

[0003] An image usually contains various different elements, such as people, landscapes, and text. Currently, the image stylization processing algorithms usually do not distinguish different elements in the image when processing the image, but process all elements in the image without discrimination to achieve the transformation of the overall style of the image.

[0004] However, this non-discriminatory processing method may cause the blurring and distortion of the text in the image, reducing the readability and clarity of the text in the processed image. Summary of the Invention

[0005] This specification provides a model training method, apparatus, device, and readable storage medium to partially solve the above problems existing in the prior art.

[0006] This specification adopts the following technical solutions:

[0007] This specification provides a model training method, including:

[0008] Obtain a sample image containing text and a reference text for describing the style of a reference image;

[0009] Input the sample image and the reference text into an image processing model to be trained, and obtain a generated image that conforms to the style of the reference image;

[0010] Determine a text loss according to the similarity between the first text included in the sample image and the second text included in the generated image, and at least train the image processing model according to the text loss; wherein, the text loss is inversely proportional to the similarity between the first text and the second text.

[0011] This specification provides a model training apparatus, including:

[0012] An obtaining module, configured to obtain a sample image containing text and a reference text for describing the style of a reference image;

[0013] An image generation determination module, configured to input the sample image and the reference text into an image processing model to be trained, and obtain a generated image that conforms to the style of the reference image;

[0014] A training module, configured to determine a text loss according to the similarity between the first text included in the sample image and the second text included in the generated image, and at least train the image processing model according to the text loss; wherein, the text loss is inversely proportional to the similarity between the first text and the second text.

[0015] This specification provides a computer-readable storage medium, where the storage medium stores a computer program, and when the computer program is executed by a processor, the above model training method is implemented.

[0016] This specification provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above model training method is implemented.

[0017] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:

[0018] In the model training method provided in this specification, a sample image containing text and a reference text for describing the style of a reference image are obtained, the sample image and the reference text are input into an image processing model to be trained, and a generated image that conforms to the style of the reference image is obtained. At least according to the similarity between the first text included in the sample image and the second text included in the generated image, a text loss is determined, and the image processing model is trained according to the text loss. It can be seen that by comparing the text information of the sample image with the text information of the generated image after image style transformation, and training the image processing model with this, the text fidelity of the image stylization algorithm is improved, and the text information in the image before and after the style transformation is ensured to be consistent. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of this specification, and constitute a part of this specification. The illustrative embodiments of this specification and their descriptions are used to explain this specification and do not constitute an improper limitation of this specification. In the attached

[0020] In the figure:

[0021] Figure 1 is a schematic flowchart of a model training method in this specification;

[0022] Figure 2 is a schematic diagram of an image processing model in this specification;

[0023] Figure 3It is a schematic flowchart of a model training method in this specification;

[0024] Figure 4 It is a schematic diagram of a noise prediction network in this specification;

[0025] Figure 5 It is a schematic diagram of a model training device provided in this specification;

[0026] Figure 6 corresponding to the Figure 1 in this specification is a schematic diagram of an electronic device. Specific Embodiments

[0027] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all of them. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of this specification.

[0028] In addition, it should be noted that all actions of obtaining signals, information, or data in this specification are carried out on the premise of complying with the corresponding data protection regulations and policies of the location and with the authorization given by the owner of the corresponding device.

[0029] It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0030] As mentioned above, in the training process of current image stylization algorithms, the goals of algorithm training or optimization are usually global, aiming to make the generated image globally match the target style. Therefore, the image stylization algorithm will uniformly perform undifferentiated style processing on various different elements such as people, landscapes, and texts included in the input image using the same rules or methods to achieve the transformation of the overall image style. However, this may cause varying degrees of deformation or distortion to the texts in the image, which makes the text information in the image lose its original readability and accuracy during the stylization process, thereby affecting the transmission effect of the text information in the processed image.

[0031] For example, in the field of advertising design, designers can use image stylization algorithms to convert promotional posters containing products and product description texts into artistic styles that match the brand image, such as vintage style, modern style, etc., to better enhance the brand recognition and attractiveness, thereby attracting the attention of the target audience. However, the product description texts in the promotional posters may be deformed or distorted during the style conversion, resulting in poor readability of the product description texts, thus reducing the ability of the promotional posters to convey product information and weakening the promotional effect.

[0032] Another example is in deep learning. When the amount of sample data for model training is insufficient, data augmentation methods can be adopted to reasonably transform the existing training samples to generate new training samples, thereby expanding the sample data volume to improve the generalization ability of the model. These transformations usually keep the labels unchanged. For example, in the intelligent voucher business, single voucher images can be extended into multiple different-style images through image stylization, which can then be used for the training of various voucher tasks, enabling the model to recognize the information in the voucher images under different color and light conditions. However, voucher images usually contain a lot of relatively important text information. The text information in the voucher images processed by existing stylization algorithms is very likely to have problems of blurring and distortion. If the voucher images with blurred and distorted text information are directly used as training samples, it will greatly reduce the accuracy and training efficiency of the model.

[0033] Based on this, this specification provides an image processing method, which compares the text information of the sample image with the text information of the generated image after image style transformation, and trains the image processing model with this, so as to improve the text fidelity of the image stylization algorithm and ensure that the text information in the images before and after the style transformation remains consistent.

[0034] The following will, with reference to the accompanying drawings, detail the technical solutions provided by the embodiments of this specification.

[0035] Figure 1 It is a schematic flowchart of a model training method provided by this specification.

[0036] S100: Obtain a sample image containing text and a reference text for describing the reference image style.

[0037] In the embodiments of this specification, a model training method is provided. The execution process of this model training method can be executed by an electronic device such as a server for model training. In addition, after the image processing model training is completed, the electronic device for processing the image to be processed based on the trained image processing model and the electronic device for executing the model training method can be the same or different. This specification does not make a limitation on this.

[0038] In this specification, an image processing model is used to stylize the input image according to the image style described in the input text, so as to obtain an image that conforms to the image style described in the input text. Among them, the input image to the image processing model contains text. In order to prevent the text in the image stylized by the image processing model from being blurred and distorted due to style transformation, in this specification, during the training process of the image processing model, according to the text information contained in the sample image input to the image processing model to be trained and the text information contained in the generated image output by the image processing model to be trained, a text loss is determined, and the image processing model is trained according to the text loss, so that the image processing model can distinguish the text part and the non-text part in the image during image stylization, overcome the deficiencies of the current image stylization algorithm in terms of image text fidelity, and ensure that the text content can be relatively clearly and completely retained after image stylization.

[0039] Therefore, a sample image containing text is obtained as the training sample of the image processing model, and a reference text for describing the reference image style is obtained as the guiding condition for the image processing model to learn image stylization.

[0040] The sample image contains one or more text characters, and the text can be located at any position in the sample image. Moreover, this specification does not limit the proportion of the text occupying the sample image. Other elements such as people, animals, objects, and landscapes can also be included in the sample image.

[0041] The reference image style described by the reference text can be any existing type of artistic style or visual effect, such as the oil painting style, cartoon style, sketch style, etc. in the painting style, or the warm color style, cold color style, etc. in the color style, or the modern style and retro style in the design style, etc. This specification does not make any limitations in this regard. The reference text can be composed of several sentences describing the reference image style. The reference text can be constructed based on prior knowledge in the field of artistic style or visual effect, or can be extracted based on other images conforming to the reference image style. This specification does not make any limitations in this regard.

[0042] Generally speaking, the reference image style described by the reference text is different from the initial image style that the sample image conforms to. This difference can be different image styles within the same style category. For example, the reference image style is the oil painting style, and the initial image style that the sample image conforms to is the sketch style. It can also be different image styles belonging to different style categories. For example, the reference image style is the sketch style, and the initial image style that the sample image conforms to is the modern style.

[0043] S102: Input the sample image and the reference text into the image processing model to be trained, and obtain a generated image that conforms to the reference image style.

[0044] The image processing model adopted in this specification can be constructed by using one or more of any existing types of network structures for image processing, such as neural networks, diffusion models, generative adversarial networks, etc. The network structure included in the image processing model can be flexibly determined according to different application scenarios, and this specification does not limit this.

[0045] During the training process, both the sample image and the reference text are used as inputs to the image processing model. The network for feature extraction in the image processing model can extract the image features of the sample image and the text features of the reference text respectively. The image features can represent different elements and the structure of the elements captured from the sample image, and the text features can represent the style attributes such as color and texture in the reference image style extracted from the reference text. Thus, the network for generating images in the image processing model can output a new generated image based on the image features under the constraint of the style attributes of the reference image style represented by the text features.

[0046] S104: Determine the text loss at least according to the similarity between the first text included in the sample image and the second text included in the generated image, and train the image processing model according to the text loss, where the text loss is inversely proportional to the similarity between the first text and the second text.

[0047] In this specification, the purpose of the image processing model is to perform style conversion on the image input into the model, but does not change the content presented by the image. Therefore, in order to achieve styleization without changing the content presented by the image, the content loss and the style loss can be defined. Among them, the content loss can ensure the similarity in content between the generated image and the sample image, that is, the content presented by the generated image output by the image processing model and the content presented by the sample image. The style loss is used to ensure that the generated image conforms to the reference image style described in the reference text.

[0048] For example, if the sample image is a promotional poster of coffee, then the sample image contains coffee and the promotional text "Enjoy the mellow fragrance every moment", and the reference text is "Vintage style". Then, input the promotional poster and the reference text into the image processing model, and it is expected that the generated image output by the image processing model can be a coffee promotional poster that conforms to the vintage style, and the coffee content in the poster remains unchanged, and the promotional text part also remains unchanged.

[0049] For another example, in the field of deep learning, the sample image is a voucher image with a cold color tone, which contains voucher text, such as the accounting time "X year Y month Z day" and the accounting amount "100 yuan". If the reference text is in a "warm color tone style", then the voucher image and the reference text are input into an image processing model, and it is expected that the generated image output by the image processing model can be a voucher image with a warm color tone, and the voucher text contained in the voucher image remains unchanged.

[0050] On this basis, in order to improve the fidelity of the text contained in the generated image output by the image processing model and reduce the difference between the text information contained in the generated image and the text information contained in the sample image, the text loss can also be determined based on the similarity between the first text contained in the sample image and the second text contained in the generated image, and the text loss is introduced into the aforementioned style loss and content loss, and the image processing model is trained by combining the three. Thus, during the training process of the image processing model, it can not only learn the image processing ability to generate a generated image that is similar in content to the sample image and conforms to the reference image style described by the reference text, but also ensure that the text information in the generated image is similar to the text information in the sample image, improve the text clarity and readability in the image after image stylization processing, achieve text fidelity in image style processing, and expand the application scenarios of the image processing model in multiple fields.

[0051] Among them, the similarity between the first text contained in the sample image and the second text contained in the generated image can be determined based on the similarity calculation method after respectively extracting the first text and the second text through text recognition technologies such as Optical Character Recognition (OCR), or can be determined by a pre-trained image-text matching model, or can also be determined by any other type of image text recognition and similarity calculation method, and this specification does not limit this.

[0052] The similarity between the first text contained in the sample image and the second text contained in the generated image is in an inverse relationship. That is, the greater the similarity between the first text and the second text, the smaller the text loss; conversely, the smaller the similarity between the first text and the second text, the greater the text loss. Thus, when training the image processing model, the training objective can be set to minimize the text loss, so as to maximize the similarity between the first text and the second text as the training objective, and train the image processing model, so that the trained image processing model can ensure that the text information in the image before and after style transformation remains consistent.

[0053] In one or more embodiments of this specification, if the image processing model has been pre-trained based on other sample images and other texts, that is, the image processing model can output a generated image that conforms to the image style and content described in the text and is consistent with the input image, then it can be re-trained only based on the text loss to have the ability in text fidelity. If the image processing model has not been pre-trained, it can be trained based on the text loss, content loss, and style loss to simultaneously learn the abilities in text fidelity, style transformation, and content consistency.

[0054] In the model training method provided in this specification, a sample image containing text and a reference text for describing the style of a reference image are obtained, the sample image and the reference text are input into the image processing model to be trained, a generated image that conforms to the style of the reference image is obtained, the text loss is determined at least according to the similarity between the first text contained in the sample image and the second text contained in the generated image, and the image processing model is trained according to the text loss.

[0055] It can be seen that by comparing the text information of the sample image with the text information of the generated image after image style transformation and training the image processing model with this, the text fidelity of the image stylization algorithm is improved, and the text information in the image before and after style transformation is ensured to be consistent.

[0056] In one or more embodiments of this specification, Figure 1 The image processing model adopted in the illustrated embodiment can be constructed based on a diffusion model. Among them, the diffusion model is a generative model that can generate synthetic images. The diffusion model usually introduces noise (such as Gaussian noise) into the image and then trains with the goal of minimizing the noise to try to denoise the image to generate an image. The diffusion model can include the Stable Diffusion model (SD), Latent Diffusion Models (LDMs), etc. For the convenience of description, the following takes the construction of an image processing model based on the latent diffusion model as an example to illustrate the specific technical solution. Thus, Figure 1 The model architecture of the image processing model in the illustrated embodiment can be as Figure 2 shown. The image processing model includes an encoder, a noise prediction network, a text feature extraction network, and a decoder. Based on Figure 2 the image processing model shown, the steps shown in the foregoing step S102 can be implemented as follows. As Figure 3 shown:

[0057] S200: Input the sample image into the encoder to obtain the initial image features of the sample image.

[0058] In this specification, the encoder and decoder in the image processing model can be the encoder and decoder in a Variational Auto-encoder (VAE). The encoder can map the input image to the latent space to obtain initial image features, and the decoder can map the input features back to the original data space, that is, map the features in the latent space to the pixel space to obtain the generated image.

[0059] Thus, when the sample image is input into the image processing model, the sample image can be directly input into the encoder, and the encoder can extract the feature vector of the sample image, use it as the initial image features of the sample image, and perform subsequent steps.

[0060] S202: Determine the reference noise, and add noise to the initial image features according to the reference noise to obtain the first image features.

[0061] Specifically, the working process of the latent diffusion model is to gradually add noise to the initial image features after mapping the input image to the latent space through the encoder to obtain image features with noise. Then, the noise prediction network predicts the noise from the image features with noise, thereby realizing denoising of the image features. The denoised image features are input into the decoder to obtain the generated image.

[0062] Therefore, in this step, determine the reference noise, and add noise to the initial image features of the sample image according to the reference noise to obtain the first image features with noise. Among them, the reference noise can be sampled from the noise that follows a normal distribution. In addition, the diffusion steps can also be sampled to determine the number of times of adding noise to the initial image features. The reference noise can be Gaussian noise, and this specification does not limit this.

[0063] S204: Input the reference text into the text feature extraction network to obtain the reference style features.

[0064] The text feature extraction network is used to map the reference text to the latent space, map the reference image style described by the reference text to the reference style features, and introduce the reference style features into the process of the noise prediction network outputting the predicted noise. By introducing the information of the reference image style, it can be controlled that the generated image obtained based on the second image features in the subsequent steps can conform to the reference image style, thereby realizing the conversion of a specific image style.

[0065] The text feature extraction network in this specification can be composed of networks such as Word Embedding, Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), and Transformer. This specification does not limit this. The text feature extraction network can be jointly trained with other network structures in the image processing model based on the sample image and the reference text, or can be pre-trained based on a general corpus. This specification does not limit this.

[0066] S206: According to the first image feature and the reference style feature, obtain the predicted noise through the noise prediction network.

[0067] After that, input the first image feature and the reference style feature into the noise prediction network to obtain the predicted noise output by the noise prediction network. Among them, the noise prediction network can be constructed based on the U-Net architecture. The noise prediction network can include multiple downsampling layers and multiple upsampling layers. The downsampling layer is used to gradually reduce the resolution of the image feature to capture the high-level abstract features of the noise from the first image feature containing noise; the upsampling layer is used to gradually restore the resolution of the image feature. At the same time, the skip connections between the downsampling layer and the upsampling layer introduce the features output by the downsampling layer into the upsampling, which further enhances the model's ability to capture fine-grained features, thereby improving the accuracy of the predicted noise.

[0068] Through the predicted noise, the model can gradually remove the noise in the first image feature. The denoised second image feature is given to the decoder to obtain the generated image, rather than directly obtaining the final generated image. This method is easier to control the generation process because the denoising amplitude is smaller each time, and the generated image will be more stable.

[0069] In addition, the diffusion steps sampled when determining the reference noise in S202 can also be input into the noise prediction network, so that the noise prediction network outputs the predicted noise at the diffusion steps, enabling it to learn the denoising strategy at different diffusion steps.

[0070] S208: According to the predicted noise, perform denoising processing on the first image feature to obtain a second image feature.

[0071] Since the predicted noise obtained in S206 is predicted by the noise prediction network based on the first image feature containing noise under the guidance and constraint of the reference style feature, thus, the part of the predicted noise in the first image feature is removed, that is, denoising processing is performed on the first image feature, thereby obtaining a second image feature. Using the second image feature as the restored image feature without noise, the generated image is obtained through the decoder.

[0072] Since the second image noise is obtained based on the predicted noise output by the noise prediction network, and the noise prediction network introduces the reference style feature in the process of outputting the predicted noise, the generated image obtained based on the second image feature can be a generated image that has undergone style conversion processing and conforms to the style of the reference image.

[0073] S210: Input the second image feature into the decoder to obtain a generated image that conforms to the style of the reference image.

[0074] Based on Figure 3 In the embodiment shown, an image processing model is constructed based on the latent diffusion model. By mapping the sample image from the pixel space to the latent space, adding reference noise to the initial image feature to obtain the first image feature, and under the guidance of the reference style feature corresponding to the reference text, obtaining the predicted noise from the first image feature through the noise prediction network, denoising the first image feature with the predicted noise to obtain the second image feature, and thus obtaining a generated image that conforms to the style of the reference image. This diffusion operation performed in the latent space can greatly reduce the computational amount and generation time compared with directly performing the diffusion process in the high-dimensional pixel space. Moreover, the features learned in the latent space are more abstract and general, so it has better generalization ability when generating new images. Through step-by-step denoising and conditional input, the generated image is not only clear but also meets the style requirements of the reference image style, realizing precise control of the image style that the generated image conforms to.

[0075] In an optional embodiment of this specification, Figure 2 In the image processing model shown, the encoder, the text feature extraction network, and the decoder can be pre-trained based on general samples. Thus, in the process of training the image processing model, in fact, only the network parameters of the noise prediction network can be optimized, thereby saving the training time of the image processing model and improving the online speed of the image processing model. In this case, when performing the foregoing step S104, it can be specifically implemented through the following embodiments:

[0076] The first step: Input the sample image into a pre-trained text recognition model to obtain the first text included in the sample image.

[0077] Among them, the text recognition model is suitable for converting the text information in the image into a machine-readable text character sequence. The text recognition model can include a feature extraction layer for extracting global or local features in the image, a recurrent layer for capturing character sequence information, and an output layer for outputting the sequence information into the final text sequence. The training samples of the text recognition model can be synthetic images with text character labels or real image data containing text information in the real world. This specification does not limit this.

[0078] In summary, the text recognition model can extract text characters from the image input to the model and output a text sequence. Therefore, by inputting the sample image into the pre-trained text recognition model, the first text output by the model is obtained, and the first text is the text contained in the sample image.

[0079] Step 2: Input the generated image into the pre-trained text recognition model to obtain the second text contained in the generated image.

[0080] Similarly, input the generated image output by the image processing model into the same pre-trained text recognition model to obtain the second text output by the model, and the second text is the text contained in the generated image.

[0081] Step 3: Determine the similarity between the first text and the second text, and substitute the similarity between the first text and the second text into the first preset loss function to obtain the text loss, where the text loss is inversely proportional to the similarity between the first text and the second text.

[0082] After that, through the text similarity calculation algorithm, determine the similarity between the first text and the second text. The text similarity calculation algorithm can be a combination of any existing type of text feature extraction model such as the Bag of Words (BoW) and Word Embeddings and any existing type of similarity calculation algorithm such as cosine similarity, Euclidean distance, and edit distance. This specification does not make any limitations in this regard.

[0083] Obtain the first preset loss function. The first preset loss function can be any existing loss function such as the CTC (Connectionist Temporal Classification) loss, Mean Squared Error (MSE), cross-entropy loss, absolute error, hinge loss, and log loss, and can be determined according to the specific application scenario. This specification does not make any limitations in this regard.

[0084] The text loss is inversely proportional to the similarity between the first text and the second text. That is, the greater the similarity between the first text and the second text, the smaller the text loss. Conversely, the smaller the similarity between the first text and the second text, the greater the text loss. Thus, when training the noise prediction network according to the text loss, the minimization of the text loss can be used as the training objective, that is, the maximization of the similarity between the first text and the second text can be used as the training objective, so that the second image feature obtained after denoising processing based on the predicted noise obtained by the trained noise prediction network can generate a generated image with the same text information as the text information contained in the sample image.

[0085] Fourth step: Train the noise prediction network at least according to the text loss.

[0086] Similar to the aforementioned S104, details are not described here.

[0087] Optionally, in addition to using the solutions of the first to third steps of the foregoing embodiments to determine the similarity between the first text extracted from the sample image and the second text included in the generated image, computer vision technology can also be used to determine whether the second text in the generated image is consistent with the first text extracted from the sample image through feature matching.

[0088] In one or more embodiments of this specification, Figure 2 The shown image processing model is constructed based on the latent diffusion model, and the working principle of the latent diffusion model is actually to gradually add noise to the initial image features in the latent space, and use the noise prediction network to gradually remove the noise to generate new image features, and then obtain the generated image. Therefore, based on Figure 2 the model architecture of the shown image processing model and the foregoing training method for the noise prediction network, the noise prediction network also needs to learn how to gradually remove the noise added to the initial image features. Thus, the output of the noise prediction network is the predicted noise. Therefore, the loss used when training the noise prediction network also needs to introduce a loss related to the noise prediction accuracy, so that the noise prediction network simultaneously has the capabilities of noise prediction and text fidelity. These capabilities together enable the image processing model to perform high-quality image style conversion while ensuring the clarity and readability of the text in the image.

[0089] Specifically, after the third step of the foregoing embodiment:

[0090] First, substitute the difference between the predicted noise and the reference noise into the second preset loss function to obtain the noise loss.

[0091] Among them, the second preset loss function can also be any existing loss function such as Mean Squared Error (MSE), cross-entropy loss, absolute error, hinge loss, logarithmic loss, etc., which can be determined according to the specific application scenario, and this specification does not limit this. And the first preset loss function and the second preset loss function can be loss functions of the same type or different types.

[0092] The noise loss obtained by substituting the difference between the predicted noise and the reference noise into the second preset loss function is directly proportional to the difference between the predicted noise and the reference noise. That is, the greater the difference between the predicted noise and the reference noise, the greater the noise loss; conversely, the smaller the difference between the predicted noise and the reference noise, the smaller the noise loss.

[0093] Therefore, when training the noise prediction network based on the noise loss, the minimization of the noise loss can be used as the training objective to optimize the network parameters of the noise prediction network, so as to realize the training of the noise prediction network. And the minimization of the noise loss is the minimization of the difference between the predicted noise and the reference noise. Therefore, the noise prediction network can learn the ability to output accurate predicted noise, so as to improve the restoration degree of the image generated based on the denoised image features subsequently.

[0094] After that, the total loss is determined according to the text loss and the noise loss.

[0095] Specifically, the text loss and the noise loss can be added to obtain the total loss, and the noise prediction network is trained with the total loss, so that it has the ability to output accurate predicted noise and the ability to generate a generated image with the same text information as the text information included in the sample image.

[0096] In addition, the weight corresponding to the text loss can be determined, and the weight corresponding to the noise loss can be determined, and the text loss and the noise loss are weighted respectively according to the corresponding weights, and the weighted text loss and the weighted noise loss are added to obtain the total loss, where the weight corresponding to the text loss is used to adjust the proportion of the text loss in the total loss, and the weight corresponding to the noise loss is used to adjust the proportion of the noise loss in the total loss.

[0097] Therefore, the fourth step of the foregoing embodiment can specifically be: training the noise prediction network according to the total loss. That is, with the minimization of the total loss as the training objective, the network parameters of the noise prediction network are optimized to realize the training of the noise prediction network.

[0098] In the embodiments of this specification, the noise prediction network can be constructed by a U-Net architecture. Therefore, the noise prediction network includes a plurality of downsampling layers and a plurality of upsampling layers. The plurality of downsampling layers therein constitute a contraction path, and the plurality of upsampling layers constitute an expansion path, so that the noise prediction network can capture higher-level features (obtained by downsampling), and at the same time retain the original spatial information (restored by upsampling). In addition, in each downsampling layer, detailed information of the first image feature containing noise will be captured, and then these information will be transmitted to the corresponding upsampling layer through skip connections. This helps to solve the problem of gradient disappearance and allows the upsampling part to utilize the high-resolution features from the early layers, thereby improving the prediction accuracy. The number of downsampling layers and the number of upsampling layers are usually the same. The downsampling layer reduces the resolution of the input feature (image), and the upsampling layer increases the resolution of the feature (image). The plurality of downsampling layers are arranged in the order of decreasing output resolution in turn. Similarly, the plurality of upsampling layers are arranged in the order of increasing output resolution in turn.

[0099] Based on this, the above step S206 can be specifically implemented through the following embodiments:

[0100] First, input the first image feature into the first downsampling layer included in the noise prediction network to obtain the first intermediate feature output by the first downsampling layer.

[0101] Secondly, for the next downsampling layer of the first downsampling layer among the multiple downsampling layers included in the noise prediction network, take the next downsampling layer of the first downsampling layer as the current downsampling layer, input the first intermediate feature output by the previous downsampling layer of the current downsampling layer into the current downsampling layer, and obtain the first intermediate feature output by the current downsampling layer.

[0102] Among them, the resolution of the first intermediate feature output by the previous downsampling layer is greater than the resolution of the first intermediate feature output by the current downsampling layer.

[0103] After that, determine whether there is a next downsampling layer of the current downsampling. If so, take the next downsampling layer of the current downsampling as the current downsampling layer again, and re - execute the above downsampling steps until all downsampling layers are traversed, and obtain the first intermediate feature output by the last downsampling layer included in the noise prediction network.

[0104] Then, take the first intermediate feature output by the last downsampling layer included in the noise prediction network and the style feature as inputs, and input them into the first upsampling layer included in the noise prediction network to obtain the second intermediate feature output by the first upsampling layer.

[0105] Furthermore, successively for each upsampling layer after the first upsampling layer among the multiple upsampling layers included in the noise prediction network, take the next upsampling layer of the first upsampling layer as the current upsampling layer, and take the second intermediate feature output by the previous upsampling layer of the current upsampling layer, the style feature, and the first intermediate feature output by the downsampling layer corresponding to the current upsampling layer as inputs, and input them into the current upsampling layer to obtain the second intermediate feature output by the current upsampling layer.

[0106] Among them, the resolution of the second intermediate feature output by the previous upsampling layer is less than the resolution of the second intermediate feature output by the current upsampling layer, and the resolution of the first intermediate feature output by the downsampling layer corresponding to the current upsampling layer is the same as the resolution of the second intermediate feature output by the previous upsampling layer of the current upsampling layer.

[0107] Therefore, it is determined whether there is a next upsampling layer of the current upsampling layer. If so, the next upsampling layer of the current upsampling layer is re - taken as the current upsampling layer, and the above upsampling steps are re - executed until all upsampling layers are traversed, and the second intermediate feature output by the last upsampling layer included in the noise prediction network is obtained.

[0108] Finally, the predicted noise is obtained based on the second intermediate feature output by the last upsampling layer and the first image feature.

[0109] As Figure 4 The noise prediction network shown includes three downsampling layers and three upsampling layers. The downsampling layers are arranged in the order of decreasing output feature resolution, and the upsampling layers are arranged in the order of increasing output feature resolution. Assume that the resolution of the first image feature is 512×512. After passing through the first downsampling layer, the resolution of the first intermediate feature output is 256×256. Each downsampling layer is traversed in turn. Finally, the resolution of the first intermediate feature output by the last downsampling layer is 64×64. Then, the first intermediate feature output by the last downsampling layer and the reference style feature are input into the first upsampling layer, and the second intermediate feature output by the first upsampling layer is obtained, with a resolution of 128×128. For the second upsampling layer, its input includes the second intermediate feature output by the first upsampling layer (the previous upsampling layer), the reference style feature, and the first intermediate feature output by the corresponding downsampling layer, with a resolution of 128×128. Each upsampling layer is traversed in turn. Finally, the resolution of the first intermediate feature output by the last upsampling layer is 512×512, which is fused with the first image feature, and the predicted noise is determined based on the fusion result.

[0110] In one or more embodiments of this specification, based on Figure 1 the shown implementation manner, a trained image processing model can be obtained. The trained image processing model is deployed to a server. The server can respond to an image processing request and perform stylization processing on a to - be - processed image containing text based on the deployed trained image processing model, so as to achieve image style conversion while ensuring that the text included in the image remains unchanged. The specific implementation steps are as follows:

[0111] Step 1: In response to the image processing request, determine the to - be - processed image containing the target text and the description text for describing the target image style.

[0112] As described above, the server for deploying the trained image processing model may be the same as or different from the server for executing the model training method. When the server receives an image processing request, the image to be processed and the target text can be parsed from the image processing request. Among them, the image to be processed contains the target text, and the number of characters of the target text, its position in the image to be processed, and the proportion of the target text occupying the image to be processed are not limited in this specification.

[0113] The description text is used to describe the target image style, which is usually different from the original image processing style corresponding to the image to be processed.

[0114] For example, in the field of advertising design, the image to be processed is a promotional poster containing a coffee product and its promotional slogan. The original image style of the promotional poster is a modern style, then the description text can be "retro style", and the target image style it describes is the retro style, which is different from the modern style of the promotional poster.

[0115] Another example is in deep learning. The image to be processed is a voucher image containing voucher text taken under a cold light source. The original image style corresponding to the voucher image is a cold color tone. The voucher text, such as in an accounting voucher, contains text such as voucher number, amount, business summary, date, etc. The description text can be "warm color tone style", then the target image style described by this description text is different from the original cold color tone image style of the voucher image.

[0116] The second step: Input the image to be processed and the description text into the trained image processing model to obtain a target image that conforms to the target image style and contains the target text.

[0117] After that, input the image to be processed and the description text into the trained image processing model. The network for extracting features in the image processing model can extract the image features of the image to be processed and the text features of the description text respectively. The image features can represent different elements and the structure of the elements captured from the image to be processed, and the text features can represent the style attributes such as color and texture in the target image style extracted from the description text. Thus, the network for generating images in the image processing model can output a new generated image based on the image features under the constraint of the style attributes of the target image style represented by the text features. This generated image is the target image that conforms to the target image style, has the same image content as the image to be processed, and contains clear and complete target text, thereby achieving both image style conversion and text fidelity at the same time.

[0118] The third step: Execute the business according to the target image.

[0119] Furthermore, execute the corresponding business based on the target image that has undergone style conversion and contains complete and clear target text.

[0120] Taking the above as an example, in the field of advertising design, a promotional poster containing coffee products and coffee slogans is subjected to style conversion through a trained image processing model, from a modern style to a retro style, while retaining the slogans therein, so as to attract more consumers who like classics and nostalgia, make the poster more in line with the brand's retro image, and thus enhance the brand recognition and attractiveness.

[0121] For another example, in the field of deep learning, a cold-tone voucher image is converted into a warm-tone voucher image through an image processing model, so as to be used for the training of various voucher tasks. This not only expands the scale of the samples, but also enables the model to recognize the information in the voucher image under different color and light conditions.

[0122] In one or more embodiments of this specification, based on Figure 2 the shown image processing model, the above second step can be implemented through the following steps:

[0123] The first step: Input the image to be processed into the encoder of the trained image processing model to obtain the initial image features of the image to be processed.

[0124] Referring to the foregoing step S200, details are not described herein.

[0125] The second step: Determine the target noise, and perform noise addition processing on the initial image features according to the target noise to obtain the first image features of the image to be processed.

[0126] Similar to the foregoing step S202, the target noise can be sampled from the noise obeying the normal distribution. In addition, the diffusion steps can also be determined to determine the number of times of adding noise to the initial image features of the image to be processed.

[0127] The third step: Input the description text into the text feature extraction network of the image processing model to obtain the target style features.

[0128] Referring to the foregoing step S204, details are not described herein.

[0129] The fourth step: According to the first image features of the image to be processed and the target style features, obtain the predicted noise through the noise prediction network of the image processing model.

[0130] Similar to the foregoing step S206, details are not described herein.

[0131] The fifth step: According to the predicted noise, perform denoising processing on the first image features of the image to be processed to obtain the second image features of the image to be processed.

[0132] Similar to the foregoing step S208, details are not described herein.

[0133] Sixth step: Input the second image feature of the image to be processed into the decoder of the image processing model to obtain a target image that conforms to the target image style and contains the target text.

[0134] The above is the model training method provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding image processing device, as Figure 5 shown.

[0135] Figure 5 Schematic diagram of a model training device provided by this specification, specifically including:

[0136] An acquisition module 300, configured to acquire a sample image containing text and a reference text for describing the style of a reference image;

[0137] A generated image determination module 302, configured to input the sample image and the reference text into an image processing model to be trained, and obtain a generated image that conforms to the style of the reference image;

[0138] A training module 304, configured to determine a text loss according to the similarity between the first text included in the sample image and the second text included in the generated image, and at least train the image processing model according to the text loss; wherein, the text loss is inversely proportional to the similarity between the first text and the second text.

[0139] Optionally, the image processing model includes an encoder, a noise prediction network, a text feature extraction network, and a decoder;

[0140] Optionally, the generated image determination module 302 is specifically configured to input the sample image into the encoder to obtain an initial image feature of the sample image; determine a reference noise, and perform noise addition processing on the initial image feature according to the reference noise to obtain a first image feature; input the reference text into the text feature extraction network to obtain a reference style feature; obtain a predicted noise through the noise prediction network according to the first image feature and the reference style feature; perform noise removal processing on the first image feature according to the predicted noise to obtain a second image feature; and input the second image feature into the decoder to obtain a generated image that conforms to the style of the reference image.

[0141] Optionally, the encoder, the text feature extraction network, and the decoder in the image processing model are pre-trained;

[0142] Optionally, the training module 304 is specifically configured to input the sample image into a pre-trained text recognition model to obtain the first text included in the sample image; input the generated image into the pre-trained text recognition model to obtain the second text included in the generated image; determine the similarity between the first text and the second text, and substitute the similarity between the first text and the second text into a first preset loss function to obtain a text loss, where the text loss has an inverse relationship with the similarity between the first text and the second text; train the noise prediction network at least according to the text loss.

[0143] Optionally, the training module 304 is specifically configured to substitute the difference between the predicted noise and the reference noise into a second preset loss function to obtain a noise loss; determine a total loss according to the text loss and the noise loss; train the noise prediction network according to the total loss.

[0144] Optionally, the noise prediction network includes a plurality of downsampling layers and a plurality of upsampling layers;

[0145] Optionally, the generated image determination module 302 is specifically configured to iteratively execute: input the first intermediate feature output by the previous downsampling layer into the current downsampling layer to obtain the first intermediate feature output by the current downsampling layer until all downsampling layers are traversed; wherein the resolution of the first intermediate feature output by the previous downsampling layer is greater than the resolution of the first intermediate feature output by the current downsampling layer; the input of the first downsampling layer included in the noise prediction network is the first image feature; input the first intermediate feature output by the last downsampling layer included in the noise prediction network and the style feature as inputs into the first upsampling layer included in the noise prediction network to obtain the second intermediate feature output by the first upsampling layer; iteratively execute: input the second intermediate feature output by the previous upsampling layer, the style feature, and the first intermediate feature output by the downsampling layer corresponding to the current upsampling layer as inputs into the current upsampling layer to obtain the second intermediate feature output by the current upsampling layer until all upsampling layers are traversed; wherein the resolution of the second intermediate feature output by the previous upsampling layer is less than the resolution of the second intermediate feature output by the current upsampling layer, and the resolution of the first intermediate feature output by the downsampling layer corresponding to the current upsampling layer is the same as the resolution of the second intermediate feature output by the previous upsampling layer of the current upsampling layer; obtain the predicted noise according to the second intermediate feature output by the last upsampling layer included in the noise prediction network and the first image feature.

[0146] Optionally, the device further includes:

[0147] The target image determination module 308 is specifically configured to, in response to an image processing request, determine a to-be-processed image containing target text and a description text for describing the style of the target image; input the to-be-processed image and the description text into a trained image processing model to obtain a target image that conforms to the style of the target image and contains the target text; and perform operations according to the target image.

[0148] Optionally, the target image determination module 308 is specifically configured to input the to-be-processed image into an encoder of a trained image processing model to obtain initial image features of the to-be-processed image; determine target noise, and perform noise addition processing on the initial image features according to the target noise to obtain first image features of the to-be-processed image; input the description text into a text feature extraction network of the image processing model to obtain target style features; obtain predicted noise through a noise prediction network of the image processing model according to the first image features of the to-be-processed image and the target style features; perform denoising processing on the first image features of the to-be-processed image according to the predicted noise to obtain second image features of the to-be-processed image; and input the second image features of the to-be-processed image into a decoder of the image processing model to obtain a target image that conforms to the style of the target image and contains the target text.

[0149] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 model training method shown.

[0150] This specification also provides Figure 6 a schematic structural diagram of the electronic device shown. As Figure 6 described, at the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, other hardware required for other services may also be included. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1 model training method shown. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but may also be hardware or a logic device.

[0151] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL). There is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0152] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0153] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0154] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0155] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0156] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 a means for implementing the functions specified in one or more of the blocks or multiple blocks.

[0157] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 a means for implementing the functions specified in one or more of the blocks or multiple blocks.

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 a means for implementing the functions specified in one or more of the blocks or multiple blocks.

[0159] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0160] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0161] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0162] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0163] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0164] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0165] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0166] The above is only the embodiment of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included within the scope of the claims of this specification.

Claims

1. A model training method, the method comprising: Obtain a sample image containing text and a reference text for describing the style of the reference image; Inputting the sample image and the reference text into the image processing model to be trained to obtain a generated image that conforms to the style of the reference image; A text loss is determined based on the similarity between a first text included in the sample image and a second text included in the generated image, and the image processing model is trained at least based on the text loss; wherein the text loss is inversely proportional to the similarity between the first text and the second text.

2. The method according to claim 1, wherein the image processing model comprises an encoder, a noise prediction network, a text feature extraction network and a decoder; Inputting the sample image and the reference text into the image processing model to be trained to obtain a generated image that conforms to the style of the reference image, specifically includes: Inputting the sample image into the encoder to obtain initial image features of the sample image; Determining a reference noise, and performing noise addition processing on the initial image feature according to the reference noise to obtain a first image feature; Inputting the reference text into the text feature extraction network to obtain reference style features; According to the first image feature and the reference style feature, obtaining predicted noise through the noise prediction network; According to the predicted noise, denoising the first image feature to obtain a second image feature; The second image feature is input into the decoder to obtain a generated image that conforms to the style of the reference image.

3. The method according to claim 2, wherein the encoder, text feature extraction network and decoder in the image processing model are pre-trained; Determining a text loss according to a similarity between a first text included in the sample image and a second text included in the generated image, and training the image processing model at least according to the text loss, specifically comprising: Inputting the sample image into a pre-trained text recognition model to obtain a first text contained in the sample image; Inputting the generated image into the pre-trained text recognition model to obtain a second text contained in the generated image; Determine the similarity between the first text and the second text, and substitute the similarity between the first text and the second text into a first preset loss function to obtain a text loss, wherein the text loss is inversely proportional to the similarity between the first text and the second text; The noise prediction network is trained based on at least the text loss.

4. The method according to claim 3, wherein the noise prediction network is trained at least according to the text loss, and specifically comprises: Substituting the difference between the predicted noise and the reference noise into a second preset loss function to obtain a noise loss; determining a total loss based on the text loss and the noise loss; The noise prediction network is trained according to the total loss.

5. The method of claim 2, wherein the noise prediction network comprises a plurality of downsampling layers and a plurality of upsampling layers; According to the first image feature and the style feature, the predicted noise is obtained by the noise prediction network, specifically including: Iterative execution: inputting the first intermediate feature output by the previous downsampling layer into the current downsampling layer to obtain the first intermediate feature output by the current downsampling layer, until all downsampling layers are traversed; wherein the resolution of the first intermediate feature output by the previous downsampling layer is greater than the resolution of the first intermediate feature output by the current downsampling layer; the input of the first downsampling layer included in the noise prediction network is the first image feature; Taking the first intermediate feature output by the last downsampling layer included in the noise prediction network and the style feature as input, and inputting them into the first upsampling layer included in the noise prediction network, to obtain the second intermediate feature output by the first upsampling layer; Iterative execution: taking the second intermediate feature output by the previous upsampling layer, the style feature and the first intermediate feature output by the downsampling layer corresponding to the current upsampling layer as input, inputting them into the current upsampling layer, and obtaining the second intermediate feature output by the current upsampling layer, until all upsampling layers are traversed; wherein the resolution of the second intermediate feature output by the previous upsampling layer is smaller than the resolution of the second intermediate feature output by the current upsampling layer, and the resolution of the first intermediate feature output by the downsampling layer corresponding to the current upsampling layer is the same as the resolution of the second intermediate feature output by the previous upsampling layer of the current upsampling layer; The predicted noise is obtained according to the second intermediate feature output by the last upsampling layer included in the noise prediction network and the first image feature.

6. The method of claim 1, further comprising: In response to an image processing request, determining an image to be processed containing a target text and a description text for describing a style of the target image; Inputting the image to be processed and the description text into the trained image processing model to obtain a target image that conforms to the style of the target image and contains the target text; A service is executed according to the target image.

7. The method according to claim 6, wherein the image to be processed and the description text are input into a trained image processing model to obtain a target image that conforms to the style of the target image and contains the target text, specifically comprising: Inputting the image to be processed into an encoder of a trained image processing model to obtain initial image features of the image to be processed; Determine a target noise, and perform noise addition processing on the initial image feature according to the target noise to obtain a first image feature of the image to be processed; Inputting the description text into the text feature extraction network of the image processing model to obtain the target style feature; According to the first image feature of the image to be processed and the target style feature, obtaining predicted noise through the noise prediction network of the image processing model; According to the predicted noise, denoising is performed on the first image feature of the image to be processed to obtain the second image feature of the image to be processed; The second image feature of the image to be processed is input into the decoder of the image processing model to obtain a target image that conforms to the style of the target image and contains the target text.

8. A model training device, comprising: An acquisition module, used for acquiring a sample image containing text and a reference text for describing the style of the reference image; A generated image determination module, used for inputting the sample image and the reference text into the image processing model to be trained to obtain a generated image that conforms to the style of the reference image; A training module is used to determine a text loss based on the similarity between a first text contained in the sample image and a second text contained in the generated image, and to train the image processing model at least based on the text loss; wherein the text loss is inversely proportional to the similarity between the first text and the second text.

9. A computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.