Image Processing Method and Device Based on Residual Network

The residual network-based method enriches OCR model training data by rendering new text in non-text regions, improving model performance by enhancing background image diversity and relevance.

CN114155542BActive Publication Date: 2025-07-15DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111521277.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2025-07-15
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

The existing OCR algorithm cannot render new text to other locations in the image, resulting in the missing background image block information in the dataset, affecting the detection and recognition performance of the OCR model.

Method used

Using an image processing method based on the residual network, by acquiring the first text image, the second text image and the target background image, the trained residual network is used to transfer and fusion, and the target text image is generated, so as to realize text filling in any non-text area.

Benefits of technology

The proportion of background information of the data set image is increased, the problem of missing background image block information is solved, and the text detection capability of the OCR model is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155542B_ABST
    Figure CN114155542B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method and apparatus based on a residual network. Among them, the method includes: obtaining a first text image, a second text image, and a target background image, where the first text image at least includes a first text of a first style and a first background image, the second text image at least includes a second text of a second style and a second background image, and the target background image and the second background image belong to the same background image; using the trained residual network to perform style transfer on the first text image and the second text image to obtain a third text image, where the third text image at least includes the first text of the second style and the first background image; using the trained residual network to fuse the first text of the second style with the target background image to generate a target text image. The present invention solves the technical problem that the existing algorithms cannot render new texts to other positions, resulting in the lack of background image block information in the dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and image generation, and in particular, to an image processing method and apparatus based on a residual network. Background Technique

[0002] With the rapid improvement of computer performance and the advent of the big data era, OCR (Optical Character Recognition) technology, which mainly uses deep learning, has been widely applied to scenarios such as card and certificate analysis, traffic management, and bill recognition. As a data-driven technical direction, mainstream OCR algorithm models require a large amount of labeled data to ensure the improvement of model performance. However, in actual situations, there are few open-source data sets that meet the target scenarios; in addition, the method of manually labeling business data has a high cost. Compared with the method of manual labeling, the method of batch generating data using image text generation algorithms has the advantages of controllable quantity and low cost, and has been widely used in the industry.

[0003] In order to improve the detection and recognition performance of OCR algorithms, image text generation technology has become an essential pre-strategy for training OCR algorithm models. In related technologies, an image text generation method based on style transfer is provided, which can render text with a target font style onto the background image where the original text is located. However, this algorithm replaces the text at a fixed position, resulting in the loss of background image block information.

[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide an image processing method and apparatus based on a residual network to at least solve the technical problem that the existing algorithm cannot render new text to other positions, resulting in the loss of background image block information in the data set.

[0006] According to one aspect of the embodiments of the present invention, an image processing method based on a residual network is provided, including: obtaining a first text image, a second text image, and a target background image, where the first text image includes at least a first text with a first style and a first background image, the second text image includes at least a second text with a second style and a second background image, and the target background image and the second background image belong to the same background image; performing style transfer on the first text image and the second text image using a trained residual network to obtain a third text image, where the third text image includes at least the first text with the second style and the first background image; and fusing the first text with the second style with the target background image using the trained residual network to generate a target text image.

[0007] Optionally, obtaining the first text image includes: randomly obtaining a first text from a text collection; rendering the first text onto a first background image based on a first font file to generate a first text image, where the first font file is a font file corresponding to a first style.

[0008] Optionally, obtaining the target background image includes: obtaining annotation information of service data corresponding to a second text image, where the annotation information is used to characterize the position of the text included in the service data in the service data; randomly determining a target background image from the service data based on the annotation information.

[0009] Optionally, the above method further includes: constructing a training data set, where the training data set includes: a first training image, a second training image, and a third training image, where the first training image includes a first training text in a first style, the second training image includes a second training text in a second style, the third training image includes the first training text in the second style, and the background image in the third training image and the background image in the second training image are different image blocks of the same background image; training an initial residual network using the training data set to obtain a trained residual network.

[0010] Optionally, constructing the training data set includes: obtaining a first training text, a second training text, a preset background image, and a first training background image; performing block processing on the preset background image to obtain a plurality of image blocks; randomly determining a second training background image and a third training background image from the plurality of image blocks; rendering the first training text onto the first training background image based on a first font file to generate a first training image, where the first font file is a font file corresponding to a first style; rendering the second training text onto the second training background image based on a second font file to generate a second training image, where the second font file is a font file corresponding to a second style; rendering the first training text onto the third training background image based on the second font file to generate a third training image.

[0011] Optionally, the training data set further includes: a fourth training image and a first skeleton guidance image, where the fourth training image includes the first training text in the second style and the first training background image, and the first skeleton guidance image includes a text skeleton corresponding to the first training text in the second style.

[0012] Optionally, rendering the first training text onto the first training background image based on a second font file to generate a fourth training image; rendering the first training text in the first style using the second font file to obtain a rendered first training text; processing the text skeleton of the rendered first training text to generate a first skeleton guidance image.

[0013] Optionally, the initial residual network is trained using a training data set, and the trained residual network obtained includes: performing style transfer on the first training image and the second training image using the initial residual network to obtain a first generated image and a second skeleton-guided image; fusing the first generated image and the target background image using the initial residual network to generate a second generated image; determining a first loss function based on the first generated image and the fourth training image; determining a second loss function based on the second generated image and the third training image; determining a third loss function based on the first skeleton-guided image and the second skeleton-guided image; and adjusting the model parameters of the initial residual network based on the first loss function, the second loss function, and the third loss function to obtain a trained residual network.

[0014] According to another aspect of the embodiments of the present invention, there is also provided an image processing apparatus based on a residual network, including: an acquisition module, configured to acquire a first text image, a second text image, and a target background image, where the first text image at least includes a first text of a first style and a first background image, the second text image at least includes a second text of a second style and a second background image, and the second background image and the target background image belong to the same background image; a migration module, configured to perform style transfer on the first text image and the second text image using the trained residual network to obtain a third text image, where the third text image at least includes the first text of the second style and the first background image; and a fusion module, configured to fuse the first text of the second style with the target background image using the trained residual network to generate a target text image.

[0015] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned image processing method based on a residual network.

[0016] According to another aspect of the embodiments of the present invention, there is also provided a processor, where the processor is used to run a program, and when the program runs, it executes the above-mentioned image processing method based on a residual network.

[0017] In an embodiment of the present invention, by obtaining a first text image, a second text image, and a target background image, where the first text image includes at least a first text in a first style and a first background image, the second text image includes at least a second text in a second style and a second background image, and the target background image and the second background image belong to the same background image; using a trained residual network to perform style transfer on the first text image and the second text image to obtain a third text image, where the third text image includes at least the first text in the second style and the first background image; using the trained residual network to fuse the first text in the second style with the target background image to generate a target text image, the technical effect of filling text in any non-text area in the image is achieved, which can effectively increase the proportion of background information in the dataset images, and thus solves the technical problem that the existing algorithms cannot render new text to other positions, resulting in the lack of background image block information in the dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present invention and form a part of this application. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0019] Figure 1 is a schematic diagram of a text image generated by a Synthtext algorithm according to the related art;

[0020] Figure 2 is a schematic diagram of a structural diagram of an SRNet algorithm model according to the related art;

[0021] Figure 3 is a flowchart of an image processing method based on a residual network according to an embodiment of the present invention;

[0022] Figure 4 is a schematic diagram of a structural diagram of an optional residual network model according to an embodiment of the present invention;

[0023] Figure 5 is a schematic diagram of a flowchart for making an optional dataset according to an embodiment of the present invention;

[0024] Figure 6 is a structural diagram of an image processing device based on a residual network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0027] To improve the detection and recognition performance of the OCR algorithm, image text generation technology has become an essential pre-strategy for training the OCR algorithm model. Currently, the mainstream image text generation algorithms are mainly divided into the following types:

[0028] The first type can be an image text generation method based on image semantic information. The image text generation method represented by the Synthtext algorithm uses the semantic information of the image to calculate the gradient and depth information of the image, and then uses such information to determine the placement position and direction of the text in the background image. The Synthtext algorithm adopts a simple "pasting" idea and places randomly given text at a determined position in the background image. However, in actual business scenarios, the font styles such as the text color, font type, and text deformation method in the text image are determined. This type of algorithm cannot obtain the font style information in the business data through the pixel information of the image, resulting in a large visual difference between the placed text information and the semantic information of the background image. Its subjective effect is as Figure 1 shown.

[0029] The second one can be a single-character image text generation method based on the style transfer approach. Using the idea of GAN (Generative Adversarial Networks), the TET-GAN algorithm realizes single-character-based image text generation. This algorithm prepares original images and target images with different text contents, where the text style in the target image is inconsistent with that of the original text. Then, it uses the GAN training method to convert the text content in the original text image into the style of the text in the target image while ensuring that the text content remains unchanged. The TET-GAN algorithm can effectively complete the text style transfer task. However, the height of the input image of this algorithm is 256, which is quite different from the height of text images in the actual OCR business scenario (the height of text images required by general text detection and recognition models is generally 32). Therefore, its scalability is poor.

[0030] The third one can be an image text generation method based on a 3D engine. Such algorithms use a 3D engine to model the real-world scenario, then render the text into the 3D scene according to the light and shadow information in the scene, and finally restore the 3D scene containing the text to a 2D image. Since such algorithms lack 3D model files that conform to the actual scenario in applications and randomly specify the style of the rendered font, they cannot be widely used in business scenarios.

[0031] The fourth one can be a long-text image text generation method based on style transfer. Algorithms represented by SRNet (Steganalysis Residual Network) decouple the tasks of obtaining the target font style, removing the background text, and pasting the text, and can render the text with the target font style onto the background image where the original text is located, realizing the batch generation of text images at fixed positions in the background image. Figure 3 This realizes the batch generation of text images at fixed positions in the background image. Figure 2 Figure shows the schematic structure of the SRNet algorithm.

[0032] By Figure 2From the structure, it can be seen that the SRNet algorithm has completed the task of erasing the original text at fixed positions and rendering new text with different text contents and the initial text font styles at those positions onto those positions, that is, by batch replacing the text contents at fixed positions, the dataset is amplified. Since this method increases the richness of text data, it can effectively improve the recognition performance of the OCR model for text. However, since the algorithm cannot render the new text to other positions, it causes the lack of background image block information in the dataset. On the other hand, the detection sub-model of the OCR model realizes the bounding of the text content in the image by differentiating the foreground and background information in the text image. Therefore, when training the detection model, it is required that the text background images included in the training set be as rich as possible. In summary, the SRNet algorithm has model defects in improving the detection performance of the OCR model.

[0033] Embodiment 1

[0034] According to an embodiment of the present invention, an embodiment of an image processing method based on a residual network is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0035] Figure 3 is an image processing method based on a residual network according to an embodiment of the present invention, as Figure 3 shown, the method includes the following steps:

[0036] Step S302, obtain a first text image, a second text image, and a target background image;

[0037] Among them, the first text image includes at least the first text of the first style and the first background image, the second text image includes at least the second text of the second style and the second background image, and the target background image and the second background image belong to the same background image.

[0038] Specifically, the first text image may refer to a new text image of the replaced text provided by the user, and the second text image refers to the original text image in the business data that needs text replacement. The text in the text image (including the above-mentioned first text and second text) may refer to the specific content of the text in the image, and the style may refer to the format corresponding to the text content, such as text color, font type, and text deformation method, etc. The formats of the text in different text images may be the same or different, but usually they are different. The target background image may be a randomly selected new background image of the business data, and this background image is different from the background image in the original text image, but belongs to the same background image as the original text image.

[0039] In an alternative embodiment, the second text image may be a text image within a text annotation box in business data, the target background image may be randomly selected from a set of background images of the business data, and the first text image may also be generated by the user providing the first text and rendering it on a gray background image (i.e., the above-mentioned first background image).

[0040] Through the above steps, the text content, text format, and new background image of the new text image and the original text image can be obtained.

[0041] Step S304: Use the trained residual network to perform style transfer on the first text image and the second text image to obtain a third text image;

[0042] Among them, the third text image at least includes the first text in the second style and the first background image.

[0043] By observing business data, it can be seen that the area of the region containing non-text accounts for a relatively large proportion in an image. Therefore, in order to make full use of the background images in the annotation data, increase the proportion of background image information in the dataset, and further improve the text detection ability of the OCR model, the SRNet algorithm can be improved. The original SRNet algorithm is divided into three modules, namely a style transfer module for obtaining the first text in the second style, a text removal module for removing the text content in the original text image, and a fusion module for fusing the first text in the second style with a textless new background image. As Figure 4 shown, in order to increase the diversity of background image patches, the above-mentioned residual network removes the text removal module in the original SRNet algorithm; on the other hand, in order to improve the subjective consistency of the new text in the new background, the text-removed background in the input of the fusion module of the above-mentioned residual network is changed to the obtained real background.

[0044] Specifically, a above-mentioned residual network can be trained in advance using training data. As Figure 4 shown, after obtaining the new text image i_t and the original text image i_s (i.e., the original image text block i_s shown in Figure 1 ), the style transfer module in the network can be used to perform style transfer on the new text image i_t and the original text image i_s, that is, change the format of the text in the new text image i_t to the text format of the original text image i_s, to obtain a new text with the same text style as the original text image (i.e., the target style new text t_t in Figure 4 ).

[0045] It should be noted that as Figure 4As shown, the style transfer module performs style transfer on the new text image \(i_t\) and the original text image \(i_s\) to obtain the new text with the target style. Inputting the new text with the target style into the skeleton algorithm can extract the skeleton of the text in the new text with the target style \(t_t\).

[0046] Step S306: Use the trained residual network to fuse the first text in the second style with the target background image to generate a target text image.

[0047] Specifically, as Figure 4 shown, the fusion module in the above network can be used to fuse the new text after style transfer (i.e., the new text with the target style \(t_t\) in Figure 4 ) with the obtained new background image new_bg to generate a target text image (i.e., the new background new text image \(t_f\) in Figure 4 ). Specifically, the purpose of fusing the text with the background image can be achieved by rendering the new text onto the new background image.

[0048] In the embodiment of the present invention, by obtaining a first text image, a second text image, and a target background image, where the first text image at least includes a first text in a first style and a first background image, the second text image at least includes a second text in a second style and a second background image, and the target background image and the second background image belong to the same background image; using the trained residual network to perform style transfer on the first text image and the second text image to obtain a third text image, where the third text image at least includes a first text in the second style and a first background image; using the trained residual network to fuse the first text in the second style with the target background image to generate a target text image, the technical effect of filling in text for any non-text area in the image is achieved, which can effectively increase the proportion of background information in the dataset images, and further solves the technical problem that the existing algorithms cannot render new text to other positions, resulting in the lack of background image block information in the dataset.

[0049] Optionally, obtaining the first text image includes: randomly obtaining a first text from a text collection; rendering the first text onto the first background image based on a first font file to generate a first text image, where the first font file is the font file corresponding to the first style.

[0050] Specifically, in the data generation stage, first randomly select a text line text_t (i.e., the above-mentioned first text), then combine the standard font file font_t (i.e., the above-mentioned first font file) and the gray background image gray_bg (i.e., the above-mentioned first background image), and use the standard font rendering module render_t to perform rendering to obtain a new text image \(i_t\), that is, the first text image.

[0051] Optionally, obtaining the target background image includes: obtaining the annotation information of the service data corresponding to the second text image, where the annotation information is used to characterize the position of the text included in the service data in the service data; randomly determining the target background image from the service data based on the annotation information.

[0052] Specifically, the service data can be an image containing a lot of text, and the text content in the image can be boxed by humans or other algorithm models to obtain the annotation boxes of different texts, that is, the above-mentioned annotation information is obtained.

[0053] In an alternative embodiment, the background image in the service data can be obtained according to the coordinates of the annotation boxes of all the texts in the service data according to a certain logic, and then an image block can be randomly selected from the background image as the above-mentioned target background image (new_bg).

[0054] Optionally, the above method further includes: constructing a training data set, where the training data set includes: a first training image, a second training image, and a third training image, where the first training image contains the first training text of the first style, the second training image contains the second training text of the second style, the third training image contains the first training text of the second style, and the background image in the third training image and the background image in the second training image are different image blocks of the same background image; using the training data set to train the initial residual network to obtain a trained residual network.

[0055] Specifically, the first training image can be used as the new text image, and the text style of this text image is the first style; the second training image can be used as the original text image, and the text style of this text image is the second style. The third training image can be used as the new background new text image, and the background images of the third training image and the second training image belong to different image blocks in the same background image.

[0056] In order to adapt to the improvement strategy, the embodiments of the present invention can adjust the process of making the data set. Without changing the overall process, the selection of the background image can be changed in the embodiments of the present invention. In the original SRNet algorithm, the background image can be randomly selected, but in the improvement strategy, in order to ensure that all the background image blocks involved in one training process are located in one image, one background image can be divided into blocks, and two image blocks are randomly selected from the respective divided image blocks each time the background image is selected during the training phase. One is used as the old background for rendering the new text, that is, the background of the second training image, and the other is used as the new background for rendering the new text, that is, the background of the third training image.

[0057] Optionally, constructing a training dataset includes: obtaining a first training text, a second training text, a preset background image, and a first training background image; performing block processing on the preset background image to obtain a plurality of image blocks; randomly determining a second training background image and a third training background image from the plurality of image blocks; rendering the first training text onto the first training background image based on a first font file to generate a first training image, where the first font file is a font file corresponding to the first style; rendering the second training text onto the second training background image based on a second font file to generate a second training image, where the second font file is a font file corresponding to the second style; rendering the first training text onto the third training background image based on the second font file to generate a third training image.

[0058] Specifically, before constructing the training dataset, the dataset can be prepared in the above manner. Preparing the dataset requires preparing a background image bg without text, a gray background image gray_bg, a text corpus text (including the original text text_s and the new text text_t), a standard font file font_t, and a style font file font_s; in terms of code, a standard font rendering module render_t, a style font rendering module render_s, and a skeleton generation module sk need to be prepared. The difference between the two rendering modules is that the latter adds rich deformation operations (such as projective transformation, bending, etc.). After preparing the above files and modules, the dataset is made according to the Figure 5 process shown. The production process of the dataset can include four threads: the production process of the first training image, the production process of the second training image, the production of the target style new text image (i.e., the fourth training image), and the production process of the new background new text image (i.e., the third training image).

[0059] Specifically, the first training text obtained can be the first text text_t of the first text image, and the second training text can be the second text text_s of the second text image. The background image of the first training image can be the gray background image gray_bg.

[0060] Furthermore, the preset background image is subjected to block processing according to the annotation information of the business data. During the training phase, two images can be randomly selected from the above image blocks each time, one as the second training background image bg and one as the third training background image new_bg.

[0061] Specifically, the rendering module render_t is used to render the first training text text_t, the first font file font_t, and the first training background image gray_bg to generate the first training image i_t (i.e., the new text image); the rendering module render_s is used to render the second training text text_s, the style font file font_s, and the second training background image bg to generate the second training image (i.e., the original image text block); the rendering module render_s is used to render the first text text_t, the style font file font_s, and the first training background image gray_bg to generate the fourth training image (i.e., the target style new text image t_t), and at the same time, the skeleton generation module sk is used to perform skeleton extraction on the fourth training image to obtain the skeleton guidance image t_sk; the rendering module render_s is used to render the first text text_t, the style font file font_s, and the third training background image new_bg to generate the third training image (i.e., the new background new text image t_f).

[0062] Optionally, the training dataset further includes: the fourth training image and the skeleton guidance image, where the fourth training image includes the first training text of the second style and the first training background image, and the first skeleton guidance image includes the text skeleton corresponding to the first training text of the second style.

[0063] Specifically, skeleton extraction is also called binary image thinning. This algorithm can thin a connected region into a pixel width for feature extraction and target topology representation. The above-mentioned fourth training image includes the first text of the target style, that is, the first text of the first text image containing the text style of the second text image, and the skeleton guidance image includes the text skeleton corresponding to the first training text after being rendered by the second font file.

[0064] Optionally, based on the second font file, the first training text is rendered to the first training background image to generate the fourth training image; the first training text of the first style is rendered using the second font file to obtain the rendered first training text; the text skeleton of the rendered first training text is processed to generate the first skeleton guidance image.

[0065] Specifically, based on the second font file (i.e., the style font file font_s), the first training text in the new text image is rendered to the first training background image through the style font rendering module render_s to generate the fourth training image; the fourth training image is input into the skeleton algorithm to obtain the corresponding skeleton guidance image, that is, the first skeleton guidance image.

[0066] Optionally, the initial residual network is trained using a training dataset to obtain a trained residual network, including: performing style transfer on the first training image and the second training image using the initial residual network to obtain a first generated image and a second skeleton-guided image; fusing the first generated image and the target background image using the initial residual network to generate a second generated image; determining a first loss function based on the first generated image and the fourth training image; determining a second loss function based on the second generated image and the third training image; determining a third loss function based on the first skeleton-guided image and the second skeleton-guided image; adjusting the model parameters of the initial residual network based on the first loss function, the second loss function, and the third loss function to obtain a trained residual network.

[0067] Specifically, the initial residual network is trained using the above-prepared dataset. The input of the initial residual network is the first training image and the second training image, and the output is the first generated image, the second skeleton-guided image, and the second generated image. The first loss function of the initial residual network is obtained by comparing and analyzing the first generated image and the fourth training image; the second loss function of the initial residual network is obtained by comparing and analyzing the second generated image and the third training image; the third loss function of the initial residual network is obtained by comparing and analyzing the first skeleton-guided image and the second skeleton-guided image. The first, second, and third loss functions reflect the accuracy of the initial residual network. The model parameters are adjusted according to the first loss function, the second loss function, and the third loss function, and training is continued. Finally, a trained residual network model can be obtained, and this model meets the accuracy requirements.

[0068] It should be noted that the above first, second, and third loss functions can select specific function types according to actual needs, and the present invention does not make specific limitations on this.

[0069] Embodiment 2

[0070] According to another aspect of the embodiments of the present invention, there is also provided an image processing device based on a residual network. This device can execute the image processing method based on a residual network provided in the above embodiments. The specific implementation scheme and application scenario are the same as those in the above embodiments and will not be elaborated here.

[0071] Figure 6 is a schematic diagram of an image processing device based on a residual network according to an embodiment of the present invention. As Figure 6 shown, the device includes:

[0072] An acquisition module 60, configured to acquire a first text image, a second text image, and a target background image, where the first text image includes at least a first text of a first style and a first background image, the second text image includes at least a second text of a second style and a second background image, and the second background image and the target background image belong to the same background image;

[0073] A migration module 62, configured to perform style migration on the first text image and the second text image by using a trained residual network to obtain a third text image, where the third text image includes at least the first text in the second style and the first background image;

[0074] A fusion module 64, configured to fuse the first text in the second style with a target background image by using a trained residual network to generate a target text image.

[0075] Optionally, the acquisition of the first text image further includes: randomly acquiring a first text from a text set; rendering the first text onto a first background image based on a first font file to generate a first text image, where the first font file is a font file corresponding to the first style.

[0076] Optionally, the acquisition of the target background image includes: acquiring annotation information of service data corresponding to a second text image, where the annotation information is used to represent the position of the text included in the service data in the service data; randomly determining a target background image from the service data based on the annotation information.

[0077] Optionally, the above method further includes constructing a training data set, where the training data set includes: a first training image, a second training image, and a third training image, where the first training image includes a first training text in the first style, the second training image includes a second training text in the second style, the third training image includes a first training text in the second style, and the background image in the third training image and the background image in the second training image are different image blocks of the same background image; training an initial residual network by using the training data set to obtain a trained residual network.

[0078] Optionally, constructing the training data set includes: acquiring a first training text, a second training text, a preset background image, and a first training background image; performing block processing on the preset background image to obtain a plurality of image blocks; randomly determining a second training background image and a third training background image from the plurality of image blocks; rendering the first training text onto the first training background image based on a first font file to generate a first training image, where the first font file is a font file corresponding to the first style; rendering the second training text onto the second training background image based on a second font file to generate a second training image, where the second font file is a font file corresponding to the second style; rendering the first training text onto the third training background image based on the second font file to generate a third training image.

[0079] Optionally, the training data set further includes: a fourth training image and a first skeleton guidance image, wherein the fourth training image includes a first training text in a second style and a first training background image, and the first skeleton guidance image includes a text skeleton corresponding to the first training text in the second style.

[0080] Optionally, render the first training text onto the first training background image based on the second font file to generate a fourth training image; render the first training text in the first style using the second font file to obtain the rendered first training text; process the text skeleton of the rendered first training text to generate a first skeleton guidance image.

[0081] Optionally, use the training data set to train the initial residual network, and the trained residual network obtained includes: performing style transfer on the first training image and the second training image using the initial residual network to obtain a first generated image and a second skeleton guidance image; fusing the first generated image and the target background image using the initial residual network to generate a second generated image; determining a first loss function based on the first generated image and the fourth training image; determining a second loss function based on the second generated image and the third training image; determining a third loss function based on the first skeleton guidance image and the second skeleton guidance image; adjusting the model parameters of the initial residual network based on the first loss function, the second loss function, and the third loss function to obtain the trained residual network.

[0082] Embodiment 3

[0083] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, and the computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the image processing method in Embodiment 1.

[0084] Embodiment 4

[0085] According to another aspect of the embodiments of the present invention, there is also provided a processor, and the processor is used to run a program, wherein when the program runs, it executes the image processing method in Embodiment 1.

[0086] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.

[0087] In the above embodiments of the present invention, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0088] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0089] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0090] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0091] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0092] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. An image processing method based on a residual network, characterized in that Including: Obtain a first text image, a second text image, and a target background image, where the first text image at least includes a first text in a first style and a first background image, the second text image at least includes a second text in a second style and a second background image, and the target background image and the second background image both belong to the same background image; Use the trained residual network to perform style transfer on the first text image and the second text image to obtain a third text image, where the third text image at least includes the first text in the second style and the first background image; Use the trained residual network to fuse the first text in the second style with the target background image to generate a target text image; Among them, obtaining the target background image includes: Obtain the annotation information of the service data corresponding to the second text image, where the annotation information is used to characterize the position of the text included in the service data in the service data; Randomly determine the target background image from the service data based on the annotation information.

2. The method according to claim 1, wherein Obtaining the first text image includes: Randomly obtain the first text from the text set; Render the first text to the first background image based on a first font file to generate the first text image, where the first font file is the font file corresponding to the first style.

3. The method according to any one of claims 1 to 2, characterized in that The method further includes: Construct a training data set, where the training data set includes: a first training image, a second training image, and a third training image, where the first training image includes a first training text in the first style, the second training image includes a second training text in the second style, the third training image includes the first training text in the second style, and the background image in the third training image and the background image in the second training image are different image blocks of the same background image; Use the training data set to train the initial residual network to obtain the trained residual network.

4. The method according to claim 3, wherein Constructing the training data set includes: Obtain the first training text, the second training text, a preset background image, and a first training background image; Perform block processing on the preset background image to obtain a plurality of image blocks; Randomly determine a second training background image and a third training background image from the plurality of image blocks; Render the first training text to the first training background image based on a first font file to generate the first training image, where the first font file is the font file corresponding to the first style; Render the second training text to the second training background image based on a second font file to generate the second training image, where the second font file is the font file corresponding to the second style; Render the first training text to the third training background image based on the second font file to generate the third training image.

5. The method according to claim 4, wherein The training data set further includes: a fourth training image and a first skeleton guidance image, wherein the fourth training image includes the first training text in the second style and the first training background image, and the first skeleton guidance image includes a text skeleton corresponding to the first training text in the second style.

6. The method according to claim 5, wherein Constructing the training data set further includes: Rendering the first training text onto the first training background image based on the second font file to generate the fourth training image; Rendering the first training text in the first style using the second font file to obtain the rendered first training text; Processing the text skeleton of the rendered first training text to generate the first skeleton guidance image.

7. The method according to claim 5, wherein Training the initial residual network using the training data set to obtain the trained residual network includes: Performing style transfer on the first training image and the second training image using the initial residual network to obtain a first generated image and a second skeleton guidance image; Fusing the first generated image and the target background image using the initial residual network to generate a second generated image; Determining a first loss function based on the first generated image and the fourth training image; Determining a second loss function based on the second generated image and the third training image; Determining a third loss function based on the first skeleton guidance image and the second skeleton guidance image; Adjusting the model parameters of the initial residual network based on the first loss function, the second loss function, and the third loss function to obtain the trained residual network.

8. An image processing device based on a residual network, characterized in that, Including: An acquisition module, configured to acquire a first text image, a second text image, and a target background image, wherein the first text image includes at least a first text in a first style and a first background image, the second text image includes at least a second text in a second style and a second background image, and the second background image and the target background image belong to the same background image; A transfer module, configured to perform style transfer on the first text image and the second text image using the trained residual network to obtain a third text image, wherein the third text image includes at least the first text in the second style and the first background image; A fusion module, configured to fuse the first text in the second style with the target background image using the trained residual network to generate a target text image; The acquisition module is further configured to acquire annotation information of the service data corresponding to the second text image, wherein the annotation information is used to characterize the position of the text included in the service data in the service data; randomly determining the target background image from the service data based on the annotation information.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute the image processing method based on a residual network according to any one of claims 1 to 7.