Image processing method and device

By enhancing the text and non-text content features in the image and combining these features to build high-resolution images, the problem of poor display of text information in super-resolution images is solved, and higher image resolution and clearer text display are achieved.

CN120013765APending Publication Date: 2025-05-16LENOVO (BEIJING) LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510213674.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In remote office scenes such as video conferencing, the image area corresponding to the Chinese text information in the reconstructed super-resolution image is poorly displayed.

Method used

By determining the features of text content and non-text content in the image, these features are enhanced separately, and high-resolution images are constructed in combination with the enhanced features. Specific methods include processing features using matrices in text enhancement libraries and non-text enhancement libraries, and generating high-resolution images through feature fusion, interpolation, and overlay.

Benefits of technology

Improve the display effect of Chinese text content in super-resolution images, and enhance image resolution, especially the clarity of text areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013765A_ABST
    Figure CN120013765A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device. The method comprises the steps of determining a first feature of text content in a first image and a second feature of non-text content in the first image; performing enhancement processing on the first feature to obtain a third feature; performing enhancement processing on the second feature to obtain a fourth feature; and combining the third feature and the fourth feature to construct a second image, the image resolution of the second image being higher than the image resolution of the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an image processing method and device. Background Art

[0002] Image super-resolution refers to the technology of restoring a high-resolution image from a low-resolution image.

[0003] In video conferencing or other forms of remote office scenarios, after receiving the video image transmitted by the other end device, the terminal device needs to use the image super-resolution algorithm to restore the low-resolution video image to a high-resolution video image. However, the display effect of the image area corresponding to the text information in the reconstructed super-resolution image is currently poor. Summary of the invention

[0004] On the one hand, the present application provides an image processing method, comprising:

[0005] determining a first characteristic of text content in a first image and a second characteristic of non-text content in the first image;

[0006] Performing enhancement processing on the first feature to obtain a third feature;

[0007] Performing enhancement processing on the second feature to obtain a fourth feature;

[0008] The third feature and the fourth feature are combined to construct a second image, wherein the image resolution of the second image is higher than the image resolution of the first image.

[0009] In a possible implementation manner, combining the third feature and the fourth feature to construct the second image includes:

[0010] Fusing the third feature and the fourth feature, and obtaining a first target image based on the fused features;

[0011] interpolating the first image using a bicubic interpolation algorithm to obtain a second target image;

[0012] The first target image and the second target image are superimposed to obtain a second image.

[0013] In yet another possible implementation, the enhancing the first feature to obtain the third feature includes:

[0014] Processing the first feature using at least one text enhancement matrix in a text enhancement library to obtain a third feature, wherein each of the text enhancement matrices corresponds to a text image enhancement mode;

[0015] The step of enhancing the second feature to obtain a fourth feature includes:

[0016] The second feature is processed by using at least one non-text enhancement matrix in a non-text enhancement library to obtain a fourth feature, wherein each of the non-text enhancement matrices corresponds to a non-text image enhancement mode.

[0017] In yet another possible implementation, the method further includes: obtaining character mask information of the first image;

[0018] The determining a first feature of text content in the first image and a second feature of non-text content in the first image includes:

[0019] Determining a coding feature of the first image based on the first image and the character mask information;

[0020] In combination with the encoding features of the first image, a first feature of text content and a second feature of non-text content in the first image are determined.

[0021] In yet another possible implementation, determining the encoding feature of the first image based on the first image and the character mask information includes:

[0022] Determining encoding features of the first image using a multi-level feature encoding module based on the first image and the character mask information;

[0023] Each level of feature coding module determines the current level coding features and current level character mask features corresponding to the first image based on the previous level coding features and previous level character mask features output by the previous level feature coding module of the feature coding module;

[0024] When the feature encoding module is a first-level feature encoding module, the previous-level encoding feature is an image feature of the first image, and the previous-level character mask feature is a mask feature of the character mask information;

[0025] The coding features of this level output by the last level feature coding module are the coding features of the first image.

[0026] In another possible implementation, the feature encoding module at each level determines the current level encoding features and current level character mask features corresponding to the first image based on the previous level encoding features and previous level character mask features output by the previous level feature encoding module of the feature encoding module, including:

[0027] Each level of feature coding module performs multi-scale feature extraction and feature fusion processing on the previous level coding features output by the previous level feature coding module of the feature coding module to obtain multi-scale fusion features;

[0028] Based on the multi-scale fusion features and the previous-level character mask features output by the previous-level feature encoding module, current-level encoding features and current-level character mask features corresponding to the first image are determined.

[0029] In another possible implementation, determining the current-level coding features and the current-level character mask features corresponding to the first image based on the multi-scale fusion features and the previous-level character mask features output by the previous-level feature encoding module includes:

[0030] Performing feature enhancement processing on the previous-level character mask features output by the previous-level feature encoding module to obtain current-level character mask features;

[0031] The character mask features of this level and the multi-scale fusion features are fused to obtain the encoding features of this level.

[0032] In another possible implementation, the text enhancement library is obtained in the following manner:

[0033] Obtain multiple groups of text image data, each group of text image data includes: a first image sample with each text pixel point belonging to the text annotated and a text boundary map corresponding to the first image sample, the text boundary map including each boundary line belonging to the text area in the first image sample;

[0034] For each set of text image data, based on the first image sample and the text boundary map, using a pattern generation module to generate at least one text enhancement matrix;

[0035] Determining a first sample encoding feature of the first image sample based on the first image sample and a text boundary map corresponding to the first image sample;

[0036] Performing feature downsampling on the first sample coding feature to obtain a second sample coding feature;

[0037] Performing feature enhancement on the second sample encoding feature based on the at least one text enhancement matrix to obtain a third sample encoding feature;

[0038] generating a first predicted image sample based on the third sample encoding feature;

[0039] If it is determined based on the first predicted image sample and the first image sample that the set training target has not been achieved, adjust the parameters of the pattern generation module, and return to execute the operation of generating at least one text enhancement matrix based on the first image sample and the text boundary map using the pattern generation module until the set training target is achieved.

[0040] In another possible implementation, the non-text enhancement library is obtained in the following manner:

[0041] Obtaining multiple groups of non-text image data, each group of non-text image data includes: a second image sample with each non-text pixel point marked as non-text and a non-text boundary map corresponding to the second image sample, the non-text boundary map including each boundary line belonging to a text area in the second image sample;

[0042] For each set of non-text image data, generating at least one non-text enhancement matrix using a pattern generation module based on the second image sample and the non-text boundary map;

[0043] Determining a fourth sample encoding feature of the second image sample based on the second image sample and a non-text boundary map corresponding to the second image sample;

[0044] Performing feature downsampling on the fourth sample coding feature to obtain a fifth sample coding feature;

[0045] Performing feature enhancement on the fifth sample coding feature based on the at least one non-text enhancement matrix to obtain a sixth sample coding feature;

[0046] generating a second predicted image sample based on the sixth sample encoding feature;

[0047] If it is determined based on the second predicted image sample and the second image sample that the set training target has not been achieved, the parameters of the pattern generation module are adjusted, and the operation of generating at least one non-text enhancement matrix based on the second image sample and the non-text boundary map using the pattern generation module is returned to execution until the set training target is achieved.

[0048] In another aspect, the present application further provides an image processing device, comprising:

[0049] a feature determination unit, configured to determine a first feature of text content in a first image and a second feature of non-text content in the first image;

[0050] A first feature enhancement unit, used for performing enhancement processing on the first feature to obtain a third feature;

[0051] A second feature enhancement unit, used for enhancing the second feature to obtain a fourth feature;

[0052] An image construction unit is used to construct a second image by combining the third feature and the fourth feature, wherein the image resolution of the second image is higher than the image resolution of the first image. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.

[0054] Figure 1 A schematic diagram of a flow chart of the image processing method provided in this application;

[0055] Figure 2 A schematic diagram of an implementation process for constructing a text enhancement library in this application;

[0056] Figure 3 An example diagram of an implementation framework for constructing a text enhancement matrix in a text enhancement library in this application;

[0057] Figure 4 A schematic diagram of an implementation process for constructing a non-text enhancement library in this application;

[0058] Figure 5 A schematic diagram of another process flow of the image processing method provided by the present application;

[0059] Figure 6 An example diagram of an implementation principle framework of the image processing method provided in this application;

[0060] Figure 7 An example diagram of an implementation process framework of the image processing method provided in this application;

[0061] Figure 8 A schematic diagram of a composition architecture of a feature coding module provided in this application;

[0062] Fig. 9 A schematic diagram of an implementation flow of the feature coding module in this application determining the current level coding features and the current level character mask features;

[0063] Fig.10 A schematic diagram of the structure of an image processing device provided by the present application;

[0064] Fig.11 A schematic diagram of the composition architecture of an electronic device provided in this application. DETAILED DESCRIPTION

[0065] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application. The terms used in the implementation mode of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the present application. It is known to those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0066] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and need not be used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, which is only to describe the distinction mode adopted by the objects of the same attributes when describing in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.

[0067] like Figure 1 , showing a flow chart of the image processing method provided in an embodiment of the present application. The method of this embodiment can be applied to an electronic device, which may be a laptop computer, a desktop computer, a server, etc., or a node device in a cloud platform or a cluster system, etc., without limitation.

[0068] The method of this embodiment may include:

[0069] S101, determining a first feature of text content in a first image and a second feature of non-text content in the first image.

[0070] The first image is an image whose resolution is to be improved in the present application. In the present application, the first image contains text content and non-text content. The non-text content in the first image refers to the content other than the text content in the first image.

[0071] The first feature is a feature related to the text content in the first image. The first feature can reflect information about the text content in the first image, so that the text information in the first image can be taken into consideration when subsequently reconstructing a high-resolution image.

[0072] S102, enhancing the first feature to obtain a third feature.

[0073] Among them, by enhancing the first feature, the features of the text content in the first image can be enriched. For example, compared with the first feature, the third feature can show the features of the text content at more pixel points.

[0074] In the present application, there are many possibilities for the specific implementation of the enhancement processing of the first feature, which are not specifically limited.

[0075] S103, enhancing the second feature to obtain a fourth feature.

[0076] By enhancing the second feature, the features of the non-text content in the first image can be enriched. For example, compared with the second feature, the fourth feature can show the features of the non-text content on more pixels.

[0077] In the present application, there is no limitation on the specific implementation method of enhancing the second feature.

[0078] S104, constructing a second image by combining the third feature and the fourth feature.

[0079] The image resolution of the second image is higher than that of the first image. Based on this, the second image is actually a super-resolution image reconstructed based on the first image.

[0080] From the above content, it can be seen that in the process of constructing a second image with high image resolution based on the first image, the present application will respectively determine the first feature of the text content in the first image and the second feature of the non-text content, and enhance the first feature and the second feature respectively, so that the information of the text content can be fully considered in the process of constructing the second image with high image resolution, which naturally can improve the display effect of the text content in the constructed second image.

[0081] In the present application, there are many possibilities for the specific implementation of constructing the second image by combining the third feature and the fourth feature, and there is no restriction on this. For example, if the third feature and the fourth feature can respectively reflect the features of the image resolution improvement of the text content and the non-text content, then the third feature and the fourth feature can be fused, and the second image can be generated based on the fused features.

[0082] For another example, in order to reduce the complexity of processing text and non-text related features in the first image, in the present application, the third feature and the fourth feature can be used to represent the residual features required to increase the image resolution of the text content and the non-text content, respectively. Based on this, the present application can fuse the third feature and the fourth feature, and obtain the first target image based on the fused features. In addition, the bicubic interpolation algorithm is used to interpolate the first image to obtain the second target image. On this basis, the first target image and the second target image are superimposed to obtain the second image.

[0083] In the present application, there are multiple possible implementations of the specific implementation of enhancing the first feature and the second feature. In one possible implementation, the present application can pre-build a text enhancement library for enhancing features of text content and a non-text enhancement library for enhancing non-text content.

[0084] The text enhancement library includes: at least one text enhancement matrix, each text enhancement matrix corresponds to a text image enhancement mode. Based on this, different text enhancement matrices in the text enhancement library can be combined to enhance the features of the text content in different modes.

[0085] Similarly, the non-text enhancement library includes at least one non-text enhancement matrix, and each non-text enhancement matrix corresponds to a non-text image enhancement mode. Accordingly, based on different non-text enhancement matrices, the features of non-text content can be enhanced in different modes.

[0086] On this basis, various text image enhancement modes in the text enhancement library and various non-text image enhancement modes in the non-text enhancement library can be used as structural priors to guide the enhancement processing of the first feature and the second feature. Accordingly, the present application can use at least one text enhancement matrix in the text enhancement library to process the first feature to obtain the third feature. The second feature can be processed using at least one non-text enhancement matrix in the non-text enhancement library to obtain the fourth feature.

[0087] For example, since the first feature and the second feature are also essentially a matrix, each text enhancement matrix in the text enhancement library can be matrix multiplied with the first feature, and then the matrix multiplication results of each text enhancement matrix and the first feature can be added to obtain a third matrix. Similarly, each non-text enhancement matrix in the non-text enhancement library can be matrix multiplied with the second feature, and then the matrix multiplication results of each non-text enhancement matrix and the second feature can be added to obtain a fourth matrix.

[0088] In this possible implementation, the first feature determined by the present application for the text content in the first image can not only reflect the features of the text content of the first image, but also reflect the degree to which the text content of the first image is suitable for various text image enhancement modes in the text enhancement library. Similarly, the second feature can not only reflect the features of the non-text content in the second image, but also reflect the degree to which the non-text content in the first image is suitable for various non-text image enhancement modes in the non-text enhancement library.

[0089] On this basis, the third feature can be regarded as the weighted average result of enhancing the first feature using each text enhancement matrix, and the fourth feature can be regarded as the weighted average result of enhancing the second feature using each non-text enhancement matrix.

[0090] In the present application, there is no restriction on the specific implementation of constructing the text enhancement library and the non-text enhancement library. The construction process of each text enhancement matrix in the text enhancement library and the non-text enhancement matrix in the non-text enhancement library is described below in a possible implementation manner.

[0091] like Figure 2 , shows a schematic diagram of an implementation process of building a text enhancement library in this application, and this embodiment may include:

[0092] S201, obtaining multiple groups of text image data.

[0093] Each set of text image data includes: a first image sample with each text pixel belonging to the text marked and a text boundary map corresponding to the first image sample. The text boundary map includes each boundary line belonging to the text area in the first image sample.

[0094] For example, the present application can pre-screen multiple high-resolution image samples, where the number of pixels in the high-resolution image samples exceeds the set number of image samples. On this basis, the pixels belonging to the text in each image sample can be pre-labeled to obtain a first image sample with only text pixels and other pixels having values ​​of 0.

[0095] For example, the present application can obtain multiple image samples and the text classification mask and boundary map corresponding to each image sample. The text classification mask is used to distinguish the text content in the image sample. Based on this, the image sample is masked using the text classification mask to obtain the first image sample with each text pixel marked. The boundary map corresponding to the image sample is the boundary map of the first image sample corresponding to the image sample.

[0096] S202: For each set of text image data, based on the first image sample with each text pixel marked and the text boundary map, generate at least one text enhancement matrix using a pattern generation module.

[0097] The pattern generation module may use a neural network or other deep learning model, without limitation. For example, in one possible implementation, the network architecture of the pattern generation module may be a tree-like convolutional network, for example, with a depth of Tree-structured convolutional network.

[0098] In a possible implementation, in order to reduce the amount of data processing of the pattern generation module and improve the operating efficiency of the pattern generation module, before using the pattern generation module to generate a text enhancement matrix, the first image sample in which each text pixel is annotated is calculated, and the pixel average and pixel standard deviation corresponding to the first image sample are calculated. In addition, the pixel average and pixel standard deviation corresponding to the text boundary map are calculated. On this basis, the pixel average and pixel standard deviation corresponding to the first image sample and the pixel average and pixel standard deviation corresponding to the text boundary map can be input into the pattern generation module to obtain at least one text enhancement matrix generated by the pattern generation module.

[0099] The average value of the pixels corresponding to the first image sample includes the average value of each pixel in the first image sample, and the average value of the pixel is the average value between the pixel and each of its adjacent pixels. Similarly, the average value of the pixels corresponding to the text boundary map includes the average value of each pixel in the text boundary map.

[0100] The standard deviation of the pixels corresponding to the first image sample is the standard deviation calculated based on the average value of each pixel in the first image sample. The standard deviation of the pixels corresponding to the text boundary map is the standard deviation calculated based on the average value of each pixel in the text boundary map.

[0101] S203: Determine a first sample encoding feature of the first image sample based on the first image sample and a text boundary map corresponding to the first image sample.

[0102] For example, the first image sample with the text pixels marked and the corresponding text boundary map are input into the first encoder to obtain the first sample encoder feature. The first encoder may include a lightweight feature extractor and two matrix operation modules. For example, the first encoder may be a convolutional neural network model or other types of neural network models, without limitation.

[0103] S204, down-sampling the first sample coding feature to obtain a second sample coding feature.

[0104] By downsampling the first sample coding feature, the spatial resolution of the image generated based on the first sample coding feature can be reduced. Based on this, the image resolution corresponding to the image generated based on the second sample coding feature is lower than the image resolution corresponding to the image generated based on the first sample coding feature.

[0105] S205: Perform feature enhancement on the second sample encoding feature based on at least one text enhancement matrix to obtain a third sample encoding feature.

[0106] It can be understood that the purpose of performing feature enhancement on the second sample coding feature based on at least one text enhancement matrix is ​​to increase the image resolution of the image that can be generated by the enhanced second sample coding feature (i.e., the third sample coding feature) so that a high-resolution image can be generated subsequently.

[0107] Among them, the specific implementation of using the at least one text enhancement matrix to enhance the feature of the second sample encoding feature can refer to the previous introduction of using the text enhancement matrix to enhance the feature of the first feature, which will not be repeated here.

[0108] S206: Generate a first predicted image sample based on the third sample encoding feature.

[0109] For example, the third sample encoding feature is input into the first decoder, and the first predicted image sample is output.

[0110] S207, based on the first predicted image sample and the first image sample, determine whether the set training goal is achieved. If not, adjust the parameters of the pattern generation module and return to execute S202 operation; if yes, the training is completed, and the currently obtained at least one text enhancement matrix is ​​determined as at least one text enhancement matrix in the text enhancement library.

[0111] It can be understood that after downsampling the first sample coding feature of the first image sample, the purpose of feature enhancement of the second sample coding feature obtained by downsampling is to obtain a third sample coding feature that can restore the first image sample. It can be seen that if the accuracy of the pattern generation module is high, then after the second sample coding feature is enhanced by using at least one text enhancement matrix generated by the pattern generation module to obtain the third sample coding feature, the first predicted image sample generated based on the third sample coding feature should be a high-resolution image, and the first predicted image sample should be as consistent as possible with the first image sample.

[0112] Based on this, the training goal of the training pattern generation module is to minimize the gap between the first image sample and the first predicted image sample corresponding to the first image sample. For example, the set training goal can be that the number of training iterations exceeds the set number, or the gap value between each first image sample and the first predicted image sample corresponding to it is minimized, or the function value of the set first loss function converges. The smaller the gap between each first image sample and the first predicted image sample corresponding to it, the smaller the function value of the first loss function.

[0113] In order to more intuitively understand the implementation process of building the text enhancement matrix in the text enhancement library, please refer to Figure 3 . Figure 3 An example diagram of an implementation framework for constructing a text enhancement matrix in a text enhancement library in the present application is shown.

[0114] Depend on Figure 3 It can be seen that after obtaining the first image sample with each text pixel marked and the text boundary map of the first image sample, the present application can respectively calculate the pixel average and pixel standard deviation corresponding to the first image sample and the text boundary map, and input them into the pattern generation module. Accordingly, the pattern generation module can output at least one text enhancement matrix, which constitutes the text enhancement library at the current moment (such as Figure 3 Each text enhancement matrix corresponds to a text image enhancement mode. Figure 3 , generated by the pattern generation module Take a text enhancement matrix as an example. The text enhancement matrices are represented as , Text Enhancement Matrix ..., until the text enhancement matrix .

[0115] Moreover, the first pattern sample and the text boundary image are also input into the encoder to obtain the first sample coding feature. After downsampling the first sample coding feature, the second sample coding feature obtained by downsampling is subjected to feature enhancement processing using at least one text enhancement matrix in the currently obtained text enhancement library, and the third sample coding feature obtained by the feature enhancement processing is input into the decoder to obtain the first predicted image sample.

[0116] On this basis, the model parameters of the pattern generation module can be continuously optimized by combining each first image sample and the first predicted image sample corresponding to the first image sample, and then the above operations can be re-executed, and finally a text enhancement library that can accurately perform image enhancement on text content can be obtained.

[0117] The following is an introduction to the process of building a non-text enhancement library. Figure 4 This is a schematic diagram of an implementation process of constructing a non-text enhancement library in this application. This embodiment may include:

[0118] S401, obtaining multiple groups of non-text image data.

[0119] Each set of non-text image data includes: a second image sample with non-text pixels marked as non-text and a non-text boundary map corresponding to the second image sample. The non-text boundary map includes boundary lines belonging to the text area in the second image sample.

[0120] Among them, the process of obtaining each group of non-text image data is similar to the previous process of obtaining each group of text image data. For example, multiple high-resolution image samples can be pre-screened, and the pixels of each image sample that do not belong to the text can be marked to obtain a second image sample with only non-text pixels and text pixels with values ​​of 0.

[0121] For example, after obtaining multiple image samples and the non-text classification mask and boundary map corresponding to each image sample, the non-text classification mask is used to distinguish the non-text content in the image sample. Based on this, the image sample is masked using the non-text classification mask to obtain a second image sample with each non-text pixel marked, and a boundary map of the second image sample can be obtained.

[0122] S402: For each group of non-text image data, based on the second image sample and the non-text boundary map, generate at least one non-text enhancement matrix using a pattern generation module.

[0123] and Figure 2 Similar to the embodiment, in order to reduce the amount of data processing of the pattern generation model, the present application can first determine the pixel average and pixel standard deviation corresponding to the second image sample, as well as the pixel average and pixel standard deviation of the non-text boundary map. Then, the pixel average and pixel standard deviation corresponding to the second image sample and the non-text boundary map are input into the pattern generation module to obtain at least one non-text enhancement matrix.

[0124] The average value of the pixels corresponding to the second image sample includes the average value of each pixel in the second image sample, and the average value of each pixel in the second image sample is the average value between the pixel and each of its adjacent pixels. Similarly, the average value of the pixels corresponding to the non-text boundary map includes the average value of each pixel in the text boundary map.

[0125] The standard deviation of the pixels corresponding to the second image sample is the standard deviation calculated based on the average value of each pixel in the second image sample. The standard deviation of the pixels corresponding to the non-text boundary map is the standard deviation calculated based on the average value of each pixel in the non-text boundary map.

[0126] The pattern generation module in this implementation can be used with the previous Figure 2 The network architecture of the pattern generation modules in the embodiments is the same and is not specifically limited.

[0127] S403: Determine a fourth sample encoding feature of the second image sample based on the second image sample and a non-text boundary map corresponding to the second image sample.

[0128] In the present application, for the sake of distinction, the encoding feature of the first image sample is referred to as the first sample encoding feature, and the encoding feature of the second image sample is referred to as the fourth sample encoding feature.

[0129] For example, based on the second image sample and the non-text boundary map corresponding to the second image sample, a second encoder is used to determine the sample encoding feature of the second image sample. The second encoder can be an encoder of the same type as the first encoder, such as the second encoder can be a convolutional neural network model. Of course, the second encoder can be the same encoder as the first encoder, or a different encoder, without limitation.

[0130] S404, perform feature downsampling on the fourth sample coding feature to obtain a fifth sample coding feature.

[0131] Among them, the specific implementation and purpose of feature downsampling of the fourth sample coding feature can be found in the previous introduction to feature downsampling of the first sample coding feature, which will not be repeated here.

[0132] S405 , performing feature enhancement on the fifth sample coding feature based on at least one non-text enhancement matrix to obtain a sixth sample coding feature.

[0133] The fifth sample encoding feature is enhanced based on at least one non-text enhancement matrix, with the aim of increasing the image resolution of an image sample that can be generated by the sixth sample encoding feature.

[0134] The specific implementation of using the non-text enhancement matrix to perform feature enhancement on the fifth sample encoding feature can refer to the previous introduction to the enhancement processing of the second feature, which will not be repeated here.

[0135] S406: Generate a second predicted image sample based on the sixth sample encoding feature.

[0136] For example, the sixth sample encoding feature is decoded by a second decoder to obtain a second predicted image sample. The second decoder may have the same network structure as the first decoder, or may be different. Of course, the second decoder and the first decoder may also be the same decoder, without any specific limitation.

[0137] S407, based on the second predicted image sample and the second image sample, determine whether the set training goal is achieved; if not, adjust the parameters of the pattern generation module and return to the operation of step S402; if not, determine that the training is completed, and determine at least one non-text enhancement matrix currently generated by the pattern generation module as a non-text enhancement feature matrix in the non-text enhancement library.

[0138] and Figure 3 Similarly, it is more accurate to use the pattern generation module to generate at least one non-text enhancement matrix. Then, after the fifth sample encoding feature corresponding to the second image sample is feature enhanced using at least one non-text enhancement matrix, the second predicted image sample generated based on the sixth sample encoding feature obtained based on the feature enhancement should be a high-resolution image and have a smaller difference with the first image sample.

[0139] Based on this, in this embodiment, the training goal of the training pattern generation module is to minimize the gap between the second image sample and the second predicted image sample corresponding to the second image sample. For example, in this embodiment, the set training goal can be that the number of training iterations exceeds the set number, or the gap value between each second image sample and its corresponding second predicted image sample is minimized, or the function value of the set second loss function converges. The smaller the gap between each second image sample and its corresponding second predicted image sample, the smaller the function value of the second loss function.

[0140] In any of the above embodiments of the present application, in order to make the boundaries of the text content in the reconstructed second image clearer and the text sharper, the present application may also obtain the character mask information of the first image before determining the first feature and the second feature. The character mask information of the first image is used to locate the area where the text content is located in the first image, or the area where each text character is located in the first image. In an optional manner, in order to make the edges of each character in the reconstructed second image sharper (that is, the characters are sharper), the character mask information of the first image may include the character edges of each character in the first image.

[0141] On this basis, in any of the above embodiments of the present application, a first feature of the text content in the first image and a second feature of the non-text content in the first image can be determined based on the first image and the character mask information.

[0142] In a possible implementation, the present application can determine the coding feature of the first image based on the first image and the character mask information. Accordingly, the first feature of the text content and the second feature of the non-text content in the first image can be determined in combination with the coding feature of the first image.

[0143] The following takes determining the first feature and the second feature in combination with the coding feature of the first image as an example, and introduces the image processing method of the present application in combination with an implementation method.

[0144] like Figure 5 , shows another flow chart of the image processing method provided by an embodiment of the present application. The method of this embodiment may include:

[0145] S501, obtaining a first image and character mask information of the first image.

[0146] The character mask information may be information for locating text content in the first image. For example, in an optional manner, in order to improve the sharpness of characters during image super-resolution reconstruction, the character mask information of the first image may include the edges of each character in the first image, for example, the character edges of each character in the first image may be determined by canny filtering, without specific limitation.

[0147] Of course, the character mask information of the first image may also be the edge where the text content in the first image is located or other information used to locate the text content in the first image, and there is no limitation to this.

[0148] S502: Determine a coding feature of the first image based on the first image and the character mask information.

[0149] For example, based on the first image and the character mask information, a target encoder is used to determine the encoding features of the first image, and the target encoder is different from the first encoder and the second encoder mentioned above.

[0150] In particular, before determining the coding features of the first image, the image features of the first image and the mask features of the character mask information may also be determined. For example, the image features of the first image may be extracted using a convolutional neural network model, or the image features of the first image may be determined using other deep learning models. Similarly, the mask features of the character mask information may be extracted using a convolutional neural network model or other deep learning models.

[0151] In a possible implementation, the coding features of the first image can be determined by using a multi-level feature coding module based on the first image and the character mask information, wherein each level of feature coding module determines the current level coding features and current level character mask features corresponding to the first image based on the previous level coding features and previous level character mask features output by the previous level feature coding module of the feature coding module.

[0152] When the feature encoding module is a first-level feature encoding module, the previous-level encoding feature on which the feature encoding module is based is the image feature of the first image, and the previous-level character mask feature is the mask feature of the character mask information. The image feature of the first image is an image feature determined based on the first image, and as described above, the feature of the first image can be determined using a convolutional neural network or other deep learning model. The character mask information is a feature determined based on the character mask information, for example, the feature of the character mask information is determined using a convolutional neural network or other deep learning model.

[0153] Among them, the coding features of this level output by the last level feature coding module in the multi-level feature coding module are the coding features of the first image.

[0154] For the sake of distinction, the present application refers to the coding features of the first image determined by the feature coding module as the current level coding features, and the character mask features determined by the feature coding module as the current level character mask features. Correspondingly, the coding features and character mask features determined by the feature coding module at the previous level are referred to as the previous level coding features and the previous level character mask features, respectively.

[0155] S503: Determine a first feature of text content and a second feature of non-text content in the first image in combination with the encoding feature of the first image.

[0156] For example, based on the coding features of the first image, the first feature and the second feature are determined using a normalization function or a specific feature decomposition module.

[0157] Of course, there may be other ways to determine the first feature and the second feature, which are not limited to this.

[0158] S504: Process the first feature using at least one text enhancement matrix in the text enhancement library to obtain a third feature.

[0159] Among them, each text enhancement matrix corresponds to a text image enhancement mode.

[0160] S505, using at least one non-text enhancement matrix in the non-text enhancement library to process the second feature to obtain a fourth feature. Each non-text enhancement matrix corresponds to a non-text image enhancement mode.

[0161] Among them, each non-text enhancement matrix corresponds to a non-text image enhancement mode.

[0162] For the above steps S504 and S505, reference may be made to the related introduction of the previous steps, which will not be repeated here.

[0163] It should be noted that using the text enhancement matrix library and the non-text enhancement library to enhance the first feature and the second feature respectively is only one implementation method. Enhancing the first feature and the second feature by other methods is also applicable to this embodiment and is not limited to this.

[0164] S506: Construct a second image by combining the third feature and the fourth feature.

[0165] The image resolution of the second image is higher than that of the first image.

[0166] In this embodiment, since the character mask information of the first image can be used to locate the text content in the first image, the first feature of the text content in the first image determined by combining the first image and the character mask information of the first image can contain richer text information in the first image. On this basis, after respectively enhancing the first feature in the first image and the second feature of the non-text content in the first image, compared with the image super-resolution reconstruction method that does not distinguish between text areas and non-text areas in low-resolution images, the enhanced third feature and fourth feature Chinese can also contain more text information, so that the reconstructed second image can accurately restore the information of the text in the first image, and improve the clarity of the text content in the second image.

[0167] In addition, in the present embodiment, after respectively determining the first feature of the text content and the second feature of the non-text content in the first image, the first feature of the text content can be overall enhanced using at least one text enhancement matrix in the text enhancement library, and the second feature of the non-text content can be overall enhanced using at least one non-text enhancement matrix in the non-text enhancement library. This enables block-by-block reconstruction of the text content and the non-text content in the first image without the need to reconstruct the first image pixel by pixel, thereby improving the efficiency of reconstructing a high-resolution second image.

[0168] To understand this benefit, see Figure 6 , which shows a schematic diagram of an implementation principle framework of the image processing method in this application.

[0169] Depend on Figure 6 It can be seen that the present application can obtain the encoding features of the first image based on the low-resolution first image and its character mask information using an encoder. Figure 6 The encoder in can be the target encoder mentioned above, or it can include an encoder of a multi-level feature encoding module, without limitation.

[0170] After determining the first feature F1 of the text content in the first image and the second feature F2 of the non-text content in the first image based on the coding features of the first image, the present application only needs to perform two steps of feature enhancement processing, namely: using the text enhancement matrices in the text enhancement library to enhance the first feature F1, and using the non-text enhancement library to enhance the second feature F2, the reconstruction processing at the feature level can be completed without the need to continuously reconstruct the first image pixel by pixel through multiple steps through models such as implicit neural transformation. On this basis, by fusing the enhanced first feature and the enhanced second feature, and generating a high-definition image based on the fused features, the complexity of reconstructing a high-resolution image is reduced, and the steps are reduced, thereby effectively improving the efficiency of image super-resolution reconstruction.

[0171] It is understandable that in order to reduce the complexity of super-resolution image reconstruction and improve the effect of the reconstructed second image, as described above, the present application can also first fuse the third feature and the fourth feature, and obtain the first target image based on the fused features. At this time, the first target image can be considered as a residual image of the high-resolution image reconstructed from the first image. On this basis, the present application can interpolate the first image using a bicubic interpolation algorithm to obtain a second target image. Accordingly, the first target image and the second target image can be superimposed to obtain a second image.

[0172] For easier understanding, see Figure 7, which shows a schematic diagram of an implementation process framework of the image processing method provided by the present application.

[0173] Depend on Figure 7 It can be seen that the first image is input into a convolutional network model to obtain the image features of the first image; and the character mask information of the first image is input into the convolutional network model to obtain the mask features of the character mask information.

[0174] On this basis, the image features of the first image and the mask features of the character mask information are processed by a multi-level feature encoding module to obtain the encoding features of the first image.

[0175] After the encoded features of the first image are processed by the normalization function layer, the first feature of the text content in the first image can be obtained. and a second feature of the non-text content in the first image .

[0176] Depend on Figure 7 It can be seen that the text enhancement library includes text enhancement matrices, represented as text feature matrices , … The non-text feature library also includes non-text enhancement matrices, represented as non-text feature matrices , … .

[0177] As an alternative, Figure 7 As shown, in the first feature And the second feature Before feature enhancement, in order to make the dimensions of the first feature and the second feature consistent with the dimensions of the text enhancement matrix or the non-text enhancement matrix, the present application may also perform feature enhancement on the first feature and the second feature. And the second feature Perform nearest neighbor upsampling (also called nearest neighbor interpolation).

[0178] Based on the above, we can use the text enhancement library The text enhancement matrix is ​​the first feature after upsampling Processing to obtain the third feature . Using the non-text enhancement library The non-text enhancement matrix is ​​used to upsample the second feature After processing, the fourth feature is obtained . The third feature And the fourth characteristic Then, a first target image is obtained based on the fused features, and the first target image is superimposed with a second target image obtained by interpolating the first image using a bicubic interpolation algorithm, so as to obtain a reconstructed super-resolution image, that is, a second image.

[0179] It can be understood that in the above embodiments, the composition structure of each level of feature coding module can also have multiple possibilities. Accordingly, the specific implementation of the feature coding module determining the coding features and character mask features of this level can have multiple possibilities, which are not specifically limited.

[0180] In one possible implementation, for any first-level feature coding module, the feature coding module performs multi-scale feature extraction and feature fusion processing on the previous-level coding features output by its previous-level feature coding module to obtain multi-scale fusion features. On this basis, the feature coding module can determine the current-level coding features and current-level character mask features corresponding to the first image based on the multi-scale fusion features and the previous-level character mask features output by the previous-level feature coding module.

[0181] It can be understood that since the coding features of this level output by the last level feature coding module will be used as the coding features of the first image, based on this, in order to make the coding features of the first image more comprehensively cover the relevant features of the text content and non-text content in the first image, when the feature coding module determines the coding features of this level and the character mask features of this level based on the multi-scale fusion features and the character mask features of the previous level, it can be: first, the character mask features of the previous level output by the feature coding module of the previous level are enhanced to obtain the character mask features of this level. Then, the character mask features of this level and the multi-scale fusion features are fused to obtain the coding features of this level.

[0182] It can be seen that the coding features at this level are actually obtained by integrating the previous level coding features of the previous level coding feature module and the current level character mask features generated by the current level feature coding module, so that the coding features at this level can cover the features related to the text content and non-text content in the first image, thereby enabling the subsequently generated coding features of the first image to include richer text and non-text related information.

[0183] There are many possibilities for implementing feature enhancement processing on the previous level character mask features output by the previous level feature encoding module, which are not specifically limited. For example, each level feature encoding module can determine the primary character mask features based on the previous level character mask features using a convolutional neural network model, and then activate the primary character mask features to obtain the current level character mask features.

[0184] The following describes how the feature coding module determines the current level coding features and the current level character mask features corresponding to the first image in combination with a composition architecture of the feature coding module. Figure 8 A schematic diagram of the composition architecture of the feature encoding module in this application is shown. Figure 8 Based on the composition architecture of the feature encoding module shown in Fig. 9 The flowchart of the embodiment of the present invention illustrates the implementation process of the feature encoding module determining the encoding feature of the present level and the character mask feature of the present level, and the implementation process may include:

[0185] S901, obtaining the previous level coding feature and the previous level character mask feature output by the previous level feature coding module of the feature coding module.

[0186] S902, performing multi-scale feature extraction and feature fusion processing on the previous level coding features to obtain multi-scale fusion features.

[0187] like Figure 8 As shown, the previous level coding feature After being processed by the layer normalization module, it is input into a spatial enhancement module. The spatial enhancement module uses a pyramid attention structure to perform multi-scale feature extraction and feature fusion processing on the feature encoding features of the previous level after layer normalization.

[0188] like Figure 8 The spatial enhancement module is divided into four branches according to the channel dimension. These four branches perform sub-sampling of the previous level coding features after layer normalization at different scales. The numbers marked in these four branches represent the scales. For example, taking the sub-sampling operation of scale s as an example, the previous level coding features after layer normalization are first down-sampled by s times, then extracted by deep convolution features, and then up-sampled by s times. Figure 8 It can be seen that the value of s can be 1, 2, 4 and 8.

[0189] On this basis, the features after sub-sampling operations at multiple scales are merged, the merged features are processed by convolution layer and activation function, and then dot-multiplied with the previous level encoding features that have only been processed by layer normalization to obtain multi-scale fusion features.

[0190] S903, performing feature enhancement processing on the previous-level character mask features output by the previous-level feature encoding module to obtain the current-level character mask features.

[0191] like Figure 8 It can be seen that the mask feature of the previous level character After a single layer of convolution and activation processing, the mask feature is enhanced to obtain the mask feature of the character at this level. .exist Figure 8The activation function is taken as Gaussian Error Linear Unit (GELU) as an example.

[0192] The character mask features at this level may include features of the text area in the first image and the characters in the text area.

[0193] exist Figure 8 The middle layer 1 convolution represents a single layer of convolution and the convolution kernel is . And 3 layers of convolution represent convolution.

[0194] S904, based on the multi-scale fusion features and the mask features of the characters at this level, a regional enhancement module is used to determine a first fusion enhancement feature.

[0195] Depend on Figure 8 It can be seen that the character mask feature at this level The multi-scale fusion features output by the spatial enhancement structure module are processed by the regional enhancement module. The regional enhancement module can decompose the features of the text content area and the non-text content area in the first image, and enhance the features of the text content area.

[0196] Depend on Figure 8 From the structure of the regional enhancement module in , it can be seen that the scale fusion feature can be connected to the character mask feature of the current level through the regional enhancement module, and a spatial attention feature is generated through a convolution layer and a sigmoid activation function. Then, based on the spatial attention feature, the enhanced text content feature and non-text content feature are determined, and the enhanced text content feature and non-text content feature are fused to obtain the first fused feature.

[0197] S905, adding the first fusion enhancement feature and the previous level coding feature to obtain a residual compensation feature.

[0198] like Figure 8 , the first fusion feature and the previous level encoding feature Adding together, we can get the residual compensation feature .

[0199] S906, using the feature fusion module to process the residual compensation features to obtain the coding features of this level.

[0200] Depend on Figure 8 It can be seen that the residual compensation feature It will be input into the Feature Fusion Module (FFM). In the FFM module, the residual compensation features will be layer normalized, Convolution and multiple layers After the activation processing of the convolution, the encoding features of this level are obtained.

[0201] It is understandable that in order to improve the accuracy of super-resolution image reconstruction, in the present application, various models or modules involved in processing the first image may also be trained in advance.

[0202] During training, the present application may first obtain multiple image sample groups, each of which includes an image sample, a sample boundary map corresponding to the image sample, and sample character mask information. The meanings of the sample boundary map and the sample character mask are the same as those of the previous boundary map and character mask information, respectively. For the sake of easy distinction, the boundary map and character mask information of the image sample are respectively referred to as the sample boundary map and sample character mask information.

[0203] On this basis, for each image sample group, the sample coding features of the image samples can be determined based on the image samples, the sample boundary map and the sample character mask information. When determining the sample coding features of the image samples, the specific processing process can be found in Figure 7 and Figure 8 The structure of the image sample and the sample boundary map can be regarded as a whole and processed in the same way as the first image, while the sample character mask information can be processed in the same way as the character mask information, and finally the sample encoding features of the image sample can be obtained.

[0204] On this basis, based on the sample coding features of the image sample, the first sample feature of the text content and the second sample feature of the non-text content in the image sample can be determined, and the first sample feature is enhanced using the text enhancement matrix in the text enhancement library to obtain the third sample feature; similarly, the second sample feature is enhanced using the non-text enhancement matrix in the non-text enhancement library to obtain the fourth sample feature. The specific implementation of obtaining the third sample feature and the fourth sample feature is the same as the process of obtaining the third feature and the fourth feature above, and will not be repeated.

[0205] After generating the predicted image sample based on the third sample feature and the fourth sample feature, the image sample and its corresponding predicted image sample can be optimized. Figure 7 and Figure 8 The parameters of the relevant models or modules in the algorithm are adjusted until high-resolution images can be accurately generated.

[0206] On the other hand, corresponding to the image processing method provided by the present application, the present application also provides an image processing device. Fig.10 , shows a schematic diagram of a composition structure of an image processing device provided by the present application, and the device of this embodiment includes:

[0207] A feature determination unit 1001 is used to determine a first feature of text content in a first image and a second feature of non-text content in the first image;

[0208] A first feature enhancement unit 1002 is used to enhance the first feature to obtain a third feature;

[0209] A second feature enhancement unit 1003 is used to enhance the second feature to obtain a fourth feature;

[0210] The image construction unit 1004 is used to construct a second image by combining the third feature and the fourth feature, wherein the image resolution of the second image is higher than the image resolution of the first image.

[0211] In a possible implementation, the image construction unit includes:

[0212] A first image generating subunit is used to fuse the third feature and the fourth feature, and obtain a first target image based on the fused features;

[0213] A second image generating subunit, configured to interpolate the first image using a bicubic interpolation algorithm to obtain a second target image;

[0214] The image construction subunit is used to superimpose the first target image and the second target image to obtain a second image.

[0215] In yet another possible implementation, the first feature enhancement unit includes:

[0216] A first feature enhancement subunit is used to process the first feature using at least one text enhancement matrix in a text enhancement library to obtain a third feature, each of the text enhancement matrices corresponding to a text image enhancement mode;

[0217] The second feature enhancement unit comprises:

[0218] The second feature enhancement subunit is used to process the second feature using at least one non-text enhancement matrix in the non-text enhancement library to obtain a fourth feature, and each of the non-text enhancement matrices corresponds to a non-text image enhancement mode.

[0219] In any one of the above device embodiments, the device further comprises: a mask obtaining unit, configured to obtain character mask information of the first image;

[0220] The feature determination unit comprises:

[0221] An image encoding subunit, configured to determine encoding features of the first image based on the first image and the character mask information;

[0222] The feature decomposition subunit is used to determine a first feature of text content and a second feature of non-text content in the first image in combination with the encoding feature of the first image.

[0223] In yet another possible implementation, the image encoding subunit includes:

[0224] a multi-level encoding subunit, configured to determine the encoding features of the first image using a multi-level feature encoding module based on the first image and the character mask information;

[0225] Each level of feature coding module determines the current level coding features and current level character mask features corresponding to the first image based on the previous level coding features and previous level character mask features output by the previous level feature coding module of the feature coding module;

[0226] When the feature encoding module is a first-level feature encoding module, the previous-level encoding feature is an image feature of the first image, and the previous-level character mask feature is a mask feature of the character mask information;

[0227] The coding features of this level output by the last level feature coding module are the coding features of the first image.

[0228] In yet another possible implementation, the multi-stage encoding subunit includes:

[0229] The scale fusion subunit is used to perform multi-scale feature extraction and feature fusion processing on the last-level coding features output by the last-level feature coding module of the feature coding module through each level of feature coding module to obtain multi-scale fusion features;

[0230] The feature integration subunit is used to determine the current level coding features and current level character mask features corresponding to the first image based on the multi-scale fusion features and the previous level character mask features output by the previous level feature encoding module.

[0231] In yet another possible implementation, the feature integration subunit includes:

[0232] The mask enhancement subunit is used to perform feature enhancement processing on the previous-level character mask features output by the previous-level feature encoding module to obtain the current-level character mask features;

[0233] The fusion processing subunit is used to fuse the character mask features of this level and the multi-scale fusion features to obtain the coding features of this level.

[0234] In another possible implementation, the text enhancement library used in the first feature enhancement subunit is obtained in the following manner:

[0235] Obtain multiple groups of text image data, each group of text image data includes: a first image sample with each text pixel point belonging to the text annotated and a text boundary map corresponding to the first image sample, the text boundary map including each boundary line belonging to the text area in the first image sample;

[0236] For each set of text image data, based on the first image sample and the text boundary map, using a pattern generation module to generate at least one text enhancement matrix;

[0237] Determining a first sample encoding feature of the first image sample based on the first image sample and a text boundary map corresponding to the first image sample;

[0238] Performing feature downsampling on the first sample coding feature to obtain a second sample coding feature;

[0239] Performing feature enhancement on the second sample encoding feature based on the at least one text enhancement matrix to obtain a third sample encoding feature;

[0240] generating a first predicted image sample based on the third sample encoding feature;

[0241] If it is determined based on the first predicted image sample and the first image sample that the set training target has not been achieved, adjust the parameters of the pattern generation module, and return to execute the operation of generating at least one text enhancement matrix based on the first image sample and the text boundary map using the pattern generation module until the set training target is achieved.

[0242] In another possible implementation, the non-text enhancement library used by the second feature enhancement subunit is obtained in the following manner:

[0243] Obtaining multiple groups of non-text image data, each group of non-text image data includes: a second image sample with each non-text pixel point marked as non-text and a non-text boundary map corresponding to the second image sample, the non-text boundary map including each boundary line belonging to a text area in the second image sample;

[0244] For each set of non-text image data, generating at least one non-text enhancement matrix using a pattern generation module based on the second image sample and the non-text boundary map;

[0245] Determining a fourth sample encoding feature of the second image sample based on the second image sample and a non-text boundary map corresponding to the second image sample;

[0246] Performing feature downsampling on the fourth sample coding feature to obtain a fifth sample coding feature;

[0247] Performing feature enhancement on the fifth sample coding feature based on the at least one non-text enhancement matrix to obtain a sixth sample coding feature;

[0248] generating a second predicted image sample based on the sixth sample encoding feature;

[0249] If it is determined based on the second predicted image sample and the second image sample that the set training target has not been achieved, the parameters of the pattern generation module are adjusted, and the operation of generating at least one non-text enhancement matrix based on the second image sample and the non-text boundary map using the pattern generation module is returned to execution until the set training target is achieved.

[0250] The present application also provides an electronic device. Fig.11 , which shows a schematic diagram of a composition structure of the electronic device, and the electronic device at least includes a processor 1101 and a memory 1102;

[0251] The processor 1101 is used to execute the image processing method described in any one of the above embodiments;

[0252] The memory 1102 is used to store the program required by the processor to perform operations

[0253] It is understandable that the electronic device may further include a display unit 1103 and an input unit 1104 .

[0254] Of course, the electronic device may also have Fig.11 There may be more or fewer components, without limitation.

[0255] A computer program product is also provided in an embodiment of the present application, including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements any image processing method provided in the embodiment of the present application.

[0256] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any image processing method provided in the embodiment of the present application.

[0257] It should also be noted that the device embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed over multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. In addition, in the drawings of the device embodiments provided by the present application, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines.

[0258] Through the description of the above implementation mode, the technicians in the field can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by special hardware including special integrated circuits, special CPUs, special memories, special components, etc. In general, all functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be various, such as analog circuits, digital circuits or special circuits. However, for the present application, software program implementation is a better implementation mode in more cases. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer floppy disk, a U disk, a mobile hard disk, a ROM, a RAM, a disk or an optical disk, etc., including a number of instructions to enable a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0259] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.

[0260] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website site, a computer, a training device, or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website site, computer, training device, or data center. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium may be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)), etc.

Claims

1. An image processing method, comprising: determining a first characteristic of text content in a first image and a second characteristic of non-text content in the first image; Performing enhancement processing on the first feature to obtain a third feature; Performing enhancement processing on the second feature to obtain a fourth feature; The third feature and the fourth feature are combined to construct a second image, wherein the image resolution of the second image is higher than the image resolution of the first image.

2. The image processing method according to claim 1, wherein the combining the third feature and the fourth feature to construct the second image comprises: Fusing the third feature and the fourth feature, and obtaining a first target image based on the fused features; interpolating the first image using a bicubic interpolation algorithm to obtain a second target image; The first target image and the second target image are superimposed to obtain a second image.

3. The image processing method according to claim 1, wherein the step of enhancing the first feature to obtain the third feature comprises: Processing the first feature using at least one text enhancement matrix in a text enhancement library to obtain a third feature, wherein each of the text enhancement matrices corresponds to a text image enhancement mode; The step of enhancing the second feature to obtain a fourth feature includes: The second feature is processed by using at least one non-text enhancement matrix in a non-text enhancement library to obtain a fourth feature, wherein each of the non-text enhancement matrices corresponds to a non-text image enhancement mode.

4. The image processing method according to any one of claims 1 to 3, further comprising: Obtaining character mask information of the first image; The determining a first feature of text content in the first image and a second feature of non-text content in the first image includes: Determining a coding feature of the first image based on the first image and the character mask information; In combination with the encoding features of the first image, a first feature of text content and a second feature of non-text content in the first image are determined.

5. The image processing method according to claim 4, wherein determining the coding feature of the first image based on the first image and the character mask information comprises: Determining encoding features of the first image using a multi-level feature encoding module based on the first image and the character mask information; Each level of feature coding module determines the current level coding features and current level character mask features corresponding to the first image based on the previous level coding features and previous level character mask features output by the previous level feature coding module of the feature coding module; When the feature encoding module is a first-level feature encoding module, the previous-level encoding feature is an image feature of the first image, and the previous-level character mask feature is a mask feature of the character mask information; The coding features of this level output by the last level feature coding module are the coding features of the first image.

6. The image processing method according to claim 5, wherein the feature encoding module at each level determines the current level encoding features and current level character mask features corresponding to the first image based on the previous level encoding features and previous level character mask features output by the previous level feature encoding module of the feature encoding module, comprising: Each level of feature coding module performs multi-scale feature extraction and feature fusion processing on the previous level coding features output by the previous level feature coding module of the feature coding module to obtain multi-scale fusion features; Based on the multi-scale fusion features and the previous-level character mask features output by the previous-level feature encoding module, current-level encoding features and current-level character mask features corresponding to the first image are determined.

7. The image processing method according to claim 6, wherein determining the current level coding features and the current level character mask features corresponding to the first image based on the multi-scale fusion features and the previous level character mask features output by the previous level feature encoding module comprises: Performing feature enhancement processing on the previous-level character mask features output by the previous-level feature encoding module to obtain current-level character mask features; The character mask features of this level and the multi-scale fusion features are fused to obtain the encoding features of this level.

8. According to the image processing method of claim 3, the text enhancement library is obtained by: A plurality of sets of text image data are obtained, each set of text image data includes: Annotating a first image sample of each text pixel belonging to the text and a text boundary map corresponding to the first image sample, wherein the text boundary map includes each boundary line belonging to the text area in the first image sample; For each set of text image data, based on the first image sample and the text boundary map, using a pattern generation module to generate at least one text enhancement matrix; Determining a first sample encoding feature of the first image sample based on the first image sample and a text boundary map corresponding to the first image sample; Performing feature downsampling on the first sample coding feature to obtain a second sample coding feature; Performing feature enhancement on the second sample encoding feature based on the at least one text enhancement matrix to obtain a third sample encoding feature; generating a first predicted image sample based on the third sample encoding feature; If it is determined based on the first predicted image sample and the first image sample that the set training target has not been achieved, adjust the parameters of the pattern generation module, and return to execute the operation of generating at least one text enhancement matrix based on the first image sample and the text boundary map using the pattern generation module until the set training target is achieved.

9. The image processing method according to claim 3, wherein the non-text enhancement library is obtained by: A plurality of sets of non-text image data are obtained, each set of non-text image data comprising: A second image sample marking each non-text pixel belonging to non-text and a non-text boundary map corresponding to the second image sample, wherein the non-text boundary map includes each boundary line belonging to a text area in the second image sample; For each set of non-text image data, generating at least one non-text enhancement matrix using a pattern generation module based on the second image sample and the non-text boundary map; Determining a fourth sample encoding feature of the second image sample based on the second image sample and a non-text boundary map corresponding to the second image sample; Performing feature downsampling on the fourth sample coding feature to obtain a fifth sample coding feature; Performing feature enhancement on the fifth sample coding feature based on the at least one non-text enhancement matrix to obtain a sixth sample coding feature; generating a second predicted image sample based on the sixth sample encoding feature; If it is determined based on the second predicted image sample and the second image sample that the set training target has not been achieved, adjust the parameters of the pattern generation module, and return to execute the operation of generating at least one non-text enhancement matrix based on the second image sample and the non-text boundary map using the pattern generation module until the set training target is achieved.

10. An image processing device, comprising: a feature determination unit, configured to determine a first feature of text content in a first image and a second feature of non-text content in the first image; A first feature enhancement unit, used for performing enhancement processing on the first feature to obtain a third feature; A second feature enhancement unit, used for enhancing the second feature to obtain a fourth feature; An image construction unit is used to construct a second image by combining the third feature and the fourth feature, wherein the image resolution of the second image is higher than the image resolution of the first image.