Image processing method, device, electronic device and storage medium
By using pre-trained neural network models, including encoder, instance normalization module and decoder, the image is defuzzed, which solves the problem of lack of universality of image defuzzing methods in the prior art, and improves the defuzzing effect of text images in real scenes and accelerates the processing speed.
Patent Information
- Application Number
- CN202210052003.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-01-18
AI Technical Summary
Existing deep learning-based image debuffering methods perform well in specific scenarios, but lack universality in other scenarios, especially in real scenarios, text images debuffering are poor and slow processing speed.
The image defuzzing process is performed using a pre-trained neural network model. The model includes an encoder, an instance normalization module and a decoder. The feature information is extracted through the encoder, the example normalization module processes part of the feature information, and generates a defuzzing image through the decoder.
It achieves the improvement of text image debuffering effect in real scenes, and at the same time accelerates the processing speed, which is suitable for image debuffering processing in various scenes.
Smart Images

Figure CN114511458B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to an image processing method, device, electronic device and storage medium. Background Art
[0002] At present, image deblurring is one of the effective methods to improve image quality. It can make blurry images clear and help identify image content. Image deblurring is usually performed based on deep learning.
[0003] However, when deblurring images based on deep learning, a training model based on blur data in a specific scenario is required. This method is not applicable to scenarios other than the specific scenario and is not universal. In addition, the deblurring effect of text images in real scenarios is relatively poor and the processing speed is also relatively slow. Summary of the invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an image processing method, which can deblur a target image to obtain a clear target image, and is suitable for text images in real scenes, has a good deblurring effect, and a fast processing speed.
[0005] In a first aspect, an embodiment of the present disclosure provides an image processing method, including:
[0006] Get the target image;
[0007] A target image is deblurred by a pre-trained neural network model, wherein the neural network model includes at least one encoder, at least one instance normalization module and at least one decoder; a feature of the target image is extracted by the encoder to obtain first feature information; an instance normalization process is performed on part of the feature information in the first feature information by the instance normalization module to obtain second feature information; and a deblurred image corresponding to the target image is obtained by the decoder based on the second feature information.
[0008] In a second aspect, an embodiment of the present disclosure provides an image processing device, including:
[0009] An acquisition module, used for acquiring a target image;
[0010] A processing unit is used to deblur a target image through a pre-trained neural network model, wherein the neural network model includes at least one encoder, at least one instance normalization module and at least one decoder; extract features of the target image through the encoder to obtain first feature information; perform instance normalization processing on part of the feature information in the first feature information through the instance normalization module to obtain second feature information; and obtain a deblurred image corresponding to the target image based on the second feature information through the decoder.
[0011] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0012] Processor; and
[0013] Memory for storing programs,
[0014] The program includes instructions, and when the instructions are executed by a processor, the processor executes the above method.
[0015] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above method.
[0016] A fifth aspect includes a computer program, wherein the computer program implements the above method when executed by a processor.
[0017] An image processing method provided by an embodiment of the present disclosure obtains a target image, and deblurs the target image through a pre-trained neural network model, wherein the neural network model includes at least one encoder, at least one instance normalization module, and at least one decoder; the encoder performs feature extraction on the target image to obtain first feature information; the instance normalization module performs instance normalization on part of the feature information in the first feature information to obtain second feature information; the decoder obtains a deblurred image corresponding to the target image based on the second feature information, and the target image can be deblurred to obtain a clear target image, and the deblurring effect on text images in real scenes is relatively good, and the processing speed is also relatively fast. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0019] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0020] Figure 1 A schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0021] Figure 2 A flowchart of a neural network model training method provided by an embodiment of the present disclosure;
[0022] Figure 3 A schematic diagram of generating a blurred image provided by an embodiment of the present disclosure;
[0023] Figure 4 A schematic diagram of the structure of a neural network model provided in an embodiment of the present disclosure;
[0024] Figure 5 A flowchart of an image processing method provided by an embodiment of the present disclosure;
[0025] Figure 6 A schematic diagram of the structure of an image processing device provided by an embodiment of the present disclosure;
[0026] Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] The term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". Relevant definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0030] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0031] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0032] At present, image deblurring is one of the important branches of image quality improvement. Its purpose is to make blurred images clear and help identify image content. It is also one of the important image preprocessing technologies in the field of optical character recognition (OCR). Current image processing is generally implemented based on deep learning (DL) methods. Image deblurring based on deep learning is mainly divided into non-blind deblur (NBD) and blind deblur (BD). Non-blind deblurring is to build a blur dataset based on a known blur kernel for deblurring learning. Since blur is generated by blur kernel convolution, deblurring generally adopts the form of deconvolution. Blind deblurring is to generate unknown blur features, generally using blur kernel estimation or end-to-end (image-to-image) image restoration methods. However, the blur features of blurry images in natural scenes are generally complex and difficult to estimate. At present, it is mainly based on the deep learning method of blind deblurring, but the deblurring effect is relatively poor. For deep learning deblurring methods, blurry and clear image data pairs are generally required to optimize the gap between the generated image and the clear image as the label by minimizing the empirical risk. The main goal of deep learning is to effectively remove blur features and make the generated image as consistent as possible with the clear image as the label in the objective function.
[0033] In blind deblurring, the same model built based on deep learning has different effects on different data sets, especially for limited data sets. The smaller the data set, the worse the deblurring effect. Therefore, deblurring work in specific fields requires specific blur data sets. For example, the model trained based on the data set of scenario 1 is not suitable for scenario 2. Secondly, for text image deblurring, the related unsupervised image deblurring methods, although they can effectively avoid the defects of data set requirements, have poor deblurring effects on text images in real scenes, and the deblurring processing speed is also relatively slow.
[0034] In view of the above technical problems, the embodiments of the present disclosure provide an image processing method, which can perform deblurring on a target image to obtain a clear image. The deblurring effect is better for text images in real scenes, and the deblurring processing speed is also faster. The method is described in detail through one or more embodiments below.
[0035] Specifically, the image processing method can be executed by a terminal or a server. Specifically, the terminal or the server can perform deblurring on the target image through a neural network model. The execution subject of the training method of the neural network model and the execution subject of the image processing method can be the same or different.
[0036] For example, in an application scenario, such as Figure 1 As shown, the server 12 trains the neural network model. The terminal 11 obtains the trained neural network model from the server 12, and the terminal 11 deblurs the target image through the trained neural network model. The target image may be obtained by photographing the terminal 11. Alternatively, the target image is an image obtained by the terminal 11 after performing image processing on a preset image, and the preset image may be obtained by photographing the terminal 11, or the preset image may be obtained by the terminal 11 from other devices. Here, no specific limitation is made to other devices.
[0037] In another application scenario, the server 12 trains the neural network model. Further, the server 12 performs deblurring processing on the target image through the trained neural network model. The way in which the server 12 obtains the target image can be similar to the way in which the terminal 11 obtains the target image as described above, which will not be repeated here.
[0038] In another application scenario, the terminal 11 trains the neural network model. Further, the terminal 11 performs deblurring processing on the target image through the trained neural network model.
[0039] It is understandable that the neural network model training method and image processing method provided by the embodiments of the present disclosure are not limited to the several possible scenarios described above. Since the trained neural network model can be applied to the image processing method, before introducing the image processing method, the neural network model training method can be introduced below.
[0040] The following introduces a neural network model training method, that is, the training process of the neural network model, by taking the server 12 training the neural network model as an example. It is understandable that the neural network model training method is also applicable to the scenario where the terminal 11 trains the neural network model.
[0041] Figure 2 A flowchart of a neural network model training method provided in an embodiment of the present disclosure specifically includes the following steps: Figure 2 The following steps S210 to S230 are shown:
[0042] S210: Obtain a clear image and a blurred image corresponding to the clear image.
[0043] It is understandable that the server obtains a clear image with clear content and a blurred image corresponding to the clear image. The clear image may include multiple characters. For example, the clear image may be an image including characters captured from a PDF file. In the clear image, each character is relatively clear, and each character can be directly recognized based on the optical character recognition algorithm without deblurring. The blurred image corresponding to the clear image can be understood as a blurred image generated based on the clear image. It is understandable that a clear image may generate multiple blurred images, which is not limited here, but each blurred image will correspond to a clear image, which can be called a blurred-clear image pair. For example, if 1 clear image is obtained and 5 blurred images are generated based on the clear image, then there are 5 blurred-clear image pairs.
[0044] Optionally, the above S210 specifically includes: acquiring a clear image with clear content; and generating a blurred image corresponding to the clear image according to the clear image and the blurred background image.
[0045] Optionally, the generating of the blurred image corresponding to the clear image based on the clear image and the blurred background image specifically includes: blurring the clear image to obtain a blurred image; and generating the blurred image corresponding to the clear image based on the blurred image and the blurred background image.
[0046] It is understandable that the specific implementation steps of generating a blurred image based on a clear image are as follows: a clear image with clear content can be obtained, and multiple clear images including different characters can be obtained at one time. The following description takes a clear image as an example. At least one blurred background image is obtained. The blurred background image can be an image obtained by taking a photo in a real scene, and the blurred background image does not include characters. For example, a blurred background image is obtained by photographing a paper book or an electronic book in a real scene including lighting. Books of different materials, lighting, cameras used for shooting, shooting techniques, etc. make the obtained blurred background images different. There is no limitation on the method of obtaining the blurred background image. It is understandable that the order of obtaining the clear image and the blurred background image is not limited. The two can be obtained at the same time, or the blurred background image can be obtained first and then the clear image. After both the clear image and the blurred background image are acquired, the clear image is blurred to obtain a blurred image. The blurring method may be dynamic blur (motion blur) and / or Gaussian blur. The specific blurring method is not limited. Motion blur is a common blur feature in the image shooting process. By adding the noise, the learning and removal ability of the neural network for the relevant motion blur can be enhanced. Specifically, the clear image can be subjected to motion blur to generate a first blurred image. The clear image can also be subjected to Gaussian blur to generate a second blurred image. The clear image can also be subjected to Gaussian blur and motion blur to generate a third blurred image. That is, multiple blurred images can be generated based on the clear image through different blurring processes. For example, three blurred images are generated based on the clear image, which are recorded as blurred image 1, blurred image 2, and blurred image 3. Subsequently, the blurred image and the blurred background image are fused to obtain a blurred image corresponding to the clear image. The fusion method is not limited as long as image fusion can be performed. For example, the Poisson fusion method is used, that is, the blurred image containing the blurred noise of the real scene is generated by the Poisson fusion method based on the generated motion blurred image and the blurred background image of the real scene. It is understandable that the blur features in real scenes are relatively complex and cannot be simulated by effective mathematical methods at present. Therefore, the present disclosure directly adopts the form of fusion of background image (blurred background image) and text image (clear image), which can not only retain the complete real scene noise, but also obtain the clear image corresponding to the blurred image, that is, each blurred image has a one-to-one corresponding clear image.
[0047] For example, see Figure 3 , Figure 3A schematic diagram of a blurred image generation provided by an embodiment of the present disclosure. 300 includes a clear image 310, a blurred processed image 320, a blurred background image 330 and a blurred image 340. The process of generating the blurred image 340 based on the clear image 310 and the blurred background image 330 is as follows: blurring the clear image 310 to obtain the blurred processed image 320, generating the blurred image 340 based on the blurred background image 330 and the blurred processed image 320, and the blurred background image 330 can be an image of an eye protection background color. For example, one clear image 310 and three blurred background images 330 are obtained, and three blurred processed images 320 are generated according to the clear image 310. Then, the three blurred background images 330 and the three blurred processed images 320 can be arranged and combined to generate multiple blurred images 340. For example, one blurred background image 330 is fused with the three blurred processed images 320 respectively to generate three different blurred images 340, and the remaining two blurred background images 330 are also fused with the three blurred processed images 320 respectively, and finally nine different blurred images are obtained. The clear images corresponding to the nine different blurred images are all clear images 310, and nine blurred-clear image pairs can be obtained. As training samples of the neural network model, a large number of blurred images with real backgrounds can be obtained, that is, a large number of blurred sample data sets similar to real scenes can be obtained, which can minimize the influence of the sample data set on the deblurring effect of the neural network model, further improve the versatility of the neural network model, that is, increase the application scenarios of the neural network model.
[0048] S220, input the blurred image into the constructed neural network model to obtain a predicted image.
[0049] It is understandable that, based on the above S210, the multiple blurred images generated as training samples are input into the constructed neural network framework to obtain a predicted image after deblurring of the blurred image output by the neural network model. It is understandable that when the neural network model has not been trained, the deblurring effect of the output predicted image may be poorer than the clear image used as the label.
[0050] For example, see Figure 4 , Figure 4 A schematic diagram of the structure of a neural network model provided in an embodiment of the present disclosure, Figure 4The network structure of the neural network model includes an input layer 410, multiple encoders 420, multiple instance normalization modules 430, multiple decoders 440, and an output layer 450. The multiple layers in the network structure of the neural network model are symmetrical, wherein the network architectures of multiple encoders 420, multiple instance normalization modules 430, and multiple decoders 440 are the same, but the network parameters may be different. The input layer 410 and the output layer 450 may be convolutional layers, the input of the input layer 410 is a blurred image, and the output of the output layer 450 is a predicted image after deblurring, that is, an end-to-end image restoration model. The multiple encoders 420 and the multiple decoders 440 may be constructed based on a fully convolutional neural network (FCN), and Figure 3 The network structure can be regarded as symmetrical based on the third instance normalization module 430, Figure 4 The third instance normalization module 430 includes three encoders 420, one input layer 410 and two instance normalization modules 430 on the left, and three decoders 440, one input layer 450 and two instance normalization modules 430 on the right. The encoder 420 is used to extract the deep features of the blurred image input to the neural network model, and the decoder 430 restores the pixel details of the blurred image based on deconvolution and upsampling operations; the instance normalization module 430 includes a convolution layer 431, multiple sub-normalization layers 432, and an activation function 433. Figure 3Only the network structure of the first instance normalization module is shown, and the network structures of the remaining four instance normalization modules (the second instance normalization module, the third instance normalization module, and the fourth instance normalization module) are the same as the first instance normalization module, which will not be described in detail here. The instance normalization module 430 can be constructed based on the instance normalization module (random part instance normalization, RPIN). In the first instance normalization module 430, multiple sub-normalization layers 432 are at the same level, and the multiple sub-normalization layers 432 are respectively connected to the convolution layer 431 and the activation function 433, and the multiple sub-normalization layers 432 have different data to process. For example, the first instance normalization module 430 includes 5 sub-normalization layers 432, each box represents a sub-normalization layer, and the feature information output by the convolution layer 431 will be randomly distributed into 5 groups of feature information, each group of feature information will correspond to a sub-normalization layer 432, that is, 5 groups of feature information will be randomly distributed to 5 The sub-normalization layer 432, for example, the first group of feature information is assigned to the first sub-normalization layer among the five sub-normalization layers 432, and the remaining four groups of feature information are respectively assigned to the remaining sub-normalization layers. The sub-normalization layer is used to perform instance normalization on the feature information or save the context information of the feature information. For example, three of the five sub-normalization layers 432 perform instance normalization on the feature information, that is, instance normalization is performed on part of the feature information in the first feature information, and the remaining two sub-normalization layers 432 save the context information of the feature information.
[0051] S230. Update the network parameters of the neural network model according to the clear image, the predicted image and the blurred image.
[0052] Optionally, the above S230 updates the network parameters of the neural network model according to the clear image, the predicted image and the blurred image, specifically including: extracting the features of the clear image and the blurred image respectively through a pre-acquired perception model to obtain the third feature information corresponding to the clear image and the fourth feature information corresponding to the blurred image; calculating the loss value according to the clear image, the predicted image, the third feature information and the fourth feature information; and updating the network parameters of the neural network model according to the loss value.
[0053] It is understandable that, based on the above S220, the clear image and the blurred image can be input into the acquired perception model in advance, and the feature information of the clear image and the blurred image can be extracted respectively through the perception model to obtain the third feature information corresponding to the clear image and the fourth feature information corresponding to the blurred image. The perception model can be a model based on transfer learning (Visual Geometry Group Network, VGGNet), such as VGG16 or VGG19. Taking VGG19 as an example, the third feature information corresponding to the clear image and the fourth feature information corresponding to the blurred image can be the feature information output by the 14th convolution layer in VGG19. It is understandable that the perception model is not limited as long as the feature information of the image can be extracted. When the transfer model is used to extract image features, the level of the output feature information can be determined according to the application scenario and the image. After obtaining the third feature information corresponding to the clear image and the fourth feature information corresponding to the blurred image and the predicted image output by the neural network model, the clear image, the predicted image, the third feature information and the fourth feature information are input into the pre-constructed loss function to calculate the loss value. Subsequently, the network parameters of the neural network model are updated according to the calculated loss value, and the training of the neural network model can be completed after multiple cycles of iterations.
[0054] Optionally, the above-mentioned calculation of the loss value based on the clear image, the predicted image, the third characteristic information and the fourth characteristic information specifically includes: obtaining a first loss value based on the fourth characteristic information and the third characteristic information; obtaining a second loss value based on the clear image and the predicted image; and calculating the sum of the first loss value and the second loss value to obtain the loss value.
[0055] It can be understood that the method for calculating the loss value includes: obtaining a first loss value according to the third feature information and the fourth feature information, that is, calculating the 2-norm of the difference between the third feature information and the fourth feature information and performing a square calculation to obtain a first loss value, calculating the 2-norm of the difference between the pixels of the clear image and the predicted image and performing a square calculation to obtain a second loss value, and calculating the sum of the first loss value and the second loss value to obtain a loss value. Exemplarily, the loss function for calculating the loss value is shown in formula (1).
[0056] Formula (1)
[0057] Among them, L represents the loss value, α and β represent weights, y1 represents the blurred image, y2 represents the clear image, and y3 represents the predicted image. represents the fourth characteristic information, represents the third feature information, and the first loss value is recorded as , the second loss value is recorded as Understandably, in a training of a neural network model, the values of α and β remain unchanged during different iterations. In different trainings, the values of α and β can be set according to user needs. For example, in the first training, the values of α and β are 0.2 and 0.8 respectively, and the values of multiple iterations are 0.2 and 0.8. In the second training, the values of α and β may be 0.3 and 0.7, and the values of multiple iterations are 0.3 and 0.7. The sum of α and β is 1.
[0058] The embodiment of the present disclosure provides a training method for a neural network model, by obtaining a clear image and a blurred image corresponding to the clear image, that is, obtaining multiple blurred-clear data pairs as training samples, and then inputting the blurred image into the constructed neural network model to obtain a predicted image, calculating a loss value based on the clear image, the predicted image and the blurred image, and updating the network parameters of the neural network model based on the loss value. The present disclosure combines depth perception features and pixel points as a loss function, and uses the loss function to optimize the neural network model. This method uses the fine character feature information of the text and the global features of the blurred features as the learning targets of the neural network model based on the text image features in natural scenes, so that the deblurring effect of the trained neural network model is better, and the training cycle of the neural network model trained using the above loss function is short.
[0059] Based on the above embodiments, Figure 5 A flowchart of an image processing method provided by an embodiment of the present disclosure, which is applied to a terminal, specifically includes the following steps: Figure 5 The following steps S510 to S520 are shown:
[0060] S510: Acquire a target image.
[0061] It can be understood that after the terminal obtains the neural network model, it obtains the target image, which can be understood as a blurred image. The target image can be captured in a real scene. For example, the target image may be blurred due to lighting and motion during the generation process.
[0062] S520, deblurring the target image through a pre-trained neural network model, the neural network model includes at least one encoder, at least one instance normalization module and at least one decoder; extracting features of the target image through the encoder to obtain first feature information; performing instance normalization processing on part of the feature information in the first feature information through the instance normalization module to obtain second feature information; obtaining a deblurred image corresponding to the target image based on the second feature information through the decoder.
[0063] Optionally, the instance normalization module includes multiple sub-normalization layers.
[0064] Optionally, in the above S520, instance normalization processing is performed on part of the feature information in the first feature information through an instance normalization module to obtain second feature information, specifically including: determining at least one target sub-normalization layer among multiple sub-normalization layers, and performing instance normalization processing on part of the feature information of the first feature information through the target sub-normalization layer; based on other sub-normalization layers except the target sub-normalization layer among the multiple sub-normalization layers, obtaining context information of the first feature information; and obtaining the second feature information based on the partial feature information after instance normalization processing and the context information.
[0065] It can be understood that, based on the above S510, the target image is input into the pre-trained neural network model for deblurring, and the neural network model outputs the deblurred image of the target image. Specifically, the processing flow inside the neural network model includes: taking the neural network model including an encoder, an instance normalization module and a decoder as an example, the target image is convolved through the encoder to extract the depth feature to obtain the first feature information, and then the first feature information is input into the instance normalization module, at least one target sub-normalization layer is randomly determined from the multiple sub-normalization layers of the instance normalization module, the target sub-normalization layer performs instance normalization on part of the feature information in the first feature information, and the sub-normalization layers other than the target sub-normalization layer in the multiple sub-normalization layers save the context information of other feature information in the first feature information except the part of the feature information, and then the second feature information is obtained according to the part of the feature information after the instance normalization and the saved context information. The decoder performs deconvolution and sampling operations on the second feature information to obtain the deblurred image corresponding to the target image.
[0066] For example, see Figure 4 , Figure 4Any instance normalization module 430 includes 5 sub-normalization layers 432, and the first feature information can be divided into 5 groups of feature information. The 5 sub-normalization layers 432 correspond to a group of feature information respectively. At least one target sub-normalization layer 432 is randomly selected from the 5 sub-normalization layers 432. For example, 3 sub-normalization layers are randomly selected. Taking the first instance normalization module 430 as an example, the first instance normalization module 430 includes 5 sub-normalization layers from top to bottom. The 5 sub-normalization layers are respectively recorded as sub-normalization layer 1, sub-normalization layer 2, sub-normalization layer 3, sub-normalization layer 4 and sub-normalization layer 5, wherein the sub-normalization layers 432 and 432 are randomly selected. Normalization layer 1, sub-normalization layer 2, and sub-normalization layer 3 are randomly selected target sub-normalization layers. Sub-normalization layer 1 performs instance normalization processing on a set of assigned feature information. Sub-normalization layer 2 and sub-normalization layer 3 also perform instance normalization processing on a set of assigned feature information. The remaining sub-normalization layer 4 and sub-normalization layer 5 save the context information in a set of assigned feature information. Finally, according to the three sets of feature information after the strength normalization processing output by sub-normalization layer 1, sub-normalization layer 2, and sub-normalization layer 3 and the two sets of context information saved by sub-normalization layer 4 and sub-normalization layer 5, the second feature information is obtained. The method provided by the present disclosure constructs an instance normalization module 430. The instance normalization module 430 increases the randomness of the instance normalization operation based on the characteristics of instance normalization, and can also save the context information of some feature information, which effectively improves the modeling and generalization capabilities of the neural network model.
[0067] Optionally, after obtaining the deblurred image corresponding to the target image, the method further includes: recognizing characters in the deblurred image.
[0068] It can be understood that after obtaining the deblurred image corresponding to the target image, the OCR character recognition algorithm can be used to recognize the characters in the deblurred image and output the character recognition result, that is, to perform text recognition on the deblurred image, which can effectively improve the accuracy of text recognition.
[0069] An image processing method provided by an embodiment of the present disclosure obtains a target image, and deblurs the target image through a pre-trained neural network model, wherein the neural network model includes at least one encoder, at least one instance normalization module, and at least one decoder; the encoder performs feature extraction on the target image to obtain first feature information; the instance normalization module performs instance normalization on part of the feature information in the first feature information to obtain second feature information; the decoder obtains a deblurred image corresponding to the target image based on the second feature information, and the target image can be deblurred to obtain a clear target image, which can improve the visual sensory effect of the image, and has a better deblurring effect on text images in real scenes and a faster processing speed.
[0070] Figure 6 The image processing device provided by the embodiment of the present disclosure can execute the processing flow provided by the image processing method embodiment, such as Figure 6 As shown, the image processing device 600 includes:
[0071] An acquisition unit 610 is used as an acquisition module to acquire a target image;
[0072] The processing unit 620 is used to deblur the target image through a pre-trained neural network model, where the neural network model includes at least one encoder, at least one instance normalization module and at least one decoder; extract features of the target image through the encoder to obtain first feature information; perform instance normalization on part of the feature information in the first feature information through the instance normalization module to obtain second feature information; and obtain a deblurred image corresponding to the target image based on the second feature information through the decoder.
[0073] Optionally, the instance normalization module in the processing unit 620 includes multiple sub-normalization layers.
[0074] Optionally, the processing unit 620 performs instance normalization processing on part of the feature information in the first feature information through an instance normalization module to obtain second feature information, including:
[0075] Determine at least one target sub-normalization layer among the multiple sub-normalization layers, and perform instance normalization processing on part of the feature information of the first feature information through the target sub-normalization layer;
[0076] Based on other sub-normalization layers except the target sub-normalization layer in the plurality of sub-normalization layers, obtaining context information of the first feature information;
[0077] The second feature information is obtained according to the normalized partial feature information and context information of the instance.
[0078] Optionally, the third recognition submodule in the processing unit 620 includes an attention layer, a recurrent layer, and a fully connected layer.
[0079] Optionally, the apparatus 700 further includes a training unit, and the training unit is specifically configured to:
[0080] Obtain a clear image and a blurred image corresponding to the clear image;
[0081] Input the blurred image into the constructed neural network model to obtain the predicted image;
[0082] According to the clear image, the predicted image and the blurred image, the network parameters of the neural network model are updated.
[0083] Optionally, in the training unit, network parameters of the neural network model are updated according to the clear image, the predicted image and the blurred image, specifically for:
[0084] The features of the clear image and the blurred image are respectively extracted by using a pre-acquired perception model to obtain third feature information corresponding to the clear image and fourth feature information corresponding to the blurred image;
[0085] Calculate the loss value according to the clear image, the predicted image, the third feature information and the fourth feature information;
[0086] Update the network parameters of the neural network model according to the loss value.
[0087] Optionally, calculating the loss value in the training unit according to the clear image, the predicted image, the third feature information, and the fourth feature information includes:
[0088] Obtaining a first loss value according to the third feature information and the fourth feature information;
[0089] Obtaining a second loss value according to the clear image and the predicted image;
[0090] The sum of the first loss value and the second loss value is calculated to obtain the loss value.
[0091] Optionally, obtaining a clear image and a blurred image corresponding to the clear image in a training unit includes:
[0092] Get clear images with crisp content;
[0093] A blurred image corresponding to the clear image is generated according to the clear image and the blurred background image.
[0094] Optionally, in the training unit, generating a blurred image corresponding to the clear image according to the clear image and the blurred background image includes:
[0095] Performing blur processing on the clear image to obtain a blurred image;
[0096] A blurred image corresponding to the clear image is generated according to the blurred processed image and the blurred background image.
[0097] Figure 6 The image processing device of the illustrated embodiment can be used to execute the technical solution of the above-mentioned method embodiment, and its implementation principle and technical effect are similar, which will not be repeated here.
[0098] The exemplary embodiment of the present disclosure also provides an electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication. The memory stores a computer program that can be executed by the at least one processor, and the computer program is used to cause the electronic device to perform the method according to the embodiment of the present disclosure when executed by the at least one processor.
[0099] Exemplary embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method according to an embodiment of the present disclosure.
[0100] The exemplary embodiments of the present disclosure also provide a computer program product, including a computer program, wherein the computer program is used to cause the computer to perform the method according to the embodiments of the present disclosure when executed by a processor of the computer.
[0101] refer to Figure 7 , a block diagram of an electronic device 700 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0102] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0103] A plurality of components in the electronic device 700 are connected to the I / O interface 705, including: an input unit 706, an output unit 707, a storage unit 708, and a communication unit 709. The input unit 706 may be any type of device capable of inputting information to the electronic device 700, and the input unit 706 may receive input digital or character information, and generate key signal inputs related to user settings and / or function control of the electronic device. The output unit 707 may be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 704 may include, but is not limited to, a disk, an optical disk. The communication unit 709 allows the electronic device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0104] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above. For example, in some embodiments, the image processing method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 700 via the ROM 702 and / or the communication unit 709. In some embodiments, the computing unit 701 may be configured to perform the image processing method in any other appropriate manner (e.g., by means of firmware).
[0105] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0106] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0107] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0108] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0109] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0110] A computer system may include clients and servers. Clients and servers are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other.
Claims
1. An image processing method, characterized in that: include: Get the target image; Deblurring the target image using a pre-trained neural network model, wherein the neural network model includes at least one encoder, at least one instance normalization module, and at least one decoder; extracting features from the target image using the encoder to obtain first feature information; Performing instance normalization processing on part of the feature information in the first feature information by the instance normalization module to obtain second feature information; obtaining a deblurred image corresponding to the target image based on the second feature information by the decoder; Among them, the instance normalization module includes a convolution layer, multiple sub-normalization layers and an activation function. The multiple sub-normalization layers are respectively connected to the convolution layer and the activation function, and the multiple sub-normalization layers are to process different data. The feature information output by the convolution layer will be distributed into multiple groups of feature information, each group of feature information corresponds to a sub-normalization layer, and the sub-normalization layer is used to perform instance normalization processing on each group of feature information or save the context information of each group of feature information.
2. The method according to claim 1, characterized in that The step of performing instance normalization processing on part of the feature information in the first feature information by the instance normalization module to obtain second feature information includes: Determine at least one target sub-normalization layer among the multiple sub-normalization layers, and perform instance normalization processing on part of the feature information of the first feature information through the target sub-normalization layer; Based on other sub-normalization layers among the multiple sub-normalization layers except the target sub-normalization layer, acquiring context information of the first feature information; The second feature information is obtained based on the normalized portion of the feature information and the context information.
3. The method according to claim 1, characterized in that The method further comprises: Acquire a clear image and a blurred image corresponding to the clear image; Inputting the blurred image into the constructed neural network model to obtain a predicted image; The network parameters of the neural network model are updated according to the clear image, the predicted image and the blurred image.
4. The method according to claim 3, characterized in that The updating of the network parameters of the neural network model according to the clear image, the predicted image and the blurred image comprises: Extracting features of the clear image and the blurred image respectively through a pre-acquired perception model to obtain third feature information corresponding to the clear image and fourth feature information corresponding to the blurred image; Calculating a loss value according to the clear image, the predicted image, the third feature information, and the fourth feature information; Update the network parameters of the neural network model according to the loss value.
5. The method according to claim 4, characterized in that The calculating the loss value according to the clear image, the predicted image, the third feature information and the fourth feature information comprises: Obtaining a first loss value according to the third feature information and the fourth feature information; Obtaining a second loss value according to the clear image and the predicted image; The sum of the first loss value and the second loss value is calculated to obtain a loss value.
6. The method according to claim 3, characterized in that The obtaining of a clear image and a blurred image corresponding to the clear image comprises: Get clear images with crisp content; A blurred image corresponding to the clear image is generated according to the clear image and the blurred background image.
7. The method according to claim 6, characterized in that The step of generating a blurred image corresponding to the clear image according to the clear image and the blurred background image comprises: Performing blur processing on the clear image to obtain a blurred image; A blurred image corresponding to the clear image is generated according to the blurred processed image and the blurred background image.
8. An image processing device, characterized in that: include: An acquisition module, used for acquiring a target image; A processing unit, configured to perform a deblurring process on the target image by using a pre-trained neural network model, wherein the neural network model includes at least one encoder, at least one instance normalization module, and at least one decoder; extracting features from the target image by using the encoder to obtain first feature information; Performing instance normalization processing on part of the feature information in the first feature information by the instance normalization module to obtain second feature information; obtaining a deblurred image corresponding to the target image based on the second feature information by the decoder; Among them, the instance normalization module includes a convolution layer, multiple sub-normalization layers and an activation function. The multiple sub-normalization layers are respectively connected to the convolution layer and the activation function, and the multiple sub-normalization layers are to process different data. The feature information output by the convolution layer will be distributed into multiple groups of feature information, each group of feature information corresponds to a sub-normalization layer, and the sub-normalization layer is used to perform instance normalization processing on each group of feature information or save the context information of each group of feature information.
9. An electronic device, comprising: processor; as well as Memory for storing programs, The program includes instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.
Citation Information
Patent Citations
Target identification method and device thereof and storage medium
CN113515992A
Method for generating image deblurring model and iris image deblurring method
CN113643215A