Image processing method, device, electronic device and storage medium
By detecting and repairing areas of interest in the image, the problems of compressed noise and blur in the face images are solved, efficient repair and natural transition are achieved, and image quality is improved.
Patent Information
- Application Number
- CN202010863010.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2040-08-25
AI Technical Summary
Images with faces may have compressed noise and face blur, affecting image quality and visual effects.
By detecting the area of interest of the target to be repaired in the original image, extracting and repairing the image in the area, backing the repaired image to the original image, and eliminating the splicing traces.
The target is repaired separately, which improves the repair efficiency, increases detailed information, makes the blurred target clearer, and makes the repaired target natural transition to the background, avoiding splicing traces affecting the visual effect of the image.
Smart Images

Figure CN114119376B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present disclosure relate to the field of image processing technology, and in particular to an image processing method, device, electronic device, and non-transitory computer-readable storage medium. Background Art
[0002] For an image with a face, there may be problems such as compression noise and face blur that affect the image quality, which in turn affects the visual effect of the image and makes the user experience of viewing the image poor. Therefore, it is urgent to provide an image processing solution to process images with faces and thus improve the image quality. Summary of the invention
[0003] In order to solve at least one problem existing in the prior art, at least one embodiment of the present disclosure provides an image processing method, apparatus, electronic device, and non-transitory computer-readable storage medium.
[0004] In a first aspect, an embodiment of the present disclosure provides an image processing method, the method comprising:
[0005] Acquire a first image; the first image includes at least one object to be repaired;
[0006] Detecting a region of interest of the object to be repaired in the first image;
[0007] Extracting a second image in the region of interest, and repairing the second image;
[0008] Based on the coordinate position of the region of interest in the first image, pasting the restored second image back to the first image;
[0009] The splicing traces produced by the re-pasting are eliminated.
[0010] In a second aspect, an embodiment of the present disclosure further provides an image processing device, the device comprising:
[0011] An acquisition unit, configured to acquire a first image; the first image includes at least one object to be repaired;
[0012] A detection unit, used for detecting a region of interest of the object to be repaired in the first image;
[0013] an extraction and restoration unit, configured to extract a second image in the region of interest and restore the second image;
[0014] A pasting unit, configured to paste the restored second image back to the first image based on the coordinate position of the region of interest in the first image;
[0015] The trace processing unit is used to eliminate the splicing traces generated by the re-pasting.
[0016] In a third aspect, an embodiment of the present disclosure further proposes an electronic device, comprising: a processor and a memory; the processor is used to execute the steps of the method described in the first aspect by calling a program or instruction stored in the memory.
[0017] In a fourth aspect, an embodiment of the present disclosure further proposes a non-transitory computer-readable storage medium for storing a program or instruction, wherein the program or instruction enables a computer to execute the steps of the method described in the first aspect.
[0018] It can be seen that in at least one embodiment of the present disclosure, by detecting the region of interest of the target to be repaired in the original image, the image in the region of interest (including the target to be repaired) can be extracted for repair, so that the target can be repaired separately without repairing the entire image, thereby improving the repair efficiency; in addition, by adding detail information through repair, the more blurred target in the original image becomes clearer; in addition, the repaired target is pasted back to the original image and the splicing marks produced by the pasting are eliminated, so that the transition between the repaired target and the background in the original image is more natural, thereby avoiding the splicing marks from affecting the visual effect of the image. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure, and a person skilled in the art can also obtain other drawings based on these drawings.
[0020] Figure 1 is an exemplary application scenario diagram;
[0021] Figure 2 is an exemplary block diagram of an image processing device provided by an embodiment of the present disclosure;
[0022] Figure 3 is an exemplary block diagram of a separate repair module provided by an embodiment of the present disclosure;
[0023] Figure 4 is an exemplary block diagram of another separate repair module provided by an embodiment of the present disclosure;
[0024] Figure 5 is an exemplary block diagram of an electronic device provided by an embodiment of the present disclosure;
[0025] Figure 6 is an exemplary flow chart of an image processing method provided by an embodiment of the present disclosure;
[0026] Fig. 7Ais a schematic diagram of an image including an object to be repaired;
[0027] Figure 7B It is a schematic diagram of a mask matrix provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It is understood that the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. The specific embodiments described herein are only used to explain the present disclosure, rather than to limit the present disclosure. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in the field belong to the scope of protection of the present disclosure.
[0029] It should be noted that, in this document, relational terms such as “first” and “second” are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0030] The embodiments of the present disclosure provide an image processing method, device, electronic device and non-transitory computer-readable storage medium. By detecting the region of interest of the target to be repaired in the original image, the image (including the target to be repaired) in the region of interest can be extracted for repair, so as to repair the target separately without repairing the entire image, thereby improving the repair efficiency. In addition, by repairing and adding detail information, the blurry target in the original image becomes clearer. In addition, the repaired target is pasted back to the original image and the splicing traces generated by the pasting are eliminated, so that the transition between the repaired target and the background in the original image is more natural, and the splicing traces are avoided from affecting the visual effect of the image. The embodiments of the present disclosure are applicable to application scenarios of any image with a target to be repaired, wherein the target to be repaired can be any target, such as a face, a vehicle, an animal, etc. that can be distinguished from the background. In some embodiments, the embodiments of the present disclosure can also be applied to other fields, such as the field of video processing, and can realize the processing of video frames. It should be understood that the application scenarios of the embodiments of the present disclosure are only some examples or embodiments of the present disclosure. For ordinary technicians in this field, the present disclosure can also be applied to other similar scenarios without creative work.
[0031] Figure 1 is an exemplary application scenario diagram, Figure 1In the example, the quality of the image with compressed noise and human face is not high. One reason is the compression noise, and the other reason is the blur of the human face. In some embodiments, the image quality is not high for other reasons, for example, there is redundant interference information such as Gaussian noise, Poisson noise, multiplicative noise and salt and pepper noise in the image. Therefore, the image needs to be processed to improve the image quality.
[0032] Figure 1 In the embodiment, an image with compression noise and a face can be input to the image processing device 100, and the image can be processed by the image processing device 100. The image processing device 100 can remove the compression noise in the image and repair the blurred face, and then output the denoised and repaired face image to improve the image quality.
[0033] Figure 2 An exemplary block diagram of an image processing apparatus 200 provided in an embodiment of the present disclosure. In some embodiments, the image processing apparatus 200 may be implemented as Figure 1 The image processing device 100 or a part of the image processing device 100 in FIG.
[0034] like Figure 2 As shown, the image processing device 200 may be divided into a plurality of modules, for example, may include: a denoising module 201, a separate restoration module 202, an overall enhancement module 203, and some other modules that may be used for image processing.
[0035] The denoising module 201 is used to obtain an image with noise information and a target, and remove the noise information in the image to reduce the influence of the noise information on the target restoration. The noise information is, for example, compressed noise, or other types of noise information; the target is, for example, a face, or other types of targets. In some embodiments, the denoising module 201 outputs a denoised image through a first convolutional neural network, wherein the input of the first convolutional neural network is an image with noise, and the output is an image with noise removed. The use of the first convolutional neural network involves the production of training samples. A batch of high-definition images can be collected as labels for the first convolutional neural network, and random intensity compression noise is added to the high-definition images to obtain noise images of different degrees, that is, training samples, and as input for the first convolutional neural network training; the output of the first convolutional neural network training is a denoised image. In some embodiments, if the image obtained by the denoising module 201 does not contain any target, the image is directly sent to the overall enhancement module 203 for processing after denoising.
[0036] The individual repair module 202 is used to repair the face in the image from which the denoising module 201 has removed noise, instead of repairing the entire image, so as to improve the efficiency of face repair. In some embodiments, the individual repair module 202 can detect the area of the face in the image, and then extract the face in the image and repair the face individually. In some embodiments, after repairing the face, the individual repair module 202 pastes the repaired face back to the original image so as to continue processing the image, such as processing by the overall enhancement module 203.
[0037] The overall enhancement module 203 is used to perform overall enhancement on the image including the repaired face after the individual repair module 202 pastes the repaired face back to the original image, thereby improving the overall quality of the image. In some embodiments, the overall enhancement module 203 may perform overall enhancement in the form of denoising and / or sharpening, including denoising and / or sharpening, the purpose of which is to ensure that the image is denoised cleanly as a whole, and the purpose of which is to enhance the image. The denoising of the overall enhancement module 203 can be understood as slight denoising, because most of the noise has been removed in the denoising module 201. In some embodiments, the overall enhancement module 203 outputs the denoised and / or sharpened image through a third convolutional neural network, wherein the input of the third convolutional neural network is an image with noise, and the output is an image after the noise is removed and sharpened. Specifically, the overall enhancement module 203 inputs the image obtained by the elimination process into the third convolutional neural network, and uses the third convolutional neural network to perform denoising and / or sharpening. The use of the third convolutional neural network involves the problem of preparing training samples. A batch of high-definition images can be collected and sharpened as the output of the third convolutional neural network training; slight random noise can be added to the high-definition images as the input of the third convolutional neural network training.
[0038] In some embodiments, the division of each module in the image processing device 200 is only a logical function division, and there may be other division methods in actual implementation, for example, the denoising module 201, the individual restoration module 202 or the overall enhancement module 203 may also be divided into multiple units, for example, the denoising module 201 includes an acquisition unit and a denoising unit, wherein the acquisition unit is used to acquire a first image, and the first image includes at least one target to be restored; the denoising unit is used to remove noise information in the first image to obtain a denoised image, specifically, the denoising unit inputs the first image into a first convolutional neural network, uses the first convolutional neural network to remove noise information in the first image, and outputs the denoised image. At least two modules among the denoising module 201, the individual restoration module 202 and the overall enhancement module 203 may be implemented as one module, for example, the function of the denoising module 201 is performed by the individual restoration module 202, that is, the individual restoration module 202 may include the acquisition unit and the denoising unit. In some embodiments, the denoising module 201 and the overall enhancement module 203 are both optional modules, that is, the image processing device 200 only has a single restoration module 202. It is understandable that each module or unit can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application.
[0039] Figure 3 An exemplary block diagram of a separate repair module 300 provided in an embodiment of the present disclosure. In some embodiments, the separate repair module 300 may be implemented as Figure 2 A separate repair module 202 or a part of a separate repair module 202 in the system.
[0040] like Figure 3 As shown, the separate repair module 300 may include but is not limited to the following units: a detection unit 301, an extraction and repair unit 302, a repainting unit 303, and some other modules that can be used to repair an image.
[0041] The detection unit 301 is used to detect the region of interest (RoI) of the target to be repaired in the first image, wherein the first image is an image including at least one target to be repaired, and the first image may be a denoised image or an undenoised image. In some embodiments, the detection unit 301 detects the region of interest of the target to be repaired in the first image by inputting the first image into a second convolutional neural network. The use of the second convolutional neural network involves the production of training samples. A batch of images marked with RoIs may be collected as labels for the second convolutional neural network, and these images may be used as inputs for the training of the second convolutional neural network; the output of the training of the second convolutional neural network is to add RoIs to these images.
[0042] The extraction and restoration unit 302 is used to extract a second image in the region of interest, where the second image is the image of the target to be restored. The extraction and restoration unit 302 is used to restore the second image, that is, to restore the target to be restored in the first image. In some embodiments, the extraction and restoration unit 302 can extract the second image based on the coordinate position of the region of interest. In some embodiments, the extraction and restoration unit 302 can output the restored second image through a generative adversarial network (GAN), wherein the input of the generative adversarial network is the image to be restored, and the output is the restored image. Specifically, the extraction and restoration unit 302 uses the generative adversarial network to restore the second image by inputting the second image into the generative adversarial network. In some embodiments, the extraction and restoration unit 302 uses GAN to generate details for the target, so that low-definition and blurred targets become high-definition and have richer detailed textures.
[0043] The generative adversarial network includes a generator and a discriminator, and the generator and the discriminator learn and train by playing games with each other: the generator generates a restored image; the discriminator determines the probability that the restored image is a real image or a generated image, and the generator updates the parameters of the generator based on the probability, while the discriminator updates the parameters of the discriminator based on the first discriminant value (probability value) of the generated image and the second discriminant value (probability value) of the real image. Through continuous adversarial iterations, the discriminator in the generative adversarial network can more accurately determine whether the received image is a real image or a generated image, so that the generator can generate a generated image that is indistinguishable from the real image, thereby completing the training requirements.
[0044] The “training requirement” may be whether the image generated by the generator of the generative adversarial network meets the preset requirements. For example, the training requirement is that the probability that the discriminator judges that the image generated by the generator is a real image or a generated image converges. The sum of the probability that the image is a real image and the probability that the image is a generated image is 1. Due to the convergence property of the function, the training requirement of the generative adversarial network may be, for example: the probability that the repaired image is a real image or the probability that the image is a generated image is close to 0.5. When it is judged that the training requirements are not met, adversarial iterative training may be continued until the discriminator judges that the probability of being a real image or a generated image meets the requirements (both are close to 0.5). In this embodiment, the real image is a high-definition target, the generated image is an image obtained by repairing a blurred image, and the blurred image is an image obtained after blurring the high-definition image.
[0045] The pasting unit 303 is used to paste the restored second image back to the first image. In some embodiments, the pasting unit 303 pastes the restored second image back to the first image based on the coordinate position of the region of interest in the first image.
[0046] In some embodiments, the division of each unit in the separate repair module 300 is only a logical function division, and there may be other division methods in actual implementation. For example, at least two units among the detection unit 301, the extraction and repair unit 302 and the re-pasting unit 303 can be implemented as one unit; the detection unit 301, the extraction and repair unit 302 or the re-pasting unit 303 can also be divided into multiple sub-units. It is understandable that each unit or sub-unit can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application.
[0047] Figure 4 An exemplary block diagram of a separate repair module 400 provided in an embodiment of the present disclosure. In some embodiments, the separate repair module 400 may be implemented as Figure 3 A separate repair module 300 or a part of a separate repair module 300.
[0048] like Figure 4As shown, the independent repair module 400 may include but is not limited to the following units: a detection unit 401, an extraction and repair unit 402, a back-pasting unit 403, a trace processing unit 404, and other modules that can be used to repair an image. The detection unit 401, the extraction and repair unit 402, and the back-pasting unit 403 are respectively the same as the detection unit 301, the extraction and repair unit 302, and the back-pasting unit 303 of the independent repair module 300, and are not described in detail.
[0049] The trace processing unit 404 is used to eliminate the splicing traces generated by the re-pasting. The splicing traces are caused by the difference in brightness, color, etc. between the background of the restored second image and the first image, so that after the restored second image is pasted back to the first image, the splicing traces appear at the re-pasting boundary. In some embodiments, the trace processing unit 404 performs mask fusion on the splicing traces generated by the re-pasting, and forms a fusion area at the boundary of the region of interest. In some embodiments, the trace processing unit 404 performs mask fusion in a progressive manner so that the transition of the re-pasting boundary is natural. In some embodiments, the progressive manner refers to that the mask weight corresponding to each pixel in the fusion area gradually increases with the increase of the distance, and the distance is the distance between the pixel and the boundary of the region of interest. In some embodiments, the progressive increase is a linear increase or a nonlinear increase. In some embodiments, the fusion area is located in the region of interest, and the fusion area is an area distributed in a strip along the boundary of the region of interest.
[0050] For example, Fig. 7A The white area in the middle is the target to be repaired, the black area represents the background, and the boundary of the white area is the boundary of the region of interest. Figure 7B White in the middle indicates a mask weight of 0, and black indicates a mask weight of 1. Figure 7B In the figure, the box in the region of interest (for ease of understanding only, the box itself does not actually exist) represents the boundary of the fusion region, that is, the area between the boundary of the region of interest and the box is the fusion region, which is distributed in a strip and is located in the region of interest. Since the mask weight increases gradually with the distance, Figure 7B The closer to the boundary of the region of interest, the smaller the mask weight and the darker the color. Figure 7B It can be seen from the gray area that the mask weight gradually increases with the distance, and the color gradually changes from black to white.
[0051] In some embodiments, the trace processing unit 404 may determine the fusion area based on the size of the region of interest. In some embodiments, the trace processing unit 404 may determine the width of the fusion area based on the border size of the region of interest. For example, the width of the fusion area is a preset multiple of the border size of the region of interest, and the preset multiple is a positive decimal, such as Fig. 7A In the example, the region of interest is a square with a side length of 10 cm, then the vertical width of the fusion region is 0.15 times the side length, i.e., 1.5 cm, and the horizontal width of the fusion region is 0.15 times the side length, i.e., 1.5 cm. It should be noted that the preset multiple corresponding to the vertical width of the fusion region may be different from the preset multiple corresponding to the horizontal width. In some embodiments, the trace processing unit 404 may scale the region of interest in proportion, and determine the fusion region based on the scaled region and the region of interest. For example, Fig. 7A After the white area (i.e., the area of interest) in the image is scaled proportionally, the image is as follows: Figure 7B The area indicated by the box and the area between the boundary of the area indicated by the box and the white area are the fusion areas.
[0052] In some embodiments, the trace processing unit 404 may determine the mask weight corresponding to each pixel in the fusion region. The trace processing unit 404 may determine the mask weight corresponding to each pixel according to the distance between each pixel in the fusion region and the boundary of the region of interest.
[0053] In some embodiments, the trace processing unit 404 may perform AND operation processing on the pixel value of each pixel in the fusion area and the corresponding mask weight to complete the mask fusion. In some embodiments, the trace processing unit 404 may determine the mask matrix based on the mask weight corresponding to each pixel in the fusion area and the mask weight corresponding to each pixel in the non-fusion area of the first image; wherein the number of elements of the mask matrix is the same as the number of pixels in the first image, and the element values of the mask matrix are the mask weights. For example, Figure 7B In the mask matrix shown, black indicates that the mask weight is 0, corresponding to the background in the first image, specifically, corresponding to the background in the non-fused area of the first image; white indicates that the mask weight is 1, corresponding to the region of interest in the non-fused area of the first image, that is, corresponding to the region of interest, and not corresponding to the fused area. The trace processing unit 404 performs mask fusion based on the mask matrix, the first image and the restored second image. For example, mask fusion is performed by the following formula:
[0054] BI=mask×B+(1-mask)×A
[0055] Among them, BI is the fused image obtained by mask fusion, mask is the mask matrix, B is the image obtained by pasting the repaired second image back to the first image, and A is the first image or the image obtained after denoising of the first image.
[0056] In some embodiments, the division of each unit in the separate repair module 400 is only a logical function division, and there may be other division methods in actual implementation. For example, at least two units of the detection unit 401, the extraction and repair unit 402, the re-pasting unit 403 and the trace processing unit 404 can be implemented as one unit; the detection unit 401, the extraction and repair unit 402, the re-pasting unit 403 or the trace processing unit 404 can also be divided into multiple sub-units. It is understandable that each unit or sub-unit can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application.
[0057] Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Figure 5 As shown, the electronic device includes: at least one processor 501, at least one memory 502 and at least one communication interface 503. The various components in the electronic device are coupled together through a bus system 504. The communication interface 503 is used to transmit information between external devices. It can be understood that the bus system 504 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 504 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 5 Various buses are labeled as bus system 504 .
[0058] It can be understood that the memory 502 in this embodiment can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
[0059] In some implementations, the memory 502 stores the following elements, executable units or data structures, or a subset thereof, or an extended set thereof: an operating system and application programs.
[0060] The operating system includes various system programs, such as a framework layer, a core library layer, a driver layer, etc., which are used to implement various basic services and process hardware-based tasks. The application includes various application programs, such as a media player (Media Player), a browser (Browser), etc., which are used to implement various application services. The program for implementing the image processing method provided by the embodiment of the present disclosure can be included in the application.
[0061] In the embodiment of the present disclosure, the processor 501 calls the program or instructions stored in the memory 502, specifically, the program or instructions stored in the application, and the processor 501 is used to execute the steps of each embodiment of the image processing method provided in the embodiment of the present disclosure.
[0062] The image processing method provided in the embodiment of the present disclosure can be applied to the processor 501, or implemented by the processor 501. The processor 501 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in the processor 501 or an instruction in the form of software. The above-mentioned processor 501 can be a general-purpose processor, a digital signal processor (Digital Signal Processor, DSP), an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), a field programmable gate array (Field Programmable Gate Array, FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.
[0063] The steps of the image processing method provided in the embodiment of the present disclosure can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software units in the decoding processor. The software unit can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 502, and the processor 501 reads the information in the memory 502 and completes the steps of the method in combination with its hardware.
[0064] Figure 6 This is an exemplary flow chart of an image processing method provided in an embodiment of the present disclosure. The execution subject of the method is an electronic device. For ease of description, the flow of the image processing method is described in the following embodiments using the electronic device as the execution subject.
[0065] In step 601, the electronic device acquires a first image, wherein the first image includes at least one object to be repaired. The first image can be understood as Figure 1 In the image with compressed noise and human face, the target to be repaired is a human face.
[0066] Since the first image has compression noise, the electronic device removes the noise information in the first image after acquiring the first image to obtain a denoised image, thereby reducing the influence of noise on target restoration. In some embodiments, the electronic device outputs a denoised image through a first convolutional neural network, wherein the input of the first convolutional neural network is an image with noise, and the output is an image with noise removed. Specifically, the electronic device inputs the first image into the first convolutional neural network, removes the noise information in the first image using the first convolutional neural network, and outputs the denoised image. The use of the first convolutional neural network involves the production of training samples. A batch of high-definition images can be collected as markers for the first convolutional neural network, and compression noise of random intensity is added to the high-definition images to obtain noise images of different degrees, i.e., training samples, and serve as inputs for the training of the first convolutional neural network; the output of the training of the first convolutional neural network is a denoised image. In some embodiments, if the image acquired by the electronic device does not carry any target, the image is directly enhanced as a whole after the denoising process.
[0067] For the convenience of description, the first images mentioned below are all denoised first images.
[0068] In step 602, the electronic device detects the region of interest of the target to be repaired in the first image. In some embodiments, the electronic device outputs the region of interest of the target to be repaired in the first image through a second convolutional neural network, wherein the input of the second convolutional neural network is an image including the target to be repaired, and the output is the region of interest of the target to be repaired in the input image. Specifically, the electronic device detects the region of interest of the target to be repaired in the first image by inputting the first image into the second convolutional neural network. The use of the second convolutional neural network involves the production of training samples. A batch of images marked with RoI can be collected as labels for the second convolutional neural network, and these images are used as input for the training of the second convolutional neural network; the output of the second convolutional neural network training is to add RoI to these images.
[0069] In step 603, the electronic device extracts a second image in the region of interest, where the second image is the image of the target to be repaired, and then repairs the second image, that is, repairs the target to be repaired in the first image. In some embodiments, the electronic device may extract the second image based on the coordinate position of the region of interest. In some embodiments, the electronic device may output the repaired second image through a generative adversarial network, wherein the input of the generative adversarial network is the image to be repaired, and the output is the repaired image. Specifically, the electronic device inputs the second image into the generative adversarial network and repairs the second image using the generative adversarial network. In some embodiments, the electronic device uses a generative adversarial network to generate details for the target, so that low-definition and blurred targets become high-definition and have richer detail textures.
[0070] In step 604, the electronic device pastes the restored second image back to the first image. In some embodiments, the electronic device pastes the restored second image back to the first image based on the coordinate position of the region of interest in the first image.
[0071] In step 605, the electronic device removes the splicing marks generated by the pasting. The splicing marks are caused by differences in brightness, color, etc. between the restored second image and the background of the first image, resulting in splicing marks appearing at the pasting boundary after the restored second image is pasted back to the first image.
[0072] In some embodiments, the electronic device performs mask fusion on the splicing traces produced by the re-pasting, and forms a fusion area at the boundary of the region of interest. In some embodiments, the electronic device performs mask fusion in a progressive manner so that the transition of the re-pasting boundary is natural. In some embodiments, the progressive manner refers to that the mask weight corresponding to each pixel in the fusion area gradually increases with the increase of the distance, and the distance is the distance between the pixel and the boundary of the region of interest. In some embodiments, the progressive increase is a linear increase or a nonlinear increase. In some embodiments, the fusion area is located in the region of interest, and the fusion area is an area distributed in a band along the boundary of the region of interest.
[0073] For example, Fig. 7A The white area in the middle is the target to be repaired, the black area represents the background, and the boundary of the white area is the boundary of the region of interest. Figure 7B White in the middle indicates a mask weight of 0, and black indicates a mask weight of 1. Figure 7B In the figure, the box in the region of interest (for ease of understanding only, the box itself does not actually exist) represents the boundary of the fusion region, that is, the area between the boundary of the region of interest and the box is the fusion region, which is distributed in a strip and is located in the region of interest. Since the mask weight increases gradually with the distance, Figure 7B The closer to the boundary of the region of interest, the smaller the mask weight and the darker the color. Figure 7B It can be seen from the gray area that the mask weight gradually increases with the distance, and the color gradually changes from black to white.
[0074] In some embodiments, the electronic device may determine the fusion area based on the size of the region of interest, and further determine the mask weight corresponding to each pixel in the fusion area, thereby performing AND operation on the pixel value of each pixel in the fusion area and the corresponding mask weight.
[0075] In some embodiments, the electronic device may determine the width of the fusion area based on the border size of the region of interest. For example, the width of the fusion area is a preset multiple of the border size of the region of interest, and the preset multiple is a positive decimal, such as Fig. 7A In the example, the region of interest is a square with a side length of 10 cm, then the vertical width of the fusion region is 0.15 times the side length, that is, 1.5 cm, and the horizontal width of the fusion region is 0.15 times the side length, that is, 1.5 cm. It should be noted that the preset multiple corresponding to the vertical width of the fusion region may be different from the preset multiple corresponding to the horizontal width. In some embodiments, the electronic device may scale the region of interest in proportion and determine the fusion region based on the scaled region and the region of interest. For example, Fig. 7A After the white area (i.e., the area of interest) in the image is scaled proportionally, the image is as follows: Figure 7B The area indicated by the box and the area between the boundary of the area indicated by the box and the white area are the fusion areas.
[0076] In some embodiments, the electronic device may determine the mask weight corresponding to each pixel according to the distance between each pixel in the fusion area and the boundary of the region of interest.
[0077] In some embodiments, the electronic device may perform AND operations on the pixel values of each pixel in the fusion area and the corresponding mask weights to complete mask fusion. In some embodiments, the electronic device may determine a mask matrix based on the mask weights corresponding to each pixel in the fusion area and the mask weights corresponding to each pixel in the non-fusion area of the first image, and then perform mask fusion based on the mask matrix, the first image and the restored second image. The number of elements in the mask matrix is the same as the number of pixels in the first image, and the element values of the mask matrix are mask weights. For example, Figure 7BIn the mask matrix shown, black indicates that the mask weight is 0, corresponding to the background in the first image, specifically, corresponding to the background in the non-fused area of the first image; white indicates that the mask weight is 1, corresponding to the region of interest in the non-fused area of the first image, that is, corresponding to the region of interest, and not corresponding to the fused area. In some embodiments, mask fusion is performed by the following formula:
[0078] BI=mask×B+(1-mask)×A
[0079] Among them, BI is the fused image obtained by mask fusion, mask is the mask matrix, B is the image obtained by pasting the repaired second image back to the first image, and A is the first image or the image obtained after denoising of the first image.
[0080] In some embodiments, after the electronic device pastes the repaired face back to the original image, the image including the repaired face is enhanced as a whole, thereby improving the overall quality of the image. In some embodiments, the electronic device can perform overall enhancement by denoising and / or sharpening, including denoising and / or sharpening, the purpose of which is to ensure that the image is denoised cleanly as a whole, and the purpose of which is to enhance the image, wherein the denoising can be understood as slight denoising, because most of the noise has been removed after the image is acquired. In some embodiments, the electronic device outputs a denoised and / or sharpened image through a third convolutional neural network, wherein the input of the third convolutional neural network is an image with noise, and the output is an image after the noise is removed and sharpened. Specifically, the electronic device inputs the image obtained by the elimination process into the third convolutional neural network, and uses the third convolutional neural network to perform denoising and / or sharpening. Using the third convolutional neural network involves the production of training samples. A batch of high-definition images can be collected, and the high-definition images can be sharpened as the output of the third convolutional neural network training; slight random noise is added to the high-definition images as the input of the third convolutional neural network training.
[0081] Based on the image processing methods described in the above embodiments, it can be seen that the present disclosure realizes image processing by removing noise → target detection, extraction, and repair → pasting after repair → eliminating splicing traces → overall enhancement. The processed image has a natural transition and high quality, which improves the user viewing experience.
[0082] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all described as a series of action combinations, but those skilled in the art can understand that the embodiments of the present disclosure are not limited by the described action sequence, because according to the embodiments of the present disclosure, some steps can be performed in other sequences or simultaneously. In addition, those skilled in the art can understand that the embodiments described in the specification are all optional embodiments.
[0083] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium, which stores programs or instructions. The programs or instructions enable a computer to execute the steps of the various embodiments of the image processing method. To avoid repeated description, they are not repeated here.
[0084] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "includes..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0085] Those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features from different embodiments is meant to be within the scope of the present disclosure and to form different embodiments.
[0086] Those skilled in the art will appreciate that the description of each embodiment has its own emphasis, and for parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0087] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a first image, wherein the first image includes at least one object to be repaired; Detecting a region of interest of the object to be repaired in the first image; Extracting a second image in the region of interest, and repairing the second image; Based on the coordinate position of the region of interest in the first image, pasting the restored second image back to the first image; Eliminating the splicing traces produced by the re-pasting; The removing process of the splicing traces generated by the re-posting includes: Performing mask fusion on the splicing traces generated by the back pasting to form a fusion area at the boundary of the region of interest; The mask fusion includes: determining a fusion region based on a size of the region of interest; Determining a mask weight corresponding to each pixel in the fusion area; Determine a mask matrix based on the mask weights corresponding to each pixel in the fused area and the mask weights corresponding to each pixel in the non-fused area of the first image; wherein the number of elements of the mask matrix is the same as the number of pixels in the first image, and the element values of the mask matrix are mask weights; Performing mask fusion based on the mask matrix, the first image and the restored second image; The performing mask fusion based on the mask matrix, the first image and the restored second image includes: Mask fusion is performed by the following formula: BI=mask×B+(1-mask)×A Among them, BI is the fused image obtained by mask fusion, mask is the mask matrix, B is the image obtained by pasting the repaired second image back to the first image, and A is the first image.
2. The method according to claim 1, characterized in that After acquiring the first image, the method further includes: Removing noise information from the first image to obtain a denoised image; Accordingly, a region of interest of the object to be repaired in the denoised image is detected.
3. The method according to claim 2, characterized in that Removing noise information from the first image to obtain a denoised image includes: The first image is input into a first convolutional neural network, noise information in the first image is removed by using the first convolutional neural network, and the denoised image is output.
4. The method according to claim 1, characterized in that: The detecting the region of interest of the target to be repaired in the first image comprises: By inputting the first image into a second convolutional neural network, the second convolutional neural network is used to detect the region of interest of the object to be repaired in the first image.
5. The method according to claim 1, characterized in that The repairing the second image comprises: The second image is input into a generative adversarial network, and the generative adversarial network is used to repair the second image.
6. The method according to claim 1, characterized in that The mask weight corresponding to each pixel in the fusion area increases as the distance increases, and the distance is the distance between the pixel and the boundary of the region of interest.
7. The method according to claim 1, characterized in that The fusion region is located within the region of interest, and the fusion region is a region distributed in a strip shape along a boundary of the region of interest.
8. The method according to claim 1, characterized in that The mask fusion includes: determining a fusion region based on a size of the region of interest; Determining a mask weight corresponding to each pixel in the fusion area; The pixel value of each pixel in the fusion area is processed by an AND operation with the corresponding mask weight.
9. The method according to claim 1 or 8, characterized in that: The determining of the fusion region based on the size of the region of interest comprises: Determining a width of a fusion region based on a boundary size of the region of interest; or, The region of interest is scaled in equal proportion, and a fusion region is determined based on the scaled region and the region of interest.
10. The method according to claim 1, characterized in that If an element in the mask matrix corresponds to the background in the non-fused area of the first image, the element value of the element is 0; If an element in the mask matrix corresponds to a region of interest in a non-fused region of the first image, the element value of the element is 1.
11. The method according to claim 1, characterized in that: After removing the splicing traces generated by the re-posting, the method further includes: The image obtained by the elimination process is subjected to denoising and / or sharpening processing.
12. The method according to claim 11, characterized in that The performing denoising and / or sharpening processing on the image obtained by the elimination processing comprises: The image obtained by the elimination process is input into a third convolutional neural network, and the third convolutional neural network is used to perform denoising and / or sharpening processing.
13. An image processing device, characterized in that: The device comprises: An acquisition unit, configured to acquire a first image, wherein the first image includes at least one object to be repaired; A detection unit, used for detecting a region of interest of the object to be repaired in the first image; an extraction and restoration unit, configured to extract a second image in the region of interest and restore the second image; A pasting unit, configured to paste the restored second image back to the first image based on the coordinate position of the region of interest in the first image; A trace processing unit, used for removing the splicing traces generated by the re-pasting; When the trace processing unit performs the removal processing on the splicing traces generated by the re-pasting, it is specifically used to: Performing mask fusion on the splicing traces generated by the back pasting to form a fusion area at the boundary of the region of interest; The trace processing unit is specifically used for: determining a fusion region based on a size of the region of interest; Determining a mask weight corresponding to each pixel in the fusion area; Determine a mask matrix based on the mask weights corresponding to each pixel in the fused area and the mask weights corresponding to each pixel in the non-fused area of the first image; wherein the number of elements of the mask matrix is the same as the number of pixels in the first image, and the element values of the mask matrix are mask weights; Performing mask fusion based on the mask matrix, the first image and the restored second image; When the trace processing unit performs mask fusion based on the mask matrix, the first image and the restored second image, it is specifically used to: Mask fusion is performed by the following formula: BI=mask×B+(1-mask)×A Among them, BI is the fused image obtained by mask fusion, mask is the mask matrix, B is the image obtained by pasting the repaired second image back to the first image, and A is the first image.
14. An electronic device, characterized in that: include: Processor and memory; The processor is used to execute the steps of the method according to any one of claims 1 to 12 by calling the program or instruction stored in the memory.
15. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores a program or instruction, which enables a computer to execute the steps of the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Panoramic image fusion method, system and image processing equipment
CN101951487A
Image processing method and device, storage medium and electronic equipment
CN109978754A
Image processing method and device, electronic equipment and computer readable storage medium
CN111325657A