Image processing method, image processing network training method, image processing network training device and image processing network training equipment

By traversing the images based on window and step size and fusion of extended Hadamar operators, the problem of insufficient image processing efficiency and effect is solved, and more efficient and accurate image alignment is achieved.

CN120375129APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410211571.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

There are shortcomings in the improvement of image processing efficiency and effect in the prior art, especially in the fields of security monitoring, autonomous driving and medical diagnosis.

Method used

By traversing the image based on the window and step size of the preset size, the image block group corresponding to each pixel point is obtained, and the image block fusion is used to determine the aligned image.

Benefits of technology

It improves the efficiency and accuracy of image processing, reduces the computing power requirement, and improves the effect of image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375129A_ABST
    Figure CN120375129A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method, an image processing network training method, an image processing network training device and image processing network training equipment, and relates to the technical field of computers. Comprising the following steps: after obtaining a to-be-processed image set, firstly, respectively traversing a first image and a second image based on a window and a step length of a preset size to obtain an image block group corresponding to each pixel point, and then fusing a first image block and a second image block in each image block group to obtain a fused image; according to the method, the first image and the second image are integrated to obtain the pixel value after the corresponding pixel points are integrated, and finally, the third image after the first image and the second image are aligned is determined according to the pixel value after each pixel point is integrated, so that conditions are provided for improving the image processing effect, and the image processing efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to an image processing method, a training method, an apparatus, and a device for an image processing network. Background Art

[0002] With the development of deep learning and artificial intelligence technologies, great progress has been made in the fields of image processing and computer vision. In practical applications, image processing and computer vision technologies can be applied to multiple fields such as security monitoring, autonomous driving, medical diagnosis, and intelligent manufacturing. These fields all involve image processing. At present, how to improve the effect and efficiency of image processing is an urgent problem to be solved. Summary of the Invention

[0003] The present disclosure aims to at least solve one of the technical problems in the related art to some extent.

[0004] A first aspect embodiment of the present disclosure provides an image processing method, including:

[0005] Obtaining an image set to be processed, where the image set includes a first image and a second image;

[0006] Traversing the first image and the second image respectively based on a window with a preset size and a step length, and obtaining an image block group corresponding to each pixel point, where each image block group includes a first image block located in the first image and a second image block located in the second image;

[0007] Fusing the first image block and the second image block in each image block group to obtain a pixel value after fusion of the corresponding pixel point;

[0008] Determining a third image after alignment of the first image and the second image according to the pixel value after fusion of each pixel point.

[0009] A second aspect embodiment of the present disclosure provides a training method for an image processing network, including:

[0010] Obtaining a training data set, where the training data set includes a plurality of data groups, and each data group includes at least two input images, a first format target map and a second format target map corresponding to the at least two input images;

[0011] Traversing the at least two input images respectively based on a window with a preset size and a step length, and obtaining an image block group corresponding to each pixel point, where each image block group includes at least two image blocks respectively located in the input images;

[0012] Input multiple image patches in each group of image patches into an initial neural network to obtain pixel values output by the initial neural network;

[0013] Determine a loss value according to a first difference between the input image and a first-format target image, and a second difference between the output pixel values and a second-format target image;

[0014] Based on the loss value, correct the initial neural network until a trained neural network is obtained.

[0015] An embodiment of the third aspect of the present disclosure provides an image processing apparatus, including:

[0016] A first acquisition module, configured to acquire a set of images to be processed, where the set of images includes a first image and a second image;

[0017] A first traversal module, configured to traverse the first image and the second image respectively based on a window of a preset size and a step length, and acquire a group of image patches corresponding to each pixel point, where each group of image patches includes a first image patch located in the first image and a second image patch located in the second image;

[0018] A fusion module, configured to fuse the first image patch and the second image patch in each group of image patches to obtain pixel values of the corresponding pixel points after fusion;

[0019] A first determination module, configured to determine a third image after alignment of the first image and the second image according to the pixel values of each pixel point after fusion.

[0020] An embodiment of the fourth aspect of the present disclosure provides a training apparatus for an image processing network, including:

[0021] A second acquisition module, configured to acquire a training data set, where the training data set includes multiple data groups, and each data group includes at least two input images, a first-format target image and a second-format target image corresponding to the at least two input images;

[0022] A second traversal module, configured to traverse the at least two input images respectively based on a window of a preset size and a step length, and acquire a group of image patches corresponding to each pixel point, where each group of image patches includes at least two image patches respectively located in the input images;

[0023] An input module, configured to input multiple image patches in each group of image patches into an initial neural network to obtain pixel values output by the initial neural network;

[0024] A second determination module, configured to determine a loss value according to a first difference between the input image and the first-format target image, and a first difference between the output pixel value and the first-format target image;

[0025] A correction module, configured to correct the initial neural network based on the loss value until a trained neural network is obtained.

[0026] An embodiment of the fifth aspect of the present disclosure provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the image processing method proposed in the embodiment of the first aspect of the present disclosure and the training method of the image processing network proposed in the embodiment of the second aspect of the present disclosure are implemented.

[0027] An embodiment of the sixth aspect of the present disclosure provides a computer-readable storage medium, storing a computer program, which when executed by a processor, implements the image processing method proposed in the embodiment of the first aspect of the present disclosure and the training method of the image processing network proposed in the embodiment of the second aspect of the present disclosure.

[0028] The image processing method, the training method of the image processing network, the device and the equipment provided by the present disclosure have the following beneficial effects:

[0029] In the embodiment of the present disclosure, after obtaining the image set to be processed, first, based on a window with a preset size and a step size, the first image and the second image are traversed respectively to obtain an image block group corresponding to each pixel point. Then, the first image block and the second image block in each image block group are fused to obtain the pixel value after fusion of the corresponding pixel point. Finally, according to the pixel value after fusion of each pixel point, a third image after alignment of the first image and the second image is determined. Thus, conditions are provided for improving the effect of image processing, and the efficiency of image processing is improved.

[0030] The additional aspects and advantages of the present disclosure will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The above and / or additional aspects and advantages of the present disclosure will become obvious and easy to understand from the following description of the embodiments with reference to the drawings, where:

[0032] Figure 1 is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure;

[0033] Figure 2 is a schematic diagram of extracting image blocks provided by an embodiment of the present disclosure;

[0034] Figure 3 Flow diagram of an image processing method provided by another embodiment of the present disclosure;

[0035] Figure 4 Schematic diagram of the image processing effect provided by the embodiment of the present disclosure;

[0036] Figure 5 Flow diagram of an image processing method provided by another embodiment of the present disclosure;

[0037] Figure 6 Flow diagram of an image processing method provided by another embodiment of the present disclosure;

[0038] Figure 7 Flow diagram of a training method for an image processing network provided by another embodiment of the present disclosure;

[0039] Figure 8 Schematic diagram of the structure of a neural network provided by the embodiment of the present disclosure;

[0040] Figure 9 Schematic diagram of the structure of an image processing device provided by another embodiment of the present disclosure;

[0041] Figure 10 Schematic diagram of the structure of a training device for an image processing network provided by another embodiment of the present disclosure;

[0042] Figure 11 Shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure. Detailed implementation manners

[0043] The embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as a limitation of the present disclosure.

[0044] The image processing method, the training method of the image processing network, the device and the equipment of the embodiments of the present disclosure will be described below with reference to the accompanying drawings.

[0045] Figure 1 Flow diagram of an image processing method provided by the embodiment of the present disclosure.

[0046] As Figure 1 shown, the image processing method may include the following steps:

[0047] Step 101, obtaining an image set to be processed, where the image set includes a first image and a second image.

[0048] Among them, the image set can be any image set composed of consecutive multiple frames of image data of the same scene. For example, the image set can be composed of consecutive multiple frames of image data obtained by taking pictures in any scene, etc., and the present disclosure does not limit this.

[0049] Among them, the first image and the second image can be image data of different frames of the same scene respectively. For example, the first image can be the image data of the first frame in consecutive multiple frames, and the second image can be the image data of the second frame in consecutive multiple frames, etc., and the present disclosure does not limit this.

[0050] Step 102: Traverse the first image and the second image respectively based on a window of a preset size and a step size, and obtain an image block group corresponding to each pixel point.

[0051] Among them, each image block group includes a first image block located in the first image and a second image block located in the second image.

[0052] Among them, the size of the preset window is the size of the sliding window, which can be preset, and the present disclosure does not limit this.

[0053] In some embodiments, the size of the preset window can be n*n, where n is an integer greater than 1. For example, the size of the preset window can be 3*3, etc., and the present disclosure does not limit this.

[0054] Among them, the preset step size is the distance that the sliding window moves in each step of sliding, which can be preset. For example, the preset step size can be 1, etc., and the present disclosure does not limit this.

[0055] Among them, the image block group can be composed of image blocks corresponding to pixel points at the same position in images of different frames.

[0056] Among them, the image block can be an image feature block extracted by a preset window during the process of traversing the image, and the present disclosure does not limit this.

[0057] In the present disclosure, after obtaining the image set, the image can be traversed based on the size of the preset window and the step size to obtain an image block corresponding to each pixel point in the image. As Figure 2 shown, Figure 2 is a schematic diagram of extracting an image block provided by an embodiment of the present disclosure. Figure 2 In a, any matrix block is an image block extracted by a preset window. Figure 2 In b, any bar graph is obtained after unfolding the corresponding image block. For example, Figure 2 after the third image block in a is unfolded, Figure 2 the third bar graph in b is obtained, and the present disclosure does not limit this.

[0058] Step 103: Fuse the first image block and the second image block in each image block group to obtain the pixel value of the corresponding pixel point after fusion.

[0059] In the present disclosure, after obtaining the image block group corresponding to each pixel point, the pixel value F of the corresponding pixel point after fusion can be obtained by multiplying the pixel values at the corresponding positions in the first image block and the second image block of the same pixel point based on formula (1). concat Among them, formula (1) is as follows:

[0060] F concat = F1 * F2 (1)

[0061] Among them, F1 is the feature that needs to be aligned in the first image block, and F2 is the feature that needs to be aligned in the second image block.

[0062] Step 104: Determine the third image after alignment of the first image and the second image according to the pixel value of each pixel point after fusion.

[0063] In the present disclosure, after obtaining the pixel value of each pixel point after fusion, the third image after alignment of the first image and the second image can be determined based on the pixel value of each pixel point after fusion.

[0064] In the embodiment of the present disclosure, after obtaining the image set to be processed, first, based on a window of a preset size and a step size, the first image and the second image are traversed respectively to obtain the image block group corresponding to each pixel point, then the first image block and the second image block in each image block group are fused to obtain the pixel value of the corresponding pixel point after fusion, and finally, according to the pixel value of each pixel point after fusion, the third image after alignment of the first image and the second image is determined. Thus, by traversing the first image and the second image based on the sliding window and the step size, the first image block and the second image block corresponding to each pixel point are obtained, then the first image block and the second image block corresponding to each pixel point are fused to obtain the pixel value after fusion, and the aligned image is determined based on the pixel value of each pixel point after fusion, thereby providing conditions for improving the effect of image processing and improving the efficiency of image processing.

[0065] Figure 3 is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure. As Figure 3 shown, the image processing method may include the following steps:

[0066] Step 301: Obtain the image set to be processed, where the image set includes a first image and a second image.

[0067] Step 302: Traverse the first image and the second image respectively based on a window of a preset size and a step length, and obtain a group of image patches corresponding to each pixel point.

[0068] Among them, each group of image patches includes a first image patch located in the first image and a second image patch located in the second image.

[0069] Among them, for the specific implementation forms of Step 301 to Step 302, reference can be made to the detailed descriptions of other embodiments of the present disclosure, which will not be elaborated here.

[0070] Step 303: Fuse the i-th pixel value in the first image patch with the i-th pixel value in the second image patch to obtain the i-th fused pixel value, where i is a natural number less than or equal to n 2 .

[0071] In the present disclosure, after obtaining a group of image patches corresponding to each pixel point, based on the Expand Hadamard operator, the i-th pixel value in the first image patch corresponding to each pixel point can be multiplied by the i-th pixel value in the second image patch to obtain the i-th fused pixel value.

[0072] It should be noted that the Expand Hadamard operator can calculate the product of the corresponding pixel values in the first image patch and the second image patch. For example, if the preset window size is 3*3, then through the Expand Hadamard operator, the pixel values in the first image patch can be multiplied by the corresponding pixel values in the second image patch to obtain 9 fused pixel values. The present disclosure does not make any limitations in this regard.

[0073] Step 304: Determine the mean value of the sum of n 2 fused pixel values as the pixel value after pixel point fusion.

[0074] In the present disclosure, after obtaining n 2 fused pixel values, the sum of n 2 fused pixel values can be calculated first, and then the mean value of the sum of n 2 fused pixel values can be calculated. The mean value is determined as the pixel value F concat after pixel point fusion, as shown in formula (2):

[0075] Among them, f i is the i-th pixel value in the first image patch, and h i is the i-th pixel value in the second image patch.

[0076] Step 305: Determine a third image after aligning the first image and the second image according to the pixel value after pixel point fusion of each pixel point.

[0077] In this disclosure, by processing the image block group of each pixel point based on the extended Hadamard operator, the range of the receptive field can be effectively expanded, thereby improving the efficiency of image processing. As shown in Table 1, Table 1 is a table of the inference speeds when using different operators provided by the embodiments of this disclosure. Among them, the input image size is 3x128x128.

[0078] Table 1

[0079] Operator Inference speed Ordinary convolution 0.24s Self-attention operator 3.41s Extended Hadamard operator 0.08s

[0080] It can be seen from Table 1 that compared with the inference speed when using convolution or the inference speed when using the self-attention operator, the inference speed when using the extended Hadamard operator is faster and the efficiency is higher.

[0081] In addition, by processing the image with the extended Hadamard operator, the effect of image processing can also be effectively improved. For example, as Figure 4 shown, Figure 4 is a schematic diagram of the image processing effect provided by the embodiments of this disclosure. Among them, Figure 4 a, Figure 4 c, and Figure 4 e are all images obtained after processing with the attention operator, Figure 4 b, Figure 4 d, Figure 4 f are all images obtained after processing with the extended Hadamard operator.

[0082] From Figure 4 a and Figure 4 b, it can be seen that when processing the image with the extended Hadamard operator, the details of the processed image increase. From Figure 4 c and Figure 4 d, it can be seen that when processing the image with the extended Hadamard operator, the pseudo-textures of the processed image decrease. From Figure 4 e and Figure 4 f, it can be seen that when processing the image with the extended Hadamard operator, the lines of the processed image become better.

[0083] In the embodiments of this disclosure, after obtaining the image set to be processed, first, based on a window and a step size of a preset size, the first image and the second image are traversed respectively to obtain the image block group corresponding to each pixel point, then the i-th pixel value in the first image block is fused with the i-th pixel value in the second image block to obtain the i-th fused pixel value, and then n 2The mean value of the sum of the fused pixel values is determined as the pixel value after the fusion of the pixel points. Finally, based on the pixel values after the fusion of each pixel point, the third image after the alignment of the first image and the second image is determined. Thus, by determining the fused pixel value corresponding to each pixel point based on the extended Hadamard operator, and then determining the aligned image based on the pixel values after the fusion of each pixel point, the computing power during image processing is reduced, the accuracy of image alignment is improved, and the efficiency of image processing is enhanced on this basis.

[0084] Figure 5 As shown in the flowchart of an image processing method provided by an embodiment of the present disclosure, Figure 5 the image processing method may include the following steps:

[0085] Step 501: Obtain an image set to be processed, where the image set includes a first image, a second image, and at least one fourth image.

[0086] The fourth image may be image data of any different frame in the same scene as the first image and the second image. For example, when the first image is the image data of the first frame and the second image is the image data of the second frame, the fourth image may be the image of the third frame, etc. The present disclosure does not limit this.

[0087] Step 502: Traverse the first image, the second image, and at least one fourth image respectively based on a window of a preset size and a step size, and obtain an image block group corresponding to each pixel point.

[0088] Each image block group includes a first image block in the first image, a second image block in the second image, and a third image block in the fourth image.

[0089] The specific implementation forms of steps 501 to 502 may refer to the detailed description of other embodiments of the present disclosure, and will not be elaborated here.

[0090] Step 503: Fuse the first image block, the second image block, and the third image block in each image block group to obtain the pixel value after the fusion of the corresponding pixel point.

[0091] In the present disclosure, after obtaining the image block group corresponding to each pixel point, the first image block, the second image block, and the third image block in each image block group may be fused by an extended Hadamard operator to obtain the pixel value after the fusion of the corresponding pixel point. For example, taking the first image block as the reference image block, the second image block and the third image block are respectively fused with the first image block by the extended Hadamard operator to obtain the pixel value after the fusion of the corresponding pixel point. The present disclosure does not limit this.

[0092] Step 504: Determine the third image after alignment of the first image, the second image, and at least one fourth image according to the pixel values after fusion of each pixel point.

[0093] For the specific implementation form of step 504, reference may be made to the detailed description of other embodiments of the present disclosure, which will not be elaborated here.

[0094] In the embodiments of the present disclosure, after obtaining the image set to be processed, first, based on a window with a preset size and a step size, traverse the first image, the second image, and at least one fourth image respectively to obtain a group of image patches corresponding to each pixel point. Then, fuse the first image patch, the second image patch, and the third image patch in each group of image patches to obtain the pixel value after fusion of the corresponding pixel point. Finally, determine the third image after alignment of the first image, the second image, and at least one fourth image according to the pixel value after fusion of each pixel point. Thus, by fusing the first image patch, the second image patch, and the third image patch in each group of image patches based on the extended Hadamard operator, the pixel value after fusion of each pixel point is obtained, and the aligned image is determined based on the pixel value after fusion of each pixel point, thereby improving the accuracy of image processing and enhancing the effect of image processing.

[0095] Figure 6 It is a schematic flowchart of an image processing method provided by an embodiment of the present disclosure. As Figure 6 shown, the image processing method may include the following steps:

[0096] Step 601: Obtain the image set to be processed, where the image set includes a first image and a second image.

[0097] Step 602: Based on a window with a preset size and a step size, traverse the first image and the second image respectively to obtain a group of image patches corresponding to each pixel point.

[0098] Each group of image patches includes a first image patch located in the first image and a second image patch located in the second image.

[0099] For the specific implementation forms of steps 601 to 602, reference may be made to the detailed description of other embodiments of the present disclosure, which will not be elaborated here.

[0100] Step 603: Input the first image patch and the second image patch in each group of image patches into the neural network trained in advance to obtain the pixel value output by the neural network.

[0101] It should be noted that the neural network trained in advance may be any neural network for processing such as aligning and / or denoising the image, and the present disclosure does not limit this.

[0102] In the present disclosure, after obtaining the image block groups corresponding to each pixel, the first image block and the second image block in each image block group can be input into a neural network trained to generate the pixel value after fusion of each pixel output by the neural network.

[0103] Step 604: Determine a third image after alignment of the first image and the second image according to the pixel values output by the neural network.

[0104] In the present disclosure, after obtaining the pixel values output by the neural network, the third image after alignment of the first image and the second image can be determined based on the pixel values output by the neural network.

[0105] In an embodiment of the present disclosure, after obtaining an image set to be processed, first, based on a window of a preset size and a step size, the first image and the second image are traversed respectively to obtain image block groups corresponding to each pixel. Then, the first image block and the second image block in each image block group are input into a neural network trained to generate the pixel values output by the neural network. Finally, a third image after alignment of the first image and the second image is determined according to the pixel values output by the neural network. Thus, by processing the first image block and the second image block in each image block group based on the trained neural network, the pixel value after fusion of each pixel is obtained, and the aligned image is determined based on the pixel value after fusion of each pixel, thereby improving the effect of image processing and the efficiency of image processing.

[0106] Figure 7 is a schematic flowchart of a training method of an image processing network provided by an embodiment of the present disclosure. As Figure 7 shown, the training method of the image processing network may include the following steps:

[0107] Step 701: Obtain a training data set, where the training data set includes multiple data groups, and each data group includes at least two input images, a first format target image and a second format target image corresponding to the at least two input images.

[0108] It should be noted that the training data set can be preset, and the present disclosure does not limit this.

[0109] Among them, the first format can be the image format to which the input image belongs, and it can be any format. For example, the first format can be a Raw image, etc., and the present disclosure does not limit this.

[0110] Among them, the second format can be the format of the output image obtained after inputting the input image into the neural network, and it can be the same as the first format or different from the first format. For example, the second format can be an RGB image, etc., and the present disclosure does not limit this.

[0111] Among them, the first format target image can be the label image of the first format image; the second format target image can be the label image of the second format image. For example, both the first format target image and the second format target image can be images after noise reduction, etc., and the present disclosure does not limit this.

[0112] Step 702: Traverse at least two input images respectively based on a window of a preset size and a step size, and obtain an image block group corresponding to each pixel point, where each image block group includes at least two image blocks respectively located in the input images.

[0113] In the present disclosure, after obtaining at least two input images, the at least two input images can be traversed respectively based on a window of a preset size and a step size to obtain an image block group corresponding to each pixel point.

[0114] Step 703: Input multiple image blocks in each image block group into the initial neural network to obtain the pixel values output by the initial neural network.

[0115] Among them, the initial neural network can be pre-set, and the present disclosure does not limit this.

[0116] In the present disclosure, after obtaining the image block group corresponding to each pixel point of at least two input images, multiple image blocks in each image block group can be input into the initial neural network. First, each image block in each image block group is filtered through convolution and an activation function, and the filtered image blocks are processed through an extended Hadamard operator to obtain the fused pixel value of each pixel point. Then, the fused pixel value of each pixel point is processed through convolution and an activation function to output an aligned image. Finally, the aligned image is subjected to feature fusion through a fusion network to output the pixel value of the fused image.

[0117] It should be noted that in different application scenarios, the fusion network may be different or may also be the same. For example, the fusion network can be an image processing U-Net network, etc., and the present disclosure does not limit this.

[0118] The following combines Figure 8 to give an example of the initial neural network provided by the present disclosure. As Figure 8 shown, taking the image block group of each pixel point of two input images as an example, Figure 8 is a schematic structural diagram of the neural network provided by an embodiment of the present disclosure.

[0119] First, filter the image patch groups of each pixel point through convolution and activation functions. Then, based on the extended Hadamard operator, process the filtered image patch groups of each pixel point to obtain the fused pixel value of each pixel point. After that, filter the fused pixel value of each pixel point through convolution and activation functions to output the aligned image. Finally, perform feature fusion on the aligned image through a fusion network to output the pixel value of the fused image.

[0120] Step 704: Determine the loss value according to the first difference between the input image and the target image in the first format, and the second difference between the output pixel value and the target image in the second format.

[0121] Among them, the first difference can be the loss value between the input image and the target image in the first format; the second difference can be the loss value between the output pixel value and the target image in the second format.

[0122] It should be noted that in the case of at least one input image, the difference between each input image and the target image in the first format can be calculated first, and then the mean value of these at least one difference can be calculated and used as the first difference. The present disclosure does not limit this.

[0123] For example, the first format of the input image is the Raw format, and the second format of the output image of the initial neural network is the RGB format. After obtaining the pixel value output by the initial neural network, the first difference L between the input image and the target image in the first format can be calculated according to formula (3). raw And the second difference L between the output pixel value and the target image in the second format can be calculated according to formula (4). rgb Then, the loss value L can be calculated according to formula (5). total Among them, formula (3), formula (4), and formula (5) are respectively as follows:

[0124]

[0125]

[0126] L total = L rgb + αL raw (5)

[0127] Among them, T raw is the pixel value of the input image, is the pixel value of the target image in the first format, T rgb is the pixel value of the image output by the neural network, is the pixel value of the target image in the second format, and α is a coefficient used to control the ratio of the two loss functions and can be set to 1. The present disclosure does not limit this.

[0128] Step 705: Based on the loss value, correct the initial neural network until the trained neural network is obtained.

[0129] In the present disclosure, after obtaining the loss value, in order to improve the accuracy of the neural network, the initial neural network can be corrected based on the loss value until the trained neural network is obtained.

[0130] In the embodiments of the present disclosure, after obtaining the training data set, first, based on a window of a preset size and a step size, at least two input images are traversed respectively to obtain a group of image patches corresponding to each pixel point, then a plurality of image patches in each group of image patches are input into the initial neural network to obtain the pixel values output by the initial neural network, and then according to the first difference between the input image and the target image in the first format, and the second difference between the output pixel value and the target image in the second format, the loss value is determined, and finally, based on the loss value, the initial neural network is corrected until the trained neural network is obtained. Thus, by training the initial neural network based on the group of image patches corresponding to each pixel point in the input image, calculating the loss value, and correcting the initial neural network based on the loss value to obtain the trained neural network, the accuracy and reliability of the image processing network are improved, providing conditions for improving the effect of image processing.

[0131] To implement the above embodiments, the present disclosure also proposes an image processing apparatus.

[0132] Figure 9 It is a schematic structural diagram of the image processing apparatus provided in the embodiments of the present disclosure.

[0133] As Figure 9 shown, the image processing apparatus 900 may include: a first acquisition module 901, a first traversal module 902, a fusion module 903, and a first determination module 904.

[0134] The first acquisition module 901 is configured to acquire an image set to be processed, where the image set includes a first image and a second image;

[0135] The first traversal module 902 is configured to traverse the first image and the second image respectively based on a window of a preset size and a step size to obtain a group of image patches corresponding to each pixel point, where each group of image patches includes a first image patch located in the first image and a second image patch located in the second image;

[0136] The fusion module 903 is configured to fuse the first image patch and the second image patch in each group of image patches to obtain the pixel value of the corresponding pixel point after fusion;

[0137] The first determination module 904 is configured to determine a third image after alignment of the first image and the second image according to the pixel values after fusion of each pixel point.

[0138] Optionally, the size of the window is n*n, where n is an integer greater than 1.

[0139] Optionally, the fusion module 903 is further configured to:

[0140] Fuse the i-th pixel value in the first image block with the i-th pixel value in the second image block to obtain the i-th fused pixel value, where i is a natural number less than or equal to n 2 ;

[0141] Determine the mean of the sum of n 2 fused pixel values as the pixel value after fusion of the pixel point.

[0142] Optionally, the fusion module 903 is further configured to:

[0143] Fuse the first image block, the second image block, and the third image block in each image block group to obtain the pixel value after fusion of the corresponding pixel point.

[0144] Optionally, the fusion module 903 is further configured to:

[0145] Input the first image block and the second image block in each image block group into a neural network trained in advance to obtain the pixel value output by the neural network.

[0146] For the functions and specific implementation principles of the above modules in the embodiments of the present disclosure, reference may be made to the above method embodiments, and details are not described herein again.

[0147] In the image processing apparatus according to the embodiments of the present disclosure, after obtaining the image set to be processed, first, based on a window and a step size of a preset size, the first image and the second image are respectively traversed to obtain an image block group corresponding to each pixel point, then the first image block and the second image block in each image block group are fused to obtain the pixel value after fusion of the corresponding pixel point, and finally, according to the pixel value after fusion of each pixel point, a third image after alignment of the first image and the second image is determined. Thus, by traversing the first image and the second image based on the sliding window and the step size, the first image block and the second image block corresponding to each pixel point are obtained, then the first image block and the second image block corresponding to each pixel point are fused to obtain the fused pixel value, and based on the pixel value after fusion of each pixel point, the aligned image is determined, thereby providing conditions for improving the effect of image processing and improving the efficiency of image processing.

[0148] To implement the above embodiments, the present disclosure also proposes a training apparatus for an image processing network.

[0149] Figure 10 The structural schematic diagram of the training device for the image processing network provided by the embodiments of the present disclosure.

[0150] As Figure 10 shown, the training device 1000 of the image processing network may include: a second acquisition module 1001, a second traversal module 1002, an input module 1003, a second determination module 1004, and a correction module 1005.

[0151] The second acquisition module 1001 is configured to acquire a training data set, where the training data set includes a plurality of data groups, and each data group includes at least two input images, a first format target image corresponding to at least two input images, and a second format target image;

[0152] The second traversal module 1002 is configured to traverse at least two input images respectively based on a window and a step size of a preset size, and acquire an image block group corresponding to each pixel point, where each image block group includes at least two image blocks respectively located in the input images;

[0153] The input module 1003 is configured to input a plurality of image blocks in each image block group into an initial neural network to acquire pixel values output by the initial neural network;

[0154] The second determination module 1004 is configured to determine a loss value according to a first difference between the input image and the first format target image, and a second difference between the output pixel value and the second format target image;

[0155] The correction module 1005 is configured to correct the initial neural network based on the loss value until a trained neural network is acquired.

[0156] For the functions and specific implementation principles of the above-mentioned modules in the embodiments of the present disclosure, reference may be made to the above-mentioned method embodiments, and details are not described herein again.

[0157] The training device of the image processing network according to the embodiments of the present disclosure, after obtaining the training data set, first traverses at least two input images respectively based on a window of a preset size and a step size to obtain a group of image patches corresponding to each pixel point, then inputs multiple image patches in each group of image patches into the initial neural network to obtain the pixel values output by the initial neural network, then determines the loss value according to the first difference between the input image and the target image in the first format and the second difference between the output pixel value and the target image in the second format, and finally corrects the initial neural network based on the loss value until the trained neural network is obtained. Thus, by training the initial neural network based on the group of image patches corresponding to each pixel point in the input image, calculating the loss value, and correcting the initial neural network based on the loss value to obtain the trained neural network, the accuracy and reliability of the image processing network are improved, providing conditions for enhancing the effect of image processing.

[0158] To implement the above embodiments, the present disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the image processing method and the training method of the image processing network proposed in the foregoing embodiments of the present disclosure.

[0159] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the image processing method and the training method of the image processing network proposed in the foregoing embodiments of the present disclosure.

[0160] Figure 11 A block diagram of an exemplary electronic device suitable for implementing the embodiments of the present disclosure is shown. Figure 11 The illustrated electronic device 12 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0161] As Figure 11 shown, the electronic device 12 is presented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0162] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.

[0163] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including volatile and nonvolatile media, removable and non-removable media.

[0164] Memory 28 may include computer system readable media in the form of volatile memory, such as Random Access Memory (RAM) 50 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 11 not shown, typically referred to as a "hard disk drive"). Although Figure 11 not shown in the figure, a disk drive for reading and writing on a removable nonvolatile disk (such as a "floppy disk"), and an optical disk drive for reading and writing on a removable nonvolatile optical disk (such as Compact Disc Read Only Memory (CD-ROM), Digital Video Disc Read Only Memory (DVD-ROM), or other optical media) can be provided. In these cases, each drive can be connected to bus 18 through one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the various embodiments of the present disclosure.

[0165] A program / utilities 60 having a set (at least one) of program modules 42 can be stored, for example, in a memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present disclosure.

[0166] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a Local Area Network (LAN), a Wide Area Network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0167] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.

[0168] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0169] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the technical features indicated. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of the present disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0170] Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations, where functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present disclosure pertain.

[0171] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0172] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0173] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program. The said program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0174] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0175] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. An image processing method, characterized in that, including: Obtain an image set to be processed, where the image set includes a first image and a second image; Based on a window of a preset size and a step size, traverse the first image and the second image respectively to obtain an image block group corresponding to each pixel point, where each image block group includes a first image block located in the first image and a second image block located in the second image; Fuse the first image block and the second image block in each image block group to obtain the pixel value after fusion of the corresponding pixel point; Determine a third image after alignment of the first image and the second image according to the pixel value after fusion of each pixel point.

2. The method according to claim 1, characterized in that The size of the window is n*n, where n is an integer greater than 1.

3. The method according to claim 2, characterized in that The fusing the first image block and the second image block in each image block group to obtain the pixel value after fusion of the corresponding pixel point includes: Fuse the i-th pixel value in the first image block with the i-th pixel value in the second image block to obtain the i-th fused pixel value, where i is a natural number less than or equal to n 2 ; Determine the mean value of the sum of n 2 fused pixel values as the pixel value after fusion of the pixel point.

4. The method according to any one of claims 1-3, characterized in that, The image set further includes at least one fourth image, and each image block group corresponding to each pixel point further includes at least one third image block. The fusing the first image block and the second image block in each image block group to obtain the pixel value after fusion of the corresponding pixel point includes: Fuse the first image block, the second image block and the third image block in each image block group to obtain the pixel value after fusion of the corresponding pixel point.

5. The method according to any one of claims 1 to 3, characterized in that The fusing the first image block and the second image block in each image block group to obtain the pixel value after fusion of the corresponding pixel point includes: Input the first image block and the second image block in each image block group into a neural network generated by training to obtain the pixel value output by the neural network.

6. A training method for an image processing network, characterized in that, The method includes: Obtain a training data set, where the training data set includes multiple data groups, and each data group includes at least two input images, a first format target image corresponding to the at least two input images, and a second format target image; Based on a window of a preset size and a step size, traverse the at least two input images respectively to obtain an image block group corresponding to each pixel point, where each image block group includes at least two image blocks respectively located in the input images; Input the multiple image blocks in each image block group into an initial neural network to obtain the pixel value output by the initial neural network; Determine a loss value according to a first difference between the input image and the first format target image and a second difference between the output pixel value and the second format target image; Based on the loss value, correct the initial neural network until a trained neural network is obtained.

7. An image processing apparatus, characterized in that, The device includes: A first obtaining module, configured to obtain an image set to be processed, where the image set includes a first image and a second image; A first traversing module, configured to traverse the first image and the second image respectively based on a window of a preset size and a step size to obtain an image block group corresponding to each pixel point, where each image block group includes a first image block located in the first image and a second image block located in the second image; A fusion module, configured to fuse a first image patch and a second image patch in each image patch group to obtain a pixel value of a corresponding pixel point after fusion; A first determination module, configured to determine a third image after alignment of the first image and the second image according to the pixel value of each pixel point after fusion.

8. A training device for an image processing network, characterized in that, The apparatus includes: A second acquisition module, configured to acquire a training data set, where the training data set includes a plurality of data groups, and each data group includes at least two input images, a first-format target image corresponding to the at least two input images, and a second-format target image; A second traversal module, configured to traverse the at least two input images respectively based on a window of a preset size and a step length, to obtain an image patch group corresponding to each pixel point, where each image patch group includes at least two image patches respectively located in the input images; An input module, configured to input a plurality of image patches in each image patch group into an initial neural network to obtain a pixel value output by the initial neural network; A second determination module, configured to determine a loss value according to a first difference between the input image and the first-format target image and a first difference between the output pixel value and the first-format target image; A correction module, configured to correct the initial neural network based on the loss value until a trained neural network is obtained.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method according to any one of claims 1-6 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1-6 is implemented.