Image stitching method and device based on cyclic generative adversarial network

By building a loop generation adversarial network for image fusion, the problem of low image fusion efficiency in the prior art is solved, image generation with high fidelity and scalable size is achieved, and a large number of fast and effective image generation needs are met.

CN114418842BActive Publication Date: 2025-07-18GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111570884.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-07-18
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

There is a lack of fast, efficient, and extensive image fusion methods in the prior art, especially to expand the form of the target and the size or structure of the image while generating the unchanged fidelity.

Method used

Build a loop generation adversarial network, including A2B generator, B2A generator, A discriminator and B discriminator, train the network through the training data set, generate image blocks and perform fusion and splicing, and optimize generator parameters using the loss function.

Benefits of technology

The generated image is achieved with high fidelity, scalable size and structure, high generation efficiency and large quantity, meeting the needs of image fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114418842B_ABST
    Figure CN114418842B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention relate to the technical field of image stitching, and disclose an image stitching method and device based on a cyclic generative adversarial network. The method includes: constructing a data set; constructing a cyclic generative adversarial network; training the cyclic generative adversarial network using the data set; dividing a background image into a plurality of image patches with overlapping regions; setting an initial condition region in a first image patch, extracting a first conditional image from the first image patch, and obtaining a first generated image through a generator of type A2B; extracting a second conditional image from the first image patch based on the overlapping region between the first image patch and a second image patch, and obtaining a second generated image through a generator of type A2B; stitching the first generated image and the second generated image; sequentially generating generated images of other image patches based on the overlapping regions and fusing and stitching them with the previous images to obtain a stitched image. Implementing the embodiments of the present invention can quickly, effectively, and in large quantities complete image fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image stitching, and particularly relates to an image stitching method and device based on a cyclic generative adversarial network. Background Art

[0002] The use of deep learning for image target recognition has been increasingly widely applied. However, the large amount of image data required for network learning often becomes a bottleneck in application development. The generation of image data has become a hot research direction. Among them, there is a type of problem that cannot be ignored, that is, some target images have both high-definition requirements and involve unconventional image structure sizes. For example, the long and thin product defects that appear in products, etc. It is necessary to first generate local target images and then perform reasonable image fusion to generate applicable image samples.

[0003] The generative adversarial network (GAN: Generative Adversarial Network) is an implicit density generation model, and its learning nature is unsupervised. Based on the GAN network, image style transfer technology or image style mixing models have been developed, such as converting photos into paintings, converting photos taken in summer into photos taken in winter, or converting photos of horses into photos of zebras; the generation from text descriptions to images has been developed, bringing many practical applications, such as converting text-form stories into comics, etc.; image coloring, enhancement, etc. have also been developed, such as converting night images into day images, or coloring black and white images, and converting sketches into realistic photos, etc. However, the prior art does not involve how to perform fast, effective, and large-scale image fusion through network image generation technology, so that the fidelity of the target in the generated image remains unchanged, but the form of the target and the size or structure of the image can be greatly expanded. Summary of the Invention

[0004] In view of the above defects, embodiments of the present invention disclose an image stitching method and device based on a cyclic generative adversarial network, which can perform fast, effective, and large-scale image fusion through network image generation technology, so that the fidelity of the target in the generated image remains unchanged, but the form of the target and the size or structure of the image can be greatly expanded.

[0005] The first aspect of the embodiments of the present invention discloses an image stitching method based on a cyclic generative adversarial network, and the method includes:

[0006] Construct a data set, where the data set includes a target-free data set and a target data set;

[0007] Construct a cyclic generative adversarial network, which includes a generator of type A2B, a generator of type B2A, a discriminator of type A, and a discriminator of type B. The generator of type A2B is used to take an image of type A and its corresponding conditional image as inputs and output a generated image of type B. The generator of type B2A is used to take an image of type B as an input and output a generated image of type A. The discriminator of type A is used to determine the authenticity of the generated image of type A based on the image of type A. The discriminator of type B is used to determine the authenticity of the generated image of type B based on the image of type B;

[0008] Use the target dataset and the non-target dataset to train the cyclic generative adversarial network to obtain a trained cyclic generative adversarial network;

[0009] Select a background image for which a target object is to be generated, and segment the background image into a number of image patches with overlapping regions;

[0010] Select a region in the first image patch as an initial conditional region, extract a first conditional image from the first image patch according to this initial conditional region, and use the first image patch and the first conditional image as inputs to the generator of type A2B in the trained cyclic generative adversarial network to obtain a first generated image;

[0011] According to the overlapping region between the first image patch and the second image patch, extract a second conditional image from the first image patch based on the overlapping region, and use the second image patch and the second conditional image as inputs to the generator of type A2B in the trained cyclic generative adversarial network to obtain a second generated image;

[0012] Fuse and splice the first generated image and the second generated image; sequentially generate generated images of other image patches based on the overlapping regions and fuse and splice them with the previously generated images to obtain a final spliced image.

[0013] As a preferred embodiment, in the first aspect of the embodiments of the present invention, construct a dataset, the dataset including a non-target dataset and a target dataset, including:

[0014] Select an image with a target object, and perform a first manual annotation on the region where the target object is located with a bounding box of a fixed size;

[0015] From the image dataset with the first manual annotation, crop out the annotated region according to the bounding box as the target dataset;

[0016] Select a background image without a target, and perform a second manual annotation at a random position with a bounding box of the same size to give the position of the generated target;

[0017] From the second manually annotated image dataset, crop out the annotated area according to the annotation box as the target-free dataset.

[0018] As a preferred embodiment, in the first aspect of the embodiments of the present invention, use the target dataset and the target-free dataset to train the cycle generative adversarial network to obtain the trained cycle generative adversarial network, including:

[0019] Randomly select N images x i , i = 1, 2…, N, from the target-free dataset as class A images, and randomly select N images y i from the target dataset as class B images;

[0020] Randomly select a rectangular area as the conditional area F, and extract the conditional image F(x i ) from the class A image x i , where F(·) represents the extraction operation;

[0021] Use the class A image x i and the conditional image F(x i ) as the input of the A2B class generator to obtain the class B generated image

[0022] Use the class B generated image as the input of the B2A class generator to obtain the restored image of the class A image x i

[0023] Extract the conditional image F(y i ) from the class B image y i ;

[0024] Use the class B image y i as the input of the B2A class generator to obtain the class A generated image

[0025] Use the class A generated image and the conditional image F(y i ) as the input of the A2B class generator to obtain the restored image of the class B image y i

[0026] Calculate the loss function Lz:

[0027] L z = L GB + L GA + α[L cx + L cy + βL cyc

[0028] Among them, α and β are control parameter constants; L GB is the loss function of discriminator of type B, L GA is the loss function of discriminator of type A, L cx is the first conditional loss function, L cy is the second conditional loss function, L cyc is the restoration loss function;

[0029]

[0030] Among them, D B (·) represents the output of the discriminator of type B;

[0031]

[0032] Among them, D A (·) represents the output of the discriminator of type A;

[0033]

[0034]

[0035]

[0036] According to the loss function L z , use the Adam optimization algorithm to update the parameters of the generator until the number of training times reaches the threshold M.

[0037] The second aspect of the embodiments of the present invention discloses an image splicing device based on a cyclic generative adversarial network, which includes:

[0038] The first construction unit is used to construct a data set, and the data set includes a targetless data set and a targeted data set;

[0039] The second construction unit is used to construct a cyclic generative adversarial network, and the cyclic generative adversarial network includes a generator of type A2B, a generator of type B2A, a discriminator of type A, and a discriminator of type B. The generator of type A2B is used to take an image of type A and its corresponding conditional image as inputs and output a generated image of type B. The generator of type B2A is used to take an image of type B as an input and output a generated image of type A. The discriminator of type A is used to discriminate the authenticity of the generated image of type A according to the image of type A, and the discriminator of type B is used to discriminate the authenticity of the generated image of type B according to the image of type B;

[0040] The training unit is used to train the cyclic generative adversarial network using the targeted data set and the targetless data set to obtain a trained cyclic generative adversarial network;

[0041] A selection unit, configured to select a background image for generating a target object, and segment the background image into a plurality of image patches with overlapping regions;

[0042] A first generation unit, configured to select a region in a first image patch as an initial condition region, extract a first conditional image from the first image patch according to the initial condition region, and use the first image patch and the first conditional image as inputs to the A2B type generator in the trained cycle generative adversarial network to obtain a first generated image;

[0043] A second generation unit, configured to, according to the overlapping region between the first image patch and a second image patch, extract a second conditional image from the first image patch based on the overlapping region, and use the second image patch and the second conditional image as inputs to the A2B type generator in the trained cycle generative adversarial network to obtain a second generated image;

[0044] A splicing unit, configured to fuse and splice the first generated image and the second generated image; generate generated images of other image patches based on the overlapping regions in sequence, and fuse and splice them with the previously generated images to obtain a final spliced image.

[0045] As a preferred embodiment, in the second aspect of the embodiments of the present invention, the first construction unit includes:

[0046] A first annotation subunit, configured to select an image with a target object, and perform first manual annotation on the region where the target object is located with a fixed-size annotation box;

[0047] A first cropping subunit, configured to crop out the annotated region from the first manually annotated image dataset according to the annotation box as a dataset with targets;

[0048] A second annotation subunit, configured to select a background image without a target, and perform second manual annotation at a random position with an annotation box of the same size to give the position of the generated target;

[0049] A second cropping subunit, configured to crop out the annotated region from the second manually annotated image dataset according to the annotation box as a dataset without targets.

[0050] As a preferred embodiment, in the second aspect of the embodiments of the present invention, the training unit includes:

[0051] A selection subunit, configured to randomly select N images x i , i = 1, 2..., N, as type A images, and randomly select N images y i from the dataset with targets as type B images;

[0052] The first extraction subunit is used to randomly select a rectangular area as the conditional region F, and extract the conditional image F(x i from the type-A image x i ), where F(·) represents the extraction operation;

[0053] The first generation subunit is used to use the type-A image x i and the conditional image F(x i ) as the input of the A2B type generator to obtain the type-B generated image

[0054] The second generation subunit is used to use the type-B generated image as the input of the B2A type generator to obtain the restored image of the type-A image x i ;

[0055] The second extraction subunit is used to extract the conditional image F(y i ) from the type-B image y i ;

[0056] The third generation subunit is used to use the type-B image y i as the input of the B2A type generator to obtain the type-A generated image

[0057] The fourth generation subunit is used to use the type-A generated image and the conditional image F(y i ) as the input of the A2B type generator to obtain the restored image of the type-B image y i ;

[0058] The calculation subunit is used to calculate the loss function Lz:

[0059] L z = L GB + L GA + α[L cx + L cy + βL cyc

[0060] where α and β are control parameter constants; L GB is the type-B discriminator loss function, L GA is the type-A discriminator loss function, L cx is the first conditional loss function, L cy is the second conditional loss function, L cyc is the restoration loss function;

[0061]

[0062] where DB (·) represents the output of the discriminator of Class B;

[0063]

[0064] where D A (·) represents the output of the discriminator of Class A;

[0065]

[0066]

[0067]

[0068] An update subunit, configured to update the parameters of the generator according to the loss function L z using the Adam optimization algorithm until the number of training times reaches the threshold M.

[0069] A third aspect of the embodiments of the present invention discloses an electronic device, including: a memory storing executable program code; a processor coupled to the memory; the processor invoking the executable program code stored in the memory for executing an image stitching method based on a cyclic generative adversarial network disclosed in the first aspect of the embodiments of the present invention.

[0070] A fourth aspect of the embodiments of the present invention discloses a computer-readable storage medium storing a computer program, wherein the computer program causes a computer to execute an image stitching method based on a cyclic generative adversarial network disclosed in the first aspect of the embodiments of the present invention.

[0071] A fifth aspect of the embodiments of the present invention discloses a computer program product, which when running on a computer causes the computer to execute an image stitching method based on a cyclic generative adversarial network disclosed in the first aspect of the embodiments of the present invention.

[0072] A sixth aspect of the embodiments of the present invention discloses an application publishing platform for publishing a computer program product, wherein when the computer program product runs on a computer, it causes the computer to execute an image stitching method based on a cyclic generative adversarial network disclosed in the first aspect of the embodiments of the present invention.

[0073] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0074] The main purpose of generating images in the embodiments of the present invention is to increase the number of sample images. On this premise, through the improvement of the network structure, the generated images are automatically located and fused during the network generation process, further changing the size structure of the target and increasing the diversity of the target images. The main advantages are:

[0075] 1. The size and structure of the generated image can be increased;

[0076] 2. The generated image has a high degree of realism;

[0077] 3. The efficiency of generating images is high and the quantity is large. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for use in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0079] Figure 1 It is a schematic flowchart of an image stitching method based on a cyclic generative adversarial network disclosed in an embodiment of the present invention;

[0080] Figure 2 It is a schematic structural diagram of a cyclic generative adversarial network disclosed in an embodiment of the present invention;

[0081] Figure 3 It is a schematic structural diagram of an A2B type generator disclosed in an embodiment of the present invention;

[0082] Figure 4 It is a schematic flowchart of the training process of a cyclic generative adversarial network disclosed in an embodiment of the present invention;

[0083] Figure 5 It is a schematic structural diagram of background image segmentation disclosed in an embodiment of the present invention;

[0084] Figure 6 It is a schematic structural diagram of an image stitching device based on a cyclic generative adversarial network disclosed in an embodiment of the present invention;

[0085] Figure 7 It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0086] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0087] It should be noted that the terms "first", "second", "third", "fourth", etc. in the description and claims of the present invention are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "having" in the embodiments of the present invention and any variations thereof are intended to cover non-exclusive inclusion. Exemplarily, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0088] The embodiments of the present invention disclose an image stitching method and device based on a cyclic generative adversarial network, which performs fast, effective, and large-scale image fusion through network-generated image technology, achieving an unchanged fidelity of the target in the generated image, but the form of the target and the size or structure of the image can be greatly expanded. The following is a detailed description with reference to the accompanying drawings.

[0089] Embodiment 1

[0090] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of an image stitching method based on a cyclic generative adversarial network disclosed in the embodiments of the present invention. As Figure 1 shown, the image stitching method based on the cyclic generative adversarial network includes the following steps:

[0091] S110, constructing a data set, where the data set includes a target-free data set and a target-containing data set.

[0092] Select images with target objects, and perform first manual annotation on the area where the target object is located with a fixed-size annotation box; the annotation box used in the present invention is a 256×256 rectangular box. From the first manually annotated image data set, cut out the annotated area according to the annotation box as the target-containing data set.

[0093] Select background images without targets, and perform second manual annotation at random positions with the same-size annotation box to give the position of the generated target; from the second manually annotated image data set, cut out the annotated area according to the annotation box as the target-free data set.

[0094] S120, constructing a cyclic generative adversarial network.

[0095] Please refer to Figure 2 shown, the cyclic generative adversarial network mainly includes a generator of type A2B, a generator of type B2A, a discriminator of type A, a discriminator of type B, and a conditional regional consistency analysis module. Among them, the structures of the generator of type B2A, the discriminator of type A, and the discriminator of type B are the same as those of the standard cyclic generative adversarial network structure. The generator of type A2B is as Figure 3As shown, it first extracts the features of Class A images and conditional images through two convolutional networks with the same weights respectively, then concatenates the two types of features, and finally generates Class B images through a deconvolution network. All network weights are randomly initialized using a Gaussian distribution.

[0096] The conditional region consistency analysis module is an image similarity calculation module. In this embodiment, the L1 norm of the difference between two images is used.

[0097] The A2B type generator is used to take a Class A image and its corresponding conditional image as inputs and output a generated Class B image (if a restored Class A image and its corresponding conditional image are used as inputs, then a restored Class B image is output). The B2A type generator is used to take a Class B image as an input and output a generated Class A image (if a restored Class B image is used as an input, then a restored Class A image is output). The Class A discriminator is used to discriminate the authenticity of the generated Class A image according to the Class A image, and the Class B discriminator is used to discriminate the authenticity of the generated Class B image according to the Class B image.

[0098] S130. Use the target dataset and the non-target dataset to train the cyclic generative adversarial network to obtain a trained cyclic generative adversarial network.

[0099] Please refer to Figure 4 As shown, it specifically includes the following steps:

[0100] S131. Randomly select N images \(x_i\), \(i = 1, 2, \ldots, N\) from the non-target dataset as Class A images, and randomly select N images \(y_i\), \(i = 1, 2, \ldots, N\) from the target dataset as Class B images. In this embodiment, \(N = 20\). i i

[0101] S132. Randomly select a rectangular region as the conditional region F, and extract the conditional image \(F(x_i)\) from the Class A image \(x_i\), where \(F(\cdot)\) represents the extraction operation. It first generates a zero-filled image of the same size as the original image, and then copies the image from the original image region F to the same position of the zero-filled image. i i

[0102] S133. Take the Class A image \(x_i\) and the conditional image \(F(x_i)\) as inputs of the A2B type generator to obtain the generated Class B image i i

[0103] S134. Take the image as the input of the B2A type generator to obtain the restored image of the Class A image \(x_i\) i

[0104] S135. Extract the conditional image F(y i ) from the Class B image y i . Take the conditional region F in step S132 as the conditional region of the Class B image y i , and extract the conditional image F(y i ) from the Class B image y i .

[0105] S136. Use the image y i as the input of the B2A class generator to obtain the Class A generated image

[0106] S137. Use the image and the conditional image F(y i ) as the input of the A2B class generator to obtain the restored image of the Class B image y i

[0107] S138. Calculate the loss function

[0108] L z = L GB + L GA + α[L cx + L cy + βL cyc

[0109] where α and β are control parameters. In this embodiment, α = 10 and β = 10;

[0110] L GB is the loss function of the Class B discriminator:

[0111]

[0112] where D B (·) represents the output of the Class B discriminator;

[0113] L GA is the loss function of the Class A discriminator:

[0114]

[0115] D A (·) represents the output of the Class A discriminator;

[0116] L cx is the first conditional loss function, L cy is the second conditional loss function, and L cyc is the restoration loss function, which can be calculated by the conditional region consistency analysis module:

[0117] ​

[0118]

[0119]

[0120] S139, according to the loss function L z , use the Adam optimization algorithm to update the parameters of the generator, and repeat steps S131 to S138 until the number of training times reaches the threshold M. In this embodiment, M = 10000.

[0121] S140, select a background image for generating the target object, and segment the background image into several image patches with overlapping regions (denoted as the first image patch, the second image patch, the third image patch, etc.), as Figure 5 shown.

[0122] S150, select a region in the first image patch as the initial condition region, extract the first conditional image from the first image patch according to this initial condition region, and use the first image patch and the first conditional image as the inputs of the A2B type generator in the trained cyclic generative adversarial network to obtain the first generated image.

[0123] Denote the first image patch image as x1, select a region in the first image patch as the initial condition region, and extract the conditional image F(x1) from x1 according to this region; use x1 and F(x1) as the inputs of the A2B type generator to obtain the generated image

[0124] S160, based on the overlapping region between the first image patch and the second image patch, extract the second conditional image from the first image patch based on the overlapping region, and use the second image patch and the second conditional image as the inputs of the A2B type generator in the trained cyclic generative adversarial network to obtain the second generated image.

[0125] According to the position of the overlapping region in the first image patch and the second image patch, extract the conditional image from Extract the conditional image Denote the second image patch as x2, and use x2 and as the inputs of the A2B type generator to obtain the generated image

[0126] S170, fuse and splice the first generated image and the second generated image; sequentially generate the generated images of other image patches based on the overlapping region and fuse and splice them with the previously generated images to obtain the final spliced image.

[0127] Fuse and splice with to complete the splicing of the first image patch and the second image patch.

[0128] Using the same fusion and splicing method, a third image block with an overlapping area with the second image block is used to obtain a generated image based on the overlapping area between the two. (Through the third image block and the third conditional image obtained by extracting the overlapping area in the second image block, and then inputting them into the A2B type generator), and are fused and spliced, and so on until all the image blocks are spliced. If the third image block has an overlapping area with the first image block instead of the second image block, the third image block with an overlapping area with the first image block is used to obtain a generated image based on the overlapping area between the two. The and are fused and spliced.

[0129] Embodiment 2

[0130] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of an image splicing device based on a cyclic generative adversarial network disclosed in an embodiment of the present invention. As Figure 6 shown, the image splicing device based on the cyclic generative adversarial network may include:

[0131] A first construction unit 210 for constructing a data set, where the data set includes a targetless data set and a targeted data set;

[0132] A second construction unit 220 for constructing a cyclic generative adversarial network, where the cyclic generative adversarial network includes an A2B type generator, a B2A type generator, an A type discriminator, and a B type discriminator. The A2B type generator is used to take an A type image and its corresponding conditional image as inputs and output a B type generated image. The B2A type generator is used to take a B type image as an input and output an A type generated image. The A type discriminator is used to discriminate the authenticity of the A type generated image according to the A type image, and the B type discriminator is used to discriminate the authenticity of the B type generated image according to the B type image;

[0133] A training unit 230 for training the cyclic generative adversarial network using the targeted data set and the targetless data set to obtain a trained cyclic generative adversarial network;

[0134] A selection unit 240 for selecting a background image to generate a target object and segmenting the background image into several image blocks with overlapping areas;

[0135] The first generation unit 250 is configured to select a region in the first image block as the initial condition region, extract the first conditional image from the first image block according to this initial condition region, and use the first image block and the first conditional image as the inputs of the A2B type generator in the trained cyclic generative adversarial network to obtain the first generated image;

[0136] The second generation unit 260 is configured to, according to the overlapping region of the first image block and the second image block, extract the second conditional image from the first image block based on the overlapping region, and use the second image block and the second conditional image as the inputs of the A2B type generator in the trained cyclic generative adversarial network to obtain the second generated image;

[0137] The splicing unit 270 is configured to fuse and splice the first generated image and the second generated image; generate the generated images of other image blocks based on the overlapping region in sequence, and fuse and splice them with the previously generated images to obtain the final spliced image.

[0138] Preferably, the first construction unit 210 includes:

[0139] The first annotation subunit is configured to select an image with a target object, and perform the first manual annotation on the region where the target object is located with a fixed-size annotation box;

[0140] The first cropping subunit is configured to crop out the annotated region from the first manually annotated image dataset according to the annotation box as the target dataset;

[0141] The second annotation subunit is configured to select a background image without a target, and perform the second manual annotation at a random position with an annotation box of the same size to give the position of the generated target;

[0142] The second cropping subunit is configured to crop out the annotated region from the second manually annotated image dataset according to the annotation box as the non-target dataset.

[0143] Preferably, the training unit 230 includes:

[0144] The selection subunit is configured to randomly select N images x i , i = 1, 2..., N, from the non-target dataset as type A images, and randomly select N images y i from the target dataset as type B images;

[0145] The first extraction subunit is configured to randomly select a rectangular region as the conditional region F, and extract the conditional image F(x i ) from the type A image x i , where F(·) represents the extraction operation;

[0146] The first generation subunit is used to use the Class A image x i and the conditional image F(x i ) as the input of the A2B class generator to obtain the generated Class B image

[0147] The second generation subunit is used to use the generated Class B image as the input of the B2A class generator to obtain the restored image of the Class A image x i

[0148] The second extraction subunit is used to extract the conditional image F(y i from the Class B image y i );

[0149] The third generation subunit is used to use the Class B image y i as the input of the B2A class generator to obtain the generated Class A image

[0150] The fourth generation subunit is used to use the generated Class A image and the conditional image F(y i ) as the input of the A2B class generator to obtain the restored image of the Class B image y i

[0151] The calculation subunit is used to calculate the loss function L z :

[0152] L z = L GB + L GA + α[L cx + L cy + βL cyc

[0153] where α and β are control parameter constants; L GB is the Class B discriminator loss function, L GA is the Class A discriminator loss function, L cx is the first conditional loss function, L cy is the second conditional loss function, L cyc is the restoration loss function;

[0154]

[0155] where D B (·) represents the output of the Class B discriminator;

[0156]

[0157] where D A ​​(·) represents the output of the discriminator of Class A;

[0158]

[0159]

[0160]

[0161] An update subunit, configured to update the parameters of the generator according to the loss function L z , and use the Adam optimization algorithm to update the parameters of the generator until the number of training times reaches the threshold M.

[0162] Embodiment III

[0163] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device disclosed in an embodiment of the present invention. As Figure 7 shown, the electronic device may include:

[0164] A memory 310 storing executable program code;

[0165] A processor 320 coupled to the memory 310;

[0166] Wherein, the processor 320 calls the executable program code stored in the memory 310 and executes some or all of the steps in the method for image stitching based on a cyclic generative adversarial network in Embodiment I.

[0167] An embodiment of the present invention discloses a computer-readable storage medium, which stores a computer program, wherein the computer program enables a computer to execute some or all of the steps in the method for image stitching based on a cyclic generative adversarial network in Embodiment I.

[0168] An embodiment of the present invention also discloses a computer program product, wherein when the computer program product runs on a computer, it enables the computer to execute some or all of the steps in the method for image stitching based on a cyclic generative adversarial network in Embodiment I.

[0169] An embodiment of the present invention also discloses an application publishing platform, wherein the application publishing platform is used to publish a computer program product, and when the computer program product runs on a computer, it enables the computer to execute some or all of the steps in the method for image stitching based on a cyclic generative adversarial network in Embodiment I.

[0170] In various embodiments of the present invention, it should be understood that the magnitude of the sequence numbers of the various processes does not necessarily mean the order of execution, and the execution order of the various processes should be determined according to their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0171] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0172] In addition, each functional unit in the embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0173] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-accessible memory. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests for causing a computer device (which can be a personal computer, a server, or a network device, etc., specifically, the processor in the computer device) to execute some or all of the steps of the methods described in the various embodiments of the present invention.

[0174] In the embodiments provided by the present invention, it should be understood that "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0175] Those of ordinary skill in the art can understand that some or all of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, magnetic tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0176] The above has introduced in detail a method and apparatus for image stitching based on a cyclic generative adversarial network disclosed in the embodiments of the present invention. Specific examples are used herein to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An image stitching method based on a cyclic generative adversarial network, characterized in that, Including: Construct a data set, the data set including a targetless data set and a targeted data set; Construct a cyclic generative adversarial network, the cyclic generative adversarial network including a generator of type A2B, a generator of type B2A, a discriminator of type A, and a discriminator of type B. The generator of type A2B is used to take an image of type A and its corresponding conditional image as inputs and output a generated image of type B. The generator of type B2A is used to take an image of type B as an input and output a generated image of type A. The discriminator of type A is used to discriminate the authenticity of the generated image of type A according to the image of type A. The discriminator of type B is used to discriminate the authenticity of the generated image of type B according to the image of type B; Use the targeted data set and the targetless data set to train the cyclic generative adversarial network to obtain a trained cyclic generative adversarial network; Select a background image for which a target object is to be generated, and segment the background image into a number of image patches with overlapping regions; Select a region in the first image patch as an initial conditional region, extract a first conditional image from the first image patch according to this initial conditional region, take the first image patch and the first conditional image as inputs to the generator of type A2B in the trained cyclic generative adversarial network to obtain a first generated image; According to the overlapping region between the first image patch and the second image patch, extract a second conditional image from the first image patch based on the overlapping region, take the second image patch and the second conditional image as inputs to the generator of type A2B in the trained cyclic generative adversarial network to obtain a second generated image; Fuse and splice the first generated image and the second generated image; Generate generated images of other image patches in sequence based on the overlapping regions, and fuse and splice them with the previously generated images to obtain a final spliced image.

2. The image stitching method based on a cyclic generative adversarial network according to claim 1, characterized in that Construct a data set, the data set including a targetless data set and a targeted data set, including: Select an image with a target object, and perform a first manual annotation on the region where the target object is located with a bounding box of a fixed size; From the image data set of the first manual annotation, crop out the annotated region according to the bounding box as the targeted data set; Select a targetless background image, and perform a second manual annotation at a random position with a bounding box of the same size to give the position of the generated target; From the image data set of the second manual annotation, crop out the annotated region according to the bounding box as the targetless data set.

3. The image stitching method based on a cyclic generative adversarial network according to claim 2, wherein Use the targeted data set and the targetless data set to train the cyclic generative adversarial network to obtain a trained cyclic generative adversarial network, including: Randomly select N images x from the targetless dataset i , where i = 1, 2…, N, as class A images, and randomly select N images y from the targeted dataset i , as class B images; Randomly select a rectangular area as the conditional area F, and extract the conditional image F(x i ) from the type-A images x i , where F(·) represents the extraction operation; Take the Class A image x i and the conditional image F(x i ) as the input of the Class A2B generator to obtain the generated Class B image Generate the B-class image As the input of the B2A-class generator, obtain the A-class image x i Restored image Extract the conditional image F(y i ) from the class B image y i ); Take the Class B image y i as the input of the B2A class generator to obtain the Class A generated image Generate an image of Class A and the conditional image F(y i ) as the input to the Class A2B generator to obtain the restored image of the Class B image y i ​ Calculate the loss function L z : L z = L GB + L GA + α[L cx + L cy + βL cyc where α and β are control parameter constants; L GB is the loss function of discriminator of class B, L GA is the loss function of discriminator of class A, L cx is the first conditional loss function, L cy is the second conditional loss function, L cyc is the restoration loss function; Among them, D B (·) represents the output of the discriminator of type B; Among them, D A (·) represents the output of discriminator of type A; According to the loss function L z , update the parameters of the generator using the Adam optimization algorithm until the number of training times reaches the threshold M.

4. An image stitching device based on a cyclic generative adversarial network, characterized in that, It includes: A first construction unit for constructing a data set, the data set including a targetless data set and a targeted data set; A second construction unit for constructing a cyclic generative adversarial network, which includes a generator of type A2B, a generator of type B2A, a discriminator of type A, and a discriminator of type B. The generator of type A2B is used to take an image of type A and its corresponding conditional image as inputs and output a generated image of type B. The generator of type B2A is used to take an image of type B as an input and output a generated image of type A. The discriminator of type A is used to determine the authenticity of the generated image of type A based on the image of type A. The discriminator of type B is used to determine the authenticity of the generated image of type B based on the image of type B; A training unit for training the cyclic generative adversarial network using the target dataset and the non-target dataset to obtain a trained cyclic generative adversarial network; A selection unit for selecting a background image for generating a target object and dividing the background image into a plurality of image patches with overlapping regions; A first generation unit for selecting a region in a first image patch as an initial conditional region, extracting a first conditional image from the first image patch according to this initial conditional region, taking the first image patch and the first conditional image as inputs to the generator of type A2B in the trained cyclic generative adversarial network, and obtaining a first generated image; A second generation unit for, according to the overlapping region between the first image patch and a second image patch, extracting a second conditional image from the first image patch based on the overlapping region, taking the second image patch and the second conditional image as inputs to the generator of type A2B in the trained cyclic generative adversarial network, and obtaining a second generated image; A splicing unit for fusing and splicing the first generated image and the second generated image; Generating generated images of other image patches based on the overlapping regions in sequence and fusing and splicing them with the previously generated images to obtain a final spliced image.

5. The image stitching device based on the cyclic generative adversarial network according to claim 4, characterized in that, The first construction unit includes: A first annotation subunit for selecting an image with a target object and performing a first manual annotation on the region where the target object is located with a fixed-size annotation box; A first cropping subunit for cropping out the annotated region from the first manually annotated image dataset according to the annotation box as the target dataset; A second annotation subunit for selecting a background image without a target and performing a second manual annotation at a random position with an annotation box of the same size to give the position of the generated target; A second cropping subunit for cropping out the annotated region from the second manually annotated image dataset according to the annotation box as the non-target dataset.

6. The image stitching device based on a cyclic generative adversarial network according to claim 5, characterized in that, The training unit includes: Select sub-units for randomly selecting N images x from the dataset without targets i , i = 1, 2…, N, as class A images, and randomly select N images y from the dataset with targets i , as class B images; The first extraction subunit is used to randomly select a rectangular area as the conditional area F, and extract the conditional image F(x i from the type-A image x i ), where F(·) represents the extraction operation; The first generation subunit is used to take the Class-A image x i and the conditional image F(x i ) as the input of the A2B class generator to obtain the Class-B generated image The second generation subunit is used to generate an image of class B as the input of the B2A class generator to obtain an image of class A, x i restored image A second extraction subunit, configured to extract a conditional image F(y i ) from a Class B image y i ); The third generation subunit is used to take the Class B image y i as the input of the Class B2A generator to obtain the Class A generated image The fourth generation subunit is used to generate the Class-A generated image and the conditional image F(y i ) as the input of the Class-A2B generator to obtain the restored image of the Class-B image y i ​ A calculation subunit for calculating a loss function L z : L z = L GB + L GA + α[L cx + L cy + βL cyc Among them, α and β are control parameter constants; L GB is the loss function of discriminator of class B, L GA is the loss function of discriminator of class A, L cx is the first conditional loss function, L cy is the second conditional loss function, L cyc is the restoration loss function; Among them, D B (·) represents the output of the discriminator of class B; Among them, D A (·) represents the output of the discriminator of Class A; An update subunit, configured to update the parameters of the generator according to the loss function L z , using the Adam optimization algorithm until the number of training times reaches the threshold M.

7. An electronic device, characterized in that, including: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute an image splicing method based on a cyclic generative adversarial network according to any one of claims 1 to 3.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program causes a computer to execute an image splicing method based on a cyclic generative adversarial network according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • License plate image generation model construction method and device and license plate image generation method and device

    CN112102424A

  • Cambered surface defect image generation method based on cyclic generative adversarial network

    CN113011480A