A method and system for splicing images of goods on a shelf

Through the depth homography network, the homography between shelf images is predicted and deformation spliced ​​is performed, which solves the artifact and stretching problems in large parallax image splicing, and efficient splicing of images of any size is achieved.

CN115115522BActive Publication Date: 2025-05-30ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210976559.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-15
Publication Date
2025-05-30
Estimated Expiration
2042-08-15

AI Technical Summary

Technical Problem

When taking pictures of shelf products, images with large parallax will be obtained due to different shooting angles, resulting in artifacts and stretching after image stitching. In addition, the existing image stitching method based on deep neural networks is difficult to process images of any size.

Method used

The depth homography network is used to predict the homography between two images, reduce artifacts through deformation and stitching and fusion techniques, and enable the model to process input images of any size through offset adjustment.

Benefits of technology

It effectively reduces the artifact phenomenon in the stitched image, improves the image quality, and solves the problem of stitching images in different sizes, achieving efficient shelf image stitching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115522B_ABST
    Figure CN115115522B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for splicing images of shelf commodities. First, two images of shelf commodities A and B are input into a trained deep homography estimation network to obtain an estimated homography matrix H. According to the homography matrix H, the image of shelf commodity B is deformed to obtain a corresponding deformed image C. Then, the obtained deformed image C is spliced and fused with the image of shelf commodity A, and finally, feature optimization is performed to enhance the image quality and obtain a high-resolution spliced image E. The present invention uses a deep homography estimation network composed of a feature extraction module, a feature correlation layer, and a regression module to predict the homography between two images, greatly reducing the artifact phenomenon in the spliced image and improving the image quality. The model of the present invention has the function of splicing input images of any size, solving the problem of diverse sizes of input images of the shelf.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image stitching in computer vision, and specifically to a method and system for stitching shelf commodity images based on a deep neural network. Background Art

[0002] In the retail industry, in order to better understand the display situation of commodities on the shelves and thus better insight into the market to make decisions on marketing management, consumer goods manufacturers often analyze the types and placement positions of the commodities placed on the shelves by taking images of store shelves, obtain information such as the stocking rate and the number of facings of different commodities in the store, and determine whether it meets the requirements of the manufacturer. For a scene like a shelf where the shooting space is relatively narrow, with a wide viewing angle and a large amount of content, it is difficult to include all commodities by taking a single image. Therefore, it is very difficult to directly obtain an ultra-wide viewing angle and high-resolution image using a lens, and image stitching technology is required to complete it.

[0003] Image stitching refers to seamlessly connecting multiple images with overlapping regions under similar viewing angles and combining them into an ultra-wide viewing angle image. In recent years, due to its powerful feature extraction ability, deep neural networks have developed rapidly in the field of computer vision, and more and more researchers have applied deep neural networks to image stitching. However, during the process of shooting shelf commodity images, due to different shooting angles, images with large parallax will be obtained, and artifacts and stretching phenomena will occur after stitching such large-parallax images; in addition, the models obtained by training in the currently popular image stitching methods based on deep neural networks are only suitable for processing images with the same size as the training set, and for images of any size, the effect during the inference stage is not satisfactory. Summary of the Invention

[0004] In the scenario of shelf image stitching, due to shooting angle problems, images with large parallax will be obtained, and artifacts and stretching phenomena will occur after stitching such large-parallax images. In view of the above problems, the present invention provides a method for stitching shelf commodity images based on a deep learning model, which is used to process the shelf image stitching task and solve the problem of blurred stitched images caused by shooting angles, making the stitching process more convenient and efficient.

[0005] The technical solution adopted by the present invention to solve its technical problems is:

[0006] A method for stitching shelf commodity images, comprising:

[0007] (1) Obtain two shelf commodity images A and B to be stitched;

[0008] (2) Input two shelf commodity images A and B into a trained deep homography estimation network to obtain an estimated homography matrix H; deform the shelf commodity image B according to the homography matrix H to obtain a corresponding deformed image C; the deep homography estimation network consists of n sequentially connected feature processing modules, a global feature correlation layer, and a regression network module, where: the feature processing module is used to extract features based on the input image; the global feature correlation layer is used to calculate the correlation of each feature point for the feature maps corresponding to the shelf commodity images A and B output by the nth feature processing module; the regression network module is used to predict the x and y coordinate offsets of two vertices each where the shelf commodity images A and B overlap based on the correlation of each feature point output by the global feature correlation layer, and then obtain the estimated homography matrix H according to the eight obtained offset coordinates and the projection transformation factor 1; the trained deep homography estimation network is trained based on the obtained training dataset, with each sample pair of the training dataset as input, aiming to minimize the error between the predicted homography matrix H and the ground truth.

[0009] (3) Stitch and fuse the deformed image C obtained in step (2) with the shelf commodity image A to obtain a fused image D;

[0010] (4) Optimize the features of the image D obtained in step (3) to enhance the image quality and obtain a high-resolution stitched image E.

[0011] Further, the deep homography estimation network further includes n - 2 local feature correlation layers and n - 2 regression network modules, where the (n - 1)th regression network module is used to predict the x and y coordinate offsets of two vertices each where the shelf commodity images A and B overlap based on the correlation of each feature point output by the global feature correlation layer, and then obtain the estimated homography matrix H according to the eight obtained offset coordinates and the projection transformation factor 1; the (n - 2)th local feature correlation layer is used to calculate the correlation of each feature point for the feature map of the shelf commodity image A output by the (n - 2)th feature processing module and the feature map F after deforming the feature map corresponding to the shelf commodity image B output by the nth feature processing module according to the homography matrix H output by the nth regression network module. B ’ 1 / (2^(n-2)) Calculate the correlation of each feature point; the (n - 2)th regression network module is used to predict the x and y coordinate offsets of two vertices each where the shelf commodity images A and B overlap based on the correlation of each feature point output by the (n - 2)th local feature correlation layer, and then obtain the estimated homography matrix H according to the eight obtained offset coordinates and the projection transformation factor 1; and so on, until the 1st regression network module predicts the x and y coordinate offsets of two vertices each where the shelf commodity images A and B overlap based on the correlation of each feature point output by the 1st local feature correlation layer, and then obtain the estimated homography matrix H according to the eight obtained offset coordinates and the projection transformation factor 1.

[0012] Further, step (3) is specifically as follows:

[0013] Input the deformed image C obtained in step (2) and the shelf commodity image A into a trained encoder-decoder network for splicing and fusion to obtain a fused image D; the encoder-decoder network includes an encoder and a decoder. The encoder is used to reconstruct the features of the overlapping regions in the two images based on the input deformed image C and shelf commodity image A obtained in step (2). The decoder is used to decode according to the features output by the encoder and simultaneously restore the non-overlapping regions to obtain the fused image D.

[0014] Further, step (4) is specifically as follows:

[0015] Input the image D obtained in step (3) into a trained optimization branch for feature optimization to enhance the image quality and obtain a high-resolution spliced image E; the optimization branch is composed of a first convolutional layer, a plurality of deep residual blocks, and a plurality of second convolutional layers connected in sequence.

[0016] Further, in the regression network module, the x and y coordinate offsets of two vertices overlapping between the shelf commodity images A and B are predicted according to the correlation of each feature point output by the global feature correlation layer, and the predicted x and y coordinate offsets are adjusted according to the size ratio of the two shelf commodity images to be spliced and the images in the training dataset, specifically as follows:

[0017]

[0018] Among them, σw = w / W, σh = h / H, where w and h respectively represent the width and height of the image in the training dataset, and W and H respectively represent the width and height of the two shelf commodity images to be spliced; ΔU i and ΔV i (i = 1, 2, 3, 4) respectively represent the x coordinate and y coordinate offsets of the four vertices of the overlapping region in the same coordinate system when performing homography estimation on the image in the training dataset; σwΔU i and σhΔV i (i = 1, 2, 3, 4) respectively represent the x coordinate and y coordinate offsets of the four vertices of the overlapping region in the same coordinate system when performing homography estimation on the two shelf commodity images to be spliced.

[0019] A shelf commodity image splicing system for implementing the above method, comprising:

[0020] The homography estimation module is used to input two shelf commodity images A and B into a trained deep homography estimation network to obtain an estimated homography matrix H; deform the shelf commodity image B according to the homography matrix H to obtain a corresponding deformed image C;

[0021] The stitching and fusion module is used to stitch and fuse the obtained deformed image C with the shelf commodity image A to obtain a fused image D;

[0022] The feature optimization module is used to optimize the features of the obtained image D, enhance the image quality, and obtain a high-resolution stitched image E.

[0023] The beneficial effects of the present invention are mainly reflected in:

[0024] The present invention preferably solves the problem that it is difficult for the lens to accommodate the entire content of the shelf when shooting shelf images. By using the excellent feature extraction ability of the convolutional neural network, a deep learning model stitching technology for solving the shelf image scene is proposed. During the process of shooting shelf commodity images, due to different shooting angles, images with large parallax are obtained. The homography between such large-parallax images is difficult to predict, so artifacts and stretching phenomena will occur in the stitched image. To address this problem, the present invention uses a deep homography estimation network composed of a feature extraction module, a feature correlation layer, and a regression module to predict the homography between two images, greatly reducing the artifact phenomenon in the stitched image and improving the image quality. In addition, the present invention also provides a method based on offset adjustment to enable the model to have the function of stitching input images of any size. Furthermore, the shelf image does not need to be cropped to the same size as the training image before stitching, solving the problem of diverse sizes of shelf input images. Description of the Drawings

[0025] Figure 1 is the implementation flowchart of the present invention;

[0026] Figure 2 is the structure diagram of the deep homography estimation network;

[0027] Figure 3 is the schematic diagram of the warped image;

[0028] Figure 4 is the structure diagram of the image feature fusion network (encoder-decoder network);

[0029] Figure 5 is the structure diagram of the image feature optimization network (optimization branch)

[0030] Figure 6 is the structure diagram of the image feature optimization residual module. Detailed Embodiments

[0031] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.

[0032] A method for stitching images of shelf commodities provided by the present invention, as Figure 1 shown, includes the following steps:

[0033] (1) Obtain two images of shelf commodities A and B to be stitched;

[0034] The method of the present invention first deforms one image according to the other image, and then stitches and fuses the deformed image with the original image. For the convenience of distinction, in this embodiment, the undeformed and deformed images of shelf commodities in the two images are respectively defined as reference image A and target image B, and the size of both is W×H, where W represents the width of the image and H represents the height of the image;

[0035] (2) Input the two images described in step (1) into a trained deep homography estimation network for deep homography estimation, as Figure 3 shown, the homography between images refers to the projection relationship of the position coordinates between the overlapping parts of the images obtained when the camera takes pictures of the same object from different positions, and can be expressed as:

[0036]

[0037] where [u', v'] and [u, v] represent the coordinates of the images taken at different positions, H represents a 3×3 homography matrix, and [H 11 H 12 ; H 21 H 22 represents the rotation parameter, [H 13 H 23 represents the translation parameter, [H 31 H 32 represents the position parameter of the intersection points of the image with the two coordinate axes, and H 33 represents a projection transformation factor, usually 1.

[0038] Among them, the depth homography estimation network consists of n feature processing modules, a global feature correlation layer, and a regression network module connected in sequence. After inputting the two images described in step (1) into a trained depth homography estimation network, the input images are first subjected to feature processing through n feature processing modules composed of convolutional layers and max-pooling layers. The feature processing modules extract features based on the input images. The processed features are sent to the feature correlation layer to calculate the correlation of each feature point for the feature maps corresponding to the shelf commodity images A and B output by the nth feature processing module, match the features between the two images, and then sent to the regression network module. According to the correlation of each feature point output by the global feature correlation layer, the x and y coordinate offsets of two vertices (the two right vertices in the reference image and the two left vertices in the target image) where the shelf commodity images A and B overlap are obtained, and the homography matrix H between the two images is predicted based on the eight obtained offset coordinates and the projection transformation factor 1. Finally, the target image B is warped according to the homography between the two images to obtain the final warped image C.

[0039] Furthermore, the depth homography estimation network further includes n - 2 local feature correlation layers and n - 2 regression network modules. Among them, the (n - 1)th regression network module is used to predict the x and y coordinate offsets of two vertices where the shelf commodity images A and B overlap according to the correlation of each feature point output by the global feature correlation layer, and then obtain the estimated homography matrix H based on the eight obtained offset coordinates and the projection transformation factor 1. The (n - 2)th local feature correlation layer is used to perform the correlation calculation of each feature point on the feature map F obtained by warping the feature map corresponding to the shelf commodity image B output by the nth feature processing module according to the homography matrix H output by the nth regression network module and the feature map corresponding to the shelf commodity image A output by the (n - 2)th feature processing module. B ’ 1 / (2^(n-2)) Perform the correlation calculation of each feature point. The (n - 2)th regression network module is used to predict the x and y coordinate offsets of two vertices where the shelf commodity images A and B overlap according to the correlation of each feature point output by the (n - 2)th local feature correlation layer, and then obtain the estimated homography matrix H based on the eight obtained offset coordinates and the projection transformation factor 1. And so on, until the first regression network module predicts the x and y coordinate offsets of two vertices where the shelf commodity images A and B overlap according to the correlation of each feature point output by the first local feature correlation layer, and then obtain the estimated homography matrix H based on the eight obtained offset coordinates and the projection transformation factor 1.

[0040] Exemplarily, taking the depth homography estimation network composed of four feature processing modules, three feature correlation layers, and three regression network modules as an example, the homography estimation is actually a process of predicting the homography matrix H, as Figure 2As shown, each feature processing module processes the input images A and B through two 3x3 convolutional layers and a max pooling layer:

[0041] (2a) The input images A and B respectively pass through the convolutional layer conv1, the convolutional layer conv2, and the max pooling layer maxpooling1 of the first feature processing module to obtain the feature maps F 1 A and F 1 B , both with a size of W / 2 * H / 2 * 64;

[0042] (2b) The obtained feature maps F 1 A and F 1 B are respectively input into the second feature processing module to obtain the feature maps F A 1 / 2 and F B 1 / 2 , both with a size of W / 4 * H / 4 * 128;

[0043] (2c) Then the obtained feature maps F A 1 / 2 and F B 1 / 2 are input into the third feature processing module to obtain the feature maps F A 1 / 4 and F B 1 / 4 , both with a size of W / 8 * H / 8 * 256;

[0044] (2d) Finally, the obtained feature maps F A 1 / 4 and F B 1 / 4 are input into the fourth feature processing module to obtain the feature maps F A 1 / 8 and F B 1 / 8 , with a size of W / 16 * H / 16 * 512;

[0045] (2e) The feature maps F A 1 / 8 and F B 1 / 8 obtained in step 2d are jointly input into a global feature correlation layer to match the features of these two feature maps. The feature correlation between the two feature maps can be expressed as:

[0046]

[0047] Where c represents feature correlation, F A l and F B l (l=1, 1 / 2, 1 / 4, 1 / 8) represent the feature maps obtained from the reference image A and the target image B respectively, x l A , x l B Represent the feature map F A l and F B l The two-dimensional spatial position of the corresponding feature, F A l (x l A ) and F B l (x l B ) represent the feature maps F A l and F B l Up, x l A , x l B The characteristics of the location, <F A l (x l A ), F B l (x l B )> represents the dot product of two features, |F A l (x l A )||F B l (x l B )| represents the product of the moduli of the two features, c(x l A , x l B ) value, the larger the feature matching is.

[0048] Then the homography between the two is estimated through the third regression network module. The third regression network module consists of three convolutional layers and two fully connected layers to predict the eight coordinate offsets of the homography (predicted by the coordinates of the feature points with the best feature match between the four vertices of image A and image B), that is, the x and y coordinate offsets of the two vertices of the overlapping shelf product images A and B. The required homography matrix H can be obtained based on the obtained eight offset coordinates and the projection transformation factor 1; for the feature map F B1 / 4 Perform the warp operation;

[0049] (2f) The warped feature map obtained in step 2e is then combined with the feature map F A 1 / 4 and both are input into the second local feature correlation layer to match the features of these two feature maps. Then, through the second regression network module, the homography between these two layers is estimated, and accordingly, the feature map F B 1 / 2 is warped;

[0050] (2g) Similar to step 2e, the obtained feature map is correspondingly input into the first local feature correlation layer to match the features of these two feature maps. Further, through the first regression network module, the homography is estimated, and then, according to the homography obtained from this layer, the target image B is warped to obtain the final warped image C.

[0051] (3) The image C obtained in step (2) and the reference image A are subjected to an image fusion operation to obtain the fused image D;

[0052] In the above step (3), image fusion means fusing the feature information of the reference image A and the warped target image C onto one image to achieve the purpose of stitching. This process can be realized by finding and matching feature points or by deep learning. In this embodiment, mainly, an encoder-decoder network composed of convolutional layers and pooling layers is used to learn the fusion rules of image features, and skip connections are used to connect low-level and high-level features with the same resolution, thereby obtaining a fused image D. As Figure 4 shown, the specific operation steps are as follows:

[0053] (3a) Input the warped target image C obtained in step (2) and the reference image A;

[0054] (3b) The two images are sequentially subjected to feature encoding through four encoders, where only the overlapping regions of the two images are concerned, and the features of the non-overlapping regions are all suppressed (the overlapping regions of the two images are obtained through the homography matrix obtained in step (2), and the pixels of the non-overlapping regions are suppressed to make their pixel values 0). Each encoder consists of two 3×3 convolutional layers and one max pooling layer, and the corresponding number of filters is 64, 128, 256, and 512 in sequence;

[0055] (3c) The encoded features obtained in step (3b) are input into three decoders for feature decoding, and the pixel values of the non-overlapping regions are restored. Each decoder consists of three 3×3 transposed convolutional layers, and the corresponding number of filters is 256, 128, and 64 in sequence;

[0056] (3d) Through the feature encoding and feature decoding processes in steps 3a and 3b, the fused image D can be obtained.

[0057] The encoder-decoder network constructs a perceptual loss function and is trained in an unsupervised manner, which is expressed as follows:

[0058] Among them, the perceptual loss function compares the features learned by each layer of the network with the features of the input image, and the loss function can be expressed as follows:

[0059]

[0060] Among them, j represents the index of the layer number of the encoder-decoder network, C, H, and W respectively represent the number of channels, height, and width of the image, C j H j W j represents the size of the feature map on this layer of the network, represents the feature output of the input image y (A or B) on the j-th layer of the encoder network, represents the feature output learned by the j-th layer of the encoder network (i.e., the feature output on the j-th layer of the decoder network). ||*|| 2 represents the L2 norm; by minimizing the squared difference of the two eigenvalue on the same layer of the network, the content and global structure of the two features are made close.

[0061] (4) The resolution of the stitched image D obtained in step (3) is relatively small and relatively blurred, and the image features need to be further optimized to enhance the image quality to obtain a high-resolution stitched image E;

[0062] In this embodiment, the optimization of the image features is mainly achieved through a trained optimization branch composed of a convolutional layer and a deep residual block. Among them, the convolutional block is used to optimize the basic pixel features of the image, and the residual block is used to optimize the visual perception features of the image, so that the visual perception effect of the stitched image is better. As Figure 5 shown, the specific steps are implemented as follows:

[0063] (4a) The rough stitched image obtained in step (4) is first input into the first 3*3 convolutional layer for pixel feature optimization, and the number of filters in this layer is 64;

[0064] (4b) The optimized image features are then sequentially optimized through eight deep residual modules for visual perception features. Each module consists of the same five parts, namely a convolutional block, a RELU activation function, a convolutional block, a summation block, and a RELU activation function, as Figure 6 shown;

[0065] (4c) The optimized image features in step 4b finally pass through two convolutional layers with sizes of 3*3*6 and 3*3*3 respectively to obtain a high-resolution spliced shelf image E. In addition, skip connections are used to connect the first convolutional layer and the second convolutional layer to prevent the loss of feature information.

[0066] Furthermore, if the size of the image to be spliced is inconsistent with the size of the training images in the training dataset, the position offset change during image warping is obtained through the width ratio (σw) and height ratio (σh) of the two images, so that the network model can process input images of any size during the inference stage. Furthermore, the shelf images do not need to be cropped to the same size as the training images before splicing. Specifically:

[0067] (5a) During the inference stage, after the image is input into the homography estimation network, the corresponding width ratio σw and height ratio σh are first obtained according to the size of the input image and the size of the training images during model training. The calculation formulas are as follows:

[0068] σw = w / W σh = h / H

[0069] where w and h respectively represent the width and height of the images during model training, and W and H respectively represent the width and height of the image to be spliced during the actual application stage;

[0070] (5b) According to the obtained σw and σh, the coordinate offsets of the four vertices of the overlapping area of the image to be spliced are adjusted, and then homography estimation is completed through this new coordinate offset, and then the warp operation is performed. This process can be expressed as:

[0071]

[0072] where, ΔU i and ΔV i (i = 1, 2, 3, 4) respectively represent the x-coordinate and y-coordinate offsets of the four vertices of the overlapping area in the same coordinate system during homography estimation of the training images; σwΔU i and σwΔV i (i = 1, 2, 3, 4) respectively represent the x-coordinate and y-coordinate offsets of the four vertices of the overlapping area in the same coordinate system during homography estimation of the test images.

[0073] Corresponding to the foregoing embodiments of a method for splicing shelf commodity images, the present invention also provides an embodiment of a system for splicing shelf commodity images.

[0074] A system for splicing shelf commodity images, used to implement the method described in any one of the above, includes:

[0075] The homography estimation module is used to input two shelf commodity images A and B into a trained deep homography estimation network to obtain an estimated homography matrix H; deform the shelf commodity image B according to the homography matrix H to obtain a corresponding deformed image C;

[0076] The stitching and fusion module is used to stitch and fuse the obtained deformed image C with the shelf commodity image A to obtain a fused image D;

[0077] The feature optimization module is used to optimize the features of the obtained image D, enhance the image quality, and obtain a high-resolution stitched image E.

[0078] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The system embodiment described above is only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0079] Obviously, the above embodiments are only examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. The obvious changes or modifications derived therefrom are still within the protection scope of the present invention.

Claims

1. A method for stitching images of goods on a shelf, characterized in that, it includes: (1) Obtain two images of goods on a shelf, A and B, to be stitched; (2) Input the two images of goods on a shelf, A and B, into a trained deep homography estimation network to obtain an estimated homography matrix H; deform the image of goods on a shelf B according to the homography matrix H to obtain a corresponding deformed image C; the deep homography estimation network is composed of n feature processing modules, a global feature correlation layer, and a regression network module connected in sequence, where: the feature processing module is used to extract features based on the input image; the global feature correlation layer is used to calculate the correlation of each feature point for the feature maps corresponding to the images of goods on a shelf A and B output by the nth feature processing module; the regression network module is used to predict the x and y coordinate offsets of two vertices overlapping between the images of goods on a shelf A and B according to the correlation of each feature point output by the global feature correlation layer, and then obtain the estimated homography matrix H according to the eight obtained offset coordinates and the projection transformation factor 1; the trained deep homography estimation network is trained based on the obtained training dataset, with each sample pair of the training dataset as the input, and the goal is to minimize the error between the predicted homography matrix H and the ground truth; (3) Stitch and fuse the deformed image C obtained in step (2) with the image of goods on a shelf A to obtain a fused image D; (4) Optimize the features of the image D obtained in step (3) to enhance the image quality and obtain a high-resolution stitched image E; The depth homography estimation network further includes n - 2 local feature correlation layers and n - 2 regression network modules. Among them, the (n - 1)-th regression network module is used to predict the x and y coordinate offsets of two vertices of the overlapping parts of the shelf commodity images A and B according to the correlation of each feature point output by the global feature correlation layer, and then obtain the estimated homography matrix H based on the eight obtained offset coordinates and the projection transformation factor 1; the (n - 2)-th local feature correlation layer is used to perform deformation on the feature map of the shelf commodity image A output by the (n - 2)-th feature processing module and the feature map corresponding to the shelf commodity image B output by the n-th feature processing module according to the homography matrix H output by the n-th regression network module to obtain the feature map F B ’ 1 / (2^(n-2)) perform the correlation calculation of each feature point; the (n - 2)-th regression network module is used to predict the x and y coordinate offsets of two vertices of the overlapping parts of the shelf commodity images A and B according to the correlation of each feature point output by the (n - 2)-th local feature correlation layer, and then obtain the estimated homography matrix H based on the eight obtained offset coordinates and the projection transformation factor 1; and so on, until the 1st regression network module predicts the x and y coordinate offsets of two vertices of the overlapping parts of the shelf commodity images A and B according to the correlation of each feature point output by the 1st local feature correlation layer, and then obtain the estimated homography matrix H based on the eight obtained offset coordinates and the projection transformation factor 1; In the regression network module, the x and y coordinate offsets of two vertices overlapping between the images of goods on a shelf A and B are predicted according to the correlation of each feature point output by the global feature correlation layer, and the predicted x and y coordinate offsets are adjusted according to the size ratio of the two images of goods on a shelf to be stitched input and the images in the training dataset, specifically as follows: where σw = w / W, σh = h / H, w and h represent the width and height of the images in the training dataset respectively, and W and H represent the width and height of the two shelf product images to be stitched; ΔU i and ΔV i (i = 1, 2, 3, 4) represent the offsets of the x - coordinates and y - coordinates of the four vertices of the overlapping region in the same coordinate system when performing homography estimation on the images in the training dataset; σwΔU i and σwΔV i (i = 1, 2, 3, 4) represent the offsets of the x - coordinates and y - coordinates of the four vertices of the overlapping region in the same coordinate system when performing homography estimation on the two shelf product images to be stitched.

2. The method according to claim 1, characterized in that, step (3) is specifically: Input the deformed image C obtained in step (2) and the image of goods on a shelf A into a trained encoder-decoder network for stitching and fusion to obtain a fused image D; the encoder-decoder network includes an encoder and a decoder, and the encoder is used to reconstruct the features of the overlapping area in the two images according to the input deformed image C obtained in step (2) and the image of goods on a shelf A; The decoder is used to decode according to the features output by the encoder and simultaneously restore the non-overlapping area to obtain a fused image D.

3. The method according to claim 1 or 2, characterized in that, step (4) is specifically: Input the image D obtained in step (3) into a trained optimization branch for feature optimization to enhance the image quality and obtain a high-resolution stitched image E; the optimization branch is composed of a first convolutional layer, multiple deep residual blocks, and multiple second convolutional layers connected in sequence.

4. A system for stitching images of goods on a shelf, characterized in that, used to implement the method described in any one of claims 1-3, including: The homography estimation module is used to input two shelf product images A and B into a trained deep homography estimation network to obtain an estimated homography matrix H; deform the shelf product image B according to the homography matrix H to obtain a corresponding deformed image C; The stitching and fusion module is used to stitch and fuse the obtained deformed image C with the shelf product image A to obtain a fused image D; The feature optimization module is used to optimize the features of the obtained image D, enhance the image quality, and obtain a high-resolution stitched image E.

Citation Information

Patent Citations

  • Rapid panoramic stitching method and system for microscopic images

    CN111626936A

  • Multi-channel picture splicing method based on end-to-end neural network

    CN111709880A