Homography matrix estimation method, electronic device, storage medium and program product

By designing a homography matrix estimation network and using intermediate images for training, the problem of large holographic matrix estimation error in the large baseline scenarios is solved, and efficient and robust homography matrix estimation is achieved.

CN116266385BActive Publication Date: 2025-06-06YUANLI TUXIN (CHONGQING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310107729.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-31
Publication Date
2025-06-06
Estimated Expiration
2043-01-31

AI Technical Summary

Technical Problem

The existing deep learning-based homography matrix estimation method has a large error in large baseline scenarios and depends on feature matching points between images, requiring the image to contain rich textures and good lighting conditions.

Method used

By designing a homography matrix estimation method, the homography matrix between the pending source image and the pending target image is predicted using the homography matrix estimation network. The network is obtained through training, and the intermediate image is obtained using the target homography matrix and the sample source image transformation, and the input network is used as the image in the training image to calculate the loss of the predicted homography matrix and optimize it.

Benefits of technology

This method can efficiently and robustly solve the homography matrix in large baseline scenarios, reducing dependence on image texture and lighting conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116266385B_ABST
    Figure CN116266385B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides a homography matrix estimation method, electronic device, storage medium and program product. The method includes: obtaining a source image to be processed and a target image to be processed; inputting the obtained image into a homography matrix estimation network to obtain a homography matrix between the images, and the above network is trained and obtained in the following manner: obtaining a sample source image and a sample target image; for each group of target homography matrices in at least one group of target homography matrices, using each target homography matrix therein, performing image distortion on the distorted image to obtain an intermediate state image; for each group of training images in at least one group of training images, inputting each image pair in the group of training images into the above network to obtain a predicted homography matrix corresponding to the image pairs in the group of training images one by one; calculating the total prediction loss; and optimizing the parameters in the network based on the total prediction loss. The scheme can accurately solve the homography matrix in a large baseline scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine learning, and more specifically to a homography matrix estimation method, electronic device, storage medium and computer program product. Background Art

[0002] Homography matrix estimation is a basic and important computer vision task, and has been widely used in fields such as high dynamic range (HDR) imaging, image stitching, and video stabilization. Early traditional methods mainly rely on feature extraction and feature matching, and solve the homography matrix of two images through the direct linear transformation (DLT). However, traditional methods rely heavily on feature matching points between images, requiring images to contain rich textures and good lighting conditions. Deep learning-based methods do not rely on matching feature points between images, but directly learn the homography matrix of two images through convolutional neural networks. However, due to the large parallax transformation of large baseline scenes and the large non-overlap rate between source and target images, the existing deep learning-based methods have large errors in solving the homography matrix in large baseline scenes. Summary of the invention

[0003] The present invention is proposed in view of the above problems. The present invention provides a homography matrix estimation method, electronic device, storage medium and computer program product.

[0004] According to one aspect of the present invention, a homography matrix estimation method is provided, comprising: obtaining a source image to be processed and a target image to be processed; inputting the source image to be processed and the target image to be processed into a homography matrix estimation network to obtain a homography matrix between the source image to be processed and the target image to be processed, wherein the homography matrix estimation network is trained and obtained in the following manner: obtaining a sample source image and a sample target image; for each set of target homography matrices in at least one set of target homography matrices, performing image distortion on an image to be distorted by using each target homography matrix in the set of target homography matrices to obtain an intermediate state image, wherein each set of target homography matrices in the at least one set of target homography matrices includes at least one homography matrix, and in each set of target homography matrices in the at least one set of target homography matrices, the image to be distorted corresponding to the first target homography matrix is ​​the sample source image, and the image to be distorted corresponding to each of the remaining target homography matrices is an intermediate state image obtained by distortion of the previous target homography matrix; for each set of training images in at least one set of training images, Each image pair in the group of training images is input into a homography matrix estimation network to obtain a predicted homography matrix corresponding to the image pairs in the group of training images, wherein the at least one group of training images corresponds to the at least one group of target homography matrices in a one-to-one manner, each group of training images in the at least one group of training images includes a first image pair and a second image pair and / or includes at least one third image pair, the first image pair includes the sample source image and the sample target image, the second image pair includes the sample target image and an intermediate state image corresponding to the last target homography matrix in the corresponding group of target homography matrices, the at least one third image pair corresponds to at least one target homography matrix in the corresponding group of target homography matrices, and each third image pair includes an image to be distorted and an intermediate state image corresponding to the corresponding target homography matrix; based on at least part of the predicted homography matrices corresponding to the image pairs in each group of training images in the at least one group of training images, a total prediction loss of the homography matrix estimation network is calculated; and parameters in the homography matrix estimation network are optimized based on the total prediction loss.

[0005] Exemplarily, the calculating the total prediction loss of the homography matrix estimation network based on at least a portion of the predicted homography matrices corresponding to each of at least some of the image pairs in each of the at least one group of training images comprises: for each of the at least one group of training images, calculating the first prediction loss corresponding to the group of training images based on at least the predicted homography matrices corresponding to each of the first image pairs and the second image pairs in the group of training images, and / or, calculating the second prediction loss corresponding to the group of training images based on the predicted homography matrices corresponding to each of at least one third image pair in the group of training images and the target homography matrix; calculating the total prediction loss based on the first prediction loss and / or the second prediction loss corresponding to each of the at least one group of training images.

[0006] Exemplarily, the calculating of the first prediction loss corresponding to the group of training images at least based on the predicted homography matrices corresponding to the first image pair and the second image pair in the group of training images includes: calculating the first prediction loss corresponding to the group of training images based on the difference between the first matrix multiplication result and the second matrix multiplication result, wherein the first matrix multiplication result is the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images, and the second matrix multiplication result is the matrix cumulative multiplication result of the predicted homography matrices corresponding to at least one third image pair in the group of training images.

[0007] Exemplarily, the first prediction loss corresponding to the group of training images is calculated based on the predicted homography matrices corresponding to the first image pair and the second image pair in the group of training images, including: calculating the first prediction loss corresponding to the group of training images based on the difference between the first matrix multiplication result and the second matrix multiplication result, wherein the first matrix multiplication result is the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images, and the second matrix multiplication result is the matrix accumulation result of a group of target homography matrices corresponding to the group of training images.

[0008] Exemplarily, the method of calculating the second prediction loss corresponding to the group of training images based on the predicted homography matrix and the target homography matrix corresponding to each of at least one third image pairs in the group of training images includes: calculating the second prediction loss corresponding to the group of training images based on the difference between the predicted homography matrix and the target homography matrix corresponding to each of at least one third image pairs in the group of training images.

[0009] Exemplarily, the homography matrix estimation network includes an encoder, a global association layer and a homography estimation module, and the step of inputting the source image to be processed and the target image to be processed into the homography matrix estimation network to obtain the homography matrix between the source image to be processed and the target image to be processed comprises: inputting the source image to be processed and the target image to be processed into the encoder to obtain at least one group of first image features corresponding to the source image to be processed and at least one group of second image features corresponding to the target image to be processed, wherein the at least one group of first image features corresponds to the at least one group of second image features in one-to-one correspondence; inputting each group of first image features and the corresponding group of second image features into the global association layer to obtain global association features corresponding to the group of first image features; inputting the global association features corresponding to each of the at least one group of first image features into the homography estimation module to obtain the homography matrix between the source image to be processed and the target image to be processed.

[0010] Exemplarily, the homography estimation module includes at least one motion estimation module corresponding to the at least one group of first image features one by one, and the global correlation features corresponding to each of the at least one group of first image features are input into the homography estimation module to obtain the homography matrix between the source image to be processed and the target image to be processed, including: inputting the global correlation features corresponding to the i-th group of first image features into the i-th motion estimation module to obtain the i-th homography matrix in the form of a matrix stream; inputting the i-th homography matrix in the form of a matrix stream with the global correlation features corresponding to the i+1-th group of first image features; The local correlation features are combined to obtain combined correlation features; the combined correlation features are input into the i+1th motion estimation module to obtain the i+1th homography matrix in the form of a matrix stream; the homography matrix in the form of a matrix stream output by the last motion estimation module is directly linearly transformed to obtain the homography matrix in the form of a matrix as the homography matrix between the source image to be processed and the target image to be processed; wherein i=1,2,3,…,N-1, N is the number of the at least one group of first image features, and the dimension of the i-th group of first image features is smaller than the dimension of the i+1-th group of first image features.

[0011] According to another aspect of the present invention, an electronic device is provided, comprising a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used to execute the above-mentioned homography matrix estimation method when the processor is executed.

[0012] According to another aspect of the present invention, a storage medium is provided, on which program instructions are stored, and the program instructions are used to execute the above homography matrix estimation method when running.

[0013] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program, wherein the computer program is used to execute the above homography matrix estimation method when running.

[0014] According to the homography matrix estimation method, electronic device, storage medium and computer program product of the embodiment of the present invention, a homography matrix estimation network is used to predict the homography matrix between the source image to be processed and the target image to be processed. The homography matrix estimation network is obtained by training in the following manner. That is, through the target homography matrix, an intermediate state image is obtained based on the sample source image conversion, and at least an image pair in the training image is composed based on the intermediate state image. Subsequently, each image pair can be input into the homography matrix estimation network to obtain a predicted homography matrix, and the loss is calculated based on at least the predicted homography matrix to optimize the homography matrix estimation network. This scheme can efficiently and robustly realize the training of the homography matrix estimation network by dividing the large baseline scene into multiple small baseline scenes using the intermediate state image. Therefore, when the homography matrix estimation network trained by the above training method is applied to the actual homography matrix estimation, the homography matrix under the large baseline scene can be accurately solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 A schematic block diagram showing an example electronic device for implementing a homography matrix estimation method according to an embodiment of the present application;

[0017] Figure 2 A schematic flow chart showing a homography matrix estimation network training method according to an embodiment of the present application is shown;

[0018] Figure 3 A schematic diagram showing how an image to be distorted generates an intermediate image according to a target homography matrix according to an embodiment of the present application;

[0019] Figure 4 A schematic diagram showing a homography matrix estimation network training method according to an embodiment of the present application;

[0020] Figure 5 A schematic diagram showing a homography matrix estimation network training method according to another embodiment of the present application;

[0021] Figure 6A schematic diagram showing a homography matrix estimation network according to an embodiment of the present application is shown;

[0022] Figure 7 A schematic diagram showing a homography estimation module according to an embodiment of the present application is shown;

[0023] Figure 8 A schematic diagram showing a coarse motion estimation module according to an embodiment of the present application is shown;

[0024] Fig. 9 A schematic diagram showing a fine motion estimation module according to an embodiment of the present application;

[0025] Fig.10 A schematic flow chart showing a homography matrix estimation method according to an embodiment of the present application;

[0026] Fig.11 A schematic block diagram showing a homography matrix estimation device according to an embodiment of the present invention; and

[0027] Fig.12 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0028] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, image processing, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.

[0029] In order to make the purpose, technical solutions and advantages of the present application more obvious, the example embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein. Based on the embodiments of the present application described in the present application, all other embodiments obtained by those skilled in the art without creative work should fall within the scope of protection of the present application.

[0030] The embodiments of the present application provide a homography matrix estimation method, electronic device, storage medium and computer program product. According to the homography matrix estimation method of the embodiments of the present application, the homography matrix in a large baseline scenario can be accurately solved. The homography matrix estimation technology according to the embodiments of the present application can be applied to any field involving homography matrix estimation.

[0031] First, refer to Figure 1 An example electronic device 100 for implementing the homography matrix estimation method according to an embodiment of the present application is described.

[0032] like Figure 1As shown, the electronic device 100 includes one or more processors 102 and one or more storage devices 104. Optionally, the electronic device 100 may also include an input device 106, an output device 108, and an image acquisition device 110, and these components are interconnected via a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structures of the electronic device 100 shown are merely exemplary and non-limiting. The electronic device may also have other components and structures as required.

[0033] The processor 102 can be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic array (PLA), and a microprocessor. The processor 102 can be a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), or one or a combination of other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 100 to perform desired functions.

[0034] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 102 may run the program instructions to implement the client functions (implemented by the processor) in the embodiments of the present application described below and / or other desired functions. Various applications and various data may also be stored in the computer-readable storage medium, such as various data used and / or generated by the application, etc.

[0035] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, and the like.

[0036] The output device 108 can output various information (such as images and / or sounds) to the outside (such as a user), and can include one or more of a display, a speaker, etc. Optionally, the input device 106 and the output device 108 can be integrated together and implemented using the same interactive device (such as a touch screen).

[0037] The image acquisition device 110 can acquire images and store the acquired images in the storage device 104 for use by other components. The image acquisition device 110 can be a separate camera or a camera in a mobile terminal, etc. It should be understood that the image acquisition device 110 is only an example, and the electronic device 100 may not include the image acquisition device 110. In this case, other devices with image acquisition capabilities can be used to acquire images and send the acquired images to the electronic device 100.

[0038] Exemplarily, the example electronic device for implementing the homography matrix estimation method according to the embodiment of the present application can be implemented on a device such as a personal computer, a terminal device, an attendance machine, a panel machine, a camera or a remote server, etc. The terminal device includes but is not limited to: a tablet computer, a mobile phone, a PDA (Personal Digital Assistant), a touch-screen all-in-one machine, a wearable device, etc.

[0039] In order to solve the above technical problems, the present application provides a homography matrix estimation method, which processes a source image to be processed and a target image to be processed by a homography matrix estimation network to obtain a homography matrix between the source image to be processed and the target image to be processed. The homography matrix estimation network can be obtained by training using a training method according to an embodiment of the present invention. For ease of understanding, the following first refers to Figure 2 Describe the training method for the homography matrix estimation network described above.

[0040] Figure 2 FIG. 1 is a schematic flow chart of a homography matrix estimation network training method according to an embodiment of the present application. Figure 2 As shown, the homography matrix estimation network training method 200 includes the following steps S210, S220, S230, S240 and S250.

[0041] In step S210, a sample source image and a sample target image are acquired.

[0042] The sample source image and the sample target image may be any image having a homography transformation relationship. The sample source image and / or the sample target image may be an original image acquired by an image acquisition device, an image acquired by preprocessing the original image acquired by the image acquisition device, or a composite image generated by computer technology. At least part of the image area in the sample source image and at least part of the image area in the sample target image correspond to the same object or scene. Exemplarily, the sample source image and the sample target image may be two images of a scene taken from different shooting angles. The sample source image and the sample target image may come from an external device, and are transmitted by the external device to the electronic device 100 for homography matrix estimation network training. In addition, the sample source image and the sample target image may also be acquired by the electronic device 100 itself. For example, the electronic device 100 may use an image acquisition device 110 (e.g., an independent camera) to acquire a sample source image and a corresponding sample target image. The image acquisition device 110 may transmit the acquired sample source image and sample target image to the processor 102, and the processor 102 may perform homography matrix estimation network training.

[0043] In step S220, for each set of target homography matrices in at least one set of target homography matrices, an image to be distorted is distorted through each target homography matrix in the set of target homography matrices to obtain an intermediate image, wherein each set of target homography matrices in at least one set of target homography matrices includes at least one homography matrix, and in each set of target homography matrices in at least one set of target homography matrices, the image to be distorted corresponding to the first target homography matrix is ​​a sample source image, and the image to be distorted corresponding to each of the remaining target homography matrices is an intermediate image obtained by distortion of the previous target homography matrix.

[0044] Exemplarily, for a set of target homography matrices, each target homography matrix may be equal or unequal. The target homography matrix may be set according to the non-overlapping ratio between the corresponding image to be distorted and the intermediate image. For example, the non-overlapping ratio may be set to 20%, and the target homography matrix may be set to ensure that the non-overlapping ratio between the intermediate image obtained by distortion of the target homography matrix and the image to be distorted is less than 20%.

[0045] Exemplarily, for a set of target homography matrices, the image to be distorted corresponding to the first target homography matrix is ​​the sample source image. The image to be distorted corresponding to the second target homography matrix is ​​the first intermediate state image obtained by distorting the sample source image through the first target homography matrix. The image to be distorted corresponding to the third target homography matrix is ​​the second intermediate state image obtained by distorting the first intermediate state image through the second target homography matrix. By analogy, the image to be distorted corresponding to the Nth target homography matrix is ​​the N-1th intermediate state image obtained by distorting the N-2th intermediate state image through the N-1th target homography matrix.

[0046] The number of homography matrices included in any two sets of target homography matrices may be equal or unequal. In one embodiment, a set of target homography matrices includes three target homography matrices, namely, matrix 1, matrix 2, and matrix 3. Among them, the image to be distorted corresponding to matrix 1 is the sample source image. The image to be distorted corresponding to matrix 2 is the first intermediate state image obtained by distorting the sample source image through matrix 1. The image to be distorted corresponding to matrix 3 is the second intermediate state image obtained by distorting the first intermediate state image through matrix 2.

[0047] Figure 3 FIG. 1 is a schematic diagram showing how an image to be distorted generates an intermediate image according to a target homography matrix according to an embodiment of the present application. Figure 3 As shown, the sample source image is I s 0 The first target homography matrix H gt 1 The corresponding image to be distorted is I s 0 The second target homography matrix H gt 2 The corresponding image to be distorted is I s 0 By H gt 1 The first intermediate image I obtained by distortion s 1 ...the i-th target homography matrix H gt i The corresponding image to be distorted is the i-2th intermediate state image I s i-2 Through the i-1th target homography matrix H gt i-1 The i-1th intermediate state image I obtained by distortion s i-1 ...the Nth target homography matrix H gt N The corresponding image to be distorted is the N-2th intermediate state image I s N-2Through the N-1th target homography matrix H gt N-1 The N-1th intermediate state image I obtained by distortion s N-1 . N is the number of target homography matrices in the current set of target homography matrices.

[0048] In step S230, for each group of training images in at least one group of training images, each image pair in the group of training images is input into a homography matrix estimation network to obtain a predicted homography matrix corresponding one-to-one to the image pairs in the group of training images, wherein at least one group of training images corresponds one-to-one to at least one group of target homography matrices, and each group of training images in at least one group of training images includes a first image pair and a second image pair and / or includes at least one third image pair, the first image pair includes a sample source image and a sample target image, the second image pair includes the sample target image and an intermediate image corresponding to the last target homography matrix in the corresponding group of target homography matrices, and the at least one third image pair corresponds one-to-one to at least one target homography matrix in the corresponding group of target homography matrices, and each third image pair includes an image to be distorted and an intermediate image corresponding to the corresponding target homography matrix.

[0049] Exemplarily, the homography matrix estimation network can be implemented using any existing or future neural network capable of performing homography matrix estimation. For example, the homography matrix estimation network can be implemented using one or more of the following neural networks: Convolutional Neural Networks (CNN), U-Net, Fully Convolutional Networks (FCN), Deep Convolutional Encoder-Decoder Structure for Image Segmentation (SegNet), Pyramid Scene Parsing Network (PSPNet), Residual Network (ResNet), etc. Of course, the above neural network models are only examples, and the homography matrix estimation network can also be implemented using other suitable network structures.

[0050] In one embodiment, a group of training images may include a first image pair and a second image pair. Alternatively, a group of training images may include at least one third image pair. Still alternatively, a group of training images may include a first image pair, a second image pair, and at least one third image pair. The types of image pairs included in any two groups of training images may be consistent or inconsistent. For example, the first group of training images may include a first image pair and a second image pair, and the second group of training images may include at least one third image pair. For another example, the first group of training images may include a first image pair, a second image pair, and at least one third image pair, and the second group of training images may also include a first image pair, a second image pair, and at least one third image pair.

[0051] In such Figure 3 In the embodiment shown, I t represents a sample target image. The first image pair includes I s 0 and I t The second image pair includes I t and I s N The third image pair includes the image to be distorted and the intermediate image corresponding to the target homography matrix. The target homography matrix can be H gt 1 , H gt 2 , H gt 3 ...H gt i ...H gt N It can be understood that at least one third image pair corresponds to at least one target homography matrix in the corresponding set of target homography matrices, that is, the number of third image pairs is the same as the number of homography matrices in the corresponding set of target homography matrices. Figure 3 In the specific embodiment shown, the number of third image pairs is N, which are respectively gt 1 , H gt 2 , H gt 3 ...H gt i ...H gt N That is, the first third image pair and H gt 1 Corresponding, including I s 0 and I s 1 ; The second third image pair and H gt 2 Corresponding, including Is 1 and I s 2 ; ...; The i-th third image pair and H gt i Corresponding, including I s i-1 and I s i ; ...; The Nth third image pair and H gt N Corresponding, including I s N-1 and I s N .

[0052] By inputting any of the above image pairs into the homography matrix estimation network, a predicted homography matrix output by the homography matrix estimation network at the output end or the intermediate position can be obtained.

[0053] In step S240, a total prediction loss of the homography matrix estimation network is calculated based on at least the predicted homography matrices corresponding to at least some of the image pairs in each training image group in at least one group of training images.

[0054] Exemplarily, the training images may be a group, and the total prediction loss of the homography estimation matrix is ​​calculated based on the predicted homography matrices corresponding to at least some of the training image pairs in the group of training images. For example, in the embodiment where the above-mentioned group of training images includes a first image pair, a second image pair, and at least one third image pair, the prediction loss corresponding to the group of training images can be calculated only by the predicted homography matrices corresponding to the first image pair and the second image pair in the group of training images, and the prediction loss corresponding to the group of training images is used as the total prediction loss. Alternatively, the training images may be multiple groups. In one embodiment, the prediction loss corresponding to each group of training images can be calculated separately, and the prediction losses corresponding to each group of training images can be summed or averaged to serve as the total prediction loss of the homography matrix estimation network. For example, there are 5 groups of training images, and the prediction losses corresponding to each group of training images are L, ... 1 , L 2 , L 3 , L 4 and L 5 The total prediction loss of the homography estimation matrix L = (L 1 +L 2 +L 3 +L 4 +L 5 ) / 5.

[0055] In step S250, the parameters in the homography matrix estimation network are optimized based on the total prediction loss.

[0056] It can be understood that the optimization process can be one, two or more times. In one embodiment, the first total prediction loss corresponding to the first group of training images can be used to optimize the parameters in the homography matrix estimation network. After the first optimization is completed, the above steps S210, S220, S230 and S240 can be repeated for the second group of training images based on the homography matrix estimation network optimized for the first time to obtain the second total prediction loss. The homography matrix estimation network optimized for the first time is optimized again based on the second total prediction loss. Repeat the above process, and iterate and optimize the homography matrix estimation network for multiple times for the third group of training images, the fourth group of training images, the fifth group of training images, etc., until the total prediction loss of the homography matrix estimation network is within a certain range (loss convergence), the optimization is completed.

[0057] The above steps S210, S220, S230, S240 and S250 may represent the training phase of the homography matrix estimation network. It is understood that when the homography matrix estimation network is used to perform actual homography matrix estimation (which may be referred to as a testing or reasoning phase), the processing steps for the source image to be processed and the target image to be processed may be implemented by referring to the processing operation of the homography matrix estimation network on any image pair in the training phase. For the sake of brevity, it will not be described here.

[0058] According to the above technical solution, an intermediate state image is obtained based on the sample source image conversion through the target homography matrix, and the image pairs in the training image are composed at least based on the intermediate state image. Subsequently, each image pair can be input into the homography matrix estimation network to obtain a predicted homography matrix, and the loss is calculated at least based on the predicted homography matrix to optimize the homography matrix estimation network. This solution can efficiently and robustly realize the training of the homography matrix estimation network by using the intermediate state image to divide the large baseline scene into multiple small baseline scenes, and the obtained homography matrix estimation network can accurately solve the homography matrix under the large baseline scene.

[0059] Exemplarily, calculating the total prediction loss of the homography matrix estimation network (step S240) based on at least some of the predicted homography matrices corresponding to each of at least some of the image pairs in each group of training images in at least one group of training images includes at least the following steps: for each group of training images in at least one group of training images, calculating the first prediction loss corresponding to the group of training images based on at least the predicted homography matrices corresponding to each of the first image pairs and the second image pairs in the group of training images, and / or calculating the second prediction loss corresponding to the group of training images based on the predicted homography matrices corresponding to each of at least one third image pair in the group of training images and the target homography matrix; calculating the total prediction loss based on the first prediction loss and / or the second prediction loss corresponding to each of the at least one group of training images.

[0060] Exemplarily, for each group of at least one group of training images, the third prediction loss corresponding to the group of training images can be calculated based on the first prediction loss and / or the second prediction loss corresponding to the group of training images, and the third prediction losses corresponding to each of the at least one group of training images can be summed or averaged to obtain a total prediction loss. Exemplarily, when calculating the third prediction loss, the third prediction loss can be calculated based solely on the first prediction loss. For example, the first prediction loss is determined to be the third prediction loss. Alternatively, the third prediction loss can be calculated based solely on the second prediction loss. For example, the second prediction loss is determined to be the third prediction loss. Yet alternatively, the third prediction loss can be calculated based on the first prediction loss and the second prediction loss. For example, the first prediction loss and the second prediction loss can be summed or averaged to determine the third prediction loss. In one embodiment, represents the first prediction loss, represents the second prediction loss. Then the third prediction loss corresponding to the i-th group of training images is The process of determining the first predicted loss and the second predicted loss is shown in the following embodiment.

[0061] According to the above technical solution, by using the first prediction loss and / or the second prediction loss, the total prediction loss can be accurately determined. Therefore, based on the accurate total prediction loss, the optimization effect of the homography matrix estimation network can be guaranteed, and the accuracy of the optimized homography matrix estimation network can be improved.

[0062] Exemplarily, based on at least the predicted homography matrices corresponding to the first image pair and the second image pair in the group of training images, the step of calculating the first prediction loss corresponding to the group of training images may include the following steps: Based on the difference between the first matrix multiplication result and the second matrix multiplication result, the first prediction loss corresponding to the group of training images is calculated, wherein the first matrix multiplication result is the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images, and the second matrix multiplication result is the matrix cumulative multiplication result of the predicted homography matrix corresponding to at least one third image pair in the group of training images.

[0063] It can be understood that, theoretically, the intermediate state image corresponding to the last image pair in at least one third image pair should be the same as the image obtained by distorting the sample target image through the inverse matrix of the homography matrix corresponding to the second image pair. According to the cyclic recombination relationship, when the estimation of the homography matrix estimation network is completely accurate, the matrix accumulation result of the predicted homography matrix corresponding to each of at least one third image pair in the group of training images and the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images should be the same. That is, the first matrix multiplication result and the second matrix multiplication result should be the same. For N predicted homography matrices, with represents the predicted homography matrix corresponding to the first image pair. Represents the predicted homography matrix corresponding to the second image pair. represents the predicted homography matrix corresponding to the i+1th third image pair. Then theoretically there exists:

[0064]

[0065] Wherein, Π· represents matrix multiplication. The difference between the first matrix multiplication result and the second matrix multiplication result can reflect the accuracy of the homography matrix estimation network. The larger the difference, the lower the accuracy of the homography matrix estimation network. Therefore, the first prediction loss can be calculated according to the difference between the first matrix multiplication result and the second matrix multiplication result.

[0066] In the embodiment where the above-mentioned set of target homography matrices includes three target homography matrices, the images distorted by the three target homography matrices are the first intermediate state image, the second intermediate state image and the third intermediate state image. The three predicted homography matrices are H 1 , H 2 and H 3 . With H 0 Represents the predicted homography matrix corresponding to the first image pair, that is, H 0 Represents the predicted homography matrix corresponding to the sample source image and the sample target image. t Represents the predicted homography matrix corresponding to the second image pair. That is, H t represents the prediction homography matrix corresponding to the third intermediate state image and the sample target image. Then the first prediction loss can be In a specific embodiment, Among them, |·| 1 represents the L1 norm.

[0067] Figure 4 FIG. 2 shows a schematic diagram of a homography matrix estimation network training method according to an embodiment of the present application. Figure 4As shown, the predicted homography matrix corresponding to the first image pair can be expressed as The predicted homography matrix corresponding to the second image pair can be expressed as The predicted homography matrix corresponding to the third image pair can be expressed as The first prediction loss can be:

[0068]

[0069] According to the above technical solution, based on the cyclic recombination relationship, the first prediction loss can be accurately determined by using the difference between the first matrix multiplication result and the second matrix multiplication result. This optimization solution based on the cyclic recombination relationship can train a homography matrix estimation network with higher accuracy.

[0070] Exemplarily, based on at least the predicted homography matrices corresponding to the first image pair and the second image pair in the group of training images, the step of calculating the first prediction loss corresponding to the group of training images may include the following steps: Calculate the first prediction loss corresponding to the group of training images based on the difference between the first matrix multiplication result and the second matrix multiplication result, wherein the first matrix multiplication result is the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images, and the second matrix multiplication result is the matrix accumulation result of a group of target homography matrices corresponding to the group of training images.

[0071] As described above, when the estimation of the homography matrix estimation network is completely accurate, the matrix multiplication result of the predicted homography matrix corresponding to at least one third image pair in the group of training images and the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images should be the same. It can be understood that when the homography matrix estimation network is accurate, the predicted homography matrix corresponding to the same third image pair should also be the same as the target homography matrix. Therefore, the matrix multiplication result of a group of target homography matrices corresponding to the group of training images can be used as the second matrix multiplication result of the above embodiment. For N target homography matrices, represents the predicted homography matrix corresponding to the first image pair. Represents the predicted homography matrix corresponding to the second image pair. represents the target homography matrix corresponding to the third image pair. Then theoretically there exists:

[0072]

[0073] The difference between the first matrix multiplication result and the second matrix multiplication result can reflect the accuracy of the homography matrix estimation network. The larger the difference, the lower the accuracy of the homography matrix estimation network. Therefore, the first prediction loss can be calculated according to the difference between the first matrix multiplication result and the second matrix multiplication result.

[0074] According to the above technical solution, based on the cyclic recombination relationship, the first matrix multiplication result and the second matrix multiplication result are used to accurately calculate the first prediction loss. During training, the solution only needs to obtain the prediction homography matrix corresponding to the first image pair and the prediction homography matrix corresponding to the second image pair through the homography matrix estimation network to perform loss calculation, which occupies less resources and is highly efficient.

[0075] Exemplarily, based on the predicted homography matrix and the target homography matrix respectively corresponding to at least one third image pair in the group of training images, the step of calculating the second prediction loss corresponding to the group of training images may include the following steps: Based on the difference between the predicted homography matrix and the target homography matrix respectively corresponding to at least one third image pair in the group of training images, the second prediction loss corresponding to the group of training images is calculated.

[0076] As described above, when the estimation of the homography matrix estimation network is completely accurate, the predicted homography matrix and the target homography matrix corresponding to the same third image pair should be the same. Therefore, the second prediction loss corresponding to the set of training images can be calculated by the difference between the predicted homography matrix and the target homography matrix corresponding to the same third image pair.

[0077] It can be understood that the number of third image pairs can be selected according to the calculation accuracy. The more the number of third image pairs, the higher the calculation accuracy. Otherwise, vice versa. In one embodiment, the number of target homography matrices is N. Accordingly, the number of third image pairs is N. The second prediction loss can be calculated based on any one or several of the N third image pairs. For example, the second prediction loss can be calculated based on the difference between the predicted homography matrix corresponding to a third image pair and the target homography matrix. Specifically, for example, the absolute value of the difference between the predicted homography matrix corresponding to the third image pair and the target homography matrix can be determined as the second prediction loss. For example, the second prediction loss represents the i+1th predicted homography matrix corresponding to the i+1th image pair, and represents the i+1th target homography matrix. Alternatively, the differences between the predicted homography matrices and the target homography matrices corresponding to the N third image pairs may be summed or averaged to determine the second prediction loss. In a specific embodiment, the absolute values ​​of the differences between the predicted homography matrices and the target homography matrices corresponding to the N third image pairs are summed to determine the second prediction loss. For example, the second prediction loss

[0078] Figure 5 A schematic diagram of a homography matrix estimation network training method according to an embodiment of the present application is shown. Figure 5 The figure shows a set of target homography matrices. When there are more than one set of target homography matrices, the processing method for each set of target homography matrices is the same and will not be described in detail. Figure 5 As shown, a set of target homography matrices includes N target homography matrices. First, the image to be distorted is distorted by using the N target homography matrices in the set of target homography matrices to obtain an intermediate image. The specific method for obtaining the intermediate image has been described in detail in the above embodiment, and for the sake of brevity, it will not be repeated here. Then, the first image pair, the second image pair and the N third image pairs are input into the homography matrix estimation network to obtain multiple predicted homography matrices. Among them, For the first image pair, For the second image pair, Corresponding to the third image pair, i∈[0, N-1]. Then, the predicted homography matrices corresponding to multiple third image pairs can be multiplied to obtain Finally, calculate and The difference between them is used to determine the total prediction loss, and the total prediction loss is used to optimize the homography matrix estimation network.

[0079] According to the above technical solution, the second prediction loss can be accurately determined by using the difference between the predicted homography matrix and the target homography matrix corresponding to the third image pair, thereby improving the optimization effect of the homography matrix estimation network based on the accurate second prediction loss.

[0080] Exemplarily, the homography matrix estimation network includes an encoder, a global association layer and a homography estimation module. In the training stage of the homography matrix estimation network, for each group of training images in at least one group of training images, each image pair in the group of training images is input into the homography matrix estimation network, and the step of obtaining a predicted homography matrix corresponding to the image pairs in the group of training images can specifically include the following steps: for each image pair in each group of training images in at least one group of training images, the image pair is input into the encoder, and at least one group of first image features corresponding to the first image in the image pair and at least one group of second image features corresponding to the second image in the image pair are obtained, and at least one group of first image features corresponds to at least one group of second image features; each group of first image features and the corresponding group of second image features are input into the global association layer to obtain global association features corresponding to the group of first image features; the global association features corresponding to each of the at least one group of first image features are input into the homography estimation module to obtain the predicted homography matrix corresponding to the image pair.

[0081] It can be understood that an image pair refers to two images corresponding to a target homography matrix. The two images are the image to be distorted and the intermediate image corresponding to the target homography matrix. That is, the first image can be one of the image to be distorted and the intermediate image, and the second image is the other of the image to be distorted and the intermediate image.

[0082] Exemplarily, the encoder module may be implemented using any downsampling module, which may include, for example, a convolutional layer, a pooling layer, etc., so that the encoder module may obtain multi-layer features corresponding to the image pair.

[0083] It can be understood that since the first image and the second image in the image pair are input into the same encoder, the number of features of the first image is the same as the number of features of the second image, and the encoding features of the first image and the second image at the same layer correspond. In one embodiment, the first image and the second image both obtain four layers of encoding features through the encoder, and the first layer of encoding features of the first image corresponds to the first layer of encoding features of the second image. The second layer of encoding features of the first image corresponds to the second layer of encoding features of the second image. The third layer of encoding features of the first image corresponds to the third layer of encoding features of the second image. The fourth layer of encoding features of the first image corresponds to the fourth layer of encoding features of the second image. Each layer of encoding features can be a set of image features.

[0084] After the first image features and the second image features are acquired, each group of first image features and the corresponding group of second image features are input into the global association layer to obtain the global association features corresponding to the group of first image features.

[0085] Exemplarily, the number of global correlation features is the same as the number of first image features. As described above, the number of first image features is the same as the number of second image features. In the above embodiment where the coding features are four layers, the number of first image features and the number of second image features are both four groups. The first image features are sequentially recorded as and The second image features are recorded as and Then and Input them into the global correlation layer respectively to obtain the global correlation features and

[0086] The global correlation features corresponding to at least one set of first image features are input into a homography estimation module to obtain a predicted homography matrix corresponding to the image pair.

[0087] In the above, we get the global correlation features and In an embodiment of the present invention, one of the global correlation features can be input into the homography estimation module. Alternatively, multiple global correlation features can be input into the homography estimation module. The number of global correlation features input into the homography estimation module can be selected according to the estimation accuracy. The more the number of global correlation features input, the higher the estimation accuracy; otherwise, vice versa.

[0088] According to the above technical solution, by utilizing the first image features and the second image features in a set of image pairs, and obtaining the global correlation features corresponding to each first image feature in a set of first image features through a global correlation layer, the global correlation features can be used to accurately determine the homography matrix between the first image and the second image.

[0089] Exemplarily, the encoder is a multi-scale convolutional encoder, and before inputting the image pair into the encoder for each image pair in each group of training images in at least one group of training images to obtain at least one group of first image features corresponding to the first image in the image pair and at least one group of second image features corresponding to the second image in the image pair, method 200 may further include the following steps: for each image pair in each group of training images in at least one group of training images, reducing the image pair into at least one small-resolution image pair corresponding to at least one target resolution. For each image pair in each group of training images in at least one group of training images, inputting the image pair into the encoder to obtain at least one group of first image features corresponding to the first image in the image pair and at least one group of second image features corresponding to the second image in the image pair, including: for each image pair in each group of training images in at least one group of training images, inputting the image pair and the corresponding at least one small-resolution image pair into the multi-scale convolutional encoder to obtain at least one group of original resolution features generated based on the image pair and at least one group of small-resolution features generated based on at least one small-resolution image pair, wherein the at least one group of first image features corresponding to the image pair includes at least one group of original resolution features and at least one group of small-resolution features.

[0090] It is understandable that the resolution of each set of training images may be the same or different. For example, the resolution of the first set of training images is 100PPI, and the resolution of the second set of training images is 150PPI. Since the computational complexity of the global correlation layer in calculating the global correlation features is related to the resolution of the image, in order to improve processing efficiency and accuracy and adapt to the processing of images with different resolutions, the image can be compressed into a small resolution image with a target resolution. For example, the resolution of the first set of training images and the second set of training images can be compressed to 50PPI.

[0091] After the image is reduced to a small-resolution image, features of the image pair and the corresponding small-resolution image can be extracted respectively. In other words, for each image pair in each training image in at least one set of training images, the image pair is input into an encoder to obtain at least one set of first image features corresponding to the first image in the image pair and at least one set of second image features corresponding to the second image in the image pair, which may include the following steps.

[0092] For each image pair in each group of training images in at least one group of training images, the image pair and the corresponding at least one small-resolution image pair are input into a multi-scale convolutional encoder to obtain at least one group of original-resolution features generated based on the image pair and at least one group of small-resolution features generated based on at least one small-resolution image pair, wherein the at least one group of first image features corresponding to the image pair includes at least one group of original-resolution features and at least one group of small-resolution features.

[0093] Exemplarily, after obtaining at least one set of original resolution features generated based on the image pair and at least one set of small resolution features generated based on at least one small resolution image pair, some of the small resolution features in the small resolution image can be used to replace the corresponding original resolution features in the image pair. For example, in the embodiment in which the encoder extracts four layers of coding features, the last two layers of image features of the small resolution image can be used to replace the original resolution features of the last two layers in the image pair. In the first image feature above, it is recorded as and In the example, the features of the small-resolution image of the first image are recorded as and After feature replacement, the first image feature input to the global correlation layer can be and

[0094] Figure 6 FIG. 4 shows a schematic diagram of a homography matrix estimation network according to an embodiment of the present application. Figure 6 As shown, first, the source image I s and the target image I t Reduce them to fixed small resolution images respectively to get the corresponding images and Will I s ,I t , and Input them into the multi-scale convolution encoder respectively to obtain the four-layer features of the source image and the target image. That is, the corresponding features of the source image are obtained Corresponding features of the target image in, For I s Features, for characteristics. For I t Features, for Then, the above features are input into the global correlation layer to obtain the global correlation features. Finally, the global correlation features are input into the homography estimation module to obtain the source image I s and the target image I t The homography estimation matrix corresponding to the two.

[0095] According to the above technical solution, by reducing the image to a small resolution image in advance, the computational complexity of the global correlation feature can be reduced and the computational efficiency can be improved. The solution can adapt to the operation of images with different resolutions.

[0096] Exemplarily, the homography estimation module includes at least one motion estimation module corresponding one-to-one to at least one set of first image features. The step of inputting the global correlation features corresponding to each of at least one group of first image features into the homography estimation module to obtain the predicted homography matrix corresponding to the image pair may include the following steps: inputting the global correlation features corresponding to the i-th group of first image features into the i-th motion estimation module to obtain the i-th homography matrix in the form of a matrix stream; combining the i-th homography matrix in the form of a matrix stream with the global correlation features corresponding to the i+1-th group of first image features to obtain the combined correlation features; inputting the combined correlation features into the i+1-th motion estimation module to obtain the i+1-th homography matrix in the form of a matrix stream; determining the homography matrix in the form of a matrix stream output by the last motion estimation module as the predicted homography matrix corresponding to the image pair, or performing a direct linear transformation on the homography matrix in the form of a matrix stream output by the last motion estimation module to obtain the homography matrix in the form of a matrix as the predicted homography matrix corresponding to the image pair; wherein i=1,2,3,…,N-1, N is the number of groups of at least one group of first image features, and the dimension of the i-th group of first image features is smaller than the dimension of the i+1-th group of first image features.

[0097] In the above, we get the global correlation features and In the embodiment of the present invention, the global correlation feature corresponding to the first group of first image features is The global correlation feature corresponding to the first image feature of the second group is The global correlation feature corresponding to the first image feature of the third group is The global correlation feature corresponding to the first image feature of the fourth group is In one embodiment, first, The first homography matrix is ​​input into the first motion estimation module to obtain the first homography matrix in the form of a matrix stream. Then, the first homography matrix in the form of a matrix stream is combined with the global correlation feature corresponding to the first image feature of the second group. Combine to obtain the combined correlation feature. Input the combined correlation feature into the second motion estimation module to obtain the second homography matrix in the form of matrix flow. Then, the second homography matrix in the form of matrix flow is combined with the global correlation feature corresponding to the third group of first image features. Combine to obtain the combined correlation feature. Input the combined correlation feature into the third motion estimation module to obtain the third homography matrix in the form of a matrix flow. After obtaining the third homography matrix, the third homography matrix in the form of a matrix flow is combined with the global correlation feature corresponding to the fourth group of first image features. Combine to obtain the combined correlation features. Input the combined correlation features into the fourth motion estimation module to obtain the fourth homography matrix in the form of a matrix stream. Finally, the fourth homography matrix in the form of a matrix stream is determined as the predicted homography matrix corresponding to the image pair.

[0098] Exemplarily, after obtaining the fourth homography matrix in the form of a matrix flow, in order to avoid local motion disturbances that may exist in the matrix flow, the fourth homography matrix can be directly linearly transformed to obtain a homography matrix in the form of a matrix as the predicted homography matrix corresponding to the image pair. It can be understood that in the training stage of the homography matrix estimation network, the homography matrix in the form of a matrix flow can be directly applied as the predicted homography matrix without converting it into a homography matrix in the form of a matrix, thereby saving computing resources and improving training efficiency.

[0099] exist Figure 6 In the embodiment shown, the same layer features in the source image features and the target image features are associated in the global association layer. and Input the global correlation layer separately to obtain the global correlation features and Figure 7 FIG. 2 shows a schematic diagram of a homography estimation module according to an embodiment of the present application. Figure 7 As shown, first, Input into the first motion estimation module (shown as the first motion estimation module) to obtain the first homography matrix in the form of matrix flow Then, the first homography matrix in the form of matrix flow is and Combine to obtain the combined associated features. Input the combined associated features into the second motion estimation module (shown as the second motion estimation module) to obtain the second homography matrix in the form of a matrix flow Next, the second homography matrix in the form of matrix flow and Combine to obtain the combined associated features. Input the combined associated features into the third motion estimation module (shown as the third motion estimation module) to obtain the third homography matrix in the form of matrix flow: After obtaining the third homography matrix, the third homography matrix in the form of matrix flow is and Combine to obtain the combined associated features. Input the combined associated features into the fourth motion estimation module (shown as the fourth motion estimation module) to obtain the fourth homography matrix in the form of a matrix flow: Finally, the fourth homography matrix in the form of a matrix flow is The homography matrix is ​​obtained by direct linear transformation in matrix form

[0100] In such Figure 7 In the embodiment shown, the second homography matrix can also be By direct linear transformation, we get This matrix represents the homography matrix between the small-resolution image of the source image and the small-resolution image of the target image.

[0101] According to the above technical solution, by using multiple motion estimation modules to obtain the homography matrix and using the combined correlation features between two adjacent groups of global correlation features, the accuracy of the obtained predicted homography matrix can be improved.

[0102] Exemplarily, the last motion estimation module is a fine motion estimation module, and at least part of the remaining motion estimation modules are coarse motion estimation modules. The coarse motion estimation module includes a first residual module, a first subspace constraint module, and a second residual module connected in sequence. The fine motion estimation module is sequentially connected to a third residual module, a second subspace constraint module, and a fourth residual module, a merging module, a first convolution module, a third subspace constraint module, and a second convolution module, wherein the input feature of the fine motion estimation module is jump-connected to the input end of the merging module, and in the merging module, the output feature of the fourth residual module is merged with the input feature of the jump connection.

[0103] In such Figure 7 In the illustrated embodiment, the first motion estimation module and the third motion estimation module may both be coarse motion estimation modules (CME), and the second motion module and the fourth motion module may both be fine motion estimation modules (Fine Motion Estimator, FME).

[0104] Exemplarily, the sizes of the first residual module and the second residual module may be the same or different. In one embodiment, the size of the first residual module may be set to include 32 residual blocks, and the size of the second residual module may be set to include 64 residual blocks, so as to obtain output values ​​of different scales.

[0105] Exemplarily, the sizes of the third residual module and the fourth residual module may be the same or different. The specific size setting method is similar to the setting method of the first residual module and the second residual module, and will not be repeated here for the sake of brevity.

[0106] Figure 8 FIG. 4 shows a schematic diagram of a coarse motion estimation module according to an embodiment of the present application. Figure 8As shown, after the features are input into CME, the first residual module extracts the features, the first subspace constraint module processes the extracted features, and the second residual module extracts the features output by the first subspace constraint module again, and then outputs the homography matrix in the form of a matrix flow.

[0107] Fig. 9 FIG. 4 is a schematic diagram showing a fine motion estimation module according to an embodiment of the present application. Fig. 9 First, the input feature is input into FME, and the features are extracted and processed by the third residual module, the second subspace constraint module and the fourth residual module in sequence, and then input into the merging module to be merged with the input feature. Then, the merged feature is processed by the first convolution module, the third subspace constraint module and the second convolution module in sequence, and the homography matrix in the form of matrix flow is output.

[0108] According to the above technical solution, by constructing a coarse motion estimation module and a fine motion estimation module and combining the two modules, the homography matrix corresponding to each image pair can be accurately estimated. The module structure of this solution is simple and the result is reliable.

[0109] According to another aspect of the present application, a homography matrix estimation method is provided. Fig.10 FIG. 1 is a schematic block diagram of a homography matrix estimation method 1000 according to an embodiment of the present application. Fig.10 As shown, the homography matrix estimation method 1000 may include step S1010 and step S1020.

[0110] In step S1010, a source image to be processed and a target image to be processed are obtained.

[0111] The methods of obtaining the source image to be processed and the target image to be processed are similar to those of the sample source image and the sample target image, and will not be described in detail.

[0112] In step S1020, the source image to be processed and the target image to be processed are input into the above-mentioned homography matrix estimation network to obtain the homography matrix between the source image to be processed and the target image to be processed.

[0113] In step S1020, the above-mentioned Figure 6-9 Preferably, in step S1020, the homography matrix can be obtained in matrix form through the homography matrix estimation network, for example, the homography matrix obtained is obtained from Figure 7 The direct linear transformation module below outputs the homography matrix in matrix form.

[0114] According to the above-mentioned homography matrix estimation method, a homography matrix estimation network is used to predict the homography matrix between the source image to be processed and the target image to be processed. The homography matrix estimation network is obtained by training in the following manner. That is, through the target homography matrix, an intermediate state image is obtained based on the sample source image conversion, and at least an image pair in the training image is composed based on the intermediate state image. Subsequently, each image pair can be input into the homography matrix estimation network to obtain a predicted homography matrix, and the loss is calculated based on at least the predicted homography matrix to optimize the homography matrix estimation network. This scheme can efficiently and robustly realize the training of the homography matrix estimation network by dividing the large baseline scene into multiple small baseline scenes using the intermediate state image. Therefore, when the homography matrix estimation network trained by the above training method is applied to the actual homography matrix estimation, the homography matrix under the large baseline scene can be accurately solved.

[0115] Exemplarily, a homography matrix estimation network includes an encoder, a global association layer and a homography estimation module. A source image to be processed and a target image to be processed are input into the homography matrix estimation network to obtain a homography matrix between the source image to be processed and the target image to be processed, including: inputting the source image to be processed and the target image to be processed into the encoder to obtain at least one group of first image features corresponding to the source image to be processed and at least one group of second image features corresponding to the target image to be processed, wherein at least one group of first image features corresponds to at least one group of second image features; inputting each group of first image features and the corresponding group of second image features into the global association layer to obtain global association features corresponding to the group of first image features; inputting the global association features corresponding to each of the at least one group of first image features into the homography estimation module to obtain a homography matrix between the source image to be processed and the target image to be processed.

[0116] When describing the homography matrix estimation network training method 200 above, the structure and working principle of the encoder, global correlation layer and homography estimation module in the homography matrix estimation network have been described. You can refer to the above text about the working method of each module in the network when any image pair in the training image is input into the homography matrix estimation network to understand the working method of each module in the network when the source image to be processed and the target image to be processed are input into the homography matrix estimation network. No further details will be given here.

[0117] According to the above technical scheme, by utilizing the first image features of the source image to be processed and the second image features of the target image to be processed, and obtaining the global correlation features corresponding to each first image feature in a set of first image features through the global correlation layer, the global correlation features can be used to accurately determine the homography matrix between the source image to be processed and the target image to be processed.

[0118] Exemplarily, the homography estimation module includes at least one motion estimation module corresponding to at least one group of first image features one by one, and the global correlation features corresponding to each of the at least one group of first image features are input into the homography estimation module to obtain the homography matrix between the source image to be processed and the target image to be processed, including: inputting the global correlation features corresponding to the i-th group of first image features into the i-th motion estimation module to obtain the i-th homography matrix in the form of a matrix stream; combining the i-th homography matrix in the form of a matrix stream with the global correlation features corresponding to the i+1-th group of first image features to obtain the combined correlation features; inputting the combined correlation features into the i+1-th motion estimation module to obtain the i+1-th homography matrix in the form of a matrix stream; performing direct linear transformation on the homography matrix in the form of a matrix stream output by the last motion estimation module to obtain the homography matrix in the form of a matrix as the homography matrix between the source image to be processed and the target image to be processed; wherein, i=1,2,3,…,N-1, N is the number of groups of at least one group of first image features, and the dimension of the i-th group of first image features is smaller than the dimension of the i+1-th group of first image features.

[0119] When describing the homography matrix estimation network training method 200 above, the structure and working principle of the homography estimation module have been described. You can refer to the above description of the working method of the homography estimation module when any image pair in the training image is input into the homography matrix estimation network to understand the working method of the homography estimation module when the source image to be processed and the target image to be processed are input into the homography matrix estimation network. It will not be repeated here.

[0120] According to the above technical solution, by using multiple motion estimation modules to obtain a homography matrix and using the combined correlation features between two adjacent groups of global correlation features, the accuracy of the obtained homography matrix can be improved.

[0121] Exemplarily, the encoder is a multi-scale convolutional encoder. Before inputting the source image to be processed and the target image to be processed into the encoder to obtain at least one set of first image features corresponding to the source image to be processed and at least one set of second image features corresponding to the target image to be processed, method 1000 may further include the following steps. Reduce the source image to be processed and the target image to be processed to at least one small-resolution image pair corresponding to at least one target resolution. In this embodiment, inputting the source image to be processed and the target image to be processed into the encoder to obtain at least one set of first image features corresponding to the source image to be processed and at least one set of second image features corresponding to the target image to be processed may include the following steps. Inputting the source image to be processed and the target image to be processed and the corresponding at least one small-resolution image pair into the multi-scale convolutional encoder to obtain at least one set of original resolution features generated based on the source image to be processed and the target image to be processed and at least one set of small-resolution features generated based on at least one small-resolution image pair, the at least one set of first image features corresponding to the source image to be processed and the target image to be processed includes at least one set of original resolution features and at least one set of small-resolution features.

[0122] This step is similar to the step performed after the image pair is compressed into a small-resolution image pair in the above training stage. The specific details have been described in detail in the above embodiment and will not be repeated here for the sake of brevity.

[0123] According to the above technical solution, by reducing the image to a small resolution image in advance, the computational complexity of the global correlation feature can be reduced and the computational efficiency can be improved. The solution can adapt to the operation of images with different resolutions.

[0124] Exemplarily, the homography matrix estimation method according to the embodiment of the present application can be implemented in a device, apparatus or system having a memory and a processor.

[0125] According to the embodiment of the present application, the homography matrix estimation method can be deployed at the image acquisition end, for example, it can be deployed at a personal terminal or a server end.

[0126] Alternatively, the homography matrix estimation method according to the embodiment of the present application can also be deployed in a distributed manner on the server side (or cloud) and the personal terminal. For example, an image can be acquired on the client side, and the client side transmits the acquired image to the server side (or cloud), and the server side (or cloud) performs homography matrix estimation.

[0127] According to another aspect of the present application, a homography matrix estimation device is provided. Fig.11 FIG. 1 is a schematic block diagram of a homography matrix estimation device 1100 according to an embodiment of the present application.

[0128] like Fig.11As shown, the homography matrix estimation device 1100 according to the embodiment of the present application includes an acquisition module 1110 and an input module 1120. Each module can respectively execute the above Fig.10 The following describes the various steps of the homography matrix estimation method. The following only describes the main functions of the various components of the homography matrix estimation device 1100, and omits the details described above.

[0129] The acquisition module 1110 is used to acquire a source image to be processed and a target image to be processed.

[0130] The input module 1120 is used to input the source image to be processed and the target image to be processed into the above-mentioned homography matrix estimation network to obtain the homography matrix between the source image to be processed and the target image to be processed.

[0131] Fig.12 A schematic block diagram of an electronic device 1200 according to an embodiment of the present application is shown. The electronic device 1200 includes a memory 1210 and a processor 1220 .

[0132] The memory 1210 stores computer program instructions for implementing corresponding steps in the homography matrix estimation method according to the embodiment of the present application.

[0133] The processor 1220 is used to run the computer program instructions stored in the memory 1210 to perform corresponding steps of the homography matrix estimation method according to the embodiment of the present application.

[0134] Exemplarily, the electronic device 1200 may further include an image acquisition device 1230. The image acquisition device 1230 is used to acquire a source image to be processed and a target image to be processed. The image acquisition device 1230 is optional, and the electronic device 1200 may not include the image acquisition device 1230. At this time, the processor 1220 may obtain the source image to be processed and the target image to be processed by other means, such as obtaining the source image to be processed and the target image to be processed from an external device or from the memory 1210.

[0135] In addition, according to an embodiment of the present application, a storage medium is also provided, on which program instructions are stored, and when the program instructions are executed by a computer or a processor, the corresponding steps of the homography matrix estimation method of the embodiment of the present application are used to execute, and the corresponding modules in the homography matrix estimation device according to the embodiment of the present application are implemented. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.

[0136] In addition, according to an embodiment of the present application, a computer program product is also provided. The computer program product includes a computer program, and the computer program is used to execute the above-mentioned homography matrix estimation method when running.

[0137] Each module in the electronic device according to the embodiment of the present application can be implemented by running computer program instructions stored in a memory by a processor of the electronic device that implements the homography matrix estimation method according to the embodiment of the present application, or can be implemented when computer instructions stored in a computer-readable storage medium of a computer program product according to the embodiment of the present application are executed by a computer.

[0138] In addition, according to an embodiment of the present application, a computer program is also provided, which is used to execute the above-mentioned homography matrix estimation method when running.

[0139] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present application to this. Those of ordinary skill in the art may make various changes and modifications therein without departing from the scope and spirit of the present application. All these changes and modifications are intended to be included within the scope of the present application as required by the appended claims.

[0140] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0141] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0142] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0143] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various application aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present application should not be interpreted as reflecting the following intention: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the corresponding claims, the application point is that the corresponding technical problem can be solved with less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiment are hereby explicitly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.

[0144] It will be understood by those skilled in the art that, except for mutually exclusive features, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this specification may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0145] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.

[0146] The various component embodiments of the present application can be implemented in hardware, or implemented in software modules running on one or more processors, or implemented in a combination thereof. It should be understood by those skilled in the art that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all functions of some modules in the homography matrix estimation network training device or the homography matrix estimation device according to the embodiment of the present application. The application can also be implemented as a device program (for example, a computer program and a computer program product) for performing a part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0147] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and that those skilled in the art may design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets should not be constructed as a limitation to the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of multiple such elements. The present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim that lists several devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names.

[0148] The above is only a specific implementation method or description of a specific implementation method of the present application, and the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. The protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A homography matrix estimation method, include: Obtaining a source image to be processed and a target image to be processed; The source image to be processed and the target image to be processed are input into a homography matrix estimation network to obtain a homography matrix between the source image to be processed and the target image to be processed, wherein the homography matrix estimation network is trained in the following manner: Obtain a sample source image and a sample target image; For each set of target homography matrices in at least one set of target homography matrices, perform image warping on the image to be distorted through each target homography matrix in the set of target homography matrices to obtain an intermediate state image, wherein each set of target homography matrices in the at least one set of target homography matrices includes at least one homography matrix, and in each set of target homography matrices in the at least one set of target homography matrices, the image to be distorted corresponding to the first target homography matrix is ​​the sample source image, and the image to be distorted corresponding to each of the remaining target homography matrices is an intermediate state image obtained by warping the previous target homography matrix; For each group of training images in at least one group of training images, each image pair in the group of training images is input into a homography matrix estimation network to obtain a predicted homography matrix corresponding one-to-one to the image pairs in the group of training images, wherein the at least one group of training images corresponds one-to-one to the at least one group of target homography matrices, and each group of training images in the at least one group of training images includes a first image pair and a second image pair and / or includes at least one third image pair, the first image pair includes the sample source image and the sample target image, the second image pair includes the sample target image and an intermediate state image corresponding to the last target homography matrix in the corresponding group of target homography matrices, the at least one third image pair corresponds one-to-one to at least one target homography matrix in the corresponding group of target homography matrices, and each third image pair includes an image to be distorted and an intermediate state image corresponding to the corresponding target homography matrix; Calculating a total prediction loss of the homography matrix estimation network based at least on predicted homography matrices corresponding to at least some of the image pairs in each of the at least one group of training images; Parameters in the homography matrix estimation network are optimized based on the total prediction loss.

2. The method according to claim 1, in, The calculating the total prediction loss of the homography matrix estimation network based on the predicted homography matrices corresponding to at least some of the image pairs in each group of training images in the at least one group of training images comprises: For each set of training images in the at least one set of training images, Calculating a first prediction loss corresponding to the group of training images based on at least a predicted homography matrix corresponding to each of a first image pair and a second image pair in the group of training images, and / or calculating a second prediction loss corresponding to the group of training images based on a predicted homography matrix and a target homography matrix corresponding to each of at least one third image pair in the group of training images; The total prediction loss is calculated based on the first prediction loss and / or the second prediction loss corresponding to each of the at least one set of training images.

3. The method according to claim 2, in, The step of calculating the first prediction loss corresponding to the group of training images based on at least the prediction homography matrices corresponding to the first image pair and the second image pair in the group of training images, comprises: Based on the difference between the first matrix multiplication result and the second matrix multiplication result, the first prediction loss corresponding to the group of training images is calculated, wherein the first matrix multiplication result is the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images, and the second matrix multiplication result is the matrix cumulative multiplication result of the predicted homography matrices corresponding to at least one third image pair in the group of training images.

4. The method according to claim 2, in, The step of calculating the first prediction loss corresponding to the group of training images based on at least the prediction homography matrices corresponding to the first image pair and the second image pair in the group of training images, comprises: Based on the difference between the first matrix multiplication result and the second matrix multiplication result, the first prediction loss corresponding to the group of training images is calculated, wherein the first matrix multiplication result is the cross product result between the inverse matrix of the predicted homography matrix corresponding to the second image pair in the group of training images and the predicted homography matrix corresponding to the first image pair in the group of training images, and the second matrix multiplication result is the matrix accumulation result of a group of target homography matrices corresponding to the group of training images.

5. The method according to claim 2, in, The calculating the second prediction loss corresponding to the group of training images based on the prediction homography matrix and the target homography matrix corresponding to each of at least one third image pair in the group of training images comprises: Based on the difference between the predicted homography matrix and the target homography matrix corresponding to at least one third image pair in the group of training images, a second prediction loss corresponding to the group of training images is calculated.

6. The method according to any one of claims 1 to 5, in, The homography matrix estimation network includes an encoder, a global association layer and a homography estimation module, and the step of inputting the source image to be processed and the target image to be processed into the homography matrix estimation network to obtain the homography matrix between the source image to be processed and the target image to be processed includes: Inputting the source image to be processed and the target image to be processed into the encoder, obtaining at least one set of first image features corresponding to the source image to be processed and at least one set of second image features corresponding to the target image to be processed, wherein the at least one set of first image features corresponds to the at least one set of second image features in one-to-one correspondence; Inputting each group of first image features and the corresponding group of second image features into the global association layer to obtain global association features corresponding to the group of first image features; The global correlation features corresponding to each of the at least one group of first image features are input into the homography estimation module to obtain a homography matrix between the source image to be processed and the target image to be processed.

7. The method according to claim 6, in, The homography estimation module includes at least one motion estimation module corresponding to the at least one group of first image features one by one, and the global correlation features corresponding to each of the at least one group of first image features are input into the homography estimation module to obtain a homography matrix between the source image to be processed and the target image to be processed, including: Inputting the global correlation features corresponding to the i-th group of first image features into the i-th motion estimation module to obtain the i-th homography matrix in the form of a matrix flow; Combining the i-th homography matrix in the matrix flow form with the global correlation feature corresponding to the i+1-th group of first image features to obtain a combined correlation feature; Inputting the combined correlation features into the i+1th motion estimation module to obtain the i+1th homography matrix in the form of a matrix flow; Performing a direct linear transformation on the homography matrix in the form of a matrix stream output by the last motion estimation module to obtain a homography matrix in the form of a matrix as a homography matrix between the source image to be processed and the target image to be processed; Wherein, i=1, 2, 3, ..., N-1, N is the number of the at least one group of first image features, and the dimension of the i-th group of first image features is smaller than the dimension of the i+1-th group of first image features.

8. An electronic device comprising a processor and a memory, in, The memory stores computer program instructions, which are used by the processor to execute the homography matrix estimation method according to any one of claims 1 to 7 when the processor is running the computer program instructions.

9. A storage medium having program instructions stored thereon, in, The program instructions are used to execute the homography matrix estimation method according to any one of claims 1 to 7 when running.

10. A computer program product, the computer program product comprising a computer program, in, The computer program is used to execute the homography matrix estimation method according to any one of claims 1 to 7 when running.

Citation Information

Patent Citations

  • Image registration method and device, electronic equipment and computer readable storage medium

    CN113516697A

  • Image alignment method and device, computer readable storage medium and electronic equipment

    CN113962846A