Image stitching methods and apparatus, storage media and electronic devices
By calculating the intermediate optical flow and the optical flow inversion network, combined with the mask calculation network, the image is directly mapped and stitched, which solves the problems of large computational load and low efficiency in the existing technology, and realizes efficient and artifact-free image stitching.
Patent Information
- Application Number
- CN202110597189.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-28
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-05-28
AI Technical Summary
Existing image stitching technologies are computationally intensive, inefficient, and difficult to achieve real-time stitching. Furthermore, the stitching results contain artifacts and distortions.
By calculating the first and second intermediate optical flows, and using optical flow calculation networks and optical flow inversion networks, combined with mask calculation networks, the image is directly mapped and stitched, avoiding iterative calculation of multiple homography matrices.
It improves the efficiency and quality of image stitching, reduces artifacts and distortion, and achieves high-quality image stitching.
Smart Images

Figure CN113469880B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and more specifically, to an image stitching method and apparatus, a storage medium and an electronic device. Background Technology
[0002] Image stitching refers to the process of combining multiple images with overlapping areas to obtain a seamless panoramic image. In recent years, image stitching technology has been widely used in fields such as aerospace, minimally invasive medical surgery, medical microscopic observation, and geological surveys.
[0003] In existing technologies, image stitching can be achieved through camera calibration, but this method is computationally intensive, requiring the iterative calculation of multiple homography matrices, which makes the stitching process inefficient. Summary of the Invention
[0004] The purpose of this application is to provide an image stitching method, apparatus, storage medium, and electronic device to improve the above-mentioned technical problems.
[0005] In a first aspect, embodiments of this application provide an image stitching method, comprising: acquiring a first image and a second image; calculating a first intermediate optical flow and a second intermediate optical flow based on the first image and the second image, wherein the first intermediate optical flow is the optical flow between the intermediate image and the first image, the second intermediate optical flow is the optical flow between the intermediate image and the second image, the intermediate image is an image with a viewing angle between the first image and the second image, the size of the first intermediate optical flow is the same as the size of the first image, and the size of the second intermediate optical flow is the same as the size of the second image; and calculating a stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image.
[0006] The image stitching method described above is simple in its steps, eliminating the need for complex iterative calculations of multiple homography matrices to complete the image stitching. This improves efficiency and makes the stitching process more real-time, thus enhancing its practicality. Furthermore, the method calculates the same size for both the first and second intermediate optical flows as for the first and second images, facilitating direct mapping between them and enabling rapid image stitching.
[0007] In one implementation of the first aspect, calculating the first intermediate optical flow and the second intermediate optical flow based on the first image and the second image includes: respectively cropping image regions containing a common scene in the first image and the second image to obtain a first screenshot and a second screenshot; inputting the first screenshot and the second screenshot into an optical flow calculation network to obtain a first screenshot optical flow and a second screenshot optical flow, wherein the first screenshot optical flow is the optical flow between the intermediate image and the first screenshot, and the second screenshot optical flow is the optical flow between the intermediate image and the second screenshot; upsampling the first screenshot optical flow to the size of the first image to obtain the first intermediate optical flow, and upsampling the optical flow of the second screenshot to the size of the second image to obtain the second intermediate optical flow.
[0008] For image regions in the first and second images that share a common area, the intermediate optical flow can be estimated relatively accurately due to the correspondence between pixels. However, for image regions in the first and second images that do not share a common area, the intermediate optical flow is difficult to estimate because there is no correspondence between pixels. Therefore, directly using the complete first and second images for intermediate optical flow estimation may yield inaccurate results.
[0009] In the above implementation, optical flow estimation is first performed using the first and second screenshots (corresponding to image regions in the first and second images that share a common view). Then, the estimated small-scale optical flow is upsampled to obtain the required intermediate optical flow, which improves the accuracy of the intermediate optical flow. Furthermore, since the motion of a target in an image is consistent globally and locally in most cases, the target motion pattern reflected by the small-scale local optical flow is the same as that reflected by the large-scale global optical flow (i.e., the intermediate optical flow). Therefore, the validity of the new optical flow value generated during the upsampling process can be guaranteed.
[0010] In one implementation of the first aspect, calculating a stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image includes: mapping a first intermediate image based on the first intermediate optical flow and the first image, and mapping a second intermediate image based on the second intermediate optical flow and the second image; calculating a target optical flow based on the first intermediate optical flow and the second intermediate optical flow; mapping a first stitched image based on the target optical flow and the first intermediate image, and mapping a second stitched image based on the target optical flow and the second intermediate image; and stitching the first stitched image and the second stitched image together to obtain the stitched image.
[0011] The images to be stitched are typically images captured from different viewpoints targeting the same object (if the first and second images are completely unrelated, there is generally no need to stitch them together). The optical flow between the two images can be considered a quantitative representation of the motion of the object within the image. This motion includes both the motion of the object itself and the movement of the camera position (including the shooting angle). Therefore, using the intermediate image as a reference, the first intermediate optical flow (representing the motion of the first image relative to the intermediate image) corresponds to the viewpoint of the first image, and the second intermediate optical flow (representing the motion of the second image relative to the intermediate image) corresponds to the viewpoint of the second image.
[0012] Furthermore, since the target optical flow is generated at least based on the fusion of the first and second intermediate optical flows, it can be considered that the target optical flow also corresponds to a specific viewpoint. The image acquired from this viewpoint incorporates information from the first and second images, thus reflecting the state of the acquired target under different viewpoints, making it a relatively ideal stitched image. This stitched image can be calculated using the intermediate images and the target optical flow, exhibiting high stitching quality and improving upon the artifacts, distortions, and difficulty in aligning the images to be stitched that exist in traditional image stitching methods.
[0013] In one implementation of the first aspect, the step of calculating the target optical flow based on the first intermediate optical flow and the second intermediate optical flow includes: interpolating and calculating at least one transition optical flow based on the first intermediate optical flow and the second intermediate optical flow; and fusing the first intermediate optical flow, the at least one transition optical flow, and the second intermediate optical flow to obtain the target optical flow.
[0014] Based on the above explanation, the optical flow between two images can be considered a quantitative representation of the motion of the target within the image. Therefore, using the intermediate image as a reference, the first intermediate optical flow corresponds to the viewing angle of the first image, and the second intermediate optical flow corresponds to the viewing angle of the second image. Each transitional optical flow obtained by interpolating the first and second intermediate optical flows corresponds to a transitional viewing angle located between the first and second images. The image acquired at the transitional viewing angle is called the transitional image (no actual transitional image is acquired; the concept of a transitional image is introduced here simply to facilitate the explanation of the scheme's principle).
[0015] A virtual image acquisition process can be considered: first image is acquired from a certain perspective, then the camera is moved to each transitional perspective to acquire transitional images, and finally the second image is acquired. In this process, since the parallax between adjacent perspectives is small, the first intermediate optical flow changes smoothly into each transitional optical flow, and finally changes into the second intermediate optical flow (also known as the smooth transition of optical flow).
[0016] Furthermore, since the target optical flow in the above implementation is generated by fusing the first intermediate optical flow, at least one transitional optical flow, and the second intermediate optical flow, it contains optical flow information from various viewpoints. Therefore, the target optical flow reflects the gradual change in optical flow during the virtual image acquisition process. Thus, it can be considered that the target optical flow also corresponds to a special gradual viewpoint. The image acquired under this gradual viewpoint incorporates information from the first image, the second image, and at least one transitional image, thereby comprehensively reflecting the state of the acquired target from various viewpoints, resulting in a relatively ideal stitched image. This stitched image can be calculated using the intermediate images and the target optical flow. As mentioned above, since this stitched image reflects the entirety of the acquired target, it has high stitching quality, improving upon the artifacts, distortion, and difficulty in aligning the images to be stitched that exist in traditional image stitching methods.
[0017] In one implementation of the first aspect, the step of interpolating and calculating at least one transition optical flow based on the first intermediate optical flow and the second intermediate optical flow includes: obtaining at least one weight value; and performing a weighted summation of the first intermediate optical flow and the second intermediate optical flow based on each weight value to obtain the at least one transition optical flow.
[0018] The first and second intermediate optical flows can be considered as the two endpoints of an interpolation operation. The interpolation operation estimates the value at at least one position between these two endpoints to achieve a smooth transition of the optical flow. The weighted summation operation described above is a linear interpolation, but nonlinear interpolation (e.g., quadratic, cubic, reciprocal interpolation) can also be used. Linear interpolation has the advantage of being computationally simple, and in most cases, linear motion is sufficient to describe the motion of the target between the first and second images, and the results of linear interpolation are also good enough.
[0019] In one implementation of the first aspect, the magnitude of the weight value is related to the spectral position of the transition optical flow corresponding to the weight value.
[0020] In the above implementation, the viewpoint position of the transition optical flow corresponding to the weight value is considered when setting the weight value, so that the transition optical flow calculated using the weight value will be consistent with its viewpoint position.
[0021] In one implementation of the first aspect, the sum of the weighting coefficients of the first intermediate optical flow and the second intermediate optical flow is 1, the weight value is the weighting coefficient of the first intermediate optical flow, and the magnitude of the weight value is positively correlated with the proximity between the viewing angle position of the transition optical flow corresponding to the weight value and the viewing angle position of the first intermediate optical flow.
[0022] When the sum of the weighting coefficients of the first intermediate optical flow and the second intermediate optical flow is 1, the weighting coefficient of the first intermediate optical flow can be regarded as the weight value (at which point the weighting coefficient of the second intermediate optical flow is 1 minus the weight value), or the weighting coefficient of the second intermediate optical flow can be regarded as the weight value (at which point the weighting coefficient of the first intermediate optical flow is 1 minus the weight value). There is no substantial difference between the two schemes.
[0023] Taking the former as an example, the closer the viewing angle of the transition optical flow is to the viewing angle of the first intermediate optical flow, the larger the weight value should be. That is, increase the weighting coefficient of the first intermediate optical flow while decreasing the weighting coefficient of the second intermediate optical flow. In this way, the value of the transition optical flow will be more influenced by the first intermediate optical flow, and will be consistent with its viewing angle. Furthermore, if all weight values are set according to this rule, it can be guaranteed that the calculated transition optical flows are gradual.
[0024] In one implementation of the first aspect, the at least one weight value is uniformly distributed within the interval (0,1).
[0025] In the above implementation, since the weight values are uniformly distributed in the interval (0,1), the distribution of the viewpoint position of the transition optical flow calculated using these weight values between the first and second images is also relatively uniform. Such a viewpoint position distribution allows the transition image to fully describe the overall picture of the acquired target between the viewpoints corresponding to the first and second images. Therefore, the stitched image that incorporates the transition image information can be regarded as having high quality.
[0026] In one implementation of the first aspect, the step of calculating the target optical flow based on the first intermediate optical flow and the second intermediate optical flow includes: calculating a preliminary optical flow based on the first intermediate optical flow and the second intermediate optical flow; and inputting the preliminary optical flow into an optical flow inversion network to obtain the target optical flow.
[0027] Optical flow inversion differs from simple matrix inversion, and the calculation process is more complex. In the above implementation, a neural network is used to perform optical flow inversion operations. On the one hand, this simplifies the calculation and improves the efficiency of optical flow inversion; on the other hand, the learning ability of the neural network can be used to improve the accuracy of optical flow inversion.
[0028] In one implementation of the first aspect, the step of fusing the first intermediate optical flow, the at least one transition optical flow, and the second intermediate optical flow to obtain the target optical flow includes: obtaining N+2 weight matrices, where N is the total number of transition optical flows; performing a weighted summation on the first intermediate optical flow, the N transition optical flows, and the second intermediate optical flow based on the N+2 weight matrices to obtain the preliminary optical flow; the preliminary optical flow is the target optical flow, or the preliminary optical flow is input into an optical flow inversion network to obtain the target optical flow.
[0029] In the above implementation, the transition optical flow is weighted and summed using a weight matrix. Unlike one-dimensional weight values, the weight matrix is two-dimensional, which allows for more flexible combination of information from different optical flows in the initial optical flow. This enables the initial optical flow to reflect the gradual change from the first intermediate optical flow to the second intermediate optical flow, and naturally, the target optical flow obtained based on the initial optical flow can also reflect this gradual change.
[0030] In one implementation of the first aspect, the position of the maximum value of the elements in the weight matrix is related to the viewpoint position of the optical flow corresponding to the weight matrix.
[0031] In the above implementation, taking the case where the position of the maximum value of an element in a weight matrix coincides with the viewing angle position of its corresponding optical flow (which could be the first intermediate optical flow, transitional optical flow, or second intermediate optical flow) as an example, the optical flow (also a matrix) contributes the most to the calculation of the initial optical flow at its viewing angle position, while the optical flow values at other positions contribute relatively less. Furthermore, if all weight matrices are set according to this rule, the optical flow value corresponding to each viewing angle position in the initial optical flow can be mainly contributed by the optical flow corresponding to that viewing angle position. This allows the initial optical flow to reflect the gradual transition from the first intermediate optical flow to the second intermediate optical flow, and naturally, the target optical flow obtained based on the initial optical flow can also reflect this gradual transition.
[0032] In one implementation of the first aspect, the step of stitching the first stitched image and the second stitched image to obtain the stitched image includes: inputting the first stitched image and the second stitched image into a mask computing network to obtain a stitching mask; and stitching the first stitched image and the second stitched image together based on the stitching mask to obtain the stitched image.
[0033] In the above implementation, a stitching mask is used to achieve a smooth transition between the first stitched image and the second stitched image at the stitching point. Furthermore, the stitching mask is not pre-set, but is learned by the mask computing network, which can further improve the quality of the stitched image.
[0034] In one implementation of the first aspect, the first intermediate optical flow and the second intermediate optical flow are calculated using the first screenshot optical flow and the second screenshot optical flow output by the optical flow calculation network; the target optical flow is calculated using the optical flow inversion network; and the stitched image is calculated using the stitching mask output by the mask calculation network. The method further includes: acquiring the first real screenshot optical flow, the second real screenshot optical flow, the real target optical flow, and the real stitched image; calculating the optical flow prediction loss based on the first screenshot optical flow, the second screenshot optical flow, the first real screenshot optical flow, and the second real screenshot optical flow; calculating the optical flow inversion loss based on the target optical flow and the real target optical flow; calculating the image stitching loss based on the stitched image and the real stitched image; calculating the total loss based on the optical flow prediction loss, the optical flow inversion loss, and the image stitching loss; and updating the parameters of the optical flow calculation network, the optical flow inversion network, and the mask calculation network based on the total loss.
[0035] The above implementation provides an end-to-end model training method that can be used to train an image stitching model. The image stitching model includes an optical flow calculation network, an optical flow inversion network, and a mask calculation network. When calculating the loss, the losses corresponding to the three networks are considered simultaneously: optical flow prediction loss, optical flow inversion loss, and image stitching loss. That is, by training, the accuracy of intermediate optical flow prediction, target optical flow prediction, and stitching mask prediction of the model are improved at the same time, so that the final image stitching model can achieve high-quality image stitching.
[0036] In one implementation of the first aspect, the target optical flow is obtained by fusing a first intermediate optical flow, at least one transition optical flow, and a second intermediate optical flow. Acquiring the first image and the second image includes: calculating the first image based on the intermediate image and a homography matrix, and calculating the second image based on the intermediate image and the inverse of the homography matrix, wherein the intermediate image is a real image. Acquiring the first real screenshot optical flow, the second real screenshot optical flow, the real target optical flow, and the real stitched image includes: calculating the first real intermediate optical flow based on the homography matrix, and calculating the first real screenshot optical flow based on the first real intermediate optical flow; calculating the second real intermediate optical flow based on the inverse of the homography matrix, and calculating the second real screenshot optical flow based on the second real intermediate optical flow; interpolating at least one transition matrix based on the homography matrix and the inverse of the homography matrix; fusing the homography matrix, the at least one transition matrix, and the inverse of the homography matrix to obtain a target matrix, and calculating the real target optical flow based on the target matrix; and calculating the real stitched image based on the intermediate image and the target matrix.
[0037] In the above implementation, the intermediate image is a real image, and the homography matrix can be predefined. Using the intermediate image and the homography matrix, the training supervision signals—real crop optical flow, real target optical flow, and the real stitched image—can be calculated. If a set of images to be stitched (including the first and second images) and their corresponding supervision signals are considered as a training sample, since the homography matrix can be arbitrarily specified, this implementation can quickly generate a large number of training samples using a small number of real images. Furthermore, these samples can cover different scenes, thus enabling the trained image stitching model to have good generalization ability.
[0038] Secondly, embodiments of this application provide an image stitching device, comprising: an image acquisition module for acquiring a first image and a second image; an intermediate optical flow calculation module for calculating a first intermediate optical flow and a second intermediate optical flow based on the first image and the second image, wherein the first intermediate optical flow is the optical flow between the intermediate image and the first image, the second intermediate optical flow is the optical flow between the intermediate image and the second image, the intermediate image is an image with a viewing angle between the first image and the second image, the size of the first intermediate optical flow is the same as the size of the first image, and the size of the second intermediate optical flow is the same as the size of the second image; and an image stitching module for calculating a stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image.
[0039] Thirdly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when read and executed by a processor, perform the method provided in the first aspect or any possible implementation thereof.
[0040] Fourthly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the memory stores computer program instructions, and the computer program instructions are read and executed by the processor to perform the method provided in the first aspect or any possible implementation of the first aspect. Attached Figure Description
[0041] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1This illustrates one possible flow of the image stitching method provided in an embodiment of this application;
[0043] Figure 2 This illustrates one possible data flow of the image stitching method provided in an embodiment of this application;
[0044] Figure 3 This illustrates the working principle of the image stitching method provided in the embodiments of this application;
[0045] Figure 4 This demonstrates the process of acquiring images using a fisheye camera and obtaining images to be stitched together.
[0046] Figure 5 This illustrates a possible flow of the model training method provided in an embodiment of this application;
[0047] Figure 6 This illustrates one possible way of generating training samples in the model training method provided in the embodiments of this application;
[0048] Figure 7 This invention illustrates one possible structure of the image stitching apparatus provided in an embodiment of this application;
[0049] Figure 8 This application illustrates one possible structure of an electronic device provided in an embodiment of the present application. Detailed Implementation
[0050] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. It should be noted that similar reference numerals and letters in the following drawings indicate similar items; therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0051] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0052] The terms “first,” “second,” etc., are used only to distinguish one entity or operation from another, and should not be construed as indicating or implying relative importance, nor as requiring or implying any such actual relationship or order between these entities or operations.
[0053] Figure 1This paper illustrates a possible flow of the image stitching method provided in an embodiment of this application. Figure 2 This illustrates one possible data flow during the execution of an image stitching method, for reference when explaining the steps of the method. The image stitching method can, but is not limited to, [the following]... Figure 8 The electronic device shown performs the following instructions; for details regarding the structure of this electronic device, please refer to the section on... Figure 8 The explanation. (Refer to...) Figure 1 The method includes:
[0054] Step S110: Obtain the first image and the second image.
[0055] The first and second images are the images to be stitched together. Figure 2 The two images are denoted as I0 and I1, respectively. The sources of the first and second images are not limited; for example, they can be images captured by a camera or images generated by a computer vision algorithm, etc. The following text mainly uses the case of camera capture as an example. The first and second images have the same image size, or the first and second images were the same size at the time of capture, or the first and second images had different sizes at the time of capture but were processed to the same size after capture.
[0056] The image stitching method proposed in this application does not, in principle, limit the content of the first and second images. However, considering the practical applications of image stitching, it is advisable to take as an example that the first and second images are images taken from different perspectives targeting the same target (e.g., Figure 3 (I0 and I1 in the image are captured at viewpoint 0 and viewpoint 1, respectively). Here, "target" refers to any object that can be photographed, such as people, animals, plants, landscapes, etc.
[0057] Since the stitched image (hereinafter referred to as the stitched image) is usually larger than the first and second images, to facilitate the calculations during the image stitching process, some implementations first roughly align the original first and second images (image alignment is an operation where, ideally, pixels corresponding to the same target position in the two aligned images can overlap). Then, as needed, zeros are padded around the roughly aligned image (i.e., pixels with a value of 0) to make it the same size as the stitched image. The zero-padded image is then used as the first and second images for subsequent image stitching. For example, Figure 2 The black parts in I0 and I1 (the left and right sides and the bottom of I0, and the left side of I1) are the parts filled with zeros.
[0058] The following examples illustrate the use of wide-angle and fisheye cameras for image acquisition:
[0059] For the former, a first wide-angle image and a second wide-angle image are first captured by a wide-angle camera. Then, a global homography matrix is calculated according to the camera calibration method. The first wide-angle image and the second wide-angle image are then roughly aligned according to this homography matrix. Finally, zeros are added around the roughly aligned images to obtain the final first image and the second image.
[0060] For the latter, a first fisheye image and a second fisheye image are first obtained by taking pictures with a fisheye camera. These two images are taken from the same position but in opposite directions. Figure 4 As shown in the left column. Based on the unfolding parameters, the first fisheye image and the second fisheye image can be unfolded separately to obtain the unfolded first and second unfolded images, as shown below. Figure 4 As shown in the column on the right. Zero padding has already been performed during the unfolding process, so there's no need for additional zero padding. Furthermore, the first and second unfolded images are roughly aligned. The first unfolded image is divided into two parts: A1 on the left and B1 on the right. The second unfolded image is also divided into two parts: A2 on the left and B2 on the right. A1 and A2 form one set of first and second images, while B1 and B2 form another set of first and second images.
[0061] Step S120: Calculate the first intermediate optical flow and the second intermediate optical flow based on the first image and the second image.
[0062] The optical flow between two images can be considered a quantitative representation of the motion of a target within the images. This motion includes both the motion of the target itself and the movement of the camera capturing the two images (including changes in the shooting angle). For the latter, because motion is relative, the movement of the camera changes the position of the target in the image, which is equivalent to the motion of the target itself. Specifically, the motion of the target causes a point on the target to correspond to a pixel at a different location in the two images. The coordinate offset (a vector) between these two pixels is the optical flow value at one of the pixel locations. For two images of the same size, the optical flow between them can also be considered as an optical flow map, which can be the same size as the two images, and each pixel value in the map is the optical flow value mentioned above.
[0063] In step S120, the intermediate image is a virtual image captured from an intermediate viewpoint (e.g., Figure 3 Image I acquired from viewpoint m mThe intermediate viewpoint is a perspective between the first and second images. The acquisition of the intermediate image here should be understood as a virtual acquisition; that is, if the camera is placed at the intermediate viewpoint to photograph the target, an intermediate image can be acquired, but in reality, this image acquisition behavior is not performed. The size of the intermediate image is the same as that of the first and second images.
[0064] It should be noted that in certain applications of image stitching methods (e.g., in the training phase of image stitching models, as detailed later), the intermediate image is not necessarily a virtual image; it may also be a real, captured image. However, in the explanation... Figure 1 When performing the steps, you can temporarily understand the intermediate image as a virtual image.
[0065] The first intermediate optical flow refers to the optical flow between the intermediate image and the first image, and the second intermediate optical flow refers to the optical flow between the intermediate image and the second image. Sometimes, for simplicity, both are referred to as intermediate optical flow. The first intermediate optical flow has two directions: one is the intermediate image I... m Optical flow to the first image I0, one is the optical flow from the first image I0 to the intermediate image I0. m The optical flow is denoted as F. m→0 and F 0→m When performing image stitching, only one direction of optical flow needs to be used. Figure 2 F is used m→0 Similarly, the second intermediate optical flow also has two directions, one being the intermediate image I. m The optical flow to the second image I1 is one type of optical flow from the second image I1 to the intermediate image I1. m The optical flow is denoted as F. m→1 and F 1→m When performing image stitching, only one direction of optical flow needs to be used. Figure 2 F is used m→1 (to be with F) m→0 To maintain consistency, the optical flow always originates from I. m Set off).
[0066] In the scheme of this application, the size of the first intermediate optical flow is the same as the size of the first image, and the size of the second intermediate optical flow is the same as the size of the second image, so as to facilitate the direct mapping of the first image and the second image using the first intermediate optical flow and the second intermediate optical flow in subsequent steps (the significance of mapping will be explained later), thereby quickly realizing image stitching.
[0067] In some implementations, a pre-trained neural network can be used to estimate the first intermediate optical flow and the second intermediate optical flow, taking the first image and the second image as input.
[0068] However, because the first and second images were captured from different perspectives, only a portion of the image area in each image contains the same scene. For example, Figure 2 In the image, I0 and I1 only share the front of the vehicle, while the body and rear of the vehicle exist only in I0. According to the definition of optical flow mentioned earlier, optical flow estimation can only be effectively performed if both images contain pixels corresponding to the same point on the target. Therefore, for image regions in the first and second images that share a common area, the intermediate optical flow can be estimated relatively accurately due to the correspondence between pixels. For image regions in the first and second images that do not share a common area, the intermediate optical flow is difficult to estimate because there is no correspondence between pixels (or, although the optical flow value can be calculated, the calculated value is inaccurate).
[0069] Therefore, directly using a neural network to perform intermediate optical flow estimation based on the complete first and second images may yield poor estimation results. Thus, in some implementations, a region containing common elements in the first and second images can be extracted first. Optical flow estimation is then performed only within this region using a neural network to obtain a small-scale optical flow with higher accuracy. This small-scale optical flow is then upsampled to obtain a larger-scale optical flow, i.e., the intermediate optical flow. This improves the estimation accuracy of the intermediate optical flow. In particular, for images captured by wide-angle and fisheye cameras, the overlapping areas are often large, while the non-overlapping areas are relatively small; this method is more conducive to obtaining high-quality optical flow estimation results. The following is a detailed explanation:
[0070] Step a: Extract the image regions containing the same scene from the first image and the second image respectively to obtain the first screenshot and the second screenshot.
[0071] For example, a rectangular frame can be used to crop an image region that shares a common element with the first and second images. The cropped image regions are referred to as the first screenshot and the second screenshot, respectively. Figure 2 In the image, the first and second screenshots are labeled overlap-0 and overlap-1, respectively.
[0072] Step b: Input the first and second screenshots into the optical flow calculation network to obtain the optical flow of the first and second screenshots.
[0073] Among them, the optical flow computation network is a neural network for estimating optical flow, and its training method is described in the introduction. Figure 5 This will be explained later. The network takes the first and second screenshots as input and outputs the optical flow of the first and second screenshots. The specific structure of the network is not limited.
[0074] The first screenshot optical flow refers to the optical flow between the intermediate image (more precisely, the portion of the intermediate image corresponding to the screenshot area) and the first screenshot. The second screenshot optical flow refers to the optical flow between the intermediate image (more precisely, the portion of the intermediate image corresponding to the screenshot area) and the second screenshot. Sometimes, for simplicity, both are collectively referred to as screenshot optical flow. Similar to the first intermediate optical flow, both the first and second screenshot optical flows have two directions. Figure 2 The optical flow used is from the intermediate image to the first crop and from the intermediate image to the second crop, denoted as F respectively. m→overlap-0 and F m→overlap-1 .
[0075] Step c: Upsample the optical flow of the first screenshot to the size of the first image to obtain the first intermediate optical flow, and upsample the optical flow of the second screenshot to the size of the second image to obtain the second intermediate optical flow.
[0076] Since the size of the optical flow in the first screenshot is the same as the size of the first screenshot itself, which is smaller than the size of the first image, upsampling of the optical flow in the first screenshot is required to obtain the first intermediate optical flow. Similarly, upsampling is also necessary for the optical flow in the second screenshot. Since optical flow can also be considered as a special image where each pixel value is a vector, upsampling can be performed using image interpolation algorithms, such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation. Alternatively, deep learning-based upsampling methods such as DUpsampling and Meta-Upscale can also be used.
[0077] Since the motion of a target in an image is consistent globally and locally in most cases, the target motion pattern reflected by the small-sized local optical flow (i.e., cropping optical flow) is the same as that reflected by the large-sized global optical flow (i.e., intermediate optical flow). Therefore, the validity of the new optical flow value generated by interpolation during upsampling can be guaranteed. In other words, for image regions in the first and second images that do not contain shared areas, the optical flow value calculated by upsampling is also relatively reliable. Furthermore, even if some optical flow values calculated by upsampling are not accurate enough, in some implementations with stitching masks (described later), the negative impact of inaccurate optical flow value calculation can be mitigated to some extent by changing the pixel values in the mask.
[0078] Regarding step S120, one issue needs clarification: what exactly does "intermediate viewpoint" in acquiring the intermediate image refer to? More precisely, the intermediate viewpoint refers to a specific expected viewpoint position between the corresponding viewpoint positions of the first and second images. Taking the case where the intermediate optical flow is estimated through an optical flow calculation network as an example, this expected viewpoint position is determined during the training of the optical flow calculation network. That is, the training data determines which position the optical flow calculation network should estimate the intermediate optical flow at (more precisely, it first estimates the cropped optical flow and then calculates the intermediate optical flow). The trained optical flow calculation network can then estimate the intermediate optical flow at that position, which is the viewpoint position for acquiring the intermediate image, i.e., the position of the intermediate viewpoint. For example, according to the following text... Figure 5 The explanation is that the "middle" in the intermediate perspective can be the "middle" in the sense of projection transformation, which is determined by the homography matrix, and does not refer to the exact center of the viewpoint position corresponding to the first image and the second image.
[0079] Step S130: Calculate the stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image.
[0080] Optionally, step S130 may further include the following sub-steps:
[0081] Step A: Based on the first intermediate optical flow and the first image, a first intermediate image is obtained by mapping; and based on the second intermediate optical flow and the second image, a second intermediate image is obtained by mapping.
[0082] The first and second intermediate images can be understood as parts of an intermediate image (the intermediate image is obtained by stitching the two together). Figure 2 The two are respectively denoted as I. m←0 and I m←1 As explained above, optical flow reflects the coordinate offset between the corresponding pixels of the same point on the target in two images. Therefore, given one of the two images and the optical flow between them, the other image can be estimated. This estimation method is called warping.
[0083] Specifically, the first intermediate image can be obtained by mapping the first image based on the first intermediate optical flow. Depending on the direction of the first intermediate optical flow, there are two different mapping methods. If the first intermediate optical flow is F... m→0 Then backward warping is used, and the mapping process can be represented as I m←0 =backward_warping(I0,F m→0 If the first intermediate optical flow is F; 0→m Then forward warping is used, and the mapping process can be represented as I0→m =forward_warping(I0,F 0→m In the following text, we will mainly use backward mapping, which is currently widely used, as an example for explanation. Similarly, by mapping the second image based on the second intermediate optical flow, we can obtain the second intermediate image. The mapping process can be represented as I... m←1 =backward_warping(I1,F m→1 ).
[0084] Step B: Calculate the target optical flow based on the first intermediate optical flow and the second intermediate optical flow.
[0085] Optionally, step B may further include the following sub-steps:
[0086] Step B1: Calculate at least one transition optical flow by interpolation based on the first intermediate optical flow and the second intermediate optical flow.
[0087] First, we introduce the concepts of transitional viewpoints and transitional images. Similar to intermediate images, transitional images are virtual images captured from transitional viewpoints (e.g., ...). Figure 3 Image I acquired from viewpoint v v The transitional viewpoint is the viewpoint between the acquisition viewpoints of the first image and the second image. The acquisition of the transitional image mentioned here should be understood as acquisition in a virtual sense.
[0088] Clearly, there are countless transitional perspectives in theory, for example, Figure 3 The image below shows four transitional viewpoints, named Viewpoint 0.2, Viewpoint 0.4, Viewpoint 0.6, and Viewpoint 0.8, each corresponding to a transitional image. The weight values 0.2, 0.4, 0.6, and 0.8 (the definition of weight values is explained later) roughly represent the positional relationship between the viewpoints. That is, starting from Viewpoint 0, the transition to Viewpoint 1 follows the order of "Viewpoint 0 → Viewpoint 0.2 → Viewpoint 0.4 → Viewpoint 0.6 → Viewpoint 0.8 → Viewpoint 1".
[0089] After defining the transition viewpoint and transition image, a virtual image acquisition process can be considered: at a certain initial viewpoint (e.g., Figure 3 The first image is captured from viewpoint 0, and then the camera is moved sequentially to each transitional viewpoint (e.g., ...). Figure 3 Transition images were acquired at viewing angles of 0.2, 0.4, 0.6, and 0.8, and finally, a terminating viewing angle (e.g., ...) was used. Figure 3 The second image was captured from perspective 1). This process can be visualized as follows: the photographer holds a mobile phone, moves around a target, and continuously takes pictures of the target from different angles during the movement.
[0090] Using the intermediate image as a reference, all acquired images are considered as the result of mapping between the intermediate image and optical flow. Therefore, each image acquired from a given viewpoint can be correlated with the intermediate image and the optical flow between them. For example, the first image corresponds to the first intermediate optical flow, the second image to the second intermediate optical flow, and the transition image to the transition optical flow. The transition optical flow is the optical flow between the intermediate image and the transition image. Figure 3 It can be denoted as F m→v (Of course, it could also be F) v→m Of course, since there is a correspondence between viewpoint and image, there is also a correspondence between viewpoint and optical flow.
[0091] Combining the virtual image acquisition process described above, this process can also be viewed as the process of the first image being transformed into various transitional images, and finally into the second image. During this image transition process, due to the correspondence between the acquired image and the optical flow, the first intermediate optical flow is also transformed into various transitional optical flows, and finally into the second intermediate optical flow. During this optical flow transition process, because the parallax between images acquired from adjacent viewpoints is small, the transition between different optical flows is relatively smooth, especially when many transitional viewpoints are selected.
[0092] Based on the above analysis, the transition from the first intermediate optical flow to the second intermediate optical flow is smooth. Therefore, the first and second intermediate optical flows can be regarded as two endpoints. The transition optical flow at any position between these two endpoints can be calculated by interpolation. The specific interpolation algorithm is not limited. For example, linear interpolation, quadratic interpolation, cubic interpolation, and inverse distance interpolation can be used. Linear interpolation will be used as an example later. It will not be elaborated here.
[0093] The specific locations and number of transition optical flows to be calculated can be determined based on actual needs, but at least one transition optical flow should be calculated. For example, in Figure 3 In the calculation, four transition optical flows were calculated, located at viewing angles of 0.2, 0.4, 0.6, and 0.8. However, it's important to note that the exact location of viewing angle 0.2 does not need to be determined before calculating the transition optical flow F. m→0.2 Instead of performing calculations, the transition optical flow F is directly calculated by interpolation using a weight value of 0.2 (the definition of the weight value is explained later). m→0.2 That's it, F m→0.2 The corresponding position is naturally the viewing angle 0.2. The same applies to viewing angles of 0.4, 0.6, and 0.8; the transition optical flow F can be interpolated and calculated according to the weight values of 0.4, 0.6, and 0.8 respectively. m→0.4 F m→0.6 and F m→0.8 .
[0094] Step B2: Based on the first intermediate optical flow, at least one transition optical flow, and the second intermediate optical flow, the target optical flow is obtained by fusing them.
[0095] As described in step B1, the first intermediate optical flow, at least one transitional optical flow, and the second intermediate optical flow are gradually transitioned. The "fusion" in step B2 refers to an optical flow merging operation that combines the first intermediate optical flow, at least one transitional optical flow, and the second intermediate optical flow into a single optical flow, called the target optical flow, ensuring that the target optical flow exhibits the gradual transition characteristics between the individual optical flows. Possible fusion operations include weighted summation and splicing; the following section will illustrate this using an example of optical flow fusion achieved through a weight matrix.
[0096] Furthermore, since the target optical flow contains information about the optical flow from various viewpoints and reflects the gradual change characteristics of the optical flow from each viewpoint, it can be considered that the target optical flow also corresponds to a special gradual viewpoint. The image acquired under this gradual viewpoint integrates information from the first image, the second image, and at least one transitional image, thus comprehensively reflecting the state of the acquired target from various viewpoints. That is, the image acquired under the gradual viewpoint is a relatively ideal stitching result, which is the stitched image between the first and second images to be calculated (referred to as the stitched image). Therefore, the target optical flow can also be regarded as the optical flow between the intermediate image and the stitched image, used for the calculation of the stitched image. A gradual viewpoint can be figuratively understood as follows: A photographer holds a mobile phone, starting from a starting viewpoint and ending at a ending viewpoint, moving around a target to take pictures. During the shooting process, the mobile phone continuously stitches together images acquired from different angles, without missing information from any viewpoint. The final image can reflect the entirety of the target between the starting and ending viewpoints.
[0097] In some implementations, step B2 can be divided into two sub-steps:
[0098] First, a preliminary optical flow is obtained by fusing the first intermediate optical flow, at least one transitional optical flow, and the second intermediate optical flow. Then, the preliminary optical flow is input into an optical flow inversion network to obtain the target optical flow output by the optical flow inversion network. Figure 2 These two sub-steps are shown. The reason for performing optical flow inversion is as follows:
[0099] If the first intermediate optical flow, at least one transitional optical flow, and the second intermediate optical flow all originate from the intermediate image, then the preliminary optical flow obtained by directly fusing them also originates from the intermediate image. Figure 2 The middle is denoted as F m→0~1 That is, the optical flow from the intermediate image to the stitched image. If F is directly used... m→0~1As the target optical flow, the mapping in step C (see the explanation of step C) can only be a forward mapping. However, forward mapping is not widely used due to some defects. Therefore, in the above implementation, F is obtained by inverting the optical flow. m→0~1 Transformed into a reverse optical flow F 0~1→m That is, the optical flow from stitching images to intermediate images, and F 0~1→m As the target optical flow, backward mapping can be used in step C. It should be understood that if a better forward mapping method exists, F can be directly mapped... m→0~1 It can also be used as a target optical flow.
[0100] It should be noted that optical flow inversion differs from simple matrix inversion; its computational process is more complex. Therefore, the above implementation uses a neural network for optical flow inversion. This simplifies the computation and improves the efficiency of optical flow inversion. Furthermore, the learning ability of the neural network can be leveraged to improve the accuracy of optical flow inversion. Improved accuracy obviously enhances the quality of the resulting stitched image. The training method for the optical flow inversion network will be discussed in the following section. Figure 5 This will be explained later. The specific structure of the optical flow inversion network is not limited. For example, in some simpler implementations, the optical flow inversion network can be constructed using L (L>1) consecutive convolutional layers. The first convolutional layer takes the initial optical flow as input, and the last convolutional layer outputs the target optical flow. For instance, if L=2 and the kernel size is 3×3, meaning the optical flow inversion network only includes two 3×3 convolutions, then the computation process of the optical flow inversion network can be represented as F... 0~1->m =conv(conv(F) m->0~1 ,3,3),3,3), where conv represents the convolution operation.
[0101] This simple network design is well-suited for stitching together images captured by wide-angle or fisheye cameras. Since wide-angle and fisheye cameras have a relatively large shooting range, the movement of the target in the image is relatively small. Therefore, the changes in the optical flow values of each optical flow (first intermediate optical flow, second intermediate optical flow, and transitional optical flow) are relatively gradual, and large changes in optical flow values are not likely to occur. The same applies to the initial optical flow obtained by fusion. Such initial optical flow is easy to invert, so there is no need to use a complex network structure. A simple network can actually improve the efficiency of optical flow inversion.
[0102] In some other implementations of step B (different from steps B1 and B2), the transition optical flow can be omitted. Instead, the target optical flow can be obtained directly by fusing the first intermediate optical flow and the second intermediate optical flow (similar to the above scheme, the preliminary optical flow can be obtained by fusing first, and then the preliminary optical flow can be directly used as the target optical flow or the target optical flow can be calculated using an optical flow inversion network). In this case, although the gradient characteristics of the target optical flow are not as good as the above scheme (referring to steps B1 and B2), the calculation is simpler.
[0103] Based on the above analysis, the target optical flow at this point can be considered to correspond to a degraded, gradually changing viewpoint (directly transitioning from the viewpoint of the first image to the viewpoint of the second image). The image acquired under this gradually changing viewpoint fuses information from both the first and second images. Although it does not fuse information from the transition image, it already contains all the original information (first and second images) used for image stitching, and it can also reflect the state of the acquired target under different viewpoints, thus making it a relatively ideal stitched image. This stitched image can also be calculated using the target optical flow and the intermediate image. Step C: Based on the target optical flow and the first intermediate image, map to obtain the first stitched image, and based on the target optical flow and the second intermediate image, map to obtain the second stitched image.
[0104] The first and second stitched images can be understood as parts of the stitched image to be calculated in step D (the stitched image is obtained by stitching the two together). Figure 2 The two are respectively denoted as I. 0~1←m←0 and I 0~1←m←1 I 0~1←m←0 The subscripts 0~1←m←0 in the image are abbreviations of (0~1)←(m←0), indicating that the first stitched image I is used. m←0 and target optical flow F 0~1→m The result of the backward mapping, I 0~1←m←1 The subscripts 0~1←m←1 in the image are abbreviations of (0~1)←(m←1), indicating that the second stitched image I is used. m←1 and target optical flow F 0~1→m The result of backward mapping. Forward mapping can be represented similarly, and will not be elaborated further. As mentioned in step B, the target optical flow can be considered as the optical flow between the intermediate image and the stitched image, making such mapping feasible.
[0105] Step D: Based on the first stitched image and the second stitched image, stitch together the first image and the second image to obtain a stitched image.
[0106] Step C has already yielded the first and second stitched images; stitching them together will produce the final stitched image. For example, in... Figure 2The stitched image is obtained by stitching the left image region of the first stitched image and the right image region of the second stitched image together, denoted as I. 0~1 .
[0107] If the accuracy of the target optical flow is high enough, the calculated first and second stitched images are already aligned, and they can be directly superimposed to obtain the stitched image. However, considering that the accuracy of the target optical flow calculation is affected by many factors (e.g., the prediction accuracy of the optical flow calculation network, the upsampling accuracy of the optical flow, etc.), its accuracy may not be high enough. Consequently, the first and second stitched images calculated based on the target optical flow may not be well aligned. Therefore, in some implementations, a stitching mask can be set to achieve a smooth transition between the first and second stitched images at the stitching point, thereby improving the quality of the stitched image. The specific steps are as follows:
[0108] First, the first and second stitched images are input into the mask computing network to obtain the stitching mask output by the mask computing network. Then, based on the stitching mask, the first and second stitched images are stitched together to obtain the stitched image. Figure 2 The steps for image stitching using a mask are shown, where the stitching mask is denoted as mask.
[0109] The mask computation network is a pre-trained neural network, the specific structure of which is not limited, and its training method will be introduced later. Figure 5 This will be explained later. The input to the mask computation network includes at least the first and second stitched images, but may also include other information, such as target optical flow. Utilizing the learning ability of neural networks, the mask can be predicted relatively accurately, which helps improve the quality of the stitched images.
[0110] The stitching mask can also be considered as an image, with the same size as the first stitched image (or the second stitched image). Optionally, the pixel values in the stitching mask can be values between [0,1], and the specific values are calculated by the mask calculation network. For pixel position (x,y), the pixel value p1 of the first stitched image at that position, the pixel value p2 of the second stitched image at that position, and the pixel value m of the stitching mask at that position are obtained. Then, the pixel value p of the stitched image at (x,y) can be calculated by weighting p according to the formula p = m × p1 + (1-m) × p2, where m represents the weighting coefficient of p1. Of course, if m represents the weighting coefficient of p2, the formula can be changed to p = (1-m) × p1 + m × p2, which is not essentially different from the previous formula.
[0111] For example, if the pixel values in the stitching mask represent weighting coefficients for the pixel values in the first stitched image, then a possible stitching mask would be as follows:
[0112]
[0113] The values in the left and right columns of the stitching mask are 1 and 0, respectively, while the values in the middle two columns are 0.5. This means that the left two columns of the stitched image take the pixel values from the first stitched image, the right two columns take the pixel values from the second stitched image, and the middle two columns take the average of the pixel values from the two stitched images. This value selection is reasonable, as the middle two columns are likely located in image areas where the two stitched images contain common elements. By taking the average, the first and second stitched images can transition smoothly in these areas.
[0114] Note that this is an example, so the splicing mask is only 6x6 in size and is not suitable for general use. Figure 2 Image stitching. But Figure 2 The actual stitching mask shown is very similar to this example. White on the left represents a pixel value of 1, black on the right represents a pixel value of 0, and gray in the middle represents a pixel value within the range (0,1).
[0115] It is understandable that a mask is not always necessary to stitch the first and second images together. For example, the two images can be superimposed first, and then smoothing filters can be used to improve the abrupt transitions in the image.
[0116] In some other implementations of step S130 (different from steps A to D), the first intermediate image and the second intermediate image can be calculated first (similar to step A); then the first intermediate image and the second intermediate image can be stitched together (directly stitched or stitched using a mask) to obtain an intermediate image; then the target optical flow can be calculated based on the first intermediate optical flow and the second intermediate optical flow (similar to step B); finally, the stitched image can be obtained by mapping based on the intermediate image and the target optical flow. Details of each step can be found in steps A to D and will not be repeated here.
[0117] In summary, the image stitching method provided in this application has a relatively simple calculation process. Unlike some traditional image stitching methods, it does not require complex iterations to operate on multiple homography matrices, thus making image stitching more real-time and enhancing its practicality.
[0118] In some implementations of this method, image stitching is achieved by calculating the target optical flow. Since the target optical flow is generated by fusing a first intermediate optical flow, at least one transitional optical flow, and a second intermediate optical flow (it may not include the transitional optical flow, but can be analyzed similarly), it contains optical flow information from various viewpoints. Considering the correspondence between the image and the optical flow, the stitched image calculated based on the target optical flow also incorporates information from the first image, the second image, and at least one transitional image. This allows for a comprehensive reflection of the state of the acquired target from various viewpoints between the first and second images, resulting in a relatively ideal stitching result. The stitched image obtained by this method is of high quality, improving upon the artifacts, distortion, and difficulty in aligning the images to be stitched that exist in traditional image stitching methods. It can also effectively stitch images even when there is a large parallax between the first and second images.
[0119] It should be understood that if two or more images need to be stitched together, the above image stitching method can be applied consecutively. For example, to stitch together the first image, the second image, and the third image, the method can be used first to stitch the first and second images to obtain an intermediate stitching result, and then the method can be used to further stitch the intermediate stitching result with the third image to obtain the final stitched image.
[0120] Below, based on the above embodiments, we will continue to introduce the method of calculating the transition optical flow by linear interpolation in step B1.
[0121] Given two endpoints (first intermediate optical flow and second intermediate optical flow), linear interpolation is essentially a weighted summation operation. The specific steps can be as follows:
[0122] First, obtain at least one weight value; then, based on each weight value, perform a weighted summation of the first intermediate optical flow and the second intermediate optical flow to obtain at least one transition optical flow.
[0123] The number of weight values is the same as the number of transition optical flows. For example, if there are 4 weight values, then 4 transition optical flows are calculated by weighted summation. The specific number can be determined according to the requirements. The weight values can be preset (e.g., written in a configuration file or program) and directly read and used when calculating the transition optical flows. Alternatively, the weight values can be generated by a certain algorithm when calculating the transition optical flows.
[0124] When performing weighted summation, two weighting coefficients are required: the weighting coefficient of the first intermediate optical flow and the weighting coefficient of the second intermediate optical flow. These two weighting coefficients are mutually restrictive; once one weighting coefficient is known, the other weighting coefficient can be calculated accordingly.
[0125] For example, the weighting coefficients are limited to the interval (0,1), and the sum of the two weighting coefficients is 1. In this case, the weight value w within the interval (0,1) can be used as the weighting coefficient for one of the intermediate optical flows, for example, the weighting coefficient for the first intermediate optical flow. Then the weighting coefficient for the second intermediate optical flow is 1-w. The formula for calculating the transition optical flow can be expressed as:
[0126] F m->v =w×F m->0 +(1-w)×F m->1
[0127] Understandably, if the weight value w is used as the weighting coefficient for the second intermediate optical flow, then the weighting coefficient for the first intermediate optical flow is 1-w. The formula for calculating the transition optical flow can be expressed as:
[0128] F m->v = (1-w)×F m->0 +w×F m->1
[0129] The two formulas above are not substantially different. Linear interpolation has the advantage of being computationally simple, and in most cases (especially when the time interval between the acquisition of the first and second images is not too long), linear motion is sufficient to describe the motion of the target between the first and second images, and the accuracy of the transition optical flow calculated by linear interpolation is also high enough. Of course, as mentioned above, other nonlinear interpolation methods can also be used.
[0130] The following further explains some principles for setting weight values:
[0131] As a principle, the weight value can be set to be related to the viewing angle position of the transition optical flow corresponding to that weight value. This setting can ensure that the transition optical flow calculated using the weight value is consistent with its viewing angle position, thereby ensuring that the target optical flow calculated using the transition optical flow can reflect the characteristics of optical flow gradient.
[0132] For example, assuming that the sum of the weighting coefficients of the first intermediate optical flow and the second intermediate optical flow is 1, and the weight value is the weighting coefficient of the first intermediate optical flow, the weight value can be set to be positively correlated with the proximity of the viewpoint position of the transition optical flow corresponding to the weight value and the viewpoint position of the first intermediate optical flow (or, negatively correlated with the proximity of the viewpoint position of the transition optical flow corresponding to the weight value and the viewpoint position of the second intermediate optical flow, which are equivalent).
[0133] Specifically, the closer the viewing angle of the transition optical flow is to the viewing angle of the first intermediate optical flow, the larger the weight value is set. That is, the weighting coefficient of the first intermediate optical flow is increased while the weighting coefficient of the second intermediate optical flow is decreased. In this way, the value of the transition optical flow will be more affected by the first intermediate optical flow and less affected by the second intermediate optical flow, which is consistent with its viewing angle. Similarly, the farther the viewing angle of the transition optical flow is from the viewing angle of the first intermediate optical flow, the smaller the weight value is set. That is, the weighting coefficient of the first intermediate optical flow is decreased while the weighting coefficient of the second intermediate optical flow is increased. In this way, the value of the transition optical flow will be more affected by the second intermediate optical flow and less affected by the first intermediate optical flow, which is consistent with its viewing angle.
[0134] For example, in Figure 3 In the calculation, the proximity of viewing angles 0.2, 0.4, 0.6, and 0.8 to viewing angle 0 gradually decreases, so the transition optical flow F is calculated at viewing angle 0.2. m->0.2 The transition optical flow F is calculated with a weight value w = 0.8 and a viewing angle of 0.4. m->0.4 The transition optical flow F is calculated with a weight value w = 0.6 and a viewing angle of 0.6. m->0.6 The transition optical flow F is calculated with a weight value w = 0.4 and a viewing angle of 0.8. m->0.8 The weight value w = 0.2, and the corresponding transition optical flow can be calculated according to the following formula:
[0135] F m->0.2 =0.8×F m->0 +0.2×F m->1
[0136] F m->0.4 =0.6×F m->0 +0.4×F m->1
[0137] F m->0.6 =0.4×F m->0 +0.6×F m->1
[0138] F m->0.8 =0.2×F m->0 +0.8×F m->1
[0139] Imagine several transition optical flows distributed between the first intermediate optical flow and the second intermediate optical flow. If the weight values corresponding to all transition optical flows are set according to the above rule (referring to the positive correlation rule), then from the first intermediate optical flow to the second intermediate optical flow, the influence of the first intermediate optical flow on each transition optical flow will decrease from strong to weak, and the influence of the second intermediate optical flow will increase from weak to strong. This ensures that the calculated transition optical flows are gradual.
[0140] It should be noted that setting the weight value to be related to the viewing angle position of the transition optical flow corresponding to that weight value does not mean that the viewing angle position of the transition optical flow must be calculated precisely before determining the corresponding weight value. It simply means that the viewing angle position of the transition optical flow should be considered when setting the weight value. For example, when setting the weight value w = 0.8, it is not necessary to quantitatively calculate the position of viewing angle 0.2. It is only necessary to know that viewing angle 0.2 is closer to viewing angle 0 than viewing angles 0.4, 0.6, and 0.8, and that the angle between viewing angle 0.2 and viewing angle 0 is approximately 20% of the angle between viewing angle 0 and viewing angle 1. The weight value w = 1 - 20% = 0.8.
[0141] As another principle, at least one weight value can be set to be uniformly distributed within the interval (0,1). For example, if there is only one weight value, it can be 0.5, and the interval between this weight value and both 0 and 1 is 0.5, which is a uniform distribution. If there are M (M>1) weight values, these weight values can be i / (M+1) (i is an integer between 1 and M), the interval between any two weight values is 1 / (M+1), and the interval between the first weight value and 0, and the interval between the Mth weight value and 1 are also 1 / (M+1), which is a uniform distribution. For example, Figure 3 This refers to the case where M=4.
[0142] In these implementations, since the weight values are uniformly distributed within the interval (0,1), the viewpoint positions of the transition optical flow calculated using these weight values are also uniformly distributed between the first and second images, which allows the acquired transition images to fully describe the overall picture of the acquired target between the corresponding viewpoints in the first and second images (without favoring certain viewpoints). Since the stitched image can be regarded as a fusion of information from all transition images, the resulting stitched image has high quality.
[0143] Of course, it is not mandatory for the weight values to be evenly distributed in the interval (0,1). For example, the weight values can be set to be more densely distributed in some intervals of (0,1) and more sparsely distributed in other intervals.
[0144] It should be understood that the above two principles for setting weight values can also be used in combination.
[0145] Below, based on the above embodiments, we will continue to describe the method for fusing the target optical flow according to the first intermediate optical flow, at least one transition optical flow and the second intermediate optical flow in step B2.
[0146] In some implementations, optical flow fusion can be achieved using a weight matrix. The specific steps are as follows:
[0147] First, obtain N+2 weight matrices, where N is the total number of transition optical flows (N≥1); then, based on the N+2 weight matrices, perform a weighted summation on the first intermediate optical flow, the N transition optical flows, and the second intermediate optical flow to obtain the preliminary optical flow; finally, depending on the implementation method, the preliminary optical flow can be directly used as the target optical flow or inverted and used as the target optical flow.
[0148] Taking the case of inverting optical flow as an example, the formula can be used. This represents the initial optical flow calculation process. Where F... t The optical flow to be fused can be a first intermediate optical flow, a transitional optical flow, or a second intermediate optical flow, W. t F represents t The corresponding weight matrix, i.e., the t-th weight matrix, where × denotes matrix multiplication. Each element in the weight matrix can be considered as a weight value obtained by weighted summation. Optionally, the elements in the weight matrix take values within the interval [0,1], and the elements in each weight matrix satisfy the following relationship: W t (i,j) represents the element in the i-th row and j-th column of the t-th weight matrix.
[0149] The transition optical flow is weighted and summed using a weight matrix. Unlike weighted summation using one-dimensional weight values (e.g., calculating the transition optical flow), the weight matrix is two-dimensional. This allows for more flexible combination of information from different optical flows to be fused in the initial optical flow, enabling the initial optical flow to reflect the gradual change from the first intermediate optical flow to the second intermediate optical flow. Furthermore, it allows the target optical flow calculated by inverting the optical flow to also reflect this gradual change.
[0150] For example, the elements in the weight matrix can be set according to the following rule: the position of the maximum value of the element in the weight matrix is related to the viewpoint position of the optical flow corresponding to the weight matrix.
[0151] For example, the "correlation" mentioned above could mean that the position of the maximum value of an element in the weight matrix is consistent with the position of the viewpoint of the optical flow to be fused within the entire viewpoint range (referring to the region between the viewpoint of the first image and the viewpoint of the second image).
[0152] Combination Figure 3 To explain the meaning of "consistency," and for simplicity, let's assume F... t The matrix has only 6 columns:
[0153] F1( Figure 3 F in m->0The first intermediate optical flow is angle 0, which is located at the leftmost edge of the entire angle range (from angle 0 to angle 1). Therefore, W1, set according to the above rules, takes the maximum value in the leftmost column, and the values in the other columns can be gradually decreased, or set in other ways. For example, the following two W1 values both satisfy the above rules:
[0154]
[0155]
[0156] F2( Figure 3 F in m->0.2 The transition optical flow (F1) corresponds to the viewing angle 0.2, which is located approximately 20% from the left of the entire viewing angle range. Therefore, according to the rules described above, W2 has its maximum value in the second column from the left, while the values in the remaining columns can be gradually decreased or set in other ways. For example, the following two W2 values both satisfy the above rules:
[0157]
[0158]
[0159] For F3, F4, F5, F6 ( Figure 3 F in m->0.4 F m->0.6 F m->0.8 F m->1 The settings for the corresponding weight matrices W3, W4, W5, and W6 are similar and will not be elaborated further.
[0160] Optionally, the maximum value of each element in the weight matrices can be kept uniform. For example, in W1 and W2 above, the maximum value of each element could be either 1 or 0.8. The above example can be easily generalized to cases where optical flow contains more columns, in which case the element with the maximum value in the weight matrix might be from one or more columns. For example, for... Figure 3 In the scene, W2 (where the optical flow is no longer limited to 6 columns) is the maximum value of one or several nearby columns at the position where the total number of columns in the matrix is multiplied by 20%.
[0161] If the position of the maximum value of an element in a weight matrix coincides with the viewpoint position of the corresponding optical flow to be fused (which is also a matrix), then after weighted calculation using this weight matrix, the optical flow values corresponding to their viewpoint positions in the optical flow matrix of the optical flow to be fused will contribute the most to the calculation of the optical flow value at the same position in the initial optical flow, while the optical flow values corresponding to other positions will contribute relatively less. If all weight matrices are set according to the above rules and the maximum values of the elements in each weight matrix are consistent, the optical flow value corresponding to each viewpoint position in the optical flow matrix of the initial optical flow can be mainly contributed by the optical flow to be fused corresponding to that viewpoint position. This allows the initial optical flow to reflect the gradual change from the first intermediate optical flow to the second intermediate optical flow, and further allows the target optical flow calculated by optical flow inversion to also reflect this gradual change.
[0162] For example, still considering the case where the optical flow to be fused has only 6 columns, W t (t is an integer from 1 to 6) is set so that all elements in column t are 1, and all other elements are 0. Then, after weighted summation, the initial optical flow F is obtained. m->0~1 In the diagram, column t comes from F. t Or rather, each F t Contributed F m->0~1 The column of element values corresponding to its viewpoint position, therefore F m->0~1 This is sufficient to reflect the gradual change in optical flow from F1 to F6. In contrast, if a one-dimensional weight value (let's assume it's denoted as w) is used... t The weighted summation of the fused optical flow is performed, since w t Acting indiscriminately on F t Each optical flow value in the table cannot accurately reflect the F-value. t Optical flow values at different locations in F m->0~1 The difference in optical flow values between the two values in the calculation of their contribution will lead to the difference in the obtained F. m->0~1 It cannot accurately reflect the gradual change in optical flow from F1 to F6.
[0163] It should be noted that the "correlation" mentioned above can also be defined in other ways. The position of the maximum value of an element in the weight matrix and the position of the corresponding optical flow to be fused within the entire field of view may not be strictly consistent, but only roughly consistent, and so on.
[0164] Consider the following scenario: the image stitching method described above is based on a neural network model called the image stitching model. This model includes three sub-networks: an optical flow calculation network, an optical flow inversion network, and a mask calculation network. The optical flow calculation network estimates the optical flow of the first and second screenshots, and then calculates the first and second intermediate optical flows (see step b). The optical flow inversion network calculates the inverse optical flow of the initial optical flow to obtain the target optical flow (see step B2). The mask calculation network calculates the stitching mask, and then calculates the stitched image (see step D).
[0165] The functions of the three sub-networks have been described in detail above and will not be repeated here. Below, based on the above embodiments, the training method for the image stitching model is introduced. The possible process is as follows: Figure 5 As shown. The model training method can be, but is not limited to, by... Figure 8 The electronic device shown performs the following instructions; for details regarding the structure of this electronic device, please refer to the section on... Figure 8 The explanation. (Refer to...) Figure 5 The method includes:
[0166] Step S210: Calculate the stitched image of the first image and the second image using the image stitching model.
[0167] Step S220: Obtain the first real screenshot optical flow, the second real screenshot optical flow, the real target optical flow, and the real stitched image.
[0168] Step S230: Calculate the optical flow prediction loss based on the optical flow of the first screenshot, the optical flow of the second screenshot, the optical flow of the first real screenshot, and the optical flow of the second real screenshot.
[0169] Step S240: Calculate the optical flow inversion loss based on the target optical flow and the actual target optical flow.
[0170] Step S250: Calculate the image stitching loss based on the stitched image and the real stitched image.
[0171] Step S260: Calculate the total loss based on the optical flow prediction loss, optical flow inversion loss, and image stitching loss, and update the parameters of the optical flow calculation network, optical flow inversion network, and mask calculation network based on the total loss.
[0172] Step S210 can be implemented using the image stitching method provided in the embodiments of this application (the corresponding step requires the use of the three sub-networks in the image stitching model), and will not be described again. However, it should be noted that, due to Figure 8 This is a training method applied during the model training phase. Therefore, the first and second images in step S210 are training images, while step S110 does not specify whether the first and second images are training images or inference images.
[0173] In step S220, the first real screenshot optical flow, the second real screenshot optical flow, the real target optical flow, and the real stitched image are all supervision signals. Specifically, the first and second real screenshot optical flows are used in step S230 to calculate the optical flow prediction loss, thereby supervising the training of the optical flow computation network; the real target optical flow is used in step S240 to calculate the optical flow inversion loss, thereby supervising the training of the optical flow inversion network; and the real stitched image is used in step S250 to calculate the image stitching loss, thereby supervising the training of the mask computation network. Of course, the losses calculated for later sub-networks in the image stitching model may also have a supervisory effect on earlier sub-networks in the image stitching model. For example, the optical flow inversion loss may also supervise the optical flow computation network, and the image stitching loss may also supervise both the optical flow computation network and the optical flow inversion network.
[0174] The execution order of steps S220 to S250 is relatively flexible, as explained below:
[0175] The three monitoring signals in step S220 can be acquired together or separately. "Acquiring" here includes both calculation and direct reading. For separate acquisition, the corresponding loss can be calculated immediately after acquiring a monitoring signal. For example, after acquiring the optical flow of the first and second real-image screenshots, step S230 can be executed to calculate the optical flow prediction loss (provided that the optical flow of the first and second screenshots has already been calculated). It is not necessary to wait until all three monitoring signals have been acquired before executing step S230.
[0176] The timing of step S220 is not necessarily related to that of step S210; it can be executed before or after step S210. Figure 8 (As shown in the example), it can also be executed in parallel with step S210.
[0177] The timing of calculating the three losses in steps S230 to S250 is not limited. For example, the three losses can be calculated after step S210 is completed. At this time, the optical flow of the first screenshot, the optical flow of the second screenshot, the target optical flow, and the stitched image are all calculated during the execution of step S210, so the loss can be calculated, and the order of calculation of the three losses is irrelevant. Alternatively, the losses can be calculated during step S210. For example, after calculating the optical flow of the first and second screenshots, step S230 can be executed (provided that the optical flow of the first and second real screenshots has been obtained), without having to wait until the stitched image is calculated.
[0178] In step S260, the total loss is calculated based on the optical flow prediction loss, optical flow inversion loss, and image stitching loss. This can be achieved using methods such as weighted summation or product calculation. Taking weighted summation as an example, the total loss can be expressed as: Loss = α × Loss1 + β × Loss2 + γ × Loss3. Here, α, β, and γ are weighting coefficients, and Loss1, Loss2, and Loss3 are the optical flow prediction loss, optical flow inversion loss, and image stitching loss, respectively. Backpropagation based on the total loss allows for parameter updates to the image stitching model.
[0179] The above model training method considers the losses corresponding to the three sub-networks simultaneously: optical flow prediction loss, optical flow inversion loss, and image stitching loss. That is, by training, the accuracy of intermediate optical flow prediction, target optical flow prediction, and stitching mask prediction of the model are improved at the same time, so that the final image stitching model can achieve high-quality image stitching.
[0180] It should be understood that in some implementations, even if the image stitching model includes the three sub-networks mentioned above, only one or two losses need to be calculated, instead of all three. For example, since the image stitching loss can also serve as a supervisory function for the optical flow calculation network and the optical flow inversion network, only the image stitching loss needs to be calculated. Obviously, if some losses are not calculated, the corresponding supervisory signals do not need to be acquired.
[0181] It should also be understood that some steps in image stitching methods may not be implemented using neural networks, and therefore the image stitching model may not necessarily include the three sub-networks mentioned above. For example, optical flow inversion can also be performed using other methods, in which case the image stitching model will not include the optical flow inversion network, and the corresponding optical flow inversion loss will naturally not need to be calculated.
[0182] In some implementations, a real image can be used as an intermediate image to generate the first and second images in step S210. The real image referred to here can be a real-world image or an image generated by a computer vision algorithm. In short, the intermediate image is a known, actually existing image, which differs from the previous explanation of the intermediate image (previously, when introducing step S120, the intermediate image was considered a virtual image). The method for generating a set of first and second images is as follows:
[0183] The first image is calculated based on the intermediate image and a certain homography matrix, and the second image is calculated based on the intermediate image and the inverse of the homography matrix. Figure 6 This illustrates the image generation process, where h represents a homography matrix, h -1 Representing the inverse matrix of h, based on the intermediate image, using h and h respectively. -1 By performing a projection transformation, we can obtain the first image and the second image.
[0184] Clearly, for a given intermediate image, simply changing h yields a set of different first and second images. Thus, this method can "create" a large number of training images from a small number of real images. The homography matrix h can be pre-set or generated temporarily by the algorithm.
[0185] Furthermore, the three monitoring signals in step S220 can be generated in the following way:
[0186] (1) Generation of optical flow from real screenshots
[0187] Calculate the first true intermediate optical flow based on h, and calculate the first true screenshot optical flow based on the first true intermediate optical flow, and, based on h -1 Calculate the second true intermediate optical flow, and then calculate the second true screenshot optical flow based on the second true intermediate optical flow. The true screenshot optical flow should be understood as the screenshot optical flow estimation result under ideal conditions.
[0188] The projection transformation represented by the homography matrix gives the correspondence between pixels of two images. Therefore, given the homography matrix, it is easy to calculate the optical flow between two images.
[0189] Therefore, using the homography matrix h between the intermediate image and the first image, the optical flow between them can be calculated, called the first true intermediate optical flow. After obtaining the first true intermediate optical flow, the first true crop optical flow can be obtained by taking a screenshot (cropping the image area containing the common scene in the first and second images). Similarly, using the homography matrix h between the intermediate image and the second image... -1 The optical flow between the two can be calculated, which is called the second true intermediate optical flow. After obtaining the second true intermediate optical flow, the second true screenshot optical flow can be obtained by taking a screenshot.
[0190] (2) Generation of optical flow of real target
[0191] According to h and h -1 Interpolation is used to calculate at least one transition matrix, based on h, at least one transition matrix, and h. -1 The target matrix is obtained by fusing the two matrices, and the true target optical flow is calculated based on the target matrix. The true target optical flow should be understood as the inverse result of the preliminary optical flow under ideal conditions.
[0192] The interpolation method used here is similar to that used in calculating the transition optical flow (e.g., weighted summation using weight values). The fusion method used here is similar to that used in calculating the target optical flow (e.g., weighted summation using weight matrices). Although the previous calculation methods targeted optical flow, mathematically, the homography matrix and optical flow are both matrices and have no fundamental difference. Therefore, the previous methods for calculating optical flow can be applied here. It should be noted that here, h, at least one transition matrix, and h are fused. -1 After that, the target matrix can be obtained directly without calculating the inverse matrix (the target optical flow is obtained by inverting the initial optical flow). This is because for the fusion of homography matrices, inversion needs to be performed on the matrix itself before fusion, but inverting h is simply h. -1 , for h -1 Finding the inverse is simply h, just with the two swapped, so the steps for finding the inverse of a matrix can be omitted.
[0193] Since the target matrix can also be regarded as the homography matrix between the intermediate image and the real stitched image (e.g. Figure 6 As shown in the figure, the optical flow between the two can be calculated based on the target matrix, which is the real target optical flow.
[0194] Additionally, if a certain alternative solution does not require calculating the transition optical flow, then the transition matrix above also does not need to be calculated; it can be directly calculated based on h and h. -1 Calculating the target matrix is sufficient.
[0195] (3) Real spliced images
[0196] The true stitched image can be calculated from the intermediate image and the target matrix, similar to the calculation of the first and second images, and will not be repeated here. The true stitched image should be understood as the ideal result of stitching the first and second images together.
[0197] In the above implementation, if a set of images to be stitched (including the first image and the second image) and their corresponding supervision signals are regarded as a training sample, since the homography matrix can be arbitrarily specified, this implementation can quickly generate a large number of training samples using a small number of real images (which can be located in the training set). Furthermore, by selecting different homography matrices, these samples can cover different scenes, thereby enabling the trained image stitching model to have good generalization ability.
[0198] Having introduced the generation of training data, we can now provide a more precise interpretation of the term "middle" in the intermediate image. Here, "middle" can refer to the "middle" in the sense of projection transformation, as determined by the homography matrix. That is, the image captured at this "middle" position can be transformed into the first image using a certain homography matrix, and the second image can be transformed using the inverse of that homography matrix. Understandably, if other methods are used to generate the training data, the definition of "middle" will change accordingly.
[0199] Figure 7 The structure of the image stitching apparatus 300 provided in an embodiment of this application is shown. (Refer to...) Figure 7 The image stitching device 300 includes:
[0200] Image acquisition module 310 is used to acquire a first image and a second image;
[0201] The intermediate optical flow calculation module 320 is used to calculate a first intermediate optical flow and a second intermediate optical flow based on the first image and the second image. The first intermediate optical flow is the optical flow between the intermediate image and the first image, and the second intermediate optical flow is the optical flow between the intermediate image and the second image. The intermediate image is an image with a viewing angle between the first image and the second image. The size of the first intermediate optical flow is the same as the size of the first image, and the size of the second intermediate optical flow is the same as the size of the second image.
[0202] The image stitching module 330 is used to calculate a stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image and the second image.
[0203] In one implementation of the image stitching device 300, the intermediate optical flow calculation module 320 calculates a first intermediate optical flow and a second intermediate optical flow based on the first image and the second image, including: respectively cropping image regions containing common scenes in the first image and the second image to obtain a first screenshot and a second screenshot; inputting the first screenshot and the second screenshot into an optical flow calculation network to obtain a first screenshot optical flow and a second screenshot optical flow, wherein the first screenshot optical flow is the optical flow between the intermediate image and the first screenshot, and the second screenshot optical flow is the optical flow between the intermediate image and the second screenshot; upsampling the first screenshot optical flow to the size of the first image to obtain the first intermediate optical flow, and upsampling the second screenshot optical flow to the size of the second image to obtain the second intermediate optical flow.
[0204] In one implementation of the image stitching device 300, the image stitching module 330 calculates a stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image, including: mapping a first intermediate image based on the first intermediate optical flow and the first image, and mapping a second intermediate image based on the second intermediate optical flow and the second image; calculating a target optical flow based on the first intermediate optical flow and the second intermediate optical flow; mapping a first stitched image based on the target optical flow and the first intermediate image, and mapping a second stitched image based on the target optical flow and the second intermediate image; and stitching the first stitched image and the second stitched image together to obtain the stitched image.
[0205] In one implementation of the image stitching device 300, the image stitching module 330 calculates the target optical flow based on the first intermediate optical flow and the second intermediate optical flow, including: interpolating and calculating at least one transition optical flow based on the first intermediate optical flow and the second intermediate optical flow; and fusing the target optical flow based on the first intermediate optical flow, the at least one transition optical flow, and the second intermediate optical flow.
[0206] In one implementation of the image stitching device 300, the image stitching module 330 interpolates and calculates at least one transition optical flow based on the first intermediate optical flow and the second intermediate optical flow, including: obtaining at least one weight value; and performing a weighted summation of the first intermediate optical flow and the second intermediate optical flow based on each weight value to obtain the at least one transition optical flow.
[0207] In one implementation of the image stitching device 300, the magnitude of the weight value is related to the viewing angle position of the transition optical flow corresponding to the weight value.
[0208] In one implementation of the image stitching device 300, the sum of the weighting coefficients of the first intermediate optical flow and the second intermediate optical flow is 1, the weight value is the weighting coefficient of the first intermediate optical flow, and the magnitude of the weight value is positively correlated with the proximity between the viewing angle position of the transition optical flow corresponding to the weight value and the viewing angle position of the first intermediate optical flow.
[0209] In one implementation of the image stitching device 300, the at least one weight value is uniformly distributed within the interval (0,1).
[0210] In one implementation of the image stitching device 300, the image stitching module 330 calculates the target optical flow based on the first intermediate optical flow and the second intermediate optical flow, including: calculating a preliminary optical flow based on the first intermediate optical flow and the second intermediate optical flow; and inputting the preliminary optical flow into an optical flow inversion network to obtain the target optical flow.
[0211] In one implementation of the image stitching device 300, the image stitching module 330 fuses the first intermediate optical flow, the at least one transition optical flow, and the second intermediate optical flow to obtain the target optical flow, including: obtaining N+2 weight matrices, where N is the total number of transition optical flows; performing a weighted summation on the first intermediate optical flow, the N transition optical flows, and the second intermediate optical flow based on the N+2 weight matrices to obtain the preliminary optical flow; the preliminary optical flow is the target optical flow, or the preliminary optical flow is input into an optical flow inversion network to obtain the target optical flow.
[0212] In one implementation of the image stitching device 300, the position of the maximum value of the element in the weight matrix is related to the viewing angle position of the optical flow corresponding to the weight matrix.
[0213] In one implementation of the image stitching device 300, the image stitching module 330 stitches together the first stitched image and the second stitched image to obtain the stitched image, including: inputting the first stitched image and the second stitched image into a mask computing network to obtain a stitching mask; and stitching the first stitched image and the second stitched image together based on the stitching mask to obtain the stitched image.
[0214] In one implementation of the image stitching device 300, the first intermediate optical flow and the second intermediate optical flow are calculated using the first and second screenshot optical flows output by an optical flow calculation network; the target optical flow is calculated using an optical flow inversion network; and the stitched image is calculated using a stitching mask output by a mask calculation network. The device further includes:
[0215] The supervision signal acquisition module is used to acquire the optical flow of the first real screenshot, the optical flow of the second real screenshot, the optical flow of the real target, and the real stitched image;
[0216] The loss calculation module is used to calculate optical flow prediction loss based on the first screenshot optical flow, the second screenshot optical flow, the first real screenshot optical flow, and the second real screenshot optical flow; and to calculate optical flow inversion loss based on the target optical flow and the real target optical flow; and to calculate image stitching loss based on the stitched image and the real stitched image.
[0217] The parameter update module is used to calculate the total loss based on the optical flow prediction loss, the optical flow inversion loss, and the image stitching loss, and update the parameters of the optical flow calculation network, the optical flow inversion network, and the mask calculation network based on the total loss.
[0218] In one implementation of the image stitching device 300, the target optical flow is obtained by fusing the first intermediate optical flow, at least one transition optical flow, and the second intermediate optical flow. The image acquisition module 310 acquires the first image and the second image, including: calculating the first image based on the intermediate image and a specified homography matrix, and calculating the second image based on the intermediate image and the inverse matrix of the homography matrix, wherein the intermediate image is a real image.
[0219] The monitoring signal acquisition module acquires a first real screenshot optical flow, a second real screenshot optical flow, a real target optical flow, and a real stitched image, including: calculating a first real intermediate optical flow based on the homography matrix, and calculating the first real screenshot optical flow based on the first real intermediate optical flow; calculating a second real intermediate optical flow based on the inverse of the homography matrix, and calculating the second real screenshot optical flow based on the second real intermediate optical flow; interpolating at least one transition matrix based on the homography matrix and the inverse of the homography matrix, fusing the homography matrix, the at least one transition matrix, and the inverse of the homography matrix to obtain a target matrix, and calculating the real target optical flow based on the target matrix; and calculating the real stitched image based on the intermediate image and the target matrix.
[0220] The image stitching device 300 provided in this application embodiment has been described in the foregoing method embodiment in terms of its implementation principle and the resulting technical effects. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the method embodiment.
[0221] Figure 8 This illustration shows a possible structure of the electronic device 400 provided in an embodiment of this application. (Refer to...) Figure 8 The electronic device 400 includes a processor 410, a memory 420, and a communication interface 430. These components are interconnected and communicate with each other via a communication bus 440 and / or other forms of connection mechanism (not shown).
[0222] The processor 410 includes one or more (only one is shown in the figure), which can be an integrated circuit chip with signal processing capabilities. The processor 410 can be a general-purpose processor, including a Central Processing Unit (CPU), a Microcontroller Unit (MCU), a Network Processor (NP), or other conventional processors; it can also be a special-purpose processor, including a Graphics Processing Unit (GPU), a Neural-network Processing Unit (NPU), a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. Furthermore, when there are multiple processors 410, some can be general-purpose processors and others can be special-purpose processors.
[0223] The memory 420 includes one or more (only one is shown in the figure), which may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0224] Processor 410 and other possible components can access memory 420, read and / or write data therein. In particular, memory 420 can store one or more computer program instructions, which processor 410 can read and execute to implement the image stitching method provided in the embodiments of this application.
[0225] Communication interface 430 includes one or more (only one is shown in the figure) that can be used to communicate directly or indirectly with other devices for data exchange. Communication interface 430 may include interfaces for wired and / or wireless communication.
[0226] Understandable. Figure 8 The structure shown is for illustrative purposes only; the electronic device 400 may also include more than [other components]. Figure 8 The more or fewer components shown, or having the same Figure 8 The different configurations shown. Figure 8 The components shown can be implemented using hardware, software, or a combination thereof. Electronic device 400 may be a physical device, such as a PC, laptop, tablet, mobile phone, server, or smart wearable device, or a virtual device, such as a virtual machine or virtualization container. Furthermore, electronic device 400 is not limited to a single device; it can also be a combination of multiple devices or a cluster of numerous devices.
[0227] This application also provides a computer-readable storage medium storing computer program instructions. These instructions are read and executed by a computer's processor to perform the image stitching method provided in this application. For example, the computer-readable storage medium can be implemented as follows: Figure 8 The memory 420 in the electronic device 400.
[0228] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. An image stitching method, characterized in that, include: Acquire the first image and the second image; Based on the first image and the second image, a first intermediate optical flow and a second intermediate optical flow are calculated. The first intermediate optical flow is the optical flow between the intermediate image and the first image, and the second intermediate optical flow is the optical flow between the intermediate image and the second image. The intermediate image is an image with a viewing angle between the first image and the second image. The size of the first intermediate optical flow is the same as the size of the first image, and the size of the second intermediate optical flow is the same as the size of the second image. Based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image, a stitched image of the first image and the second image is calculated; The step of calculating the stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image includes: A first intermediate image is obtained by mapping based on the first intermediate optical flow and the first image, and a second intermediate image is obtained by mapping based on the second intermediate optical flow and the second image; a target optical flow is calculated based on the first intermediate optical flow and the second intermediate optical flow; a first stitched image is obtained by mapping based on the target optical flow and the first intermediate image, and a second stitched image is obtained by mapping based on the target optical flow and the second intermediate image; the stitched image is obtained by stitching based on the first stitched image and the second stitched image. or, A first intermediate image is obtained by mapping based on the first intermediate optical flow and the first image, and a second intermediate image is obtained by mapping based on the second intermediate optical flow and the second image; the intermediate image is obtained by stitching based on the first intermediate image and the second intermediate image; a target optical flow is calculated based on the first intermediate optical flow and the second intermediate optical flow; and the stitched image is obtained by mapping based on the intermediate image and the target optical flow.
2. The image stitching method according to claim 1, characterized in that, The step of calculating the first intermediate optical flow and the second intermediate optical flow based on the first image and the second image includes: The image regions containing the same scene in the first image and the second image are respectively cropped to obtain the first screenshot and the second screenshot; The first screenshot and the second screenshot are input into the optical flow calculation network to obtain the optical flow of the first screenshot and the optical flow of the second screenshot. The optical flow of the first screenshot is the optical flow between the intermediate image and the first screenshot, and the optical flow of the second screenshot is the optical flow between the intermediate image and the second screenshot. The first screenshot optical flow is upsampled to the size of the first image to obtain the first intermediate optical flow, and the second screenshot optical flow is upsampled to the size of the second image to obtain the second intermediate optical flow.
3. The image stitching method according to claim 1, characterized in that, The step of calculating the target optical flow based on the first intermediate optical flow and the second intermediate optical flow includes: Based on the first intermediate optical flow and the second intermediate optical flow, at least one transition optical flow is calculated by interpolation; The target optical flow is obtained by fusing the first intermediate optical flow, the at least one transition optical flow, and the second intermediate optical flow.
4. The image stitching method according to claim 3, characterized in that, The step of interpolating and calculating at least one transition optical flow based on the first intermediate optical flow and the second intermediate optical flow includes: Obtain at least one weight value; The first intermediate optical flow and the second intermediate optical flow are weighted and summed based on each weight value to obtain the at least one transition optical flow.
5. The image stitching method according to claim 4, characterized in that, The magnitude of the weight value is related to the spectral position of the transition optical flow corresponding to that weight value.
6. The image stitching method according to claim 5, characterized in that, The sum of the weighting coefficients of the first intermediate optical flow and the second intermediate optical flow is 1. The weight value is the weighting coefficient of the first intermediate optical flow. The magnitude of the weight value is positively correlated with the proximity between the viewing angle position of the transition optical flow corresponding to the weight value and the viewing angle position of the first intermediate optical flow.
7. The image stitching method according to any one of claims 4-6, characterized in that, The at least one weight value is uniformly distributed within the interval (0,1).
8. The method according to claim 1, characterized in that The step of calculating the target optical flow based on the first intermediate optical flow and the second intermediate optical flow includes: The preliminary optical flow is calculated based on the first intermediate optical flow and the second intermediate optical flow; The preliminary optical flow is input into the optical flow inversion network to obtain the target optical flow.
9. The image stitching method according to claim 3, characterized in that, The step of fusing the first intermediate optical flow, the at least one transition optical flow, and the second intermediate optical flow to obtain the target optical flow includes: Obtain N+2 weight matrices, where N is the total number of transition optical flows; Based on the N+2 weight matrices, the first intermediate optical flow, the N transition optical flows, and the second intermediate optical flow are weighted and summed to obtain the preliminary optical flow. The preliminary optical flow is the target optical flow, or the preliminary optical flow is input into an optical flow inversion network to obtain the target optical flow.
10. The image stitching method according to claim 9, characterized in that, The position of the maximum value of an element in the weight matrix is related to the spectral position of the optical flow corresponding to the weight matrix.
11. The image stitching method according to claim 1, characterized in that, The step of stitching the first stitched image and the second stitched image to obtain the stitched image includes: The first stitched image and the second stitched image are input into the mask calculation network to obtain the stitching mask; Based on the stitching mask, the first stitched image and the second stitched image are stitched together to obtain the stitched image.
12. The image stitching method according to claim 1, characterized in that, The first intermediate optical flow and the second intermediate optical flow are calculated using the first screenshot optical flow and the second screenshot optical flow output by the optical flow calculation network. The target optical flow is calculated using the optical flow inversion network. The stitched image is calculated using the stitching mask output by the mask calculation network. The method further includes: Obtain the optical flow of the first real screenshot, the optical flow of the second real screenshot, the optical flow of the real target, and the real stitched image; Calculate the optical flow prediction loss based on the first screenshot optical flow, the second screenshot optical flow, the first real screenshot optical flow, and the second real screenshot optical flow; Calculate the optical flow inversion loss based on the target optical flow and the actual target optical flow; Calculate the image stitching loss based on the stitched image and the actual stitched image; The total loss is calculated based on the optical flow prediction loss, the optical flow inversion loss, and the image stitching loss, and the parameters of the optical flow calculation network, the optical flow inversion network, and the mask calculation network are updated based on the total loss.
13. The image stitching method according to claim 12, characterized in that, The target optical flow is obtained by fusing the first intermediate optical flow, at least one transition optical flow, and the second intermediate optical flow. Acquiring the first image and the second image includes: The first image is calculated based on the intermediate image and the homography matrix, and the second image is calculated based on the intermediate image and the inverse of the homography matrix, wherein the intermediate image is the real image; The acquisition of the first real screenshot optical flow, the second real screenshot optical flow, the real target optical flow, and the real stitched image includes: The first true intermediate optical flow is calculated based on the homography matrix, and the first true screenshot optical flow is calculated based on the first true intermediate optical flow. The second true intermediate optical flow is calculated based on the inverse of the homography matrix, and the second true screenshot optical flow is calculated based on the second true intermediate optical flow. Based on the homography matrix and its inverse, at least one transition matrix is interpolated and calculated. Based on the homography matrix, the at least one transition matrix, and the inverse of the homography matrix, a target matrix is obtained by fusing them. The true target optical flow is then calculated based on the target matrix. The true stitched image is calculated based on the intermediate image and the target matrix.
14. An image stitching device, characterized in that, include: The image acquisition module is used to acquire the first image and the second image; An intermediate optical flow calculation module is used to calculate a first intermediate optical flow and a second intermediate optical flow based on the first image and the second image. The first intermediate optical flow is the optical flow between the intermediate image and the first image, and the second intermediate optical flow is the optical flow between the intermediate image and the second image. The intermediate image is an image with a viewing angle between the first image and the second image. The size of the first intermediate optical flow is the same as the size of the first image, and the size of the second intermediate optical flow is the same as the size of the second image. An image stitching module is used to calculate a stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image. The image stitching module calculates a stitched image of the first image and the second image based on the first intermediate optical flow, the second intermediate optical flow, the first image, and the second image, including: A first intermediate image is obtained by mapping based on the first intermediate optical flow and the first image, and a second intermediate image is obtained by mapping based on the second intermediate optical flow and the second image; a target optical flow is calculated based on the first intermediate optical flow and the second intermediate optical flow; a first stitched image is obtained by mapping based on the target optical flow and the first intermediate image, and a second stitched image is obtained by mapping based on the target optical flow and the second intermediate image; the stitched image is obtained by stitching based on the first stitched image and the second stitched image. or, A first intermediate image is obtained by mapping based on the first intermediate optical flow and the first image, and a second intermediate image is obtained by mapping based on the second intermediate optical flow and the second image; an intermediate image is obtained by stitching based on the first intermediate image and the second intermediate image; a target optical flow is calculated based on the first intermediate optical flow and the second intermediate optical flow; and the stitched image is obtained by mapping based on the intermediate image and the target optical flow.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when read and executed by a processor, perform the method as described in any one of claims 1-13.
16. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer program instructions, which are read and executed by the processor to perform the method according to any one of claims 1-13.
Citation Information
Patent Citations
Image merging method and apparatus
CN106803899A
Image stitching method and image stitching device
CN106997579A