Image stitching method, device and electronic device
By using the first neural network model to fuse the target overlapping area during the image stitching process, the problem of low image stitching efficiency in the prior art is solved, and more efficient image stitching processing is achieved.
Patent Information
- Application Number
- CN202110990427.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-08-26
AI Technical Summary
In the process of image stitching, the prior art performs fusion calculations on each pixel based on the CPU, resulting in a longer time, which reduces the fusion processing efficiency and the processing efficiency of image stitching.
The first neural network model is used to fuse the target overlapping region in the initial stitching image to obtain the fused overlapping region, and then determine the target overlapping image. After determining the initial stitching image, the method performs fusion processing through the neural network model, avoiding the calculation of each pixel one by one.
It saves time to perform fusion calculations on all pixels in the initial stitching image, improves the fusion processing efficiency, and thus improves the processing efficiency of image stitching.
Smart Images

Figure CN113902657B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to an image stitching method, device and electronic equipment. Background Art
[0002] Image stitching is the process of combining multiple images with overlapping fields of view to produce an image with a larger field of view and higher resolution. In the image stitching process, it is usually necessary to perform fusion processing on the images to be stitched. In related technologies, it is usually based on the CPU (central processing unit) to perform fusion-related calculations on each pixel in the stitched image after the initial stitching. Since the CPU needs a certain amount of processing time to calculate each pixel, it takes a long time to perform fusion calculations on all pixels in the stitched image after the initial stitching, which reduces the efficiency of the fusion processing and thus reduces the processing efficiency of image stitching. Summary of the invention
[0003] The object of the present invention is to provide an image stitching method, device and electronic device to improve the processing efficiency of image stitching.
[0004] The present invention provides an image stitching method, which includes: acquiring a first image and a second image to be stitched; determining an initial stitching image of the first image and the second image; using a first neural network model to perform a fusion process on a target overlapping area in the initial stitching image to obtain a fused overlapping area corresponding to the target overlapping area; and determining a target stitching image corresponding to the first image and the second image based on the fused overlapping area and the initial stitching image.
[0005] Furthermore, before determining the initial stitched image of the first image and the second image, the method also includes: performing illumination compensation on the second image; correspondingly, determining the initial stitched image of the first image and the second image includes: determining the initial stitched image based on the first image and the illumination-compensated second image.
[0006] Furthermore, the step of performing illumination compensation on the second image includes: performing illumination compensation on the second image based on the second neural network model and the first image.
[0007] Furthermore, based on the second neural network model and the first image, illumination compensation is performed on the second image, including: determining a projection transformation matrix based on the first image and the second image; determining an initial overlapping area between the first image and the second image based on the projection transformation matrix; wherein the initial overlapping area includes: a first sub-overlapping area corresponding to the first image and a second sub-overlapping area corresponding to the second image; inputting the first sub-overlapping area and the second sub-overlapping area into the second neural network model, and determining a mapping relationship between a first pixel value of each pixel in the first sub-overlapping area and a second pixel value of a pixel at the same position in the second sub-overlapping area through the second neural network model; obtaining the mapping relationship output by the second neural network model; and for each color channel, matching the pixel value of each pixel point in the second image in the color channel with the pixel value of each pixel point in the first image in the color channel based on the mapping relationship, so as to perform illumination compensation on the second image.
[0008] Furthermore, based on the projection transformation matrix, the step of determining the initial overlapping area of the first image and the second image includes: obtaining the boundary coordinates of the second image; wherein the boundary coordinates are used to indicate the image area of the second image; based on the projection transformation matrix and the boundary coordinates of the second image, determining the boundary coordinates after the projection transformation; based on the boundary coordinates after the projection transformation, determining the second image after the projection transformation; and determining the overlapping image area of the second image after the projection transformation and the first image as the initial overlapping area.
[0009] Furthermore, based on the first image and the second image, the step of determining the projection transformation matrix includes: extracting at least one first feature point in the first image and at least one second feature point in the second image; based on the at least one first feature point and at least one second feature point, determining at least one matching feature point pair; based on the at least one matching feature point pair, determining the projection transformation matrix.
[0010] Furthermore, based on the second neural network model and the first image, the step of performing illumination compensation on the second image includes: inputting the first image and the second image into the second neural network model, performing illumination compensation on the second image based on the first image through the second neural network, and obtaining the illumination compensated second image.
[0011] Furthermore, the target overlapping area includes: a third sub-overlapping area corresponding to the first image and a fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a stitching model and a fusion model; the step of using the first neural network model to fuse the target overlapping area in the initial stitching image to obtain the fused overlapping area corresponding to the target overlapping area includes: inputting the third sub-overlapping area and the fourth sub-overlapping area into the stitching model, searching for the seam between the third sub-overlapping area and the fourth sub-overlapping area through the stitching model, and obtaining the seam area corresponding to the third sub-overlapping area and the fourth sub-overlapping area; based on the seam area, fusing the third sub-overlapping area and the fourth sub-overlapping area to obtain an initial fused overlapping area; inputting the initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area into the fusion model, and fusing the initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area through the fusion model to obtain a fused overlapping area.
[0012] Furthermore, the target overlapping area includes: a third sub-overlapping area corresponding to the first image and a fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a stitching model; the target overlapping area in the initial stitched image is fused using the first neural network model to obtain a fused overlapping area corresponding to the target overlapping area, including: inputting the third sub-overlapping area and the fourth sub-overlapping area into the stitching model to obtain a stitching area corresponding to the third sub-overlapping area and the fourth sub-overlapping area; feathering the stitching area to obtain a feathered overlapping area, and determining the feathered overlapping area as the fused overlapping area.
[0013] Furthermore, the fusion model is trained in the following manner: obtaining a first image; translating and / or rotating the first image to obtain a second image; fusing the first image and the second image to obtain an initial fused image; and training the fusion model based on the first image, the second image and the initial fused image.
[0014] Furthermore, based on the fused overlapping area and the initial stitched image, the step of determining the target stitched image corresponding to the first image and the second image includes: replacing the target overlapping area in the initial stitched image with the fused overlapping area to obtain the target stitched image corresponding to the first image and the second image.
[0015] Furthermore, the color channels of the first image and the second image are arranged in RGGB.
[0016] The present invention provides an image stitching device, which includes: an acquisition module, used to acquire a first image and a second image to be stitched; a first determination module, used to determine an initial stitched image of the first image and the second image; a fusion module, used to use a first neural network model to perform a fusion process on a target overlapping area in the initial stitched image to obtain a fused overlapping area corresponding to the target overlapping area; and a second determination module, used to determine the target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image.
[0017] An electronic device provided by the present invention comprises a processing device and a storage device, wherein the storage device stores a computer program, and the computer program executes any of the above-mentioned image stitching methods when the processing device is running.
[0018] The present invention provides a machine-readable storage medium, which stores a computer program. When the computer program is run by a processing device, any of the above-mentioned image stitching methods is executed.
[0019] The image stitching method, device and electronic device provided by the present invention first obtain the first image and the second image to be stitched; determine the initial stitching image of the first image and the second image; then use the first neural network model to fuse the target overlapping area in the initial stitching image to obtain the fused overlapping area corresponding to the target overlapping area; finally, based on the fused overlapping area and the initial stitching image, determine the target stitching image corresponding to the first image and the second image. In this method, after determining the initial stitching image of the first image and the second image, the first neural network model is used to fuse the target overlapping area in the initial stitching image to obtain the corresponding fused overlapping area, and there is no need to perform fusion-related calculations on each pixel in the initial stitching image based on the CPU, which saves the time of performing fusion calculations on all pixels in the initial stitching image, improves the efficiency of fusion processing, and thus improves the processing efficiency of image stitching. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 A schematic diagram of the structure of an electronic system provided by an embodiment of the present invention;
[0022] Figure 2 A flowchart of an image stitching method provided by an embodiment of the present invention;
[0023] Figure 3 A flowchart of another image stitching method provided by an embodiment of the present invention;
[0024] Figure 4 A flowchart of another image stitching method provided by an embodiment of the present invention;
[0025] Figure 5 A flowchart of another image stitching method provided by an embodiment of the present invention;
[0026] Figure 6 A flow chart of a RAW domain image stitching method provided by an embodiment of the present invention;
[0027] Figure 7 A schematic structural diagram of an image stitching device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0029] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.
[0030] Image stitching is the process of combining multiple images with overlapping fields of view to produce a larger field of view and high-resolution images. The field of view of a single eye of an average person is about 120°, and the field of view of two eyes is generally about 160-220°, while the FOV (Field Of Vision) of an ordinary camera is generally only about 40°-60°, making it difficult to achieve high-definition imaging of detailed objects while ensuring a large imaging field of view. Image stitching technology can combine multiple cameras with small field of view into a multi-camera with a large field of view, which has important application value in security, remote conferencing, sports event directing and other fields. In the related art, image stitching processing is usually performed based on RGB (R: Red; G: Green; B: Blue) domain images, while the current relatively novel image processing and analysis processes are usually implemented based on RAW domain images. For example, the detection, segmentation or recognition of RAW domain images is used to solve the scene problems such as dark light and backlight that cannot be solved by RGB domain images. In the case where image processing and analysis are mainly concentrated on RAW domain images, when the corresponding RGB domain images are stitched, since the RGB domain images lack some image details compared with the RAW domain images, the stitched images in the RGB domain lack the corresponding image analysis results, which is not conducive to the subsequent image detection, segmentation and other processing; In the related art, each pixel in the stitched image after the initial stitching is usually subjected to fusion-related calculation processing based on the CPU (central processing unit). Since the CPU needs a certain processing time to calculate each pixel, it takes a long time to perform fusion calculations on all pixels in the stitched image after the initial stitching, which reduces the fusion processing efficiency, thereby reducing the processing efficiency of image stitching. Based on this, an embodiment of the present invention provides an image stitching method, device and electronic device. The technology can be applied to applications that require image stitching. The embodiment of the present invention is described in detail below.
[0031] First, refer to Figure 1 An exemplary electronic system 100 for implementing the image stitching method, apparatus, and electronic device according to the embodiments of the present invention is described.
[0032] like Figure 1 The electronic system 100 includes one or more processing devices 102, one or more storage devices 104, an input device 106, an output device 108, and one or more image acquisition devices 110. These components are interconnected via a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1The components and structure of the electronic system 100 shown are merely exemplary and non-limiting. The electronic system may also have other components and structures as required.
[0033] The processing device 102 may be a gateway, or an intelligent terminal, or a device including a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities. It may process data of other components in the electronic system 100, and may also control other components in the electronic system 100 to perform desired functions.
[0034] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processing device 102 may run the program instructions to implement the client functions and / or other desired functions in the embodiments of the present invention (implemented by the processing device) described below. Various applications and various data, such as various data used and / or generated by the application, may also be stored in the computer-readable storage medium.
[0035] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, and the like.
[0036] The output device 108 may output various information (eg, images or sounds) to the outside (eg, a user), and may include one or more of a display, a speaker, and the like.
[0037] The image acquisition device 110 may acquire preview video frames or image data, and store the acquired preview video frames or image data in the storage device 104 for use by other components.
[0038] Exemplarily, the components in the example electronic system for implementing the image stitching method, device, and electronic device according to the embodiment of the present invention can be integrated or dispersed, such as integrating the processing device 102, the storage device 104, the input device 106, and the output device 108 into one body, and setting the image acquisition device 110 at a specified position where the target image can be acquired. When the components in the above electronic system are integrated, the electronic system can be implemented as an intelligent terminal such as a camera, a smart phone, a tablet computer, a computer, a vehicle-mounted terminal, etc., or a server.
[0039] This embodiment provides an image stitching method, which is executed by a processing device in the above electronic system; the processing device can be any device or chip with data processing capability. Figure 2 As shown, the method comprises the following steps:
[0040] Step S202: acquiring a first image and a second image to be stitched.
[0041] In the above-mentioned first image and second image, the first image can be used as the reference image and the second image can be used as the image to be registered, or the second image can be used as the reference image and the first image can be used as the image to be registered. The specific selection can be made according to actual needs; for the convenience of explanation, taking the first image as the reference image and the second image as the image to be registered as an example, the first image is usually located in the reference image coordinate system, and the first image can be represented by a tar image, etc.; the second image is usually located in the image coordinate system to be registered, and the second image can be represented by a src image, etc.; the first image and the second image can be RAW domain images or RGB domain images, etc. In actual implementation, when it is necessary to stitch the images, it is usually necessary to first obtain the first image and the second image to be stitched, and the first image and the second image are usually images with a partially overlapping field of view.
[0042] Step S204: determining an initial stitched image of the first image and the second image.
[0043] The acquired first image and the second image are initially stitched together. For example, if the first image is used as a reference image and the second image is used as an image to be registered, the second image can be projected onto the first image for initial stitching to obtain an initial stitched image of the first image and the second image.
[0044] Step S206, using the first neural network model to perform fusion processing on the target overlapping area in the initial stitched image to obtain a fused overlapping area corresponding to the target overlapping area.
[0045] The above-mentioned first neural network model can be implemented by a variety of convolutional neural networks, such as residual networks, VGG networks, etc.; since the first image and the second image are images with a portion of overlapping fields of view, the target overlapping area can be understood as the overlapping area after the first image and the second image are initially stitched together; in actual implementation, after determining the above-mentioned initial stitched image, the target overlapping area where the first image and the second image overlap can be determined from the initial stitched image, and the first neural network is used to fuse the target overlapping area to obtain the corresponding fused overlapping area.
[0046] Step S208: determining a target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image.
[0047] In actual implementation, after determining the above-mentioned fusion overlapping area, a target stitched image of the first image and the second image can be obtained based on the fusion overlapping area and the initial stitched image; the target stitched image can be two or more images with overlapping parts stitched together into a seamless panoramic image or high-resolution image; wherein, the two or more images with overlapping parts may be images obtained at different times, different perspectives or different sensors.
[0048] The above-mentioned image stitching method first obtains the first image and the second image to be stitched; determines the initial stitching image of the first image and the second image; then uses the first neural network model to fuse the target overlapping area in the initial stitching image to obtain the fused overlapping area corresponding to the target overlapping area; finally, based on the fused overlapping area and the initial stitching image, determines the target stitching image corresponding to the first image and the second image. In this method, after determining the initial stitching image of the first image and the second image, the first neural network model is used to fuse the target overlapping area in the initial stitching image to obtain the corresponding fused overlapping area. It is not necessary to perform fusion-related calculations on each pixel in the initial stitching image based on the CPU, which saves the time of performing fusion calculations on all pixels in the initial stitching image, improves the efficiency of fusion processing, and thus improves the processing efficiency of image stitching.
[0049] The embodiment of the present invention also provides another image stitching method, which is implemented on the basis of the method in the above embodiment; Figure 3 As shown, the method comprises the following steps:
[0050] Step S302: acquiring a first image and a second image to be stitched.
[0051] Step S304: performing illumination compensation on the second image.
[0052] The step S304 can be specifically implemented by the following step 1:
[0053] Step 1: Based on the second neural network model and the first image, perform illumination compensation on the second image.
[0054] The second neural network model can be implemented by a variety of convolutional neural networks, such as a residual network, a VGG network, etc.; the second neural network model can be a network model different from the first neural network model, or it can be the first neural network model, except that the process of illumination compensation is performed by a sub-model or sub-module in the first neural network model. In actual implementation, considering that due to light reasons, the second image captured may have color deviation caused by unbalanced light. In order to offset the color deviation in the second image, illumination compensation can be performed on the second image based on the pre-trained second neural network model and the first image.
[0055] This step 1 can be specifically implemented by the following steps A to D:
[0056] Step A: determining a projection transformation matrix based on the first image and the second image.
[0057] In actual implementation, the projection transformation matrix can be determined according to the acquired first image and the second image, wherein the first image and the second image usually have the same number of color channels, for example, the first image has four color channels of R, G, G and B, and the second image also has four color channels of R, G, G and B; specifically, step A can be implemented through steps a to c:
[0058] Step a: extracting at least one first feature point in the first image and at least one second feature point in the second image.
[0059] In order to better perform image matching, it is usually necessary to select representative areas in the image, such as corner points, edge points, bright spots in dark areas or dark spots in bright areas, etc. in the image. The above-mentioned first feature points can be corner points, edge points, bright spots in dark areas or dark spots in bright areas, etc. extracted from the first image, and the number of the first feature points can be one or more; the above-mentioned second feature points can be corner points, edge points, bright spots in dark areas or dark spots in bright areas, etc. extracted from the second image, and the number of the second feature points can be one or more.
[0060] In the present embodiment, in the color channel of each pixel of the first image, each color channel is preset with a corresponding first weight value; in the color channel of the second image, each color channel is preset with a corresponding second weight value; in actual implementation, the above-mentioned first weight value and second weight value are usually pre-set fixed values, which can be set according to actual needs and are not limited here. For example, taking the first image and the second image both having four color channels of R, G, G and B as an example, in the first image, the first weight value corresponding to the R channel is 0.3, the first weight values corresponding to the two G channels are both 0.2, and the first weight value corresponding to the B channel is 0.3; in the second image, the second weight value corresponding to the R channel is 0.3, the second weight values corresponding to the two G channels are both 0.2, and the second weight value corresponding to the B channel is 0.3.
[0061] In the above step a, the step of extracting at least one first feature point in the first image may include steps a0 to a3:
[0062] Step a0: for each pixel in the first image, multiply the component of each color channel in the pixel by the first weight value corresponding to the color channel to obtain a first calculation result for each color channel.
[0063] The above color channels can be understood as channels for storing image color information. For example, an RGB image has three color channels, namely, an R channel, a G channel, and a B channel. The components of each color channel can be understood as the brightness value of each color channel. In actual implementation, the first image usually includes multiple pixels, and each pixel usually has multiple color channels. For each pixel, the component of each color channel in the pixel can be multiplied by the corresponding first weight value to obtain the first calculation result of the color channel. For example, the first image has four color channels, namely, R, G, G, and B, wherein the first weight value corresponding to the R channel is 0.3, the first weight values corresponding to the two G channels are both 0.2, and the first weight value corresponding to the B channel is 0.3. For example, the component corresponding to the R channel of a specified pixel is 150, the components corresponding to the two G channels are both 100, and the component corresponding to the B channel is 80. For the specified pixel, the first calculation result of the R channel is 150*0.3=45, the first calculation results of the two G channels are 100*0.2=20, and the first calculation result of the B channel is 80*0.3=24.
[0064] Step a1, adding the first calculation results of each color channel in the pixel to obtain a first grayscale value corresponding to the pixel.
[0065] After obtaining the first calculation result of each color channel in each pixel in the first image through the above step a0, the first calculation result of each color channel in the pixel can be added to obtain the first grayscale value of the pixel. For example, still taking the specified pixel in the above step a0 as an example, the first calculation result of each color channel in the specified pixel can be added to obtain the first grayscale value corresponding to the specified pixel, that is, 45+20+20+24=109.
[0066] Step a2: determining a grayscale image of the first image based on a first grayscale value corresponding to each pixel in the first image.
[0067] After the first grayscale value corresponding to each pixel in the first image is obtained through the above steps a0 and a1, a grayscale image of the first image can be obtained according to the obtained multiple first grayscale values.
[0068] Step a3: extract at least one first feature point from the grayscale image of the first image.
[0069] In actual implementation, one or more first feature points can be extracted from the grayscale image of the obtained first image based on algorithms such as SIFT (Scale-Invariant Feature Transform) or SuperPoint. For details, please refer to the relevant technology, and the process of extracting feature points using methods such as SIFT or SuperPoint will not be repeated here; among them, SIFT has scale invariance, can detect key points in the image, and is a local feature descriptor; SuperPoint is a feature point detection and descriptor extraction method based on self-supervised training.
[0070] In the above step a, the step of extracting at least one second feature point in the second image may include steps a4 to a7:
[0071] Step a4: for each pixel in the second image, multiply the component of each color channel in the pixel by the second weight value corresponding to the color channel to obtain a second calculation result for each color channel.
[0072] In actual implementation, the second image usually also includes multiple pixels, and each pixel usually has multiple color channels. For each pixel, the component of each color channel in the pixel can be multiplied by the corresponding second weight value to obtain the second calculation result of the color channel. For example, the second image still has four color channels, namely R, G, G and B, wherein the first weight value corresponding to the R channel is 0.3, the first weight values corresponding to the two G channels are both 0.2, and the first weight value corresponding to the B channel is 0.3. For example, the component corresponding to the R channel of a target pixel is 100, the components corresponding to the two G channels are both 120, and the component corresponding to the B channel is 150. Then, in the target pixel, the second calculation result of the R channel is 100*0.3=30, the second calculation results of the two G channels are 120*0.2=24, and the second calculation result of the B channel is 150*0.3=45.
[0073] Step a5, adding the second calculation results of each color channel in the pixel to obtain a second grayscale value corresponding to the pixel.
[0074] After obtaining the second calculation result of each color channel in each pixel in the second image through the above step a4, the second calculation results of each color channel in the pixel can be added to obtain the second grayscale value of the pixel. For example, still taking the target pixel in the above step a4 as an example, the second calculation results of each color channel in the target pixel can be added to obtain the second grayscale value corresponding to the target pixel, that is, 30+24+24+45=123.
[0075] Step a6: determining a grayscale image of the second image based on the second grayscale value corresponding to each pixel in the second image.
[0076] After the second grayscale value corresponding to each pixel in the second image is obtained through the above steps a4 and a5, a grayscale image of the second image can be obtained according to the obtained multiple second grayscale values.
[0077] Step a7: extract at least one second feature point from the grayscale image of the second image.
[0078] In actual implementation, one or more second feature points may be extracted from the grayscale image of the obtained second image based on algorithms such as SIFT or SuperPoint.
[0079] Step b: determining at least one matching feature point pair based on at least one first feature point and at least one second feature point.
[0080] After at least one first feature point in the first image and at least one second feature point in the second image are extracted through the above step a, feature point matching can be performed. Specifically, based on at least one first feature point and at least one second feature point, KNN (K-NearestNeighbor, K nearest neighbor classification algorithm) or RANSAC (RANdom SAmpleConsensus, random sampling consensus algorithm) and other algorithms can be used to obtain one or more matching feature point pairs in the first image and the second image. The number of matching feature point pairs is generally not less than four pairs; for details, reference can be made to the process of matching feature points using KNN or RANSAC and other methods in the relevant technology, which will not be repeated here; wherein, KNN means that each sample can be represented by its closest K neighboring values, and its core idea is that if most of the K nearest neighboring samples of a sample in the feature space belong to a certain category, then the sample also belongs to this category and has the characteristics of the samples in this category; RANSAC is random sampling consistency matching, which uses matching points to calculate the homography matrix between the two images, and then uses the reprojection error to determine whether a certain match is a correct match.
[0081] Step c: determining a projection transformation matrix based on at least one matching feature point pair.
[0082] After determining at least one matching feature point pair, the projection transformation matrix can be calculated based on the matching information of the matching feature point pair, such as the coordinates of the matching feature point pair, etc. For details, please refer to the process of calculating the projection transformation matrix based on the matching feature point pair in the related technology, which will not be repeated here. Based on the projection transformation matrix, the second image can be projected onto the first image. The projection transformation matrix works in such a way that for each color channel included in each pixel in the second image, a projection transformation is performed according to the calculated projection transformation matrix. In addition, after the second image is projected onto the first image, the corresponding coordinates transformed may be decimals. At this time, it is usually necessary to interpolate the coordinates based on the values of adjacent pixels. For details, please refer to the interpolation processing method in the related technology. For example, the Cubic bicubic interpolation method can be used to reduce the impact of resolution, where the resolution can be understood as the clarity of the second image after the projection transformation.
[0083] Step B: determining an initial overlapping area of the first image and the second image based on the projection transformation matrix; wherein the initial overlapping area includes: a first sub-overlapping area corresponding to the first image and a second sub-overlapping area corresponding to the second image.
[0084] Based on the projection transformation matrix determined in the above steps, the second image is projected onto the first image. Since the first image and the second image are images with partially overlapping fields of view, after completing the projection transformation, the initial overlapping area of the first image and the second image after the projection transformation can be obtained. The initial overlapping area includes the first sub-overlapping area of the first image and the second sub-overlapping area of the second image after the projection transformation.
[0085] The step B can be specifically implemented through steps h to k:
[0086] Step h: obtaining the boundary coordinates of the second image; wherein the boundary coordinates are used to indicate the image area of the second image.
[0087] The above boundary coordinates can be understood as edge coordinates used to indicate the overall shape of the second image, and the image area corresponding to the second image can be obtained through the edge coordinates. In actual implementation, when it is necessary to determine the initial overlapping area, it is usually necessary to first obtain the boundary coordinates of the second image, and the number of the boundary coordinates is usually multiple.
[0088] Step i: determining the boundary coordinates after the projection transformation based on the projection transformation matrix and the boundary coordinates of the second image.
[0089] When the boundary coordinates of the second image are obtained and the second image is projected onto the first image based on the projection transformation matrix, the boundary coordinates of the second image after the projection transformation can be obtained; for example, if the boundary coordinates of the second image are four corner point coordinates, after completing the projection transformation based on the projection transformation matrix, the four corner point coordinates of the second image after the projection transformation can be obtained.
[0090] Step j, determining a second image after projection transformation based on the boundary coordinates after projection transformation.
[0091] After determining the boundary coordinates of the second image after the projection transformation, the second image after the projection transformation is determined according to the image area surrounded by the boundary coordinates after the projection transformation. For example, still taking the boundary coordinates of the second image as the four corner point coordinates as an example, the image area surrounded by the four corner point coordinates of the second image after the projection transformation is the second image after the projection transformation.
[0092] Step k: determine the overlapping image area of the second image after the projection transformation and the first image as the initial overlapping area.
[0093] After the second image after the projective transformation is determined, the initial overlapping area can be obtained by taking the intersection of the second image after the projective transformation and the first image.
[0094] Step C, inputting the first sub-overlapping area and the second sub-overlapping area into the second neural network model, and determining the mapping relationship between the first pixel value of each pixel in the first sub-overlapping area and the second pixel value of the pixel at the same position in the second sub-overlapping area through the second neural network model.
[0095] In actual implementation, after determining the initial overlapping area, the pixel values of the first sub-overlapping area corresponding to the first image and the second sub-overlapping area corresponding to the second image can be extracted, and the pixel value histogram distribution of the image corresponding to the first sub-overlapping area in the first image and the pixel value histogram distribution of the image corresponding to the second sub-overlapping area in the second image are calculated respectively. The pixel value histogram distribution of the image corresponding to the second sub-overlapping area is transformed into a histogram distribution matching the pixel value histogram distribution of the image corresponding to the first sub-overlapping area based on the histogram matching method. Specifically, it can be implemented by constructing a pixel value mapping table, and the pixel value mapping table can be a LUT (Look-Up-Table) table, etc., and the number of color channels included in the first image and the second image is the same, and each color channel usually has its corresponding pixel value mapping table; for example, the first image and the second image both include four color channels of R, G, G and B, the first sub-overlapping area and the second sub-overlapping area are input into a pre-trained second neural network model, and the mapping relationship between the first pixel value of each pixel in the first sub-overlapping area and the second pixel value of the pixel at the same position in the second sub-overlapping area is determined, which can be expressed as GT (Ground Transformation). Truth, which indicates the classification accuracy of the training set of supervised learning, is used to prove or disprove a hypothesis) four LUT tables of 0-65535 are used for learning, and finally four LUT tables of 0-65535 can be calculated through the second neural network model, where the four 0-65535 LUT tables are usually different, and the number of LUT tables is related to the number of color channels. In the related art, the above-mentioned pixel value mapping table is usually calculated based on the traditional CPU method. Since it is usually necessary to calculate pixel by pixel, the calculation process takes a long time and is inefficient. This embodiment can accelerate the calculation method in the related art through NN (Neural Network).
[0096] The above-mentioned pixel value mapping table can be understood as a mapping table of pixel grayscale values. The mapping table transforms the actually sampled pixel grayscale values into another corresponding grayscale value through certain transformations, such as threshold, inversion, binarization, contrast adjustment, linear transformation, etc., so as to highlight the useful information of the image and enhance the light contrast of the image.
[0097] Step D, obtaining a mapping relationship output by the second neural network model;
[0098] Step E: for each color channel, based on the mapping relationship, the pixel value of each pixel point in the second image in the color channel is matched with the pixel value of each pixel point in the first image in the color channel to perform illumination compensation on the second image.
[0099] After obtaining the above mapping relationship output by the second neural network, for example, after obtaining the pixel value mapping table corresponding to each color channel, the pixel value mapping table can be applied to the entire second image, that is, the value range of each color channel of the second image can be mapped to a distribution similar to that of the first image through the pixel value mapping table, that is, the pixel value of each pixel point in the second image in the color channel can be matched with the pixel value of each pixel point in the first image in the color channel. For example, still taking the first image and the second image both including four color channels of R, G, G and B, and using the second neural network model to calculate four 0-65535 LUT tables corresponding to the four color channels, for each color channel, the pixel value of each pixel point in the second image in the color channel can be mapped to a distribution similar to that of the first image through the LUT table corresponding to the color channel, thereby achieving illumination compensation for the second image.
[0100] This step 1 can also be specifically implemented by the following step H:
[0101] Step H: input the first image and the second image into the second neural network model, and perform illumination compensation on the second image based on the first image through the second neural network to obtain the illumination compensated second image.
[0102] In actual implementation, the first image and the second image can be input into a pre-trained second neural network model, and the second image can be illuminated based on the first image through the second neural network. The first image and the second image after illumination compensation based on histogram matching can be used as GT to supervise the training of the second neural network model. Finally, the second image after illumination compensation is calculated by the second neural network model.
[0103] In the related art, the CPU usually calculates the corresponding compensation coefficient for each pixel in the stitched image, and then compensates each pixel. Since the CPU needs a certain amount of processing time to calculate each pixel, it takes a long time to compensate all pixels in the image, which reduces the processing efficiency of light compensation. The light compensation method in this embodiment uses a pre-trained neural network model to perform light compensation on the second image, and does not need to calculate each pixel based on the CPU, which saves the time for compensating all pixels in the image and improves the processing efficiency of light compensation.
[0104] Step S306: determining an initial stitched image based on the first image and the second image after illumination compensation.
[0105] According to the projection transformation matrix obtained above, the second image after illumination compensation is projected onto the first image and initially stitched together to obtain an initial stitched image of the first image and the second image after illumination compensation.
[0106] Step S308, using the first neural network model to perform fusion processing on the target overlapping area in the initial stitched image to obtain a fused overlapping area corresponding to the target overlapping area.
[0107] Step S310: determining a target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image.
[0108] The above-mentioned image stitching method obtains the first image and the second image to be stitched, performs illumination compensation on the second image, determines the initial stitching image based on the first image and the second image after illumination compensation, fuses the target overlapping area in the initial stitching image using the first neural network model to obtain the fused overlapping area corresponding to the target overlapping area, and determines the target stitching image corresponding to the first image and the second image based on the fused overlapping area and the initial stitching image. In this method, after determining the initial stitching image of the first image and the second image, the first neural network model is used to fuse the target overlapping area in the initial stitching image to obtain the corresponding fused overlapping area, and there is no need to perform fusion-related calculations on each pixel in the initial stitching image based on the CPU, which saves the time of performing fusion calculations on all pixels in the initial stitching image, improves the efficiency of fusion processing, and thus improves the processing efficiency of image stitching.
[0109] The embodiment of the present invention also provides another image stitching method, which is implemented on the basis of the method of the above embodiment; the method focuses on describing the specific process of using the first neural network model to fuse the target overlapping area in the initial stitched image to obtain the fused overlapping area corresponding to the target overlapping area. In the method, the target overlapping area includes: the third sub-overlapping area corresponding to the first image and the fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a stitching model and a fusion model; the stitching model can be implemented by a variety of convolutional neural networks, such as residual network, VGG network, etc.; the fusion model can also be implemented by a variety of convolutional neural networks, such as residual network, VGG network, etc.; the stitching model and the fusion model can be sub-modules or sub-models in the first neural network model, or they can be two separate neural network models; Figure 4 As shown, the method comprises the following steps:
[0110] Step S402: acquiring a first image and a second image to be stitched.
[0111] Step S404: determining an initial stitched image of the first image and the second image.
[0112] Step S406: input the third overlapping sub-region and the fourth overlapping sub-region into the stitching model, and search the stitching seam between the third overlapping sub-region and the fourth overlapping sub-region through the stitching model to obtain the stitching seam region corresponding to the third overlapping sub-region and the fourth overlapping sub-region.
[0113] In actual implementation, after the initial stitching image is determined, the third sub-overlapping area and the fourth sub-overlapping area contained in the target overlapping area can be determined according to the target overlapping area of the initial stitching image, and the third sub-overlapping area and the fourth sub-overlapping area are input into a pre-trained stitching model, and the stitching model is used to search for the stitching seam between the third sub-overlapping area and the fourth sub-overlapping area, and the searched stitching seam area is output, and the stitching seam area is usually a stitching point set or a stitching mask, etc. The stitching seam area can also be a stitching line. In the related art, graphcut, vonorio, etc. are usually used. The traditional method of seam search is used to calculate the best seam line of the third sub-overlapping area and the fourth sub-overlapping area. By searching for the best seam line, the spatial continuity of the seam connection can be reduced. Generally, it is necessary to calculate the data item by item based on the CPU, and the calculation time is relatively long. Among them, graphcut is an energy optimization algorithm, which can be used for foreground and background segmentation, stereo vision or cutout in the field of image processing. In this embodiment, the above-mentioned traditional calculation method can be distilled based on a pre-trained stitching model, and there is no need for CPU calculation, thereby achieving hardware acceleration, accelerating the processing speed, and improving the processing efficiency.
[0114] Step S408: Based on the stitching area, the third sub-overlapping area and the fourth sub-overlapping area are merged to obtain an initial merged overlapping area.
[0115] In actual implementation, after obtaining the seam area, for example, the seam area is a point set of an unfeathered seam, the third sub-overlapping area and the fourth sub-overlapping area can be directly fused based on the point set of the unfeathered seam to obtain an initial fused overlapping area.
[0116] Step S410: input the initial fused overlapping region, the third sub-overlapping region, and the fourth sub-overlapping region into a fusion model, and fuse the initial fused overlapping region, the third sub-overlapping region, and the fourth sub-overlapping region through the fusion model to obtain a fused overlapping region.
[0117] In the related art, the fusion processing of overlapping areas usually requires the CPU to calculate and fuse the data item by item, which takes a long time. In actual implementation, this method can be based on NN-blending, that is, the fusion method of neural network, to perform image fusion processing, which can improve processing efficiency. Specifically, the initial fused overlapping area obtained above and two unfused original images, namely the third sub-overlapping area and the fourth sub-overlapping area, can be sent to the fusion model together, and the initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area are fused through the fusion model. For example, after image-to-image processing, the fusion optimized result is obtained.
[0118] The above fusion model can be trained by following steps 5 to 8:
[0119] Step 5: Get the first picture.
[0120] Step six, translate and / or rotate the first image to obtain a second image.
[0121] In actual implementation, the above-mentioned first image can be any image; that is, the above-mentioned second image can be obtained by subjecting the first image to processing such as small-range translation and / or rotation; in actual implementation, the training process of the fusion model can construct training samples with a single image, and the single image is the above-mentioned first image. After the single image is subjected to small-range translation and rotation, a changed image is obtained, which is the above-mentioned second image.
[0122] Step seven, fusing the first image and the second image to obtain an initial fused image.
[0123] After obtaining the first image and the second image, the two images may be fused through randomly generated seams, for example, by using an Alpha-Blending method to obtain the initial fused image.
[0124] Step eight, training a fusion model based on the first image, the second image and the initial fusion image.
[0125] The first image, the second image and the initial fusion image are used as training samples to train the fusion model. Specifically, the training samples composed of the first image, the second image and the initial fusion image can be input into the initial fusion model to output a fusion image optimized for the initial fusion image, determine the loss value based on the fusion image and the first image, and update the weight parameters of the initial fusion model based on the loss value; continue to perform the step of obtaining the first image until the initial fusion model converges to obtain the fusion model. Among them, the loss value can be understood as the difference between the fusion image optimized for the initial fusion image and the first image output above; the weight parameters can include all parameters in the initial fusion model, such as convolution kernel parameters, etc. When training the initial fusion model, it is usually necessary to update all parameters in the initial fusion model based on the fusion image and the first image to train the initial fusion model. Then continue to perform the step of obtaining the training sample until the initial fusion model converges, or the loss value converges, and finally obtain the trained fusion model. In actual implementation, the output of the initial fusion model can be supervised by the first image, and the supervised LOSS is the L1 distance, wherein the L1-loss of the seam area can be weighted by 2x.
[0126] Step S412: determining a target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image.
[0127] The above-mentioned image stitching method obtains the first image and the second image to be stitched. Determine the initial stitching image of the first image and the second image. Input the third sub-overlapping area and the fourth sub-overlapping area into the stitching model to obtain the stitching area corresponding to the third sub-overlapping area and the fourth sub-overlapping area. Based on the stitching area, the third sub-overlapping area and the fourth sub-overlapping area are fused to obtain the initial fused overlapping area. The initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area are input into the fusion model to obtain the fused overlapping area. Based on the fused overlapping area and the initial stitching image, determine the target stitching image corresponding to the first image and the second image. In this method, the first neural network model includes a stitching model and a fusion model. The stitching model is used to search for the stitching area corresponding to the third sub-overlapping area corresponding to the first image and the fourth sub-overlapping area corresponding to the second image after illumination compensation, thereby obtaining the initial fused overlapping area, and then the fused overlapping area is obtained through the fusion model. This method only needs to use the first neural network model to complete the stitching search and fusion processing process, thereby improving the fusion processing efficiency, thereby improving the image stitching processing efficiency.
[0128] The embodiment of the present invention also provides another image stitching method, which is implemented on the basis of the method of the above embodiment; the method focuses on describing the specific process of using the first neural network model to fuse the target overlapping area in the initial stitched image to obtain the fused overlapping area corresponding to the target overlapping area, and the specific process of determining the target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image, in which the target overlapping area includes: the third sub-overlapping area corresponding to the first image and the fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a stitching model; such as Figure 5 As shown, the method comprises the following steps:
[0129] Step S502: acquiring a first image and a second image to be stitched.
[0130] In actual implementation, the first image and the second image are usually RAW images, and the format of the RAW images is usually converted first, that is, the color channels of the first image and the second image are rearranged to facilitate subsequent illumination compensation, seam search and other processing. Generally, the bayer-pattern of the original RAW domain is processed into an RGGB (R: Red, red; G: Green, green; G: Green, green; B: Blue) arrangement, that is, the color channels of the first image and the second image are arranged in RGGB. In this method, the first image and the second image can both be RAW images, and can be directly stitched based on the RAW images, without converting the first image and the second image into RGB images before stitching. Since RAW images do not lack image details compared to RGB images, the RAW image stitching result is more conducive to subsequent image detection, segmentation and other processing.
[0131] Step S504: determining an initial stitched image of the first image and the second image.
[0132] The above RAW images are usually raw data that a CMOS (Complementary Metal-Oxide-Semiconductor) or CCD (Charge Coupled Device) image sensor converts captured light source signals into digital signals. They are unprocessed. Compared with RGB images, RAW images contain complete image details and have greater advantages in post-processing, such as exposure addition and subtraction, highlight / shadow adjustment, contrast increase and decrease, and color level and curve adjustment. In actual implementation, when images need to be stitched, the first image and the second image that need to be acquired first are usually both RAW images.
[0133] Step S506: input the third overlapping sub-region and the fourth overlapping sub-region into the stitching model to obtain the stitching region corresponding to the third overlapping sub-region and the fourth overlapping sub-region.
[0134] Step S508: feathering is performed on the seam area to obtain a feathered overlapping area, and the feathered overlapping area is determined as the fused overlapping area.
[0135] The above-mentioned feathering process usually blurs the edge of the pixel selection area, blends the selected area with the surrounding pixels, that is, blurs the connecting parts inside and outside the pixel selection area, plays a role of gradient, and achieves a natural connection effect. The larger the feathering value, the wider the blurring range, that is, the softer the color transition; the smaller the feathering value, the narrower the blurring range. It can be adjusted according to the actual situation. Usually, the feathering value is set to a smaller value. Repeated feathering is a technique of feathering. In actual implementation, after obtaining the seam area, the fused overlapping area can be obtained based on the blending operation, that is, the fusion operation. Specifically, the Alpha-Blending fusion method can be used to construct a feathering effect on the seam area obtained by the seam search. The feathered pixel value can be selected from 16-22 pixels, and finally the fused overlapping area is obtained. For details, please refer to the relevant technology, the process of fusion using the Alpha-Blending method, which will not be repeated here.
[0136] Step S510: Replace the target overlapping area in the initial stitched image with the fused overlapping area to obtain a target stitched image corresponding to the first image and the second image.
[0137] In actual implementation, since the first image and the second image are usually RAW images, that is, both are images that have not been subjected to lossy compression processing, the above-mentioned target stitched image is usually also a RAW image, that is, after replacing the target overlapping area in the initial stitched image with the fused overlapping area, the stitched RAW image can be obtained, that is, the above-mentioned target stitched image.
[0138] Usually, after obtaining the above-mentioned target stitching image, the target stitching image can be subjected to analysis and processing such as detection, segmentation or recognition to obtain the corresponding analysis result; since the image stitching result is a RAW image, this method can realize the image intelligent analysis of the RAW image. The target stitching image can also be processed by using the core algorithm of ISP (Image Signal Processing) such as Demasaic to obtain the corresponding RGB image stitching result, wherein ISP is mainly used for post-processing the signal output by the front-end image sensor, and its main functions include linear correction, noise removal, bad pixel removal, interpolation, white balance, automatic exposure control, etc. ISP includes multiple core algorithms such as Demasaic. After being processed by relevant algorithms, ISP can output images in the RGB domain to obtain stitching results that are easy to visualize.
[0139] The above-mentioned image stitching method obtains the first image and the second image to be stitched. The third sub-overlapping area and the fourth sub-overlapping area are input into the stitching model to obtain the stitching area corresponding to the third sub-overlapping area and the fourth sub-overlapping area. The stitching area is feathered to obtain the overlapping area after feathering, and the overlapping area after feathering is determined as the fused overlapping area. The fused overlapping area replaces the target overlapping area in the initial stitching image to obtain the target stitching image corresponding to the first image and the second image. In this method, the first neural network model includes a stitching model, and the stitching model is used to search for the stitching area corresponding to the third sub-overlapping area corresponding to the first image and the fourth sub-overlapping area corresponding to the second image after illumination compensation, and then the fused overlapping area is obtained by feathering. This method simplifies the fusion processing process, improves the fusion processing efficiency, and thus improves the processing efficiency of image stitching.
[0140] The embodiment of the present invention further provides a method for illumination compensation, which comprises the following steps:
[0141] Step 602: Acquire a first image and a second image to be stitched.
[0142] Step 604: Perform illumination compensation on the second image based on the second neural network model and the first image.
[0143] The step 604 can be implemented by the following steps 11 to 15:
[0144] Step eleven: determine a projection transformation matrix based on the first image and the second image.
[0145] The step eleven can be specifically implemented by the following steps M to O:
[0146] Step M: extracting at least one first feature point in the first image and at least one second feature point in the second image.
[0147] Step N: determining at least one matching feature point pair based on at least one first feature point and at least one second feature point.
[0148] Step O: determining a projection transformation matrix based on at least one matching feature point pair.
[0149] Step twelve: determine the initial overlapping area of the first image and the second image based on the projection transformation matrix; wherein the initial overlapping area includes: a first sub-overlapping area corresponding to the first image and a second sub-overlapping area corresponding to the second image.
[0150] The step twelve can be specifically implemented by the following steps P to S:
[0151] Step P, obtaining the boundary coordinates of the second image; wherein the boundary coordinates are used to indicate the image area of the second image.
[0152] Step Q: determining the boundary coordinates after the projection transformation based on the projection transformation matrix and the boundary coordinates of the second image.
[0153] Step R, determining a second image after projection transformation based on the boundary coordinates after projection transformation.
[0154] Step S: determining the overlapping image area of the second image after the projection transformation and the first image as the initial overlapping area.
[0155] Step thirteen, input the first sub-overlapping area and the second sub-overlapping area into the second neural network model, and determine the mapping relationship between the first pixel value of each pixel in the first sub-overlapping area and the second pixel value of the pixel at the same position in the second sub-overlapping area through the second neural network model.
[0156] Step fourteen, obtaining the mapping relationship output by the second neural network model.
[0157] Step fifteen, for each color channel, based on the mapping relationship, the pixel value of each pixel point in the second image in the color channel is matched with the pixel value of each pixel point in the first image in the color channel to perform illumination compensation on the second image.
[0158] The step 604 can also be implemented by the following step 20:
[0159] Step 20, input the first image and the second image into the second neural network model, and perform illumination compensation on the second image based on the first image through the second neural network to obtain the illumination-compensated second image. In this embodiment, the specific implementation of each step can refer to the relevant description in the above embodiment, which will not be repeated here.
[0160] The above-mentioned illumination compensation method, after acquiring the first image and the second image to be stitched, performs illumination compensation on the second image based on the second neural network model and the first image. In this method, a pre-trained neural network model is used to perform illumination compensation on the second image, and there is no need to calculate each pixel based on the CPU, which saves the time of compensating all pixels in the image and improves the processing efficiency of illumination compensation.
[0161] To further understand the above embodiments, the following is provided: Figure 6 A flow chart of a RAW domain image stitching method is shown. For the convenience of explanation, two images are taken as an example for introduction. The two images are represented by a tar image (corresponding to the first image mentioned above) and a src image (corresponding to the second image mentioned above) respectively. Multiple images can also be stitched in this way. First, the original RAW domain bayer-pattern image is processed into an RGGB arrangement, that is, the bayer-tar is formatted and converted into an RGGB-tar image, and the bayer-src is formatted and converted into an RGGB-src image.
[0162] Next, the rggb-tar image and the rggb-src image are spatially registered. The specific process of spatial registration can be: first, feature points of the two images are extracted. The feature points can be extracted by traditional SIFT or SuperPoint based on neural network, and then the feature points are matched to obtain matching feature point pairs. Based on the feature point matching information of the matching feature point pairs, the projection transformation matrix is calculated. Based on the projection transformation matrix and the boundary coordinates of the rggb-src image, the projected area information of the rggb-src is calculated, and the overlapping area (corresponding to the above-mentioned initial overlapping area) can be obtained by intersecting with the rggb-tar image.
[0163] Then, based on the pre-trained neural network model, the rggb-src image can be illuminated and compensated. The color consistency of images from different cameras can be optimized through illumination compensation. The following methods can be used: one method is to input two small images of the overlapping area of the above-mentioned rggb-tar image and the rggb-src image into the neural network model, and learn with GT as the LUT table of 0-65535. Finally, four LUT tables of 0-65535 can be calculated through the neural network model, and illumination compensation is performed on the rggb-src image based on the LUT table; another method is to input the rggb-tar image and the rggb-src image into the neural network model, and supervise with GT as the original rggb-tar image and the rggb-src image after illumination compensation based on histogram matching. Finally, the rggb-src image after illumination compensation can be calculated through the neural network model.
[0164] After obtaining the RGGB-SRC image after illumination compensation, projection change processing can be performed on the RGGB-TAR image and the RGGB-SRC image after illumination compensation. Specifically, based on the projection transformation matrix, the RGGB-SRC image after illumination compensation is projected onto the RGGB-TAR image to obtain an initial stitched image; based on the RGGB-SRC image after illumination compensation, a new overlapping area is re-extracted, and stitching search and blending are performed based on two small images of the new overlapping area of the RGGB-TAR image and the RGGB-SRC image. Specifically, the two small images of the new overlapping area can be input into a pre-trained stitching model, and through the stitching model, a stitching mask or a stitching point set (corresponding to the above-mentioned stitching area) is output, and then a fused overlapping area (corresponding to the above-mentioned fused overlapping area) is obtained by Alpha-blending fusion or by fusion based on a pre-trained fusion model. Finally, the overlapping area replacement process is performed to replace the overlapping area in the initial stitched image with the fused overlapping area to obtain the RAW domain stitching result; optionally, after obtaining the RAW domain stitching result, the RAW domain image intelligent analysis can be performed on the RAW domain stitching result, and the ISP steps such as Demasaic can be used to obtain the RGB visualization result.
[0165] The processes of illumination compensation, seam search and fusion in the above-mentioned RAW domain image stitching method can be implemented based on a neural network, thereby achieving hardware acceleration, accelerating processing speed and improving processing efficiency. This method provides a large field-of-view solution for intelligent analysis of RAW domain images, and can also handle steps such as seam connection and illumination alignment for RAW domain images to fuse the images to be stitched with the highest possible quality. By stitching RAW domain images, original RAW domain images with a large field-of-view can be obtained, which is convenient for tasks such as dark light RAW domain detection and recognition under a large field-of-view.
[0166] The embodiment of the present invention also provides a structural schematic diagram of an image stitching device, such as Figure 7 As shown, the device includes: an acquisition module 70, used to acquire a first image and a second image to be stitched; a first determination module 71, used to determine an initial stitched image of the first image and the second image; a fusion module 72, used to use a first neural network model to perform a fusion process on a target overlapping area in the initial stitched image to obtain a fused overlapping area corresponding to the target overlapping area; a second determination module 73, used to determine the target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image.
[0167] The above-mentioned image stitching device first obtains the first image and the second image to be stitched; determines the initial stitching image of the first image and the second image; then uses the first neural network model to fuse the target overlapping area in the initial stitching image to obtain the fused overlapping area corresponding to the target overlapping area; finally, based on the fused overlapping area and the initial stitching image, determines the target stitching image corresponding to the first image and the second image. In the device, after determining the initial stitching image of the first image and the second image, the first neural network model is used to fuse the target overlapping area in the initial stitching image to obtain the corresponding fused overlapping area. It is not necessary to perform fusion-related calculations on each pixel in the initial stitching image based on the CPU, which saves the time of performing fusion calculations on all pixels in the initial stitching image, improves the efficiency of fusion processing, and thus improves the processing efficiency of image stitching.
[0168] Furthermore, the device is also used to: perform illumination compensation on the second image; correspondingly, the first determination module is also used to: determine an initial spliced image based on the first image and the illumination-compensated second image.
[0169] Furthermore, the first determination module is also used to: perform illumination compensation on the second image based on the second neural network model and the first image.
[0170] Furthermore, the first determination module is also used to: determine the projection transformation matrix based on the first image and the second image; determine the initial overlapping area of the first image and the second image based on the projection transformation matrix; wherein the initial overlapping area includes: a first sub-overlapping area corresponding to the first image and a second sub-overlapping area corresponding to the second image; input the first sub-overlapping area and the second sub-overlapping area into the second neural network model, and determine the mapping relationship between the first pixel value of each pixel in the first sub-overlapping area and the second pixel value of the pixel at the same position in the second sub-overlapping area through the second neural network model; obtain the mapping relationship output by the second neural network model; for each color channel, based on the mapping relationship, match the pixel value of each pixel point in the second image in the color channel with the pixel value of each pixel point in the first image in the color channel to perform illumination compensation for the second image.
[0171] Furthermore, the first determination module is also used to: obtain the boundary coordinates of the second image; wherein the boundary coordinates are used to indicate the image area of the second image; determine the boundary coordinates after the projection transformation based on the projection transformation matrix and the boundary coordinates of the second image; determine the second image after the projection transformation based on the boundary coordinates after the projection transformation; and determine the overlapping image area of the second image after the projection transformation and the first image as the initial overlapping area.
[0172] Furthermore, the first determination module is also used to: extract at least one first feature point in the first image and at least one second feature point in the second image; determine at least one matching feature point pair based on at least one first feature point and at least one second feature point; and determine a projection transformation matrix based on at least one matching feature point pair.
[0173] Furthermore, the first determination module is also used to: input the first image and the second image into a second neural network model, and perform illumination compensation on the second image based on the first image through the second neural network to obtain an illumination-compensated second image.
[0174] Furthermore, the target overlapping area includes: a third sub-overlapping area corresponding to the first image and a fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a stitching model and a fusion model; the fusion module is also used to: input the third sub-overlapping area and the fourth sub-overlapping area into the stitching model, search for the stitching between the third sub-overlapping area and the fourth sub-overlapping area through the stitching model, and obtain the stitching area corresponding to the third sub-overlapping area and the fourth sub-overlapping area; based on the stitching area, fuse the third sub-overlapping area and the fourth sub-overlapping area to obtain an initial fused overlapping area; input the initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area into the fusion model, and fuse the initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area through the fusion model to obtain a fused overlapping area.
[0175] Furthermore, the target overlapping area includes: a third sub-overlapping area corresponding to the first image and a fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a stitching model; the fusion module is also used to: input the third sub-overlapping area and the fourth sub-overlapping area into the stitching model to obtain the stitching area corresponding to the third sub-overlapping area and the fourth sub-overlapping area; feather the stitching area to obtain the feathered overlapping area, and determine the feathered overlapping area as the fused overlapping area.
[0176] Furthermore, the fusion module is also used to: obtain a first image; translate and / or rotate the first image to obtain a second image; fuse the first image and the second image to obtain an initial fused image; and train a fusion model based on the first image, the second image and the initial fused image.
[0177] Furthermore, the second determining module is further used to: replace the target overlapping area in the initial stitched image with the fused overlapping area to obtain the target stitched image corresponding to the first image and the second image.
[0178] Furthermore, the color channels of the first image and the second image are arranged in RGGB.
[0179] The image stitching device provided in the embodiment of the present invention has the same implementation principle and technical effects as those of the aforementioned image stitching method embodiment. For the sake of brief description, for matters not mentioned in the embodiment of the image stitching device, reference may be made to the corresponding contents in the aforementioned image stitching method embodiment.
[0180] An embodiment of the present invention further provides an electronic device, including a processing device and a storage device, wherein the storage device stores a computer program, and the computer program executes any of the above-mentioned image stitching methods when the processing device is running.
[0181] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the electronic device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0182] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processing device, the steps of the above-mentioned image stitching method are executed.
[0183] The computer program product of the image stitching method, device and electronic device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiments. The specific implementation can be referred to the method embodiments, which will not be repeated here.
[0184] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described equipment and / or device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0185] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0186] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0187] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0188] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image stitching method, characterized in that: The method comprises: Acquire a first image and a second image to be stitched; Determining an initial stitched image of the first image and the second image; Using the first neural network model to perform fusion processing on the target overlapping area in the initial spliced image to obtain a fused overlapping area corresponding to the target overlapping area; Determine a target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image; Before determining the initial stitched image of the first image and the second image, the method further includes: performing illumination compensation on the second image; Accordingly, determining an initial stitched image of the first image and the second image includes: Determining the initial stitched image based on the first image and the second image after illumination compensation; The step of performing illumination compensation on the second image comprises: Based on a second neural network model and the first image, performing illumination compensation on the second image; the second neural network model is the same as or different from the first neural network model; The performing illumination compensation on the second image based on the second neural network model and the first image includes: Determining a projection transformation matrix based on the first image and the second image; Based on the projection transformation matrix, determining an initial overlapping area of the first image and the second image; wherein the initial overlapping area includes: a first sub-overlapping area corresponding to the first image and a second sub-overlapping area corresponding to the second image; Inputting the first sub-overlapping area and the second sub-overlapping area into the second neural network model, and determining the mapping relationship between the first pixel value of each pixel in the first sub-overlapping area and the second pixel value of the pixel at the same position in the second sub-overlapping area through the second neural network model; Obtaining the mapping relationship output by the second neural network model; For each color channel, based on the mapping relationship, the pixel value of each pixel point in the second image in the color channel is matched with the pixel value of each pixel point in the first image in the color channel to perform illumination compensation on the second image.
2. The method according to claim 1, characterized in that The step of determining an initial overlapping area of the first image and the second image based on the projection transformation matrix comprises: Acquire boundary coordinates of the second image; wherein the boundary coordinates are used to indicate an image area of the second image; Determining boundary coordinates after projection transformation based on the projection transformation matrix and boundary coordinates of the second image; Determining a second image after projection transformation based on the boundary coordinates after projection transformation; An overlapping image area between the second image after the projective transformation and the first image is determined as the initial overlapping area.
3. The method according to claim 1, characterized in that The step of determining a projection transformation matrix based on the first image and the second image comprises: extracting at least one first feature point in the first image and at least one second feature point in the second image; Based on the at least one first feature point and the at least one second feature point, determining at least one matching feature point pair; Based on the at least one matching feature point pair, a projection transformation matrix is determined.
4. The method according to claim 1, characterized in that The step of performing illumination compensation on the second image based on the second neural network model and the first image comprises: The first image and the second image are input into the second neural network model, and the second neural network performs illumination compensation on the second image based on the first image to obtain the illumination-compensated second image.
5. The method according to claim 1, characterized in that The target overlapping area includes: a third sub-overlapping area corresponding to the first image and a fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a splicing model and a fusion model; the step of fusing the target overlapping area in the initial spliced image using the first neural network model to obtain a fused overlapping area corresponding to the target overlapping area includes: Inputting the third sub-overlapping area and the fourth sub-overlapping area into the stitching model, searching for the seam between the third sub-overlapping area and the fourth sub-overlapping area through the stitching model, and obtaining the seam area corresponding to the third sub-overlapping area and the fourth sub-overlapping area; Based on the stitching area, the third overlapping sub-area and the fourth overlapping sub-area are merged to obtain an initial merged overlapping area; The initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area are input into the fusion model, and the initial fused overlapping area, the third sub-overlapping area, and the fourth sub-overlapping area are fused by the fusion model to obtain the fused overlapping area.
6. The method according to claim 1, characterized in that The target overlapping area includes: a third sub-overlapping area corresponding to the first image and a fourth sub-overlapping area corresponding to the second image after illumination compensation; the first neural network model includes a splicing model; the step of fusing the target overlapping area in the initial spliced image using the first neural network model to obtain a fused overlapping area corresponding to the target overlapping area includes: Inputting the third overlapping sub-region and the fourth overlapping sub-region into the stitching model to obtain a stitching region corresponding to the third overlapping sub-region and the fourth overlapping sub-region; The seam area is subjected to feathering processing to obtain an overlapping area after feathering, and the overlapping area after feathering is determined as the fused overlapping area.
7. The method according to claim 5, characterized in that The fusion model is trained in the following way: Get the first picture; Performing translation and / or rotation processing on the first picture to obtain a second picture; Performing fusion processing on the first picture and the second picture to obtain an initial fused picture; The fusion model is trained based on the first picture, the second picture and the initial fusion picture.
8. The method according to claim 1, characterized in that The step of determining a target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image comprises: The target overlapping area in the initial stitched image is replaced by the fused overlapping area to obtain the target stitched image corresponding to the first image and the second image.
9. The method according to claim 1, characterized in that: The color channels of the first image and the second image are arranged in an RGGB arrangement.
10. An image stitching device, characterized in that: The device comprises: An acquisition module, used for acquiring a first image and a second image to be stitched; A first determining module, used to determine an initial stitched image of the first image and the second image; A fusion module, used for performing fusion processing on the target overlapping area in the initial spliced image by using the first neural network model to obtain a fused overlapping area corresponding to the target overlapping area; A second determination module, configured to determine a target stitched image corresponding to the first image and the second image based on the fused overlapping area and the initial stitched image; The device is also used for: performing illumination compensation on the second image; The first determining module is further used for: Determining the initial stitched image based on the first image and the second image after illumination compensation; The device is also used for: Based on a second neural network model and the first image, performing illumination compensation on the second image; the second neural network model is the same as or different from the first neural network model; The device is also used for: Determining a projection transformation matrix based on the first image and the second image; Based on the projection transformation matrix, determining an initial overlapping area of the first image and the second image; wherein the initial overlapping area includes: a first sub-overlapping area corresponding to the first image and a second sub-overlapping area corresponding to the second image; Inputting the first sub-overlapping area and the second sub-overlapping area into the second neural network model, and determining the mapping relationship between the first pixel value of each pixel in the first sub-overlapping area and the second pixel value of the pixel at the same position in the second sub-overlapping area through the second neural network model; Obtaining the mapping relationship output by the second neural network model; For each color channel, based on the mapping relationship, the pixel value of each pixel point in the second image in the color channel is matched with the pixel value of each pixel point in the first image in the color channel to perform illumination compensation on the second image.
11. An electronic device, characterized in that: The method comprises a processing device and a storage device, wherein the storage device stores a computer program, and when the computer program is run by the processing device, the image stitching method according to any one of claims 1 to 9 is executed.
12. A machine-readable storage medium, characterized in that: The machine-readable storage medium stores a computer program, and when the computer program is executed by a processing device, the image stitching method according to any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
Image stitching processing method, device and electronic system
CN112862685A