Intelligent panoramic photography splicing system and method for multi-modal image fusion

CN120807274APending Publication Date: 2025-10-17LISHUI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510832821.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-17

Smart Images

  • Figure CN120807274A_ABST
    Figure CN120807274A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent panoramic photography splicing system and method for multi-modal image fusion. According to the system, image data of different modals are acquired through the multi-modal image acquisition module, and after the image data are subjected to noise reduction, distortion removal and other processing through the image preprocessing module, the corresponding relation between images is established through the feature extraction and matching module. And the multi-modal image fusion module performs weighted average fusion calculation according to the feature matching result to generate a fused image. And the panoramic photography splicing module adopts an optimization-based splicing algorithm to generate a high-quality panoramic image. And the intelligent control module automatically adjusts system parameters according to a shooting environment and user requirements, so that an intelligent panoramic photography splicing process is realized. The method effectively improves the quality and information richness of the panoramic image, and is suitable for various application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of photographic imaging technology, in particular to a system and method capable of intelligent fusion processing of multi-modal images and realizing panoramic photography splicing. BACKGROUND

[0002] With the continuous development of photography technology, panoramic photography has been widely applied in many fields such as tourism, real estate, virtual reality, etc. The traditional panoramic photography splicing technology mainly processes single modal images (such as ordinary visible light images), and often faces many problems in the splicing process. For example, when the light conditions of the shooting environment are complex, the ordinary visible light images may appear overexposed or underexposed, resulting in poor quality of the spliced panoramic image; in low light environment, the noise of the image will increase significantly, affecting the splicing effect and image clarity.

[0003] At the same time, in some special scenarios, a single visible light image cannot meet the actual needs. For example, in the field of security monitoring, it is difficult to clearly obtain target information at night or in smoky environment relying on visible light images alone; in the field of geological exploration, it is necessary to combine images of different wavebands (such as infrared images, ultraviolet images, etc.) to obtain more comprehensive target features. However, the existing panoramic photography splicing system rarely involves fusion processing of multi-modal images, cannot fully utilize the advantages of different modal images, and is difficult to realize high-quality and high-information-richness panoramic image splicing. Therefore, there is an urgent need for a system and method capable of effectively fusing multi-modal images and realizing intelligent panoramic photography splicing to meet the diversified application needs.

[0004] CN108288932A (a single modal panoramic splicing system) only processes visible light images and does not involve multi-modal fusion, and the noise is significant (such as ISO 1600, noise PSNR = 22dB) in low light environment, which is improved to 28dB by infrared-visible light fusion in the present application.

[0005] US20190372681A (a multispectral image splicing system) uses fixed weight fusion and cannot adapt to environmental changes, and the splicing seam is obvious (such as error ≥3 pixels in the bright-dark boundary area) in complex lighting, which is reduced to ≤0.8 pixels by deep learning dynamic weight adjustment in the present application. SUMMARY

[0006] (I) Invention purpose

[0007] The purpose of the present application is to provide a multi-modal image fusion intelligent panoramic photography stitching system and method, which improves the quality and information richness of panoramic images by intelligently processing and fusing multi-modal images, and solves the problems of poor panoramic photography stitching effect in complex environments and the inability to effectively utilize the advantages of multi-modal images in the prior art.

[0008] (II) Technical solutions

[0009] The multi-modal image fusion intelligent panoramic photography stitching system of the present application comprises the following modules:

[0010] A multi-modal image acquisition module is used to acquire image data of different modalities, including but not limited to visible light images, infrared images, and ultraviolet images. This module can use multiple different types of sensors, such as visible light camera sensors, infrared thermal imager sensors, and ultraviolet imaging sensors, and each sensor is integrated on a rotatable or movable device to obtain multi-modal images at different angles and positions.

[0011] An image preprocessing module is used to preprocess the acquired multi-modal images, including but not limited to noise reduction processing, distortion removal processing, brightness and contrast adjustment. Different preprocessing algorithms are used according to the characteristics of different modalities. For example, for infrared images, a special infrared image noise reduction algorithm is used to remove interference caused by thermal noise and other factors; for visible light images, a distortion correction algorithm based on deep learning is used to correct lens distortion.

[0012] A feature extraction and matching module is used to extract features from the preprocessed multi-modal images using a feature extraction algorithm to obtain feature points and feature descriptors of the images. The feature extraction algorithm includes but is not limited to the Scale-Invariant Feature Transform (SIFT) algorithm, the Speeded-Up Robust Features (SURF) algorithm, and the Oriented FAST and Rotated BRIEF (ORB) algorithm. Through a feature matching algorithm, the feature points of different modal images are matched to establish a correspondence between the images. In the feature matching process, a deep learning-based matching model is introduced to improve the accuracy and efficiency of matching and reduce the occurrence of false matches.

[0013] A multi-modal image fusion module is used to fuse the multi-modal images according to the feature matching results. A weighted average-based fusion algorithm is used to assign different weights to each modality image according to the quality evaluation results of different modal images in different regions for fusion calculation. Meanwhile, a deep learning-based fusion model is introduced to learn the fusion rules of different modal images in different scenarios to improve the fusion effect. For example, in low-light environments, the weight of infrared images is increased to make the fused images display the target more clearly; in normal lighting environments, the weights are reasonably assigned according to the characteristics of visible light images and infrared images to retain more detailed information.

[0014] Panorama stitching module: using the feature matching relationship of the images and the fused images, a panorama stitching algorithm is used for stitching processing to generate a panoramic image. The panorama stitching algorithm includes but is not limited to a graph cut-based stitching algorithm and an optimization-based stitching algorithm. In the stitching process, the stitching seam is adjusted by an optimization algorithm to reduce the stitching marks and improve the visual effect of the panoramic image.

[0015] Intelligent control module: used to control the workflow and parameter settings of each module. According to the shooting environment and user needs, the acquisition parameters of the multi-modal image acquisition module (such as shooting angle, shooting interval, sensor mode, etc.) and the algorithm parameters of the image preprocessing module, multi-modal image fusion module and panorama stitching module are automatically adjusted to realize an intelligent panorama stitching process.

[0016] Based on the above system, the present application also provides an intelligent panorama stitching method for multi-modal image fusion, and the specific steps are as follows:

[0017] Image acquisition: through the multi-modal image acquisition module, image data of different modalities, angles and positions are collected.

[0018] Preprocessing: the collected image data is transmitted to the image preprocessing module for preprocessing operations such as noise reduction, distortion removal, brightness and contrast adjustment, etc.

[0019] Feature extraction and matching: using the feature extraction and matching module, the preprocessed images are subjected to feature extraction and matching to establish the corresponding relationship between the images.

[0020] Image fusion: according to the feature matching results, the multi-modal images are fused in the multi-modal image fusion module to generate fused images.

[0021] Panorama stitching: the fused images are input into the panorama stitching module, and a panorama stitching algorithm is used for stitching to generate a panoramic image.

[0022] Intelligent control: through the intelligent control module, the workflow of the entire system and the parameters of each module are intelligently adjusted and optimized according to the shooting environment and user needs.

[0023] (Three) beneficial effects

[0024] Improved image quality: through a series of processes such as preprocessing, feature extraction and matching, fusion and stitching of multi-modal images, the problem of poor image quality in traditional panoramic photography in complex environments can be effectively solved. For example, in low light environments, combining infrared images and visible light images for fusion can significantly improve image clarity and detail performance, reducing noise and exposure problems.

[0025] Rich information content: make full use of the advantages of different modal images to obtain more comprehensive target features. In security monitoring, visible light images and infrared images can be used simultaneously, not only can clearly display the appearance of the target in the daytime, but also can discover hidden heat source targets at night through infrared images; in geological exploration, combining ultraviolet images and visible light images can help to discover some special geological features and mineral compositions, providing more rich information for decision-making.

[0026] High degree of intelligence: the intelligent control module can automatically adjust system parameters and workflow according to the shooting environment and user needs, without the need for users to manually make complex settings, improving the usability and adaptability of the system, suitable for various shooting scenes and application needs.

[0027] Good splicing effect: advanced feature matching algorithm and panoramic splicing algorithm are adopted, and the splicing seam is optimized, which can effectively reduce the splicing marks, generate high-quality seamless panoramic images, and improve the user's visual experience.

[0028] Multi-modal adaptive fusion mechanism: through the dynamic weighting strategy combining quality evaluation and deep learning, the problem that the fixed weight in the prior art cannot adapt to complex scenes is solved, which is the first.

[0029] Intelligent parameter optimization closed loop: the intelligent control module monitors the environment in real time and adjusts the whole link parameters, forming a "perception-decision-execution" closed loop, improving the robustness of the system, and the prior art is mostly manual parameter setting.

[0030] Cross-modal feature matching enhancement: introduce twin network and RANSAC cascade matching to solve the matching problem of large feature difference of multi-modal images (such as visible light-infrared), the matching accuracy is improved by 23% compared with traditional SIFT. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 : System architecture diagram of the present application

[0032] Figure 2 : Method flow chart of the present application DETAILED DESCRIPTION

[0033] The present application will be further described in detail below in combination with specific embodiments.

[0034] (I) Hardware settings

[0035] The visible light camera, the infrared thermal imager, and the ultraviolet imager are integrated on a 360-degree rotatable holder device, and the holder device is connected with a data processing equipment (such as a high-performance computer) through a data line. The intelligent panoramic photography splicing system software for multi-modal image fusion of the application is installed on the data processing equipment (GPU model: NVIDIA A100)

[0036] (ii) System workflow embodiment

[0037] Image acquisition stage: the user sets the shooting parameters such as the shooting range and the shooting interval through the operation interface. The intelligent control module controls the holder device to rotate according to the user settings, and starts each sensor in the multi-modal image acquisition module to collect image data of different modalities and different angles according to the set parameters, and transmits the data to the data processing equipment.

[0038] Sensor integration details: the holder rotation accuracy (such as ±0.1°), the sensor synchronous triggering mechanism (such as hardware clock synchronization, error ≤1ms), and the field of view angle matching scheme of different modal sensors (for example, when the visible light camera FOV is 75° and the infrared thermal imager FOV is 60°, the view angles are unified through overlapping area calibration) are determined.

[0039] Dynamic adjustment logic of acquisition parameters: for example, the intelligent control module automatically switches to infrared image as the main one when the ambient light intensity is less than or equal to 50 lux, and adjusts the visible light camera ISO to 3200 and the shutter speed to 1 / 30s, and the infrared thermal imager frame rate to 25fps.

[0040] Preprocessing stage: after receiving the image data, the image preprocessing module automatically identifies the modal type of the image. For visible light images, a lens distortion correction network model based on deep learning is used for lens distortion correction, a non-local mean denoising algorithm is used for denoising processing, and the brightness and contrast are adjusted according to the histogram information of the image; for infrared images, a denoising algorithm based on median filtering is used to remove thermal noise, and an adaptive histogram equalization method is used to enhance the contrast of the image.

[0041] Denoising algorithm specific parameters:

[0042] Visible light image: non-local mean (NLM) denoising is used, the search window size is 11x11, the similar window is 5x5, and the Gaussian weight standard deviation σ is 10.

[0043] Infrared image: the median filter window size is 3x3, combined with time domain recursive filtering (α=0.7) to eliminate dynamic thermal noise.

[0044] Distortion correction algorithm implementation: distortion correction network based on deep learning (such as U-Net architecture), training data contains 100,000 sets of labeled distortion-correction image pairs, using smooth L1 loss function, the residual radial distortion after correction is less than or equal to 0.5 pixels.

[0045] Feature extraction and matching stage: the feature extraction and matching module uses ORB algorithm to extract features from pre-processed multi-modal images, and obtains feature points and descriptors of the images. Then, a deep learning-based feature matching network is used to match the feature points of different modal images, and by calculating the similarity between the feature descriptors, the correct matching point pairs are selected to establish the correspondence between the images.

[0046] Deep learning matching model architecture: using Siamese Network, the input is a 128x128 pixel feature block, the convolutional layer contains 4 residual blocks, and the output is a 128-dimensional feature vector. When training, use Triplet Loss function, and set the positive and negative sample distance threshold to 0.6.

[0047] Matching screening strategy: use RANSAC algorithm to remove false matching points, iteration times 200, inlier threshold set to 3 pixels, matching accuracy improved to more than 98.5%.

[0048] Image fusion stage: the multi-modal image fusion module first performs quality assessment on different modal images based on the feature matching results. For each pixel region, calculate the clarity, contrast and other quality indicators of the visible light image and infrared image. According to the quality assessment results, use a deep learning-based fusion model to assign weights to each modal image, perform weighted average fusion calculation, and generate the fused image.

[0049] Weighted fusion formula:

[0050] F(x,y) = ∑i=1nwi(x,y)w1(x,y)·I1(x,y) + w2(x,y)·I2(x,y) + … + wn(x,y)·In(x,y)

[0051] Where wi(x,y) is the weight of the i-th modal image at (x,y), calculated from the quality assessment indicators (clarity Qi, contrast Ci):

[0052] wi(x,y) = ∑j=1nQj(x,y)·Cj(x,y)Qi(x,y)·Ci(x,y)

[0053] Deep learning fusion model training: using cross-modal dataset (such as MSCOCO infrared-visible light pair), using encoder-decoder architecture, loss function contains pixel-level MSE (weight 0.7) and perception loss (VGG16 feature distance, weight 0.3), after training for 200 rounds, the fusion image PSNR is improved by 3-5dB.

[0054] Panoramic stitching stage: the panoramic photography stitching module uses the feature matching relationship of the image, adopts an optimized panoramic stitching algorithm, and performs stitching processing on the fused image. In the stitching process, the position of the stitching seam and the fusion mode are adjusted through the energy function optimization algorithm, so that the stitched panoramic image is more natural and seamless in vision.

[0055] Optimization algorithm details: based on the stitching seam optimization of graph cut, the energy function is defined as:

[0056] E=λ1·Ecolor+λ2·Egradient+λ3·Esmooth

[0057] Wherein, Ecolor is a color difference item, Egradient is a gradient consistency item, Esmooth is a smoothing item, weight parameters λ1=0.5, λ2=0.3, λ3=0.2.

[0058] Stitching error quantification: the geometric error of the stitched panoramic image is ≤1 pixel, the brightness difference is ≤5%, and the subjective score (5-point system) is 4.8 on average.

[0059] Intelligent control stage: in the whole process, the intelligent control module monitors the light conditions, temperature and other information of the shooting environment, and the operation instructions of the user in real time. According to the environmental changes and user needs, the shooting parameters of the multi-modal image acquisition module are automatically adjusted, such as increasing the acquisition frequency and exposure time of infrared images when the light becomes dark; at the same time, the algorithm parameters of the image preprocessing module, multi-modal image fusion module and panoramic photography stitching module are dynamically adjusted to ensure the best shooting and stitching effect.

[0060] Technical effect comparison table

[0061]

[0062] Application of the application:

[0063] In the field of security monitoring: In a large city's intelligent security project, the multi-modal image fusion intelligent panoramic photography stitching system is applied. In the city's important transportation hubs, commercial centers and other areas with large crowds and complex environments, multi-modal image acquisition equipment integrating visible light cameras and infrared thermal imagers is deployed. The system collects real-time images of different modalities, and through the multi-modal image fusion module, it fuses the details of visible light images with the temperature information of infrared images. At night or in dim light environments, when there are people or vehicle activities, the system can accurately identify targets, even if the target is in the shadow area, it can clearly capture its outline and moving track through infrared images. The panoramic image generated by the panoramic photography stitching module provides a wide, dead-angle-free monitoring view for security personnel, facilitating the timely detection of abnormal situations and effectively improving the intelligent level and monitoring efficiency of urban security.

[0064] Security monitoring scenario (data)

[0065] Test environment: night (illumination ≤ 10 lux), with moving targets of pedestrians and vehicles.

[0066] Comparison results:

[0067]

[0068] In the tourism industry: A well-known tourist attraction introduced the system to enhance the tourist experience. The multi-modal image acquisition module was used to collect natural scenery and cultural landscapes in the scenic area. Not only did it capture the beauty of the visible light images, but it also captured some plants and rock textures with special fluorescent effects through ultraviolet images. After image preprocessing, feature extraction and matching, the multi-modal image fusion module generated fusion images that highlighted the unique features of different landscapes. The panoramic photography stitching module stitched these fusion images into panoramic images, which could be viewed by tourists through the official APP of the scenic area, providing a more immersive and immersive experience. At the same time, it also provided unique and novel materials for the promotion of the scenic area.

[0069] Industrial detection scenario: In a large automobile manufacturing factory, the system is applied to the quality detection of automobile parts. The multi-modal image acquisition module captures the automobile parts on the production line to obtain visible light images to detect surface appearance defects of the parts, such as scratches, bumps, etc., and obtains X-ray images to detect internal structural defects of the parts, such as cracks, bubbles, etc. The image preprocessing module processes the collected images to improve image clarity, such as denoising and enhancement. Through the feature extraction and matching module, the correspondence of the part features in different modal images is determined, and the multi-modal image fusion module fuses the two images so that the detection personnel can observe the surface and internal conditions of the parts at the same time. The panoramic photography stitching module stitches the images of multiple parts into a production line panoramic image, which helps managers to comprehensively grasp the quality status of the production line, timely find problems and adjust the production process, greatly improving the accuracy and efficiency of industrial detection, and reducing the rate of defective products.

[0070] X-ray image and visible light fusion process:

[0071] X-ray image resolution 1200x1600, minimum internal crack detection size 0.1mm;

[0072] Visible light image resolution 4096x3072, surface defect detection accuracy 0.05mm;

[0073] After fusion, the defect recognition rate is improved to 99.2%, which is 18.7% higher than that of single mode.

[0074] Cultural heritage protection and archaeological research field: When protecting and repairing a historical ancient building, the system is used to record the current situation of the ancient building. The multi-modal image acquisition module obtains visible light images to present the appearance and color information of the ancient building, and infrared images to detect the internal structure of the ancient building wall, such as whether there are hidden dangers such as hollowing and cracking. After a series of processing and fusion, the generated panoramic image provides comprehensive and detailed ancient building data for cultural heritage protection experts, which helps to develop scientific and reasonable protection and repair schemes. In the archaeological excavation site, the system combines the soil and stratum information reflected by different waveband images through multi-modal image acquisition, and stitches the overall situation of the excavation area through panoramic photography, helping archaeologists better understand the layout of the site, accurately determine the location of cultural relics, and improve the scientificity and accuracy of archaeological work.

[0075] The above examples are only used to illustrate the technical solutions of the present application, but not to limit it; although the present application has been described in detail with reference to the above examples, those skilled in the art should understand that they can modify the technical solutions recorded in the above examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An intelligent panoramic photography stitching system for multimodal image fusion, characterized in that: include: Multimodal image acquisition module, integrating visible light camera, infrared thermal imager, ultraviolet imager, and multi-angle acquisition through 360° rotating pan-tilt, and sensor synchronous trigger error ≤1ms; Image preprocessing module, including a deep learning-based dedistortion network (U-Net architecture) and a modality-adaptive denoising algorithm, which can correct radial distortion of visible light images with a residual error of ≤0.5 pixels; The feature extraction and matching module uses a twin neural network and RANSAC cascade matching, with a matching accuracy of ≥ 98.5%; Multimodal image fusion module, based on a dynamic weighted fusion formula for quality assessment, combined with an encoder-decoder network to achieve cross-modal feature fusion; The panoramic photography stitching module uses a graph cut optimization stitching algorithm, and the energy function includes color difference, gradient consistency, and smoothness terms (weight ratio 0.5:0.3:0.2); The intelligent control module automatically switches to infrared imaging based on the ambient light intensity (≤50lux), adjusts the acquisition parameters and optimizes the algorithm parameters of each module.

2. The intelligent panoramic photography stitching system for multimodal image fusion according to claim 1, characterized in that: The weighted fusion formula: F(x,y)=∑i=1nwi(x,y)w1(x,y)·I1(x,y)+w2(x,y)·I2(x,y)+…+wn(x,y)·In(x,y) Where wi(x,y) is the weight of the i-th modality image at (x,y), which is calculated by the quality assessment indicators (clarity Qi, contrast Ci): wi(x,y)=∑j=1nQj(x,y)·Cj(x,y)Qi(x,y)·Ci(x,y).

3. The intelligent panoramic photography stitching system for multimodal image fusion according to claim 1, characterized in that: The energy function of the splicing algorithm is defined as: E=λ1·Ecolor+λ2·Egradient+λ3·Esmooth Among them, Ecolor is the color difference term, Egradient is the gradient consistency term, Esmooth is the smoothing term, and the weight parameters λ1 = 0.5, λ2 = 0.3, λ3 = 0.2; Quantification of stitching error: The geometric error of the stitched panoramic image is ≤1 pixel, the brightness difference is ≤5%, and the average subjective score (out of 5) is 4.8 points.

4. An intelligent panoramic photography stitching method based on multimodal image fusion, comprising the following specific steps: Image acquisition: Through the multimodal image acquisition module, image data of different modalities, angles and positions are collected; Preprocessing: The collected image data is transmitted to the image preprocessing module for preprocessing operations such as noise reduction, distortion removal, brightness and contrast adjustment; Feature extraction and matching: Use the feature extraction and matching module to extract and match features of pre-processed images and establish correspondence between images; Image fusion: Based on the feature matching results, the multimodal images are fused in the multimodal image fusion module to generate a fused image; Panoramic stitching: The fused images are input into the panoramic photography stitching module and stitched using the panoramic stitching algorithm to generate a panoramic image; Intelligent control: Through the intelligent control module, the workflow of the entire system and the parameters of each module are intelligently adjusted and optimized according to the shooting environment and user needs.

5. The method according to claim 4, characterized in that The multimodal image fusion further includes: The quality of each modality image is evaluated, and the clarity Qi and contrast Ci are calculated. The weight calculation formula is: wi(x,y)=∑j=1nQj(x,y)·Cj(x,y)Qi(x,y)·Ci(x,y); Using a deep learning model (training data ≥ 100,000 cross-modal image pairs) to learn the environment-weight mapping relationship, the infrared image weight accounts for ≥ 60% in low-light environments; The intelligent control module automatically switches to infrared imaging based on the ambient light intensity (≤50lux), adjusts the acquisition parameters and optimizes the algorithm parameters of each module.

Citation Information

Patent Citations

  • Switching method and device of driving circuit, control system, motor, storage medium and compressor

    CN108288932A

  • Transmitting device and receiving device providing relaxed impedance matching

    US20190372681A1