Super-photosensitive image processing method and system based on multi-modal fusion, and medium
Through multimodal image feature fusion and generative adversarial network (GAN) enhancement technology, the shortcomings of traditional low-light image processing methods are solved, and a multimodal fusion ultraphotosensitive imaging scheme with high precision, low latency and strong generalization is realized.
Patent Information
- Application Number
- CN202510603326.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional low-light image processing methods have problems in low signal-to-noise ratio, loss of details, color distortion, etc., and the existing multimodal fusion technology lacks an adaptive feature selection mechanism, cannot dynamically adjust the fusion strategy, and the algorithm has a large amount of calculation, so it cannot run in real time.
Through multimodal image feature fusion, high-fidelity image enhancement is performed using Generative Adversarial Network (GAN), and the optimal fusion algorithm is automatically selected based on scene lighting intensity, GPU computing power and real-time requirements.
It improves the accuracy of image processing in low-light environments, takes into account processing efficiency, and realizes high-fidelity image enhancement, providing reliable support for low-light vision tasks.
Smart Images

Figure CN120107090A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular, to a super-sensitive image processing method, system and medium based on multimodal fusion. Background Art
[0002] In low-light environments, traditional image acquisition technologies often face problems such as low signal-to-noise ratio, loss of details, and color distortion, resulting in a significant decrease in image quality. Traditional low-light image processing methods rely on single-mode (such as visible light) enhancement, which has problems such as noise amplification and serious loss of details. Existing multimodal fusion technologies often use fixed weight strategies (such as simple superposition of infrared and visible light), lack an adaptive feature selection mechanism, and cannot dynamically adjust the fusion strategy according to scene requirements. In addition, the algorithm has a large amount of computation and cannot run in real time on edge devices. Therefore, there is an urgent need for a multimodal fusion super-sensitive imaging solution that takes into account high precision, low latency, and strong generalization to break through the technical bottleneck in low-light scenes. Summary of the invention
[0003] The purpose of this application is to provide a super-sensitive image processing method, system and medium based on multimodal fusion, to improve the accuracy of image processing in low-light environments through multimodal image feature fusion, and to automatically select the optimal fusion algorithm according to the scene light intensity, GPU computing power and real-time requirements, so as to improve the image processing accuracy while taking into account the processing efficiency, and to achieve high-fidelity image enhancement through generative adversarial networks (GANs), providing reliable support for low-light vision tasks.
[0004] The present application also provides a super-sensitive image processing method based on multimodal fusion, comprising the following steps: Acquire multimodal images and perform image preprocessing; Extract features from preprocessed multimodal images and select suitable feature fusion algorithms; Performing multimodal image feature fusion according to the adapted feature fusion algorithm, and processing to obtain a super-sensitive image; The super-sensitive image is enhanced using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to the user end for display.
[0005] Optionally, in the super-photosensitive image processing method based on multimodal fusion described in the present application, the collecting of multimodal images and performing image preprocessing includes: Acquire multimodal image acquisition data, including visible light image data, infrared image data, and depth image data; De-noising the multimodal image acquisition data based on the preset CMID algorithm; The preset SIFT feature point matching algorithm and optical flow method are used to perform spatial registration on multimodal image acquisition data.
[0006] Optionally, in the super-sensitive image processing method based on multimodal fusion described in the present application, the feature extraction of the pre-processed multimodal image and the screening of an adaptive feature fusion algorithm include: Use the preset CNN model to extract features from the multimodal image acquisition data to obtain multimodal image feature data, including visual feature data, thermal radiation distribution feature data, and spatial feature data; Obtain response time requirement data, GPU computing power data, and light intensity data; The response time requirement data, GPU computing power data and light intensity data are input into the preset adaptive feature fusion algorithm library, and the multi-condition decision tree algorithm is used to screen the adaptation algorithm to obtain the adaptive feature fusion algorithm.
[0007] Optionally, in the super-sensitive image processing method based on multimodal fusion described in the present application, the multimodal image feature fusion is performed according to the adapted feature fusion algorithm, and the super-sensitive image is obtained by processing, including: Using the adaptive feature fusion algorithm to perform feature fusion on the multimodal image feature data and generate a fusion feature map; The deconvolution operation is used to map the fused feature map back to the image space to obtain the super-sensitive image.
[0008] Optionally, in the super-sensitive image processing method based on multimodal fusion described in the present application, the super-sensitive image is enhanced by using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to a user terminal for display, including: Inputting the super-sensitive image into a preset generator model for processing to obtain an optimized super-sensitive image; Input the optimized super-sensitive image into the preset discriminator model, obtain the probability value and feed it back to the generator model; The generator model optimizes its own parameters according to the probability value to obtain an optimized generator model; The optimized super-sensitive image is processed according to the optimized generator model to obtain an enhanced super-sensitive image.
[0009] Optionally, the super-sensitive image processing method based on multimodal fusion described in the present application further includes: Compare the enhanced super-sensitive image with the real image, and obtain PSNR data, SSIM data, and LPIPS data; After weighted processing of the PSNR data and the SSIM data, a ratio thereof to the LPIPS data is calculated to obtain similarity evaluation data; The similarity evaluation data is compared with a preset similarity evaluation data threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, a poor image processing effect feedback is given.
[0010] In a second aspect, the present application provides a super-photosensitive image processing system based on multimodal fusion, the system comprising: a memory and a processor, wherein the memory stores a program of a super-photosensitive image processing method based on multimodal fusion, and when the program of the super-photosensitive image processing method based on multimodal fusion is executed by the processor, the following steps are implemented: Acquire multimodal images and perform image preprocessing; Extract features from preprocessed multimodal images and select suitable feature fusion algorithms; Performing multimodal image feature fusion according to the adapted feature fusion algorithm, and processing to obtain a super-sensitive image; The super-sensitive image is enhanced using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to the user end for display.
[0011] Optionally, in the super-photosensitive image processing system based on multimodal fusion described in the present application, the collecting of multimodal images and performing image preprocessing includes: Acquire multimodal image acquisition data, including visible light image data, infrared image data, and depth image data; De-noising the multimodal image acquisition data based on the preset CMID algorithm; The preset SIFT feature point matching algorithm and optical flow method are used to perform spatial registration on multimodal image acquisition data.
[0012] Optionally, in the super-sensitive image processing system based on multimodal fusion described in the present application, the feature extraction of the pre-processed multimodal image and the screening of an adaptive feature fusion algorithm include: Use the preset CNN model to extract features from the multimodal image acquisition data to obtain multimodal image feature data, including visual feature data, thermal radiation distribution feature data, and spatial feature data; Obtain response time requirement data, GPU computing power data, and light intensity data; The response time requirement data, GPU computing power data and light intensity data are input into the preset adaptive feature fusion algorithm library, and the multi-condition decision tree algorithm is used to screen the adaptation algorithm to obtain the adaptive feature fusion algorithm.
[0013] In the third aspect, the present application also provides a computer-readable storage medium, which stores a super-sensitive image processing method program based on multimodal fusion. When the super-sensitive image processing method program based on multimodal fusion is executed by a processor, the steps of the super-sensitive image processing method based on multimodal fusion as described in any one of the above items are implemented.
[0014] From the above, it can be seen that the super-sensitive image processing method, system and medium based on multimodal fusion provided in this application improve the accuracy of image processing in low-light environments through multimodal image feature fusion, and automatically select the optimal fusion algorithm according to the scene light intensity, GPU computing power and real-time requirements, so as to improve the image processing accuracy while taking into account the processing efficiency, and realize high-fidelity image enhancement through generative adversarial networks (GAN), providing reliable support for low-light vision tasks.
[0015] Other features and advantages of the present application will be described in the following description, and partly become apparent from the description, or understood by practicing the embodiments of the present application. The purpose and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 A flowchart of a super-sensitive image processing method based on multimodal fusion provided in an embodiment of the present application; Figure 2 A flowchart of a screening and adaptation feature fusion algorithm for a super-sensitive image processing method based on multimodal fusion provided in an embodiment of the present application; Figure 3 A flowchart of obtaining a super-sensitive image according to the super-sensitive image processing method based on multimodal fusion provided in an embodiment of the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0019] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0020] Please refer to Figure 1 , Figure 1 : is a flowchart of a super-sensitive image processing method based on multi-modal fusion in some embodiments of the present application. The super-sensitive image processing method based on multi-modal fusion is used in a terminal device, such as a computer, a mobile phone terminal, etc. The super-sensitive image processing method based on multi-modal fusion includes the following steps: S11, collecting multimodal images and performing image preprocessing; S12, extracting features from the preprocessed multimodal images and selecting an adaptive feature fusion algorithm; S13, performing multimodal image feature fusion according to the adapted feature fusion algorithm, and processing to obtain a super-sensitive image; S14. Use a generative adversarial network to enhance the super-sensitive image, obtain an enhanced super-sensitive image, and transmit it to the user end for display.
[0021] It should be noted that this application improves the accuracy of image processing in low-light environments through multimodal image feature fusion, and automatically selects the optimal fusion algorithm according to the scene light intensity, GPU computing power and real-time requirements, so as to improve image processing accuracy while taking into account processing efficiency, and realizes high-fidelity image enhancement through generative adversarial networks (GAN), providing reliable support for low-light vision tasks.
[0022] According to an embodiment of the present invention, the acquiring of multimodal images and image preprocessing includes: Acquire multimodal image acquisition data, including visible light image data, infrared image data, and depth image data; De-noising the multimodal image acquisition data based on the preset CMID algorithm; The preset SIFT feature point matching algorithm and optical flow method are used to perform spatial registration on multimodal image acquisition data.
[0023] It should be noted that multimodal images are first collected. Visible light images can provide rich color and texture information, while infrared images can capture the thermal radiation contours of objects under low-light conditions, and depth images help to understand the spatial position relationship of objects in the scene. Then, based on the cross-modal interactive dynamic fusion (CMID) algorithm, the visible light image data, infrared image data and depth image data collected in the low-light environment are denoised, and the scale-invariant feature transform (SIFT) algorithm is used to extract feature points in the image. The correspondence between images of different modalities is determined by matching these feature points, and the accuracy of spatial registration is optimized by the optical flow method to ensure the alignment of data of different modalities.
[0024] Please refer to Figure 2 , Figure 2 The flowchart of the screening and adaptation feature fusion algorithm of the super-sensitive image processing method based on multimodal fusion in some embodiments of the present application. According to the embodiment of the present invention, the feature extraction of the pre-processed multimodal image and the screening and adaptation feature fusion algorithm include: S21. Using a preset CNN model to perform feature extraction on the multimodal image acquisition data, obtaining multimodal image feature data, including visual feature data, thermal radiation distribution feature data, and spatial feature data; S22, obtaining response time requirement data, GPU computing power data, and light intensity data; S23, inputting the response time requirement data, GPU computing power data and light intensity data into a preset adaptive feature fusion algorithm library, and using a multi-condition decision tree algorithm to screen the adaptive algorithm to obtain an adaptive feature fusion algorithm.
[0025] It should be noted that visual feature data includes color feature data, texture feature data, shape contour feature data, high-level semantic feature data and object category label data, thermal radiation distribution feature data includes temperature feature data, thermal distribution feature data and thermal contrast data, and spatial feature data includes pixel depth data, spatial relationship data and geometric feature data. The preset adaptive feature fusion algorithm library is a pre-built database containing multiple feature fusion algorithms (wavelet transform, attention mechanism, weighted average, etc.). Each algorithm has been pre-calibrated for its computational complexity, applicable lighting range and typical response time. The decision tree selects the adaptation algorithm according to the following rules: the first layer is judged as response time constraint, and the adaptation algorithm is selected according to the delay time. The second layer is judged as GPU computing power matching. For example, if the GPU video memory is less than 4GB, the high complexity algorithm is eliminated (such as the attention mechanism may be excluded). The third layer is judged as light intensity adaptation. For example, if the light intensity is 2 Lux (extremely low light), the algorithm supporting 0-50 Lux (such as wavelet fusion algorithm) is selected.
[0026] Please refer to Figure 3 , Figure 3 The present invention is a flowchart of obtaining a super-sensitive image by a super-sensitive image processing method based on multi-modal fusion in some embodiments of the present application. According to an embodiment of the present invention, the multi-modal image feature fusion is performed according to the adaptive feature fusion algorithm, and the super-sensitive image is obtained by processing, including: S31, using the adapted feature fusion algorithm to perform feature fusion on the multimodal image feature data, and generate a fusion feature map; S32. Use a deconvolution operation to map the fused feature map back to the image space to obtain a super-sensitive image.
[0027] It should be noted that the multimodal image feature fusion is performed according to the adaptive feature fusion algorithm, the fused feature map is gradually upsampled to restore to the original image size, and the feature is mapped back to the image space using the deconvolution operation.
[0028] According to an embodiment of the present invention, the method of enhancing the super-sensitive image by using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to a user terminal for display includes: Inputting the super-sensitive image into a preset generator model for processing to obtain an optimized super-sensitive image; Input the optimized super-sensitive image into the preset discriminator model, obtain the probability value and feed it back to the generator model; The generator model optimizes its own parameters according to the probability value to obtain an optimized generator model; The optimized super-sensitive image is processed according to the optimized generator model to obtain an enhanced super-sensitive image.
[0029] It should be noted that by utilizing the adversarial training mechanism of the generative adversarial network (GAN), the generated images are not only improved in terms of brightness and contrast, but also closer to high-quality images in real scenes in terms of detail retention and naturalness. The generated images are distinguished from real images under normal lighting through the discriminator. The output result of the discriminator is a probability value, which indicates the probability that the input image is a real image. The discriminator probability value can prompt the generator to continuously optimize the image generation effect, and ultimately obtain super-sensitive images with excellent visual effects and high quality.
[0030] According to an embodiment of the present invention, it also includes: Compare the enhanced super-sensitive image with the real image, and obtain PSNR data, SSIM data, and LPIPS data; After weighted processing of the PSNR data and the SSIM data, a ratio thereof to the LPIPS data is calculated to obtain similarity evaluation data; The similarity evaluation data is compared with a preset similarity evaluation data threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, a poor image processing effect feedback is given.
[0031] It should be noted that the enhanced super-sensitive image is compared with the real image, and the image processing effect is judged by the similarity comparison results. Among them, the higher the peak signal-to-noise ratio (PSNR) value, the better the image quality. The structural similarity index (SSIM) is an indicator to measure the structural similarity of two images. The closer the value is to 1, the higher the structural similarity of the two images. LPIPS is an image similarity evaluation indicator based on deep learning. The smaller the LPIPS value, the more similar the two images are in perception.
[0032] According to an embodiment of the present invention, it also includes: Obtain system performance monitoring data of the image processing system, including running time, memory usage and energy consumption data; Perform weighted average calculation based on the running time, memory occupancy rate and energy consumption data to obtain system processing efficiency evaluation data; The system processing efficiency evaluation data is compared with a preset system processing efficiency evaluation data threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, a prompt of poor system efficiency is given.
[0033] It should be noted that during the image processing process, the performance of the image processing system is monitored to timely understand the system resource usage of the current image processing task. If the resource usage is too large, the image processing task needs to be optimized.
[0034] The present invention also discloses a super-photosensitive image processing system based on multimodal fusion, comprising a memory and a processor, wherein the memory stores a super-photosensitive image processing method program based on multimodal fusion, and when the super-photosensitive image processing method program based on multimodal fusion is executed by the processor, the following steps are implemented: Acquire multimodal images and perform image preprocessing; Extract features from preprocessed multimodal images and select suitable feature fusion algorithms; Performing multimodal image feature fusion according to the adapted feature fusion algorithm, and processing to obtain a super-sensitive image; The super-sensitive image is enhanced using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to the user end for display.
[0035] It should be noted that this application improves the accuracy of image processing in low-light environments through multimodal image feature fusion, and automatically selects the optimal fusion algorithm according to the scene light intensity, GPU computing power and real-time requirements, so as to improve image processing accuracy while taking into account processing efficiency, and realizes high-fidelity image enhancement through generative adversarial networks (GAN), providing reliable support for low-light vision tasks.
[0036] According to an embodiment of the present invention, the acquiring of multimodal images and image preprocessing includes: Acquire multimodal image acquisition data, including visible light image data, infrared image data, and depth image data; De-noising the multimodal image acquisition data based on the preset CMID algorithm; The preset SIFT feature point matching algorithm and optical flow method are used to perform spatial registration on multimodal image acquisition data.
[0037] It should be noted that multimodal images are first collected. Visible light images can provide rich color and texture information, while infrared images can capture the thermal radiation contours of objects under low-light conditions, and depth images help to understand the spatial position relationship of objects in the scene. Then, based on the cross-modal interactive dynamic fusion (CMID) algorithm, the visible light image data, infrared image data and depth image data collected in the low-light environment are denoised, and the scale-invariant feature transform (SIFT) algorithm is used to extract feature points in the image. The correspondence between images of different modalities is determined by matching these feature points, and the accuracy of spatial registration is optimized by the optical flow method to ensure the alignment of data of different modalities.
[0038] According to an embodiment of the present invention, the step of extracting features from the preprocessed multimodal image and selecting an adaptive feature fusion algorithm includes: Use the preset CNN model to extract features from the multimodal image acquisition data to obtain multimodal image feature data, including visual feature data, thermal radiation distribution feature data, and spatial feature data; Obtain response time requirement data, GPU computing power data, and light intensity data; The response time requirement data, GPU computing power data and light intensity data are input into the preset adaptive feature fusion algorithm library, and the multi-condition decision tree algorithm is used to screen the adaptation algorithm to obtain the adaptive feature fusion algorithm.
[0039] It should be noted that visual feature data includes color feature data, texture feature data, shape contour feature data, high-level semantic feature data and object category label data, thermal radiation distribution feature data includes temperature feature data, thermal distribution feature data and thermal contrast data, and spatial feature data includes pixel depth data, spatial relationship data and geometric feature data. The preset adaptive feature fusion algorithm library is a pre-built database containing multiple feature fusion algorithms (wavelet transform, attention mechanism, weighted average, etc.). Each algorithm has been pre-calibrated for its computational complexity, applicable lighting range and typical response time. The decision tree selects the adaptation algorithm according to the following rules: the first layer is judged as response time constraint, and the adaptation algorithm is selected according to the delay time. The second layer is judged as GPU computing power matching. For example, if the GPU video memory is less than 4GB, the high complexity algorithm is eliminated (such as the attention mechanism may be excluded). The third layer is judged as light intensity adaptation. For example, if the light intensity is 2 Lux (extremely low light), the algorithm supporting 0-50 Lux (such as wavelet fusion algorithm) is selected.
[0040] According to an embodiment of the present invention, performing multimodal image feature fusion according to the adapted feature fusion algorithm and processing to obtain a super-sensitive image includes: Using the adaptive feature fusion algorithm to perform feature fusion on the multimodal image feature data and generate a fusion feature map; The deconvolution operation is used to map the fused feature map back to the image space to obtain the super-sensitive image.
[0041] It should be noted that the multimodal image feature fusion is performed according to the adaptive feature fusion algorithm, the fused feature map is gradually upsampled to restore to the original image size, and the feature is mapped back to the image space using the deconvolution operation.
[0042] According to an embodiment of the present invention, the method of enhancing the super-sensitive image by using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to a user terminal for display includes: Inputting the super-sensitive image into a preset generator model for processing to obtain an optimized super-sensitive image; Input the optimized super-sensitive image into the preset discriminator model, obtain the probability value and feed it back to the generator model; The generator model optimizes its own parameters according to the probability value to obtain an optimized generator model; The optimized super-sensitive image is processed according to the optimized generator model to obtain an enhanced super-sensitive image.
[0043] It should be noted that by utilizing the adversarial training mechanism of the generative adversarial network (GAN), the generated images are not only improved in terms of brightness and contrast, but also closer to high-quality images in real scenes in terms of detail retention and naturalness. The generated images are distinguished from real images under normal lighting through the discriminator. The output result of the discriminator is a probability value, which indicates the probability that the input image is a real image. The discriminator probability value can prompt the generator to continuously optimize the image generation effect, and ultimately obtain super-sensitive images with excellent visual effects and high quality.
[0044] According to an embodiment of the present invention, it also includes: Compare the enhanced super-sensitive image with the real image, and obtain PSNR data, SSIM data, and LPIPS data; After weighted processing of the PSNR data and the SSIM data, a ratio thereof to the LPIPS data is calculated to obtain similarity evaluation data; The similarity evaluation data is compared with a preset similarity evaluation data threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, a poor image processing effect feedback is given.
[0045] It should be noted that the enhanced super-sensitive image is compared with the real image, and the image processing effect is judged by the similarity comparison results. Among them, the higher the peak signal-to-noise ratio (PSNR) value, the better the image quality. The structural similarity index (SSIM) is an indicator to measure the structural similarity of two images. The closer the value is to 1, the higher the structural similarity of the two images. LPIPS is an image similarity evaluation indicator based on deep learning. The smaller the LPIPS value, the more similar the two images are in perception.
[0046] According to an embodiment of the present invention, it also includes: Obtain system performance monitoring data of the image processing system, including running time, memory usage and energy consumption data; Perform weighted average calculation based on the running time, memory occupancy rate and energy consumption data to obtain system processing efficiency evaluation data; The system processing efficiency evaluation data is compared with a preset system processing efficiency evaluation data threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, a prompt of poor system efficiency is given.
[0047] It should be noted that during the image processing process, the performance of the image processing system is monitored to timely understand the system resource usage of the current image processing task. If the resource usage is too large, the image processing task needs to be optimized.
[0048] The third aspect of the present invention provides a readable storage medium, which stores a program for a super-sensitive image processing method based on multimodal fusion. When the program for the super-sensitive image processing method based on multimodal fusion is executed by a processor, the steps of the super-sensitive image processing method based on multimodal fusion as described in any one of the above items are implemented.
[0049] The super-sensitive image processing method, system and medium based on multimodal fusion disclosed in the present invention improve the accuracy of image processing in low-light environments through multimodal image feature fusion, and automatically select the optimal fusion algorithm according to the scene light intensity, GPU computing power and real-time requirements, so as to improve the image processing accuracy while taking into account the processing efficiency, and realize high-fidelity image enhancement through generative adversarial networks (GANs), providing reliable support for low-light vision tasks.
[0050] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0051] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0052] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0053] Those skilled in the art can understand that: all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, the aforementioned program can be stored in a readable storage medium, and when the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, and other media that can store program codes.
[0054] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.
Claims
1. A super-sensitive image processing method based on multimodal fusion, characterized in that: The following steps are involved: Acquire multimodal images and perform image preprocessing; Extract features from preprocessed multimodal images and select suitable feature fusion algorithms; Performing multimodal image feature fusion according to the adapted feature fusion algorithm, and processing to obtain a super-sensitive image; The super-sensitive image is enhanced using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to the user end for display.
2. The method for processing super-sensitive images based on multimodal fusion according to claim 1, characterized in that: The collecting of multimodal images and performing image preprocessing includes: Acquire multimodal image acquisition data, including visible light image data, infrared image data, and depth image data; De-noising the multimodal image acquisition data based on the preset CMID algorithm; The preset SIFT feature point matching algorithm and optical flow method are used to perform spatial registration on multimodal image acquisition data.
3. The method for processing super-sensitive images based on multimodal fusion according to claim 2, characterized in that: The method of extracting features from the preprocessed multimodal image and selecting an adaptive feature fusion algorithm includes: Use the preset CNN model to extract features from the multimodal image acquisition data to obtain multimodal image feature data, including visual feature data, thermal radiation distribution feature data, and spatial feature data; Obtain response time requirement data, GPU computing power data, and light intensity data; The response time requirement data, GPU computing power data and light intensity data are input into the preset adaptive feature fusion algorithm library, and the multi-condition decision tree algorithm is used to screen the adaptation algorithm to obtain the adaptive feature fusion algorithm.
4. The method for processing super-sensitive images based on multimodal fusion according to claim 3, characterized in that: The step of performing multimodal image feature fusion according to the adapted feature fusion algorithm and processing to obtain a super-sensitive image includes: Using the adaptive feature fusion algorithm to perform feature fusion on the multimodal image feature data and generate a fusion feature map; The deconvolution operation is used to map the fused feature map back to the image space to obtain the super-sensitive image.
5. The method for processing super-sensitive images based on multimodal fusion according to claim 4, characterized in that: The method of using a generative adversarial network to enhance the super-sensitive image, obtain an enhanced super-sensitive image and transmit it to a user terminal for display, includes: Inputting the super-sensitive image into a preset generator model for processing to obtain an optimized super-sensitive image; Input the optimized super-sensitive image into the preset discriminator model, obtain the probability value and feed it back to the generator model; The generator model optimizes its own parameters according to the probability value to obtain an optimized generator model; The optimized super-sensitive image is processed according to the optimized generator model to obtain an enhanced super-sensitive image.
6. The method for processing super-sensitive images based on multimodal fusion according to claim 5, characterized in that: Also includes: Compare the enhanced super-sensitive image with the real image, and obtain PSNR data, SSIM data, and LPIPS data; After weighted processing of the PSNR data and the SSIM data, a ratio thereof to the LPIPS data is calculated to obtain similarity evaluation data; The similarity evaluation data is compared with a preset similarity evaluation data threshold. If the threshold comparison result does not meet the preset threshold comparison result requirement, a poor image processing effect feedback is given.
7. A super-sensitive image processing system based on multi-modal fusion, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a program of a super-photosensitive image processing method based on multimodal fusion, and when the program of the super-photosensitive image processing method based on multimodal fusion is executed by the processor, the following steps are implemented: Acquire multimodal images and perform image preprocessing; Extract features from preprocessed multimodal images and select suitable feature fusion algorithms; Performing multimodal image feature fusion according to the adapted feature fusion algorithm, and processing to obtain a super-sensitive image; The super-sensitive image is enhanced using a generative adversarial network to obtain an enhanced super-sensitive image and transmit it to the user end for display.
8. The super-sensitive image processing system based on multi-modal fusion according to claim 7 is characterized in that: The collecting of multimodal images and performing image preprocessing includes: Acquire multimodal image acquisition data, including visible light image data, infrared image data, and depth image data; De-noising the multimodal image acquisition data based on the preset CMID algorithm; The preset SIFT feature point matching algorithm and optical flow method are used to perform spatial registration on multimodal image acquisition data.
9. The super-sensitive image processing system based on multi-modal fusion according to claim 8, characterized in that: The method of extracting features from the preprocessed multimodal image and selecting an adaptive feature fusion algorithm includes: Use the preset CNN model to extract features from the multimodal image acquisition data to obtain multimodal image feature data, including visual feature data, thermal radiation distribution feature data, and spatial feature data; Obtain response time requirement data, GPU computing power data, and light intensity data; The response time requirement data, GPU computing power data and light intensity data are input into the preset adaptive feature fusion algorithm library, and the multi-condition decision tree algorithm is used to screen the adaptation algorithm to obtain the adaptive feature fusion algorithm.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a super-sensitive image processing program based on multimodal fusion. When the super-sensitive image processing program based on multimodal fusion is executed by the processor, the steps of the super-sensitive image processing method based on multimodal fusion as described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Image processing method and device, electronic equipment and computer readable storage medium
CN111080543A
Bimodal infrared image fusion algorithm selection method based on difference characteristic amplitude interval fusion validity distribution
CN111445430A
Image processing method and device based on machine vision
CN112204566A
Equipment configuration method, system and device, equipment, storage medium and access control equipment
CN115311700A
Low-light image enhancement method of image hierarchical structure network based on stream learning
CN118115378A