A focal plane recognition method and recognition system for super-resolution imaging

By fusing the features of visible light and infrared images, depth estimation and focal plane recognition are solved, and the problem that traditional super-resolution imaging methods are difficult to use multi-spectral information and distinguish focal planes at different depths is achieved, and high-precision and high-definition processing of images are achieved.

CN118918011BActive Publication Date: 2025-05-13SICHUAN NATIONAL INNOVATION VISION UHD VIDEO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410927529.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2025-05-13
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

Traditional super-resolution imaging methods are difficult to effectively utilize different spectral information and distinguish focal planes at different depths, resulting in insufficient image quality and clarity.

Method used

By acquiring visible light and infrared images, image preprocessing and feature extraction are performed, features are fused using a cross-modal attention mechanism, depth estimation and focal plane recognition are performed, and the clarity enhancement of focal planes at different depths is finally achieved.

Benefits of technology

It realizes accurate processing and clarity enhancement of focal planes at different depths in the image, improves the accuracy and detail restoration capabilities of super-resolution imaging, and enhances the overall clarity and reality of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118918011B_ABST
    Figure CN118918011B_ABST
Patent Text Reader

Abstract

The present invention provides a focal plane recognition method and recognition system for super-resolution imaging, belonging to the technical field of image processing. The present invention introduces the joint processing of infrared images and visible light images, and ensures the accuracy and consistency of the two image information through an image preprocessing step; then, a pre-trained feature extraction model is used to extract multi-scale features respectively, and a cross-modal attention mechanism is used to automatically identify and strengthen the correlation between the two features, so as to generate comprehensive multi-scale features, and then a depth estimation method is used to perform depth estimation on each pixel in the image, so as to identify focal planes of different depths; finally, according to the depth information of the focal plane, a targeted clarity enhancement processing is performed on objects on different focal planes, so as to achieve accurate processing and clarity enhancement of focal planes of different depths in the image, which not only improves the accuracy and detail restoration capability of super-resolution imaging, but also enhances the overall clarity and realism of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a focal plane recognition method and recognition system applied to super-resolution imaging. Background Art

[0002] In the field of image processing, super-resolution imaging technology is a key technical means, which aims to convert low-resolution images into high-resolution images through algorithmic processing, thereby improving the image's detail performance and clarity. However, traditional super-resolution imaging methods are often limited to a single image source (such as visible light images only), ignoring the potential contribution of different spectral information (such as infrared images) to image quality improvement. In addition, in practical applications, different objects in the image are often located at different depth planes. The clarity of these objects is affected by the focal plane in which they are located, and traditional methods find it difficult to effectively distinguish and separately process these focal planes at different depths.

[0003] Therefore, it is necessary to provide a focal plane recognition method and recognition system for super-resolution imaging to solve the above technical problems. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a focal plane recognition method and recognition system for super-resolution imaging, which realizes accurate processing and clarity enhancement of focal planes at different depths in the image by fusing multi-spectral information, performing depth estimation and focal plane recognition, thereby improving the accuracy and detail restoration capability of super-resolution imaging, and enhancing the overall clarity and realism of the image.

[0005] The present invention provides a focal plane recognition method for super-resolution imaging, the recognition method comprising the following steps:

[0006] S1: Obtaining a visible light image and an infrared image of an image to be super-resolved and performing image preprocessing, wherein the image preprocessing includes image alignment and image denoising;

[0007] S2: extracting multi-scale features from the preprocessed visible light image and infrared image using a pre-trained feature extraction model to obtain first multi-scale features and second multi-scale features corresponding to the visible light image and infrared image respectively;

[0008] S3: fusing the first multi-scale feature and the second multi-scale feature based on a cross-modal attention mechanism, and performing super-resolution processing on the fused comprehensive multi-scale features to obtain an initial super-resolution image;

[0009] S4: performing depth estimation on each pixel in the initial super-resolution image, and identifying focal planes at different depths according to the estimated depth of each pixel;

[0010] S5: Using the depth of the focal plane, the clarity of objects on different focal planes is enhanced to generate a final super-resolution image.

[0011] Preferably, step S1 specifically includes:

[0012] S101: collecting a visible light image and an infrared image of the image to be super-resolved;

[0013] S102: performing image alignment on the visible light image and the infrared image based on feature points;

[0014] S103: Using a bilateral filtering method to perform denoising on the aligned visible light image and infrared image.

[0015] Preferably, step S2 specifically includes:

[0016] S201: preprocessing the visible light image and infrared image after image alignment and denoising, and inputting the preprocessed images into a pre-trained feature extraction model;

[0017] S202: performing feature extraction on the visible light image and the infrared image using convolution layers of different levels in the feature extraction model to capture first multi-scale features of the visible light image and second multi-scale features of the infrared image;

[0018] S203: Performing feature tuning on the first multi-scale features of the visible light image and the second multi-scale features of the infrared image, wherein the feature tuning includes scale and resolution unification, channel adjustment, and semantic alignment.

[0019] Preferably, step S3 specifically includes:

[0020] S301: assigning weights to the first multi-scale feature and the second multi-scale feature based on the pre-trained attention module, and performing weighted fusion based on the respective weights to generate a comprehensive multi-scale feature;

[0021] S302: Super-resolution the comprehensive multi-scale features by combining an upsampling method and a reconstruction network based on a deep learning model to generate an initial super-resolution image.

[0022] Preferably, step S4 specifically includes:

[0023] S401: Generate an initial depth map of an initial super-resolution image using a depth estimation algorithm;

[0024] S402: Based on a preset depth threshold, using an image segmentation method to preliminarily segment the initial depth map into a plurality of initial focal plane areas of different depths;

[0025] S403: Calculate the average depth value of each focal plane area;

[0026] S404: Adjusting the boundary of the initial focal plane region based on the average depth value of the adjacent initial focal plane regions to obtain a final focal plane region.

[0027] Preferably, in step S404, specifically:

[0028] Comparing the average depth value of each of the initial focal plane regions with the average depth value of adjacent initial focal plane regions, and evaluating the difference according to the comparison result;

[0029] Adjusting the boundary of each of the initial focal plane regions based on the difference;

[0030] Each of the initial focal plane regions after boundary adjustment is smoothed to generate a final focal plane region.

[0031] Preferably, step S5 specifically includes:

[0032] S501: assigning a weight to each pixel in a final focal plane region, wherein the weight is determined based on a depth of the final focal plane region where the pixel is located and a relative distance from the pixel to other final focal plane regions;

[0033] S502: performing definition enhancement processing on pixels in each final focal plane area according to the assigned weights to obtain a focal plane image corresponding to each final focal plane area;

[0034] S503: Perform weighted fusion on all the focal plane images that have undergone clarity enhancement to generate a final super-resolution image.

[0035] The present invention also provides a focal plane recognition system for super-resolution imaging, which is applied to a focal plane recognition system for super-resolution imaging. The recognition system includes:

[0036] An image processing module, used to obtain a visible light image and an infrared image of an image to be super-resolved and perform image preprocessing, wherein the image preprocessing includes image alignment and image denoising;

[0037] A feature extraction module, configured to extract multi-scale features from the preprocessed visible light image and infrared image using a pre-trained feature extraction model, to obtain first multi-scale features and second multi-scale features corresponding to the visible light image and infrared image respectively;

[0038] A super-resolution processing module, used for fusing the first multi-scale features and the second multi-scale features based on a cross-modal attention mechanism, and performing super-resolution processing on the fused comprehensive multi-scale features to obtain an initial super-resolution image;

[0039] A focal plane recognition module, used to perform depth estimation on each pixel in the initial super-resolution image, and to recognize focal planes at different depths according to the estimated depth of each pixel;

[0040] The super-resolution image enhancement module is used to enhance the clarity of objects on different focal planes by utilizing the depth of the focal plane to generate a final super-resolution image.

[0041] Compared with the related art, the focal plane recognition method and recognition system for super-resolution imaging provided by the present invention have the following beneficial effects:

[0042] The present invention introduces the joint processing of infrared images and visible light images, and ensures the accuracy and consistency of the two image information through the image preprocessing step; then, using the pre-trained feature extraction model, multi-scale features are extracted from the two images respectively, and the cross-modal attention mechanism is used to automatically identify and strengthen the correlation between the two features, while suppressing irrelevant information, thereby generating more accurate and comprehensive comprehensive multi-scale features, and then the depth of each pixel in the image is estimated by the depth estimation method, and then the focal planes at different depths are identified; finally, according to the depth information of the focal plane, the objects on different focal planes are subjected to targeted clarity enhancement processing, thereby realizing accurate processing and clarity enhancement of focal planes at different depths in the image, which not only improves the accuracy and detail restoration capability of super-resolution imaging, but also enhances the overall clarity and realism of the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A flowchart of a focal plane recognition method applied to super-resolution imaging provided by the present invention;

[0044] Figure 2 A module structure diagram of a focal plane recognition system for super-resolution imaging provided by the present invention. DETAILED DESCRIPTION

[0045] The present invention will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention. It should also be noted that, for ease of description, only the parts related to the present invention, rather than all structures, are shown in the accompanying drawings. In addition, the embodiments of the present invention and the features in the embodiments may be combined with each other without conflict.

[0046] It should also be noted that, for ease of description, only the part relevant to the present invention but not all content is shown in the accompanying drawings. It should be mentioned before discussing exemplary embodiments in more detail that some exemplary embodiments are described as processing or methods depicted as flow charts. Although the flow chart describes each operation (or step) as sequential processing, many operations therein can be implemented in parallel, concurrently or simultaneously. In addition, the order of each operation can be rearranged. When its operation is completed, the processing can be terminated, but it can also have additional steps not included in the accompanying drawings. The processing can correspond to methods, functions, procedures, subroutines, subprograms, etc.

[0047] Embodiment 1

[0048] The present invention provides a focal plane recognition method for super-resolution imaging, referring to Figure 1 As shown, the identification method includes the following steps:

[0049] S1: Obtain a visible light image and an infrared image of an image to be super-resolved and perform image preprocessing, wherein the image preprocessing includes image alignment and image denoising.

[0050] In this embodiment, the visible light image and infrared image of the image to be super-resolved are obtained by synchronous acquisition equipment. In order to ensure the accuracy of subsequent processing, image preprocessing is a crucial first step. Among them, image alignment is achieved through rigid body transformation based on feature point matching, which ensures the precise spatial correspondence between the visible light image and the infrared image and solves the displacement problem between images of different modalities. Next, the bilateral filtering method is used for image denoising, which not only maintains the clarity of the image edge, but also effectively reduces noise interference, improves image quality, and lays a solid foundation for subsequent feature extraction.

[0051] S2: Use a pre-trained feature extraction model to extract multi-scale features from the pre-processed visible light image and infrared image to obtain first multi-scale features and second multi-scale features corresponding to the visible light image and infrared image respectively.

[0052] In this embodiment, a deep convolutional neural network (CNN) is used as a pre-trained feature extraction model. The model is trained with a large-scale data set and can automatically learn and extract complex features in images. For pre-processed visible light images and infrared images, different levels of convolutional layers of CNN are used to capture multi-scale features, where the low-level convolutional layers are responsible for capturing basic texture and edge information, while the high-level convolutional layers focus on more abstract concepts and patterns. Through this multi-scale feature extraction, not only the local details of the image are retained, but also the understanding of the global structure is enhanced, providing a rich and hierarchical information basis for subsequent feature fusion.

[0053] S3: Based on the cross-modal attention mechanism, the first multi-scale features and the second multi-scale features are fused, and the fused comprehensive multi-scale features are super-resolved to obtain an initial super-resolved image.

[0054] In this embodiment, a cross-modal attention mechanism is introduced, which allows the model to automatically assign weights based on the complementary information of visible light images and infrared images. By calculating the similarity and correlation between the features of the two modalities, the attention mechanism can intelligently select which features should be emphasized and which can be suppressed, thereby achieving effective fusion of features. The fused comprehensive multi-scale features are then super-resolved through upsampling and a deep learning-based reconstruction network to generate an initial super-resolved image with richer details and significantly improved clarity. This process not only amplifies the image resolution, but also optimizes the image quality, making the details more realistic.

[0055] S4: performing depth estimation on each pixel in the initial super-resolution image, and identifying focal planes at different depths according to the estimated depth of each pixel.

[0056] In this embodiment, depth estimation is performed by a depth estimation algorithm, which is trained based on an existing depth map database and can accurately predict the depth information of each pixel. The initial super-resolution image is input into a pre-trained depth estimation model to obtain an initial depth map. Next, the initial depth map is segmented into multiple different initial focal plane areas using an image segmentation method combined with a preset depth threshold. By calculating the average depth value of each initial focal plane area, objects of different depths can be more finely distinguished, providing key depth clues for further clarity enhancement.

[0057] S5: Using the depth of the focal plane, the clarity of objects on different focal planes is enhanced to generate a final super-resolution image.

[0058] In this embodiment, the clarity enhancement is an adaptive processing based on depth information. By analyzing the average depth value of each focal plane area, corresponding clarity enhancement levels are assigned to objects at different depths. For objects close to the focal plane, the sharpening level is increased to highlight the details; for objects far from the focal plane, the sharpening level is appropriately reduced to avoid image distortion caused by excessive enhancement. Specifically, this process is implemented through an adaptive filter to ensure that objects on each focal plane can obtain the best clarity performance, and finally generate a super-resolution image with rich details, clear layers, and excellent visual effects.

[0059] Specifically, step S1 specifically includes:

[0060] S101: Acquire a visible light image and an infrared image of the image to be super-resolved.

[0061] In this embodiment, a highly integrated dual-modal imaging device can be used, which is equipped with a highly sensitive visible light camera and an infrared thermal imager at the same time, ensuring the consistency and accuracy of image acquisition. In order to obtain the visible light image and infrared image of the image to be super-resolved, a synchronous trigger mechanism is designed so that the images of the two modes are captured at the same time, avoiding the spatial dislocation caused by the time difference. In this way, not only the real-time and consistency of the image are ensured, but also a solid foundation is laid for the subsequent image alignment and fusion, greatly improving the stability and reliability of the entire system.

[0062] S102: performing image alignment on the visible light image and the infrared image based on feature points.

[0063] In this embodiment, image alignment is achieved by detecting and matching feature points in visible light images and infrared images, using advanced feature detection algorithms, including but not limited to SIFT or SURF, which can reliably detect key points in images under a variety of lighting and viewing angle conditions. Once the feature points are detected, the RANSAC algorithm is further used to eliminate false matches and improve matching accuracy. Finally, by calculating the rigid body transformation matrix based on feature point matching, the two images are accurately aligned in the spatial coordinate system. This process not only solves the displacement problem caused by differences in device position and angle between images of different modalities, but also provides an accurate image basis for subsequent multi-scale feature extraction and fusion, significantly improving the accuracy and effect of super-resolution imaging.

[0064] S103: Using a bilateral filtering method to perform denoising on the aligned visible light image and infrared image.

[0065] In this embodiment, in order to remove noise while maintaining image edges and texture details, a bilateral filtering method is used to denoise the visible light image and infrared image after image alignment. Bilateral filtering is a nonlinear multi-scale spatial filter that not only considers the spatial proximity of pixels, but also the similarity of pixel values, so it can effectively remove noise without blurring the image boundary. By reasonably setting the parameters of bilateral filtering, it is possible to balance the denoising effect and detail retention, ensuring that the image preprocessing stage can effectively reduce noise interference while maintaining the clarity and realism of the image, providing high-quality image input for subsequent feature extraction and fusion, and significantly improving the performance and user experience of the entire super-resolution imaging system.

[0066] Specifically, step S2 specifically includes:

[0067] S201: Preprocessing the visible light image and infrared image after image alignment and denoising, and inputting the preprocessed images into a pre-trained feature extraction model.

[0068] In this embodiment, the preprocessed visible light image and infrared image are first size-standardized to ensure that they have the same resolution, which is convenient for subsequent feature extraction operations. In addition, in order to further improve the efficiency and accuracy of feature extraction, the image is normalized and the pixel values ​​are scaled to the interval [0,1], which helps the learning convergence of the model and avoids the gradient disappearance or explosion problem caused by too large or too small value range. The preprocessed image is then input into the pre-trained feature extraction model.

[0069] S202: Perform feature extraction on the visible light image and the infrared image using convolutional layers of different levels in the feature extraction model to capture first multi-scale features of the visible light image and second multi-scale features of the infrared image.

[0070] In this embodiment, the hierarchical structure of the deep convolutional neural network (CNN) is fully utilized to capture multi-scale features through convolutional layers of different depths. The low-level convolutional layers focus on capturing the basic structure of the image, such as edge, texture, and color information, which are crucial to the local details of the image. As the depth of the network increases, the high-level convolutional layers gradually focus on more complex patterns and concepts, such as the shape, contour, and structural layout of objects, which are particularly important for understanding the global content and contextual relationships of the image. In this way, the first multi-scale features are extracted from the visible light image, and the second multi-scale features are extracted from the infrared image, providing a comprehensive and rich source of information for subsequent feature fusion.

[0071] S203: Performing feature tuning on the first multi-scale features of the visible light image and the second multi-scale features of the infrared image, wherein the feature tuning includes scale and resolution unification, channel adjustment, and semantic alignment.

[0072] In this embodiment, feature tuning is a key step to ensure that the two modal features can be effectively fused. First, scale and resolution unification is performed, that is, features of different scales are adjusted to the same spatial resolution. This step eliminates the scale difference and enables the features to be compared and fused at the same spatial position. Next, channel adjustment is performed, that is, the number of feature maps is adjusted to ensure that the two modal features have the same number of channels, which is convenient for subsequent feature weighting and fusion operations. Finally, semantic alignment is performed, which is achieved by calculating the similarity matrix between the two modal features. The matrix reflects the semantic correlation between the visible light image features and the infrared image features. Based on this, the features of the two modalities can be intelligently merged by weighted averaging to maximize their complementary advantages, and finally obtain comprehensive multi-scale features. This series of feature tuning operations not only solves the incompatibility problem of features between modalities, but also enhances the representation ability of features, provides high-quality input for subsequent super-resolution reconstruction and depth estimation, and significantly improves the overall performance of the system.

[0073] Specifically, step S3 specifically includes:

[0074] S301: Assign weights to the first multi-scale feature and the second multi-scale feature based on the pre-trained attention module, and perform weighted fusion based on their respective weights to generate a comprehensive multi-scale feature.

[0075] In this embodiment, a pre-trained cross-modal attention module is used, which can automatically assign weights to the features of the two modalities according to the complementarity and correlation between the first multi-scale features of the visible light image and the second multi-scale features of the infrared image. Specifically, it is achieved by calculating the similarity matrix between the feature maps, and each element in the matrix represents the degree of similarity between the visible light feature and the infrared feature at a specific position. Based on these similarity scores, the attention module can identify which features are more critical to subsequent fusion and reconstruction tasks, and thus assign higher weights. This adaptive weight allocation mechanism enables the system to intelligently screen and fuse the most informative features, avoid interference from redundant or irrelevant features, and generate more accurate and effective comprehensive multi-scale features.

[0076] S302: Super-resolution the comprehensive multi-scale features by combining an upsampling method and a reconstruction network based on a deep learning model to generate an initial super-resolution image.

[0077] In this embodiment, the comprehensive multi-scale features are first upsampled to increase the resolution to the target level, which can be achieved by including but not limited to nearest neighbor interpolation, bilinear interpolation or deconvolution layers. Subsequently, the upsampled features are input into a reconstruction network based on deep learning, which is specially trained to recover the details of high-resolution images from low-resolution features. The reconstruction network consists of a series of convolutional layers, activation functions and upsampling modules, which can gradually reconstruct the details and structure of the image and finally output a high-resolution image. In this process, the rich information in the comprehensive multi-scale features plays a key role, helping the network to better understand and restore the details of the image, and generate an initial super-resolution image with rich details and significantly improved clarity.

[0078] Specifically, step S4 specifically includes:

[0079] S401: Using a depth estimation algorithm to generate an initial depth map of an initial super-resolution image.

[0080] In this embodiment, a depth estimation algorithm based on a convolutional neural network (CNN) framework is used to predict the depth information from a single image. Specifically, the initial super-resolution image is used as input, and the depth features of the image are extracted step by step through multi-layer convolution, pooling, and upsampling operations, and finally an initial depth map matching the size of the input image is generated. Each pixel value in the initial depth map represents the relative distance of the corresponding point in three-dimensional space. This step provides a basis for subsequent focal plane area segmentation and depth information optimization.

[0081] S402: Based on a preset depth threshold, the initial depth map is preliminarily segmented into a plurality of initial focal plane areas of different depths using an image segmentation method.

[0082] In this embodiment, the initial depth map is divided into several focal plane areas with different depth levels according to a preset depth threshold. Specifically, a set of depth intervals are defined, each interval corresponds to a focal plane area, and the pixel can be assigned to the corresponding focal plane area by comparing the depth value of each pixel with the depth threshold. This can effectively distinguish different depth levels in the scene, and provide a preliminary area division for subsequent depth information optimization and clarity enhancement.

[0083] S403: Calculate the average depth value of each focal plane area.

[0084] In this embodiment, for each focal plane area obtained by preliminary segmentation, the average depth value of all its pixels is calculated as the representative depth of the area, aiming to quantify the depth characteristics of each focal plane area and provide a basis for subsequent boundary adjustment. The average depth value can reflect the approximate distance of objects in the area and help identify and distinguish different depth levels.

[0085] S404: Adjusting the boundary of the initial focal plane region based on the average depth value of the adjacent initial focal plane regions to obtain a final focal plane region.

[0086] In this embodiment, the boundaries of the initially segmented focal plane regions are optimized and adjusted by comparing the average depth values ​​of adjacent focal plane regions. Specifically, if the difference in the average depth values ​​of two adjacent regions is less than a preset threshold, it is considered that the two regions may belong to the same depth level, and the two regions can be merged; conversely, if the difference is greater than the threshold, the boundary is kept unchanged or further refined to ensure accurate segmentation of different depth levels.

[0087] Specifically, in step S404, it is specifically as follows:

[0088] The average depth value of each of the initial focal plane regions is compared with the average depth value of adjacent initial focal plane regions, and the difference is evaluated according to the comparison result.

[0089] In this embodiment, the average depth value of each initial focal plane area and its adjacent areas is compared one by one, and the difference between the two is calculated. For example, the absolute difference between the two average depth values ​​is calculated. The purpose of evaluating the difference is to determine whether the adjacent areas belong to the same depth level or whether there is a significant depth change between them. By setting a predefined depth threshold, it can be determined whether two adjacent areas should be considered as part of the same depth level or should remain independent.

[0090] The boundary of each of the initial focal plane regions is adjusted based on the difference.

[0091] In this embodiment, based on the above calculated difference, the boundary of each initial focal plane area is dynamically adjusted. If the difference between two adjacent areas is less than the preset depth threshold, it means that they are similar in depth, and they can be considered to be merged into a larger focal plane area; on the contrary, if the difference is greater than the threshold, it means that they are significantly different in depth, and the existing boundary should be maintained, and it may even be necessary to refine the boundary to more accurately reflect the depth change. This process is carried out in an iterative manner until the boundaries of all adjacent areas are appropriately adjusted.

[0092] Each of the initial focal plane regions after boundary adjustment is smoothed to generate a final focal plane region.

[0093] In this embodiment, after completing the boundary adjustment, smoothing is performed on each focal plane area to reduce discontinuity and jagged effects on the boundary. The smoothing can be achieved by applying a filter. The filter can smooth the depth value near the boundary to make the transition more natural, thereby generating the final focal plane area, improving the overall quality and visual effect of the depth map, and ensuring the coherence and consistency of the depth information.

[0094] Specifically, step S5 specifically includes:

[0095] S501: assigning a weight to each pixel in a final focal plane region, wherein the weight is determined based on a depth of the final focal plane region where the pixel is located and a relative distance from other final focal plane regions.

[0096] In this embodiment, a weight value is assigned to each pixel point in the final focal plane area, wherein the weight calculation comprehensively considers the average depth of the focal plane area where the pixel is located, and the relative distance between the area and other focal plane areas. Specifically, pixels in areas with shallower depths (i.e., areas closer to the lens) are usually assigned higher weights, while pixels in areas with deeper depths (i.e., areas farther from the lens) have lower weights. At the same time, pixels that are closer to adjacent focal plane areas receive higher weights because they are more likely to form a continuous scene visually. In this way, it is possible to ensure that different depth levels in the image are properly valued, especially in areas where depth changes are more complex.

[0097] S502: Perform definition enhancement processing on pixels in each final focal plane area according to the assigned weights to obtain a focal plane image corresponding to each final focal plane area.

[0098] In this embodiment, the pixels within each final focal plane area are processed for clarity enhancement using the assigned weights. Specifically, a sharpening filter is applied to enhance the edges and details of the image. Since the weight of each pixel reflects its importance in the depth map, pixels with higher weights will experience stronger clarity enhancement, while pixels with lower weights will be less affected. The purpose of this is to maintain a natural transition between different depth layers while enhancing image clarity and avoiding artifacts caused by over-sharpening.

[0099] S503: Perform weighted fusion on all the focal plane images that have undergone clarity enhancement to generate a final super-resolution image.

[0100] In this embodiment, a weighted fusion method is used to integrate all the sharpness-enhanced focal plane images into a final super-resolution image. This fusion process also uses the pixel weights calculated previously to ensure that areas with shallow depth (i.e., closer to the lens) dominate the final image, while areas with deeper depth are appropriately weakened. By linearly combining the contributions of each focal plane image according to their weights, a super-resolution image with high clarity and distinct layers can be generated while retaining the details of all focal plane areas.

[0101] The working principle of a focal plane recognition method for super-resolution imaging provided by the present invention is as follows: the joint processing of infrared images and visible light images is introduced, and the accuracy and consistency of the two image information are ensured through the image preprocessing step; then, the pre-trained feature extraction model is used to extract multi-scale features from the two images respectively, and the cross-modal attention mechanism is used to automatically identify and strengthen the correlation between the two features, while suppressing irrelevant information, thereby generating more accurate and comprehensive comprehensive multi-scale features, and then the depth of each pixel in the image is estimated by the depth estimation method, so as to identify focal planes at different depths; finally, the objects on different focal planes are subjected to targeted clarity enhancement processing according to the depth information of the focal planes, so as to realize the accurate processing and clarity enhancement of focal planes at different depths in the image, which not only improves the accuracy and detail restoration capability of super-resolution imaging, but also enhances the overall clarity and realism of the image.

[0102] Embodiment 2

[0103] The present invention also provides a focal plane recognition system for super-resolution imaging, which is applied to a focal plane recognition system for super-resolution imaging, referring to Figure 2 As shown, the identification system includes:

[0104] The image processing module 100 is used to obtain a visible light image and an infrared image of an image to be super-resolved and perform image preprocessing, wherein the image preprocessing includes image alignment and image denoising.

[0105] The feature extraction module 200 is used to extract multi-scale features from the preprocessed visible light image and infrared image using a pre-trained feature extraction model to obtain a first multi-scale feature and a second multi-scale feature corresponding to the visible light image and the infrared image respectively.

[0106] The super-resolution processing module 300 is used to fuse the first multi-scale features and the second multi-scale features based on a cross-modal attention mechanism, and perform super-resolution processing on the fused comprehensive multi-scale features to obtain an initial super-resolution image.

[0107] The focal plane identification module 400 is used to perform depth estimation on each pixel in the initial super-resolution image, and identify focal planes at different depths according to the estimated depth of each pixel.

[0108] The super-resolution image enhancement module 500 is used to enhance the clarity of objects on different focal planes by utilizing the depth of the focal plane to generate a final super-resolution image.

[0109] The present application is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0110] A person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable rewritable read-only memory (EEPROM), a compact disc (CD-ROM) or other optical disc storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.

[0111] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

Claims

1. A focal plane recognition method for super-resolution imaging, characterized in that: The identification method includes the following steps: S1: Obtaining a visible light image and an infrared image of an image to be super-resolved and performing image preprocessing, wherein the image preprocessing includes image alignment and image denoising; S2: extracting multi-scale features from the preprocessed visible light image and infrared image using a pre-trained feature extraction model to obtain first multi-scale features and second multi-scale features corresponding to the visible light image and infrared image respectively; S3: fusing the first multi-scale feature and the second multi-scale feature based on a cross-modal attention mechanism, and performing super-resolution processing on the fused comprehensive multi-scale features to obtain an initial super-resolution image; S4: performing depth estimation on each pixel in the initial super-resolution image, and identifying focal planes at different depths according to the estimated depth of each pixel; S5: using the depth of the focal plane to enhance the clarity of objects on different focal planes to generate a final super-resolution image; Step S4 specifically includes: S401: Generate an initial depth map of an initial super-resolution image using a depth estimation algorithm; S402: Based on a preset depth threshold, using an image segmentation method to preliminarily segment the initial depth map into a plurality of initial focal plane areas of different depths; S403: Calculate the average depth value of each focal plane area; S404: adjusting the boundary of the initial focal plane area based on the average depth value of the adjacent initial focal plane areas to obtain a final focal plane area; In step S404, specifically: Comparing the average depth value of each of the initial focal plane regions with the average depth value of adjacent initial focal plane regions, and evaluating the difference according to the comparison result; Adjusting the boundary of each of the initial focal plane regions based on the difference; Performing smoothing processing on each of the initial focal plane regions after boundary adjustment to generate a final focal plane region; Step S5 specifically includes: S501: assigning a weight to each pixel in a final focal plane region, wherein the weight is determined based on a depth of the final focal plane region where the pixel is located and a relative distance from the pixel to other final focal plane regions; S502: performing definition enhancement processing on pixels in each final focal plane area according to the assigned weights to obtain a focal plane image corresponding to each final focal plane area; S503: Perform weighted fusion on all the focal plane images that have undergone clarity enhancement to generate a final super-resolution image.

2. The focal plane recognition method for super-resolution imaging according to claim 1, characterized in that: Step S1 specifically includes: S101: collecting a visible light image and an infrared image of the image to be super-resolved; S102: performing image alignment on the visible light image and the infrared image based on feature points; S103: Using a bilateral filtering method to perform denoising on the aligned visible light image and infrared image.

3. The focal plane recognition method for super-resolution imaging according to claim 2, characterized in that: Step S2 specifically includes: S201: preprocessing the visible light image and infrared image after image alignment and denoising, and inputting the preprocessed images into a pre-trained feature extraction model; S202: performing feature extraction on the visible light image and the infrared image using convolution layers of different levels in the feature extraction model to capture first multi-scale features of the visible light image and second multi-scale features of the infrared image; S203: Performing feature tuning on the first multi-scale features of the visible light image and the second multi-scale features of the infrared image, wherein the feature tuning includes scale and resolution unification, channel adjustment, and semantic alignment.

4. The focal plane recognition method for super-resolution imaging according to claim 3, characterized in that: Step S3 specifically includes: S301: assigning weights to the first multi-scale feature and the second multi-scale feature based on the pre-trained attention module, and performing weighted fusion based on the respective weights to generate a comprehensive multi-scale feature; S302: Super-resolution the comprehensive multi-scale features by combining an upsampling method and a reconstruction network based on a deep learning model to generate an initial super-resolution image.

5. A focal plane recognition system for super-resolution imaging, applied to a focal plane recognition method for super-resolution imaging as claimed in any one of claims 1 to 4, characterized in that: The identification system includes: An image processing module, used to obtain a visible light image and an infrared image of an image to be super-resolved and perform image preprocessing, wherein the image preprocessing includes image alignment and image denoising; A feature extraction module, configured to extract multi-scale features from the preprocessed visible light image and infrared image using a pre-trained feature extraction model, to obtain first multi-scale features and second multi-scale features corresponding to the visible light image and infrared image respectively; A super-resolution processing module, used for fusing the first multi-scale features and the second multi-scale features based on a cross-modal attention mechanism, and performing super-resolution processing on the fused comprehensive multi-scale features to obtain an initial super-resolution image; A focal plane recognition module, used to perform depth estimation on each pixel in the initial super-resolution image, and to recognize focal planes at different depths according to the estimated depth of each pixel; The super-resolution image enhancement module is used to enhance the clarity of objects on different focal planes by utilizing the depth of the focal plane to generate a final super-resolution image.

Citation Information

Patent Citations

  • Depth image super-resolution reconstruction method and system

    CN109523470A

  • Multi-mode photoelectric imaging system and method for visible light image and infrared polarization image

    CN118154428A