Lightweight super-resolution imaging method and system based on multi-mode optical flow estimation

Through acoustic injection, OIS lens jitter and multimodal optical flow estimation are controlled, combined with gyroscope data and visual information, and efficient super-resolution imaging is used to use the Swin Transformer network to solve the problems of insufficient computing complexity and accuracy on smartphones, and high-quality and low resource consumption imaging effects are achieved.

CN120238756APending Publication Date: 2025-07-01SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510302272.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The existing super-resolution imaging systems have high computational complexity, insufficient accuracy and poor device adaptability on smartphones, making it difficult to achieve high-quality imaging with high efficiency and low resource consumption.

Method used

The OIS lens regular jitter is controlled through acoustic injection, and multimodal optical flow estimation is performed by combining gyroscope data and visual data. The Swin Transformer network is used to perform high-precision optical flow calculation and feature fusion to generate high-resolution images.

Benefits of technology

It realizes sub-pixel-level optical flow estimation accuracy, reduces calculation amount and resource consumption, improves imaging speed and quality, and adapts to the real-time imaging needs of mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238756A_ABST
    Figure CN120238756A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight super-resolution imaging method and system based on multi-mode optical flow estimation, and the method comprises the following steps: controlling the regular jitter of an OIS lens through acoustic injection, and carrying out the image shooting at a specific position; the method comprises the following steps: acquiring a plurality of frames of shot low-resolution images and gyroscope data, estimating an optical flow vector between a reference frame and an offset frame by using an optical flow estimation network, and outputting high-precision optical flow information; and extracting the features of the low-resolution image by using a super-resolution network, and carrying out pixel-level alignment and fusion on the features of the multi-frame image based on the high-precision optical flow information to generate a high-resolution image. Compared with the prior art, the method has the advantages of being capable of improving optical flow estimation efficiency and super-resolution imaging efficiency, reducing resource consumption, improving imaging quality and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of super-resolution imaging technology, and in particular to a lightweight super-resolution imaging method and system based on multi-modal optical flow estimation. Background Art

[0002] With the continuous evolution of user requirements, the demand for super-resolution (SR) imaging in mobile smartphone cameras is increasing. To meet users' need for clear imaging of distant targets, smartphone manufacturers are adding telephoto lenses. However, the pursuit of a thin design conflicts with the size of telephoto lenses, thus limiting the super-resolution imaging ability. When users zoom in on an image after shooting, especially for long-distance shooting, they often encounter blurred details and noise. This problem indicates an urgent need for a super-resolution technology that can break through the physical limitations of smartphone camera hardware.

[0003] Deep learning-based super-resolution technologies have been widely studied currently. Single-frame super-resolution (SFSR) models achieve super-resolution by processing a single low-resolution (LR) image. However, SFSR methods require a large amount of computing resources and may introduce artifacts or over-smoothing, thus affecting the image quality. In contrast, multi-frame super-resolution (MFSR) models generate high-resolution images by integrating multiple low-resolution images of the same scene from different positions. These methods generate better super-resolution results by relying on actual sampling values. However, MFSR models also require a large amount of computing resources, which causes inference latency problems on mobile devices.

[0004] Mainstream MFSR methods generally include the following steps: capturing multiple frames of images, calculating the optical flow between the reference frame and adjacent frames, aligning pixels, and finally synthesizing a super-resolution image. Among these steps, optical flow estimation is crucial because it directly affects the quality of the finally generated super-resolution image. For 16-fold super-resolution imaging, optical flow calculation with an accuracy of 0.25 pixels needs to be achieved to ensure the accuracy of alignment. At this accuracy, at least four frames of images need to be merged to generate a single high-resolution image through super-resolution reconstruction.

[0005] Due to computational complexity and accuracy issues, the deployment of existing optical flow models on smartphones faces numerous challenges. In the current state-of-the-art multi-frame super-resolution (MFSR) network, the optical flow network PWCNet used has an average end-point error (i.e., pixel alignment error) of 0.65 pixels. However, the parameters of this model account for 72.41% of the resources of the entire super-resolution system, and the model size is as high as 9.37 MB. On the other hand, lightweight networks such as SpyNet have an average accuracy of 4.2 pixels, which is not sufficient to meet the requirements of high-quality super-resolution imaging. Therefore, it remains a great challenge to achieve high-precision optical flow estimation at low computational cost using only RGB information. Researchers have attempted to design multi-modal optical flow estimation methods by introducing data from non-RGB modalities (such as LiDAR, infrared cameras, and radars) to improve accuracy and robustness. However, these methods usually require additional modal data (such as point clouds), which are not common in mobile devices, resulting in an increase in system and model complexity. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems of high computational complexity, insufficient accuracy, and poor device adaptability in the existing super-resolution imaging system, and to provide a lightweight super-resolution imaging method and system based on multi-modal optical flow estimation, which can reduce resource consumption while ensuring imaging quality and efficiency.

[0007] The purpose of the present invention can be achieved through the following technical solutions:

[0008] A lightweight super-resolution imaging method based on multi-modal optical flow estimation, comprising the following steps:

[0009] Use acoustic injection to control the regular jitter of the OIS lens and take images at specific positions;

[0010] Obtain multiple frames of low-resolution images and gyroscope data taken, use the optical flow estimation network to estimate the optical flow vector between the reference frame and the offset frame, and output high-precision optical flow information;

[0011] Use the super-resolution network to extract the features of the low-resolution images, and perform pixel-level alignment and fusion on the features of multiple frames of images based on the high-precision optical flow information to generate high-resolution images.

[0012] The use of acoustic injection to control the regular jitter of the OIS lens is specifically: by playing a sound wave signal with a specific frequency, control the lens movement trajectory to be elliptical.

[0013] The specific image capture at a specific position is as follows: when the lens shakes regularly, the position of the lens is estimated according to the gyroscope data, and 4 specific lens positions are selected as the shooting points. The specific lens positions are the four vertices of an elliptical motion trajectory, which are estimated based on the change of gyroscope data starting from when the audio starts to play.

[0014] During the regular shaking of the lens, the sound frequency is adjusted according to the periodic change of the gyroscope readings to achieve stable shaking of the lens.

[0015] The optical flow estimation network includes a first feature extraction module and an optical flow estimation module. Among them, the first feature extraction module encodes the image data using a convolutional neural network, extracts a low-resolution image feature map, and performs multi-layer perceptron processing on the gyroscope data to generate an angular velocity feature. Combining with the physical model of the OIS lens movement, an association between the gyroscope data and the optical flow vector is established; the optical flow estimation module generates a high-precision optical flow vector based on multi-modal data.

[0016] The optical flow estimation module performs the following steps:

[0017] Initial optical flow estimation: Using the gyroscope data, calculate the initial optical flow caused by the lens movement through a weight matrix;

[0018] Visual correction: Generate a correlation matrix using the image feature map to correct the initial optical flow;

[0019] Multi-modal fusion: Fuse the visually corrected optical flow with the angular velocity feature of the gyroscope, and use the U-Net structure to output a high-precision optical flow vector.

[0020] The super-resolution network includes a second feature extraction module, an alignment module, a feature fusion module, an amplification module, and an output module. Among them, the second feature module uses a multi-layer convolutional neural network to extract the depth features of multiple frames of low-resolution images and generates corresponding low-dimensional feature maps; the alignment module uses high-precision optical flow information to perform pixel-level alignment on the feature maps of multiple frames of images to eliminate the inter-frame motion difference; the feature fusion module uses the window attention mechanism of the Swin Transformer framework to calculate the correlation of features within a local range, and performs weighted fusion of the features within the window based on the correlation. At the same time, a global attention mechanism is introduced to supplement the long-range dependence information to generate a fusion feature map containing rich context information; the amplification module magnifies the fusion feature map to the required size through a two-round amplification strategy. Among them, in each round of amplification, the channel dimension of the feature map is rearranged into the spatial dimension to achieve efficient upsampling, and the Swin Transformer residual block is used to improve the detail quality of the magnified image; the output module is used to generate the final high-resolution image.

[0021] The alignment module uses the optical flow vectors generated by the optical flow estimation network to align the features of the offset frame to the reference frame, and uses bilinear interpolation and feature compensation strategies to reduce the artifacts that may occur during the alignment process.

[0022] The output module post-processes the RAW image according to visualization needs and converts it into an RGB image for visualization. The post-processing includes white balance, denoising, and gamma correction.

[0023] A lightweight super-resolution imaging system based on multi-modal optical flow estimation, which includes a super-resolution imaging model for the method, and:

[0024] Model loading and initialization module: Save the super-resolution imaging model in ONNX format and load it, and initialize the inference environment. During initialization, the size and number of input frames are dynamically adjusted to adapt to the hardware capabilities of different devices;

[0025] Hardware acceleration optimization module: Used to accelerate model inference, and perform the following processes: Cut the input data into small pieces and process them in parallel, and use the FP16 format for calculation;

[0026] Real-time inference module: Used to receive user operation requests and perform super-resolution processing in real time based on the super-resolution imaging model.

[0027] Compared with the prior art, the present invention has the following beneficial effects:

[0028] (1) Improve efficiency

[0029] Improved optical flow estimation efficiency: The present invention incorporates OIS lens motion data as a new modality, combines it with visual data and gyroscope readings, and realizes multi-modal optical flow estimation. By training the network model with an artificial dataset (including simulated gyroscope readings and accurate optical flow annotations), the optical flow calculation accuracy is improved to 0.12 pixels, achieving sub-pixel level alignment accuracy. Compared with traditional optical flow estimation models (such as PWCNet with an error of 0.65 pixels), this technology reduces the error by approximately 81.5%; the number of processed frames is reduced to 4 frames, the overall computational volume is reduced by approximately 50%, and the inference time is shortened to 1.39 seconds.

[0030] Improved super-resolution imaging efficiency: The present invention uses pixel rearrangement technology and a multi-round step-by-step magnification strategy to reduce the computational cost while maintaining high-resolution imaging quality, improving the imaging speed, and reducing the average latency by 88% compared to traditional MFSR systems, meeting the real-time imaging requirements.

[0031] (2) Reduce resource consumption

[0032] Memory Usage Optimization: The present invention saves the optimized model in the ONNX format, and combines mixed-precision computing (FP16) and hardware acceleration (GPU / NPU) for inference optimization. The peak memory occupancy is reduced to 479.4MB, reducing the memory requirement by more than 85% compared to traditional methods, and avoiding lags on mobile devices due to insufficient resources.

[0033] Model File Size Optimization: The present invention saves the optimized model in the ONNX format, and the model size is reduced to 18.92% of the original, facilitating storage and operation on resource-constrained devices such as smartphones.

[0034] (3) Energy Consumption Reduction

[0035] Power Consumption Control: The present invention realizes dynamic memory allocation and energy consumption control, adapts to the mainstream smartphone environment. The power consumption for a single inference is 9.495 joules, far lower than the power consumption requirements of traditional methods (usually exceeding 30 joules), and can achieve long-term operation on smartphones with a battery capacity of 5000mAh. The dynamic power management mechanism reduces device heating and extends battery life.

[0036] (4) Image Quality Improvement

[0037] The present invention improves the fusion efficiency of multi-frame features and the image detail expressiveness by adopting a window attention mechanism and a hierarchical feature fusion technique, enabling the PSNR, SSIM, and LPIPS metrics of 16-fold super-resolution to reach 36.49, 0.8917, and 0.0687 respectively, significantly superior to the existing MFSR technology. The output images have clearer details and more natural colors, meeting the high-standard requirements of long-distance zooming and professional image processing.

[0038] (5) Ease of Operation and Use

[0039] User Friendliness: The present invention supports users to achieve real-time super-resolution imaging through simple operations (such as double-tapping the screen to zoom in), providing a smooth user experience.

[0040] Wide Compatibility: The present invention supports deployment on various mobile device platforms (such as Android and iOS), and can be extended to video enhancement and other multi-frame image processing tasks.

[0041] (6) Automatic and Efficient Shooting

[0042] The present invention monitors gyroscope data, adjusts the sound frequency of acoustic injection to ensure stable lens jitter, and determines the shooting position of the camera lens and the shooting time point by monitoring gyroscope data and the start time of jitter. By determining the shooting position, the information gap during shooting is maximized, and the image with the strongest information complementarity is used for super-resolution processing. By adjusting the sound frequency, the camera shakes stably, avoiding blurring during shooting. Brief Description of the Drawings

[0043] Figure 1 is a flowchart of the method of the present invention;

[0044] Figure 2 is a schematic structural diagram of the super-resolution imaging model of the present invention;

[0045] Figure 3 is a schematic structural diagram of the optical flow estimation network of the present invention;

[0046] Figure 4 is a schematic structural diagram of the super-resolution network of the present invention. Detailed Embodiments

[0047] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and detailed implementation manners and specific operation processes are given, but the protection scope of the present invention is not limited to the following embodiments.

[0048] This embodiment first provides a lightweight super-resolution imaging method based on multi-modal optical flow estimation, as Figure 1 and Figure 2 shown, including the following steps:

[0049] S1, using acoustic injection to control the regular jitter of the OIS lens and taking images at specific positions.

[0050] Multi-view parallax shooting is one of the necessary conditions for super-resolution imaging. By acquiring image data from different perspectives, multiple samplings of the scene can be achieved, providing a basis for subsequent image processing and analysis. In this embodiment, a camera with an OIS module is used in combination with acoustic injection technology for control, enabling precise angle adjustment and image synchronization, so as to sample additional information for the scene that requires super-resolution during the shooting process. This step uses acoustic injection to affect the OIS module to control the regular jitter of the lens and takes images at the fixed position of the lens jitter for super-resolution multi-frame fusion.

[0051] First, perform acoustic injection: By playing a sine wave of a specific frequency, such as playing 19.6 kHz on a Xiaomi 11 mobile phone, the gyroscope and accelerometer of the camera are affected to precisely control the lens movement trajectory to be elliptical, thereby enhancing the controllability of optical flow estimation.

[0052] Secondly, image capturing is performed: according to the gyroscope data, the scene is captured at specific positions. When the lens undergoes regular jitter, the position of the lens is estimated based on the gyroscope data, and 4 specific lens positions are selected as the shooting points. The specific lens positions are the four vertices of an elliptical motion trajectory (i.e., two outermost boundary points and two points on both sides of the semi - length of the ellipse). Starting from when the audio starts playing, photos are taken for sampling when the gyroscope readings reach the peak, valley, and zero point respectively. Subsequently, completely identical pictures are filtered out according to image similarity, and 4 frames of pictures after specific screening can be obtained.

[0053] Finally, feedback adjustment is required during the shooting process: the audio frequency is adjusted by changing the gyroscope data to ensure stable jitter of the lens. During the lens jitter process, if a sound signal of the same frequency is played, the lens jitter may change slightly due to the thermal effect of the gyroscope, thus affecting the shooting. Therefore, it is necessary to adjust the sound frequency according to the change in the gyroscope readings, that is, the periodic change of the gyroscope readings, to achieve stable jitter of the lens.

[0054] The core improvement points of this step are as follows, to ensure the stable movement of the camera lens and capture at specific positions, providing data support for the subsequent system:

[0055] 1. Control the gyroscope readings based on the acoustic injection method, thereby affecting the OIS module to make it perform regular jitter.

[0056] 2. Determine the approximate position of the camera based on the gyroscope readings and perform specific frame shooting.

[0057] 3. Dynamically adjust the sound frequency based on the gyroscope readings to make the lens jitter stably.

[0058] S2. Obtain multiple frames of low - resolution images and gyroscope data captured, and use the optical flow estimation network to estimate the optical flow vector between the reference frame and the offset frame, and output high - precision optical flow information.

[0059] The multi - modal optical flow estimation module is one of the core technologies of this method, aiming to achieve high - precision and low - computational - complexity optical flow estimation by fusing multi - modal data (OIS lens motion information, visual data, and gyroscope readings). Traditional optical flow estimation methods usually rely on visual data, resulting in low estimation accuracy in complex motion scenes or low - texture scenes. At the same time, these methods have high computational complexity and are difficult to be deployed in real - time on resource - constrained mobile devices. The present invention significantly improves the optical flow estimation accuracy and reduces the computational complexity by introducing new modal data related to OIS control.

[0060] First, perform data acquisition: Record the corresponding gyroscope readings during image capture. The acquired content includes: a) Image data: Multiple frames of low-resolution images, including a reference frame and multiple offset frames; b) Gyroscope data: Triaxial angular velocity information for inferring lens movement. In this step, additional gyroscope data is acquired to calculate accurate optical flow.

[0061] Secondly, perform feature extraction. As Figure 3 shown, the optical flow estimation network includes a first feature extraction module and an optical flow estimation module. Among them, the first feature extraction module encodes the image data using a convolutional neural network (CNN), extracts the low-resolution image feature map, and performs multi-layer perceptron (MLP) processing on the gyroscope data to generate angular velocity features. Combining with the physical model of OIS lens movement, an association between the gyroscope data and the optical flow vector is established; the optical flow estimation module generates a high-precision optical flow vector based on multi-modal data.

[0062] Then, use the optical flow estimation module to perform optical flow estimation and generate a high-precision optical flow vector. It includes the following steps:

[0063] a) Initial optical flow estimation: Use the gyroscope data to calculate the initial optical flow caused by lens movement through a weight matrix.

[0064] b) Visual correction: Generate a correlation matrix using the image feature map to correct the initial optical flow and reduce the impact of missing visual information on the estimation result. The specific process includes: First, use a Feature Encoder to extract features from two consecutive frames of images. This encoder shares weights for the two frames of images, and the output feature map has a resolution of 1 / 8 of the original image. Perform a full-pixel comparison on the feature maps of the first and second frames, calculate the dot product between each pixel pair, and generate a four-dimensional correlation volume with a size of [H, W, H, W], representing the similarity between each pixel in the first frame and all pixels in the second frame. Perform multi-level average pooling on the above correlation volume to construct a multi-scale correlation pyramid to capture motion information at different scales. In each iteration, use the current optical flow estimation to retrieve correlation features from the correlation pyramid, and combine them with context features to calculate the optical flow increment Δflow through a gated recurrent unit (GRU). Add this increment to the current optical flow estimation, gradually approaching the true optical flow, and finally complete the correction and upsampling after multiple iterations.

[0065] c) Multi-modal fusion: Fuse the visually corrected optical flow with the angular velocity features of the gyroscope, and use a U-Net structure to further improve the optical flow estimation accuracy and output a high-precision optical flow vector.

[0066] Next, perform model optimization and training: Generate an artificial dataset, including simulated gyroscope data, multi-frame displacement images, and optical flow annotation data, and optimize the model parameters by minimizing the optical flow estimation error (such as the endpoint error EPE).

[0067] Result verification: After training, the model achieves an optical flow estimation accuracy of 0.12 pixels in various scenarios, and the number of parameters is only 145 million.

[0068] The core improvement points of the optical flow estimation network in this embodiment are as follows:

[0069] 1. Introduce new modality data: For the first time, combine the OIS lens motion information and the gyroscope reading value as an auxiliary modality for optical flow estimation to enhance the robustness in complex scenarios.

[0070] 2. Lightweight design: Compared with traditional models (such as PWCNet), the number of parameters is reduced by about 60%, and the inference speed is significantly improved.

[0071] 3. Accuracy improvement: Through multi-modal fusion, high-precision optical flow estimation is achieved at the sub-pixel level, and the error is reduced to 0.12 pixels.

[0072] The optical flow estimation network is applied to: a) Frame alignment: Provide high-precision optical flow vectors for the subsequent frame alignment module to ensure the accurate alignment of multi-frame images. b) Resource optimization: Implement real-time optical flow estimation on mobile devices to reduce memory and computing resource consumption. c) Extended scenarios: Apply to fields such as multi-frame imaging and video super-resolution, especially for high-performance deployment on resource-constrained devices.

[0073] S3. Use the super-resolution network to extract the features of the low-resolution image, and perform pixel-level alignment and fusion on the features of multi-frame images based on the high-precision optical flow information to generate a high-resolution image.

[0074] The super-resolution network is one of the core components of this technical solution. Its function is to generate high-quality super-resolution images by extracting, aligning, fusing, and magnifying the features of multi-frame low-resolution images. Traditional super-resolution networks have problems such as large computational complexity and high memory occupancy during the inference process, especially difficult to meet the real-time and energy efficiency requirements on mobile devices. The present invention significantly improves the processing efficiency and imaging quality by introducing a super-resolution network based on Swin Transformer and combining frame alignment optimization and feature fusion technology.

[0075] As Figure 4 shown, the super-resolution network includes a second feature extraction module, an alignment module, a feature fusion module, a magnification module, and an output module.

[0076] The input of the second feature module is 4 frames of low-resolution images. A multi-layer convolutional neural network (CNN) is used to extract features from each frame of the image, obtaining the depth features of multiple frames of low-resolution images and generating corresponding low-dimensional feature maps.

[0077] The alignment module uses high-precision optical flow information to perform pixel-level alignment on the feature maps of multiple frames of images, eliminating the inter-frame motion differences. Specifically, it uses the optical flow vectors generated by the optical flow estimation network to align the features of the offset frame to the reference frame, and uses bilinear interpolation and feature compensation strategies to reduce the artifacts that may occur during the alignment process. In the specific operation, first, a pixel grid with the same size as the feature map is constructed, and the coordinates of each pixel are determined by its position in the image. Then, according to the optical flow vectors, the new position of each pixel is calculated, that is, the original pixel coordinates are added to the optical flow displacement amount to obtain the offset coordinates. Since the deformed coordinates are usually continuous values and may not be exactly aligned with the pixel grid, the bilinear interpolation method is used to perform weighted average calculation according to the values of the four nearest pixels around the target position, thus realizing the feature transformation at the sub-pixel level. In addition, in order to reduce the problem of information loss that may occur for boundary pixels after the optical flow transformation, a feature compensation strategy is introduced. When the target coordinates exceed the range of the original image, boundary filling is used, that is, the pixel value closest to the edge is used for filling. These compensation strategies can effectively avoid the loss of boundary information and improve the quality of the aligned images.

[0078] The feature fusion module generates a fused feature map containing rich context information by fusing the aligned features of multiple frames. Specifically, it uses the window attention mechanism of the Swin Transformer framework to calculate the correlation of features within a local range, performs weighted fusion of the features within the window based on the correlation, and at the same time introduces a global attention mechanism to supplement the long-range dependence information, generating a fused feature map containing rich context information. This method improves the feature expression ability through a hierarchical fusion method and introduces a residual structure to avoid information loss.

[0079] The magnification module magnifies the fused feature map to the target resolution through two rounds of magnification strategies. Among them, each round of magnification includes the following core steps:

[0080] 1) Pixel rearrangement technique: Achieve efficient upsampling by rearranging the channel dimension of the feature map into the spatial dimension.

[0081] 2) Swin Transformer residual block: Use the Swin Transformer residual block to improve the detail quality of the magnified image.

[0082] In this embodiment, the feature map is gradually magnified from 112×112 to 448×448 in two steps, avoiding the artifacts and blurring caused by single-step magnification.

[0083] The output module is used to generate a final RAW format high-resolution image with a 16-fold magnification. In addition, post-processing such as white balance, denoising, and gamma correction can be performed on the RAW image according to visualization needs, and it is converted into an RGB image for visualization display.

[0084] The core improvement points of this step are as follows:

[0085] 1. Introduce Swin Transformer: Utilize its window attention mechanism to enhance the fusion effect of multi-frame features and reduce the computational complexity of global attention.

[0086] 2. Lightweight design: By introducing pixel rearrangement technology and residual structure, reduce the number of model parameters and the amount of inference calculation.

[0087] 3. Multi-step magnification strategy: Adopt a two-round step-by-step magnification method to enhance the image detail expression while avoiding artifacts.

[0088] In this embodiment, the module parameters are designed as follows: Input resolution: 112×112 (multi-frame low-resolution images); Output resolution: 448×448 (16-fold super-resolution image); Feature dimension: The size of the encoded single-frame feature map is 96 dimensions; Transformer window size: Each window contains 7×7 feature units; Number of residual blocks: 4 residual blocks are used in the feature fusion stage and the magnification stage respectively.

[0089] According to the above settings, the experimental results and performance are as follows:

[0090] Processing time: The average processing time is 1.39 seconds; Memory usage: The running memory occupancy is approximately 479.4MB; Imaging quality: Peak signal-to-noise ratio (PSNR): 36.49. Structural similarity index (SSIM): 0.8917. Learned perceptual image patch similarity (LPIPS): 0.0687.

[0091] The super-resolution network proposed in this embodiment can be applied to: a) Mobile devices: Suitable for real-time super-resolution imaging in smartphones. b) Video enhancement: Can be extended to video super-resolution tasks. c) Professional imaging: Suitable for fields such as medical image magnification and satellite image processing.

[0092] The above is the introduction of the method embodiment. The following further illustrates the solution of the present invention through the system embodiment.

[0093] This embodiment provides a lightweight super-resolution imaging system based on multi-modal optical flow estimation. The designed supporting mobile device-related modules are aimed at realizing the efficient operation of the above method to adapt to resource-constrained hardware environments such as smartphones. The deployment of traditional super-resolution methods on mobile devices often faces problems such as high computational load, long inference latency, and excessive power consumption. Through software and hardware optimization, the present invention significantly improves the real-time performance and energy efficiency of super-resolution processing, enabling the system to run smoothly on mainstream smartphones. The system includes a super-resolution imaging model for the above method, and:

[0094] (1) Model Loading and Initialization Module

[0095] Function: Load the optimized super-resolution imaging model and initialize the inference environment.

[0096] Improvement point: Save the model in ONNX format to support cross-platform deployment. The input frame size and quantity are dynamically adjusted during the initialization process to adapt to the hardware capabilities of different devices.

[0097] (2) Hardware Acceleration Optimization Module

[0098] Function: Utilize the hardware characteristics of smartphones (such as GPU, NPU) to accelerate model inference.

[0099] Core technologies:

[0100] a. Tensor sharding: Split the input data into small pieces and process them in parallel to improve the processing speed.

[0101] b. Mixed-precision calculation: Use the FP16 format for calculation, taking into account both inference speed and accuracy.

[0102] c. Platform optimization: Special optimization is carried out for different mobile phone processors (such as Qualcomm Snapdragon, Samsung Exynos, Apple A-series chips).

[0103] (3) Real-time Inference Module

[0104] Function: Receive user operation requests and perform super-resolution processing in real time.

[0105] Process: 1. Receive an image zoom request (such as double-clicking to zoom in on a specified area). 2. Automatically select multiple frames of images and preprocess them. 3. Start super-resolution inference, generate the enlarged image, and output it to the user interface.

[0106] The average inference time of the system in this embodiment is 1.39 seconds, and delay management can also be achieved: reducing UI blocking through an asynchronous processing framework.

[0107] (4) Cross-platform Compatibility Module:

[0108] Function: Support deployment on multiple smart devices (Android, iOS).

[0109] Implementation method:

[0110] a. Build a model inference environment using cross-platform frameworks (such as TensorFlow Lite, ONNX Runtime).

[0111] b. Customize and optimize for different platforms (such as iOS Metal API and Android Vulkan API).

[0112] In summary, the core improvement points of this system except for the super-resolution imaging model are as follows:

[0113] 1. Model lightweight: The model file size (ONNX format) of the present invention is reduced by 81.08%, and the storage occupancy is significantly reduced. The peak inference memory is 479.4MB, which is suitable for most mid- to high-end mobile devices.

[0114] 2. Real-time performance: The present invention achieves near-real-time super-resolution inference, and the latency is controlled within 1.39 seconds. The GPU and NPU work together to improve the inference efficiency.

[0115] 3. Energy efficiency improvement: The power consumption of a single inference of the present invention is only 9.495 joules, which matches the battery capacity of 5000mAh of mainstream smartphones. By dynamically adjusting the operating state, the device heating and battery consumption are reduced.

[0116] 4. Compatibility and scalability: The present invention supports running on multiple processors and operating systems. It can also be extended to other multi-frame image processing tasks, such as video enhancement and real-time image restoration.

[0117] The following experimental equipment is used in this embodiment to verify the system performance:

[0118] Test equipment: Multiple mainstream Android phones (Xiaomi series).

[0119] Platform environment: Android 12, optimized version of TensorFlow Lite.

[0120] Experimental results: Average inference time: 1.39 seconds; Peak memory occupancy: 479.4MB; Imaging quality metrics: Peak signal-to-noise ratio (PSNR): 36.49, Structural similarity index (SSIM): 0.8917, Learned perceptual image patch similarity (LPIPS): 0.0687.

[0121] Accordingly, the application scenarios of the system proposed in this embodiment are as follows:

[0122] 1. Smartphones:

[0123] Provide a real-time image magnification function, which is suitable for users to view the details of distant scenes.

[0124] It can be embedded into the mobile phone camera APP to enhance the shooting effect.

[0125] 2. Online platform:

[0126] It can be extended to cloud deployment to provide image enhancement services for users through the network.

[0127] 3. Other devices:

[0128] Deployed on tablets and wearable devices for real-time image processing.

[0129] The present invention has developed a lightweight super-resolution system designed specifically for mobile devices, which supports real-time lightweight super-resolution imaging. Given the superiority of the MFSR method in terms of performance and the convenience of obtaining multiple frames of information of the same scene through a mobile camera, the MFSR technology is selected. In the present invention, a method for enhancing optical flow estimation by introducing new modal data directly related to optical flow is proposed. The design objectives of the present invention are: on the one hand, to reduce the complexity of the model, and on the other hand, to improve the accuracy of inference. Related research has demonstrated a strong correlation between the lens movement and optical flow in an optically image stabilized (OIS) camera. The present invention innovatively combines the lens movement controlled by OIS, visual data, and new modal data read by a gyroscope sensor to construct a lightweight and high-performance multimodal optical flow estimation network based on a neural network. Aiming at the problem of difficult acquisition of training data, the present invention synthesizes an artificial data set, which includes simulated gyroscope data, multi-frame displacement images, and accurate optical flow annotation data. Using this data set, a high-precision optical flow estimation network module is successfully trained. Experimental verification shows that the multimodal optical flow estimation module proposed in the present invention has significant advantages in terms of the scale of model parameters and the accuracy of optical flow estimation compared with the state-of-the-art (SOTA) optical flow estimation models. The present invention only requires 145 million parameters to achieve an optical flow estimation accuracy of 0.12 pixels, thus significantly reducing the system resource occupancy. The results show that the present invention has successfully broken through the trade-off limit between the model size and accuracy in optical flow estimation by introducing new modal data related to the lens movement controlled by OIS.

[0130] Based on the above multi-modal optical flow module, the present invention proposes a real-time 16-fold super-resolution (SR) imaging system. This system utilizes the Swin Transformer framework to fuse multiple aligned low-resolution (LR) images to generate high-resolution images. In this embodiment, a prototype of this system is implemented on an Android smartphone, and super-resolution imaging is completed by directly processing multiple frames of images captured in RAW format. Compared with compressed formats such as PNG and JPG, the RAW format has the significant advantage of being uncompressed, capable of completely retaining all data from the camera CMOS sensor without loss. This system can provide near-real-time super-resolution imaging functionality, suitable for users to zoom in and view details of distant scenes, such as double-clicking on the screen to locally magnify a specific area. Experimental verification shows that the system of the present invention achieves 16-fold super-resolution enhancement on the tested smartphone CPU, capable of magnifying an 112×112 pixel image to 448×448 pixels, with an average processing time of only 1.39 seconds. The peak memory usage of the system is approximately 479.4MB, enabling it to operate near real-time on mobile devices. Through experimental comparative analysis with existing mainstream multi-frame super-resolution (MFSR) systems, the present invention shows significant advantages in terms of model lightweight and imaging quality. Specifically, the inference latency is reduced by up to 88%, and the model file size (such as the.onnx file) is reduced by up to 81.08%. More importantly, this system provides excellent imaging quality in 16-fold super-resolution imaging, reaching a peak signal-to-noise ratio (PSNR) of 36.49, a structural similarity index (SSIM) of 0.8917, and a learned perceptual image patch similarity (LPIPS) score of 0.0687. In summary, the system of the present invention has significant advantages such as lightweight and low latency, and can be conveniently integrated into mobile devices and network applications. Considering that the battery capacity of mainstream smartphones is usually 5000mAh (about 66600 joules), the inference process of this system has significant energy efficiency.

[0131] The main contributions of the present invention include the following points:

[0132] 1. A novel acoustic injection lens control module is proposed, enabling the mobile phone camera to capture multi-view images of the same scene at a fixed position.

[0133] 2. A novel multi-modal optical flow estimation network is proposed, which fuses lens motion information as a new modal data, successfully realizing an optical flow estimation model with low computational complexity and high accuracy.

[0134] 3. Based on the proposed multi-modal optical flow estimation model, the present invention designs a lightweight super-resolution (SR) network. Based on Swin Transformer, this network can be efficiently deployed on a smartphone to achieve real-time inference for 16-fold super-resolution imaging.

[0135] 4. A prototype system was implemented and deployed on various Android smartphones. The system is designed to apply 16-fold super-resolution imaging when the user zooms in on a specific area of an image.

[0136] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A lightweight super-resolution imaging method based on multimodal optical flow estimation, characterized in that: The following steps are involved: Use acoustic injection to control the regular shaking of the OIS lens and capture images at specific locations; Obtain multiple frames of low-resolution images and gyroscope data, use the optical flow estimation network to estimate the optical flow vector between the reference frame and the offset frame, and output high-precision optical flow information; A super-resolution network is used to extract features of low-resolution images, and the features of multiple frames are aligned and fused at the pixel level based on high-precision optical flow information to generate high-resolution images.

2. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 1, characterized in that: The method of controlling the regular shaking of the OIS lens by using acoustic injection specifically includes: controlling the movement trajectory of the lens to be an ellipse by playing a sound wave signal of a specific frequency.

3. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 2, characterized in that: The image shooting at a specific position is specifically as follows: when the lens shakes regularly, the position of the lens is estimated according to the gyroscope data, and 4 specific lens positions are selected as shooting points. The specific lens positions are the four vertices of an elliptical motion trajectory, starting from when the audio starts playing, and are estimated based on the changes in gyroscope data.

4. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 1, characterized in that: During regular lens shaking, the sound frequency is adjusted according to the periodic changes in the gyroscope readings to achieve stable lens shaking.

5. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 1, characterized in that: The optical flow estimation network includes a first feature extraction module and an optical flow estimation module, wherein the first feature extraction module uses a convolutional neural network to encode image data, extracts a low-resolution image feature map, and performs multi-layer perceptron processing on gyroscope data to generate angular velocity features, and combines the physical model of OIS lens movement to establish an association between gyroscope data and optical flow vectors; the optical flow estimation module generates high-precision optical flow vectors based on multimodal data.

6. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 5, characterized in that: The optical flow estimation module performs the following steps: Initial optical flow estimation: Using gyroscope data, the initial optical flow caused by lens motion is calculated through a weight matrix; Visual correction: Use the image feature map to generate a correlation matrix to correct the initial optical flow; Multimodal fusion: The visually corrected optical flow is fused with the angular velocity features of the gyroscope, and the U-Net structure is used to output high-precision optical flow vectors.

7. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 1, characterized in that: The super-resolution network includes a second feature extraction module, an alignment module, a feature fusion module, an amplification module and an output module, wherein the second feature module uses a multi-layer convolutional neural network to extract deep features of multiple frames of low-resolution images and generate corresponding low-dimensional feature maps; the alignment module uses high-precision optical flow information to perform pixel-level alignment on the feature maps of multiple frames of images to eliminate inter-frame motion differences; the feature fusion module uses the window attention mechanism of the Swin Transformer framework to calculate the correlation of features in a local range, and weightedly fuses the features in the window based on the correlation, while introducing a global attention mechanism to supplement long-distance dependency information to generate a fused feature map containing rich contextual information; the amplification module amplifies the fused feature map to the required size through a two-round amplification strategy, wherein each round of amplification rearranges the channel dimension of the feature map into a spatial dimension to achieve efficient upsampling, and uses the Swin Transformer residual block to improve the detail quality of the amplified image; the output module is used to generate a final high-resolution image.

8. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 7, characterized in that: The alignment module uses the optical flow vector generated by the optical flow estimation network to align the features of the offset frame to the reference frame, and uses bilinear interpolation and feature compensation strategies to reduce artifacts that may be generated during the alignment process.

9. The lightweight super-resolution imaging method based on multimodal optical flow estimation according to claim 7, characterized in that: The output module performs post-processing on the RAW image according to visualization requirements and converts it into an RGB image for visualization. The post-processing includes white balance, denoising and gamma correction.

10. A lightweight super-resolution imaging system based on multimodal optical flow estimation, characterized in that: The system comprises a super-resolution imaging model for implementing the method according to any one of claims 1 to 9, and: Model loading and initialization module: using ONNX format to save the super-resolution imaging model and load it, initializing the inference environment, wherein the size and number of input frames are dynamically adjusted during the initialization process to adapt to the hardware capabilities of different devices; Hardware acceleration optimization module: used to accelerate model reasoning, performing the following process: splitting input data into small blocks and processing them in parallel, and using FP16 format for calculation; Real-time inference module: used to receive user operation requests and perform super-resolution processing in real time based on the super-resolution imaging model.