Image processing method, apparatus, device, storage medium and computer program product

By integrating spatial-temporal implicit coding and pixel-by-pixel decoding of the image processing model, RS frame correction, deblurring, and frame interpolation tasks are combined to solve the RS distortion and blurring problems of cameras in fast-moving scenes, and achieve efficient recovery of clear global shutter frames.

CN118590780BActive Publication Date: 2026-05-08HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HONG KONG UNIV OF SCI & TECH (GUANGZHOU)
Filing Date
2024-05-30
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing technologies, frames captured by consumer-grade cameras based on CMOS sensors in fast-moving scenes exhibit RS distortion and blurring. Independent processing of RS frame correction, deblurring, and frame interpolation leads to accumulated errors and human traces, reducing image clarity.

Method used

An image processing model is used for spatial-temporal implicit coding, exposure time embedding, and pixel-by-pixel decoding. RS frame correction, deblurring, and frame interpolation tasks are integrated to recover clear global shutter frames through the high spatiotemporal resolution of the event camera.

Benefits of technology

It improves the sharpness of GS frames recovered from blurred RS frames, reduces accumulated errors and human artifacts, and enhances image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118590780B_ABST
    Figure CN118590780B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method, device, equipment, storage medium and computer program product. The method comprises the following steps: acquiring a rolling shutter (RS) frame collected by an event camera and event data corresponding to the RS frame; encoding the RS frame and the event data through an encoding unit of an image processing model to obtain space-time implicit representation (STR) data corresponding to the RS frame; embedding an exposure time of a global shutter (GS) frame corresponding to the RS frame into the STR data through a time embedding unit of the image processing model to obtain a time tensor corresponding to the RS frame; and performing pixel-by-pixel decoding on the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame through a decoding unit of the image processing model to generate a GS frame. The scheme disclosed by the application can improve the definition of the GS frame recovered from the RS frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of computer vision, and in particular relates to an image processing method, apparatus, device, storage medium and computer program product. Background Technology

[0002] In related technologies, consumer cameras based on CMOS (Complementary Metal Oxide Semiconductor) sensors typically employ the RS (Rolling Shutter) mechanism. However, in fast-moving scenes, frames captured by CMOS-based consumer cameras often exhibit RS distortion and blurring. To improve the sharpness of images captured by the camera, it is necessary to recover high-frame-rate, sharp GS (Global Shutter) frames from the blurred RS frames. This is typically achieved by correcting, deblurring, and interpolating the RS frames to recover the sharp GS frames.

[0003] However, in related techniques, RS frame correction, deblurring, and frame interpolation are typically processed as independent tasks and cascaded through existing image enhancement networks. This approach increases accumulated errors and significant human artifacts, thereby reducing the sharpness of the recovered GS frames. Summary of the Invention

[0004] This application provides an image processing method, apparatus, device, storage medium, and computer program product that can improve the clarity of GS frames recovered from RS frames.

[0005] In a first aspect, embodiments of this application provide an image processing method, the method comprising: acquiring a rolling shutter RS ​​frame and event data corresponding to the RS frame captured by an event camera; encoding the RS frame and event data through an encoding unit of an image processing model to obtain spatial-temporal implicit representation (STR) data corresponding to the RS frame; embedding the exposure time of a global shutter GS frame corresponding to the RS frame into the STR data through a temporal embedding unit of the image processing model to obtain a temporal tensor corresponding to the RS frame; and decoding the temporal tensor corresponding to the RS frame and the STR data corresponding to the RS frame pixel by pixel through a decoding unit of the image processing model to generate a GS frame.

[0006] Secondly, embodiments of this application provide an image processing apparatus, comprising: a data acquisition module for acquiring a rolling shutter RS ​​frame and event data corresponding to the RS frame captured by an event camera; an encoding module for encoding the RS frame and event data through an encoding unit of an image processing model to obtain spatial-temporal implicit representation (STR) data corresponding to the RS frame; a data embedding module for embedding the exposure time of a global shutter GS frame corresponding to the RS frame into the STR data through a temporal embedding unit of an image processing model to obtain a temporal tensor corresponding to the RS frame; and a decoding module for decoding the temporal tensor corresponding to the RS frame and the STR data corresponding to the RS frame pixel by pixel through a decoding unit of an image processing model to generate a GS frame.

[0007] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement the image processing method as described in the first aspect.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the image processing method as described in the first aspect.

[0009] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the image processing method as described in the first aspect.

[0010] As described above, in this embodiment, an image processing model is used to perform spatial-temporal implicit encoding, exposure time embedding, and pixel-by-pixel decoding on blurred RS frames, thereby recovering clear GS frames from the blurred RS frames. The image processing model integrates three tasks—RS frame correction, deblurring, and frame interpolation—into a single unit, reducing the cumulative errors and human intervention caused by processing RS frames separately, thus improving the clarity and quality of the recovered GS frames. Furthermore, in this embodiment, spatial-temporal implicit representation of RS frames provides a comprehensive spatiotemporal background for RS frame recovery; exposure time embedding optimizes the temporal information representation of RS frames, improving the clarity of GS frames; and pixel-by-pixel decoding ensures the accuracy of the decoded GS frames, further enhancing their quality.

[0011] Therefore, it can be seen that the solution proposed in this application can improve the clarity of the GS frame recovered from the RS frame. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic flowchart of an image processing method provided in one embodiment of this application;

[0014] Figure 2 This is a schematic diagram of the structure of an image processing model provided in one embodiment of this application;

[0015] Figure 3 This is a schematic diagram of the framework for implementing exposure time embedding by a time embedding unit according to an embodiment of this application;

[0016] Figure 4 This is a schematic diagram of the framework for a decoding unit to implement pixel-by-pixel decoding according to an embodiment of this application;

[0017] Figure 5 This is a schematic diagram illustrating the training of an image processing model provided in one embodiment of this application;

[0018] Figure 6 This is a schematic diagram of experimental visualization results on the GevRS-long Exposure dataset provided in one embodiment of this application;

[0019] Figure 7 This is a schematic diagram of experimental visualization results on the GevRS-Real World dataset provided in one embodiment of this application;

[0020] Figure 8 This is a schematic diagram of experimental visualization results on the Fastec dataset provided in one embodiment of this application;

[0021] Figure 9 This is a schematic diagram of experimental visualization results on the GevRS dataset provided in one embodiment of this application;

[0022] Figure 10 This is a schematic diagram of the structure of an image processing apparatus provided in another embodiment of this application;

[0023] Figure 11 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0024] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0025] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0026] To facilitate understanding, before explaining the solution provided in this application, the background of the solution provided in this application will be explained first.

[0027] In related technologies, the recovery of sharp GS frames from blurred RS frames typically involves treating RS frame correction, deblurring, and frame interpolation as independent tasks, cascaded through existing image enhancement networks. This approach increases accumulated errors and significant human artifacts; for example, simply cascading frame interpolation and RS correction networks reduces the sharpness of the GS frame. While event camera methods in related technologies offer the advantage of high spatiotemporal resolution, they cannot recover sharp GS frames at arbitrary frame rates from a single blurred RS frame and suffer from performance limitations when dealing with RS frame distortion, motion blur, and temporal discontinuities.

[0028] To address the problems of the prior art, embodiments of this application provide an image processing method, apparatus, device, storage medium, and computer program product. The image processing method provided in this application embodiment will be described first below.

[0029] The image processing method proposed in this application can recover clear GS frames of arbitrary frame rates from blurred RS frames. Utilizing the high spatiotemporal resolution of an event camera, it directly maps position and time coordinates to color values ​​using spatial-temporal implicit neural representation technology, effectively solving the interleaving degradation problem in the image restoration process. This method mainly processes RS frames in three aspects: spatial-temporal implicit encoding, exposure time embedding, and pixel-by-pixel decoding. The method proposed in this application significantly outperforms existing technologies with its lightweight model (only 0.379M parameters) and high efficiency (achieving an inference speed of 2.83ms / frame with 31x frame interpolation), providing a unified and efficient solution for image processing tasks such as rolling shutter correction, deblurring, and frame interpolation.

[0030] The implementation process of the method proposed in the embodiments of this application is described below.

[0031] Figure 1 A schematic flowchart of an image processing method according to an embodiment of this application is shown. Figure 1 As shown, the method includes the following steps:

[0032] Step S101: Obtain the rolling shutter RS ​​frame and the event data corresponding to the RS frame captured by the event camera.

[0033] In step S101, the event camera is a camera that acquires images through a dynamic vision sensor. The event camera can acquire "events" (i.e. event data) during the image generation process, such as changes in pixel brightness.

[0034] Furthermore, in this embodiment, the shutter is a mechanism used by the camera to control the effective exposure time of the film. A rolling shutter RS ​​frame is an image obtained by scanning line by line with the camera's sensor until all pixels are exposed; while a global shutter GS frame is an image generated by exposing the entire scene simultaneously. Using the method proposed in this embodiment, a clear GS frame can be recovered from a blurry RS frame.

[0035] Step S102: The RS frame and event data are encoded by the encoding unit of the image processing model to obtain the spatial-temporal implicit representation STR data corresponding to the RS frame.

[0036] In step S102, the image processing model includes an encoding unit, a temporal embedding unit, and a decoding unit. The encoding unit encodes the RS frames acquired by the event camera and the corresponding event data to obtain a spatial-temporal implicit representation of the RS frames.

[0037] In one embodiment, the encoding unit can perform multi-scale feature extraction on RS frames and the event data corresponding to RS frames to better capture the spatial-temporal information of RS frames and the event data corresponding to RS frames at different scales, thereby improving the image processing model's ability to capture details and global structures in RS frames and providing a foundation for improving the clarity of GS frames.

[0038] Step S103: The exposure time of the global shutter GS frame corresponding to the RS frame is embedded into the STR data through the temporal embedding unit of the image processing model to obtain the temporal tensor corresponding to the RS frame.

[0039] In step S103, the exposure time of the GS frame refers to the duration from the start of exposure to the end of exposure of the GS frame.

[0040] In this embodiment of the application, the exposure time of the GS frame can be determined by the user according to the desired frame rate of the GS frame. The exposure time of the GS frame is embedded into the STR data to obtain a time tensor that matches the dimension of the STR data. This time tensor can reflect the impact of different exposure times on the sharpness of the RS frame.

[0041] Step S104: The decoding unit of the image processing model performs pixel-by-pixel decoding on the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame to generate the GS frame.

[0042] In step S104, the STR data can be decoded pixel by pixel based on the time tensor using a multilayer perceptron. Because a multilayer perceptron is used, the decoding module can adaptively focus on the key areas of the image, thereby improving the quality of the restored image while maintaining efficient decoding.

[0043] Based on the scheme defined in steps S101 to S104 above, it can be understood that in this embodiment, an image processing model is used to perform spatial-temporal implicit encoding, exposure time embedding, and pixel-by-pixel decoding on the blurred RS frame, thereby recovering a clear GS frame from the blurred RS frame. The image processing model integrates three tasks—RS frame correction, deblurring, and frame interpolation—into one, reducing the cumulative error and human error caused by the separate processing of RS frames by the three tasks, thus improving the clarity and quality of the recovered GS frame. Furthermore, in this embodiment, spatial-temporal implicit representation of the RS frame provides a comprehensive spatiotemporal background for RS frame recovery; exposure time embedding optimizes the temporal dimension information expression of the RS frame, improving the clarity of the GS frame; and pixel-by-pixel decoding ensures the accuracy of the decoded GS frame, further improving the quality of the GS frame.

[0044] Therefore, it can be seen that the solution proposed in this application can improve the clarity of the GS frame recovered from the RS frame.

[0045] The implementation process of the method proposed in the embodiments of this application will be further explained below.

[0046] In one embodiment, Figure 2 This diagram illustrates the structure of an image processing model, consisting of... Figure 2 As can be seen, the image processing model mainly includes an encoding unit (STE), a temporal embedding unit (ETE), and a decoding unit (PPD). The STE encodes the blurred RS frames and event data into an implicit representation that captures spatial-temporal information, namely STR data. The ETE embeds the exposure time of the GS frame into the STR data, providing the exposure time for each pixel. The PPD combines the STR data and the exposure time to generate a clear GS frame.

[0047] It should be noted that an RS frame is considered as a line-by-line combination of consecutive GS frames within the exposure time, thus allowing a sharp RS frame to be constructed line-by-line based on a series of sharp GS frames. In this embodiment, the image processing model can be represented by the function F(x,t,θ), which characterizes the mapping relationship between pixel position, timestamp, and color value in the RS frame. Here, x is the pixel position x of each pixel in the RS frame, t is the exposure timestamp of each pixel, and θ is the model parameter. That is, in this embodiment, the blurred RS frame I... rsb When the event data E is input into the image processing model, the image processing model can output a series of clear GS frames with a high frame rate.

[0048] Depend on Figure 2 It can be seen that when the blurry RS frame I is obtained rsb After the event data E, the RS frame I is processed by the encoder in the coding unit. rsb Encoding the event data E yields STR data with a height of H, a width of W, and a feature dimension of C.

[0049] In one embodiment, the encoding unit extracts features from the RS frame and event data to obtain the spatial and temporal features corresponding to the RS frame; then, the encoding unit performs sparse learning on the spatial and temporal features corresponding to the RS frame to obtain the STR data corresponding to the RS frame.

[0050] In the above embodiments, the encoding unit may have a sparse learning-based encoder architecture. In this scenario, by capturing spatial-temporal information during the RS frame exposure process using a sparse learning-based encoder, computationally intensive optical flow estimation is avoided, improving the efficiency of data representation and thus improving the efficiency of RS frame recovery.

[0051] In another embodiment, the encoding unit may be an encoder with a convolutional neural network, in which feature extraction is performed on RS frames and event data through the convolutional layers of the convolutional neural network to obtain the spatial and temporal features corresponding to the RS frames.

[0052] It should be noted that in this embodiment, standard or depth-separable convolutional layers are used instead of sparse learning encoders to extract features from blurred RS frames and event data. This approach can integrate existing efficient CNN (Convolutional Neural Networks) architectures, such as MobileNet or EfficientNet, to improve the efficiency and accuracy of feature extraction.

[0053] In another embodiment, the encoding unit may also be an encoder with a graph neural network, in which the graph neural network is used to extract features from the RS frame and event data to obtain the spatial and temporal features corresponding to the RS frame.

[0054] It should be noted that, considering the unstructured nature of event data, in this embodiment, a graph neural network is used to process the event data, which better captures the spatiotemporal relationships between events and provides a foundation for improving the clarity of GS frames.

[0055] Furthermore, after encoding is completed and STR data is generated, the exposure time is embedded into the STR data through a time embedding unit to obtain the time tensor corresponding to the RS frame.

[0056] In one embodiment, the exposure time corresponding to the GS frame and the timestamp corresponding to the RS frame are first obtained; then, a mapping relationship between the exposure time of the GS frame and the timestamp of the RS frame is constructed through a time embedding unit; then, according to the mapping relationship, the exposure time of the GS frame is embedded into the STR data to obtain the time tensor corresponding to the RS frame.

[0057] In the above embodiments, the timestamp is used to characterize the start time of RS frame exposure. The exposure time corresponding to the GS frame can be determined according to the frame rate of the GS frame expected by the user. That is, the user inputs the expected frame rate of the GS frame, and the method proposed in the embodiments of this application can restore the blurry RS frame to a GS frame with that frame rate. It can be seen that the scheme based on the embodiments of this application can restore a single blurry RS frame to a clear GS frame of any frame rate.

[0058] In one example, the time embedding unit can embed the exposure time using a multilayer perceptron. In this scenario, the embedding position corresponding to the exposure time of the GS frame is determined by the mapping relationship between the exposure time of the GS frame and the timestamp of the RS frame. Then, the exposure time of the GS frame is embedded into the embedding position by the multilayer perceptron to obtain the time tensor corresponding to the RS frame.

[0059] For example, Figure 3 This illustrates the framework for implementing exposure time embedding using a time embedding unit, consisting of... Figure 3 It can be seen that the temporal embedding unit sets the exposure time t of the GS frame. g Embedded into STR data of dimension H×W×1, a mapping is constructed between the exposure time of GS frames and the timestamp of RS frames. Then, dimensions are increased (e.g., ...) using a single-layer MLP (Multilayer Perceptron). Figure 3 This involves using an MLP×1 to effectively embed the exposure time. The above method generates a time tensor that matches the dimensions (H×W×C) of the STR data.

[0060] In one embodiment, the embedding of the exposure time of the GS frame can also be achieved by dynamically adjusting the embedding method. Specifically, firstly, the relationship between the timestamp and the exposure time embedding strategy is obtained, and the target embedding strategy corresponding to the timestamp is determined. Then, the embedding position corresponding to the exposure time is determined through the mapping relationship. Finally, the exposure time is embedded into the embedding position of the STR data through the target embedding strategy to obtain the time tensor corresponding to the RS frame.

[0061] It should be noted that by dynamically adjusting the exposure time embedding method based on the input timestamp, the STR data after embedding the exposure time can more accurately reflect the impact of different exposure times, thereby improving the clarity of the GS frame.

[0062] Furthermore, it should be noted that the aforementioned exposure time embedding methods include, but are not limited to, MLP-based embedding methods, periodic function embedding methods, and direct time encoding methods. For periodic function embedding methods, sine and cosine periodic functions can be used to encode the exposure time, similar to positional encoding, which can effectively capture the periodicity and continuity of time. For direct time encoding methods, the exposure time can be directly used as one of the inputs to the image processing model, rather than being transformed through an embedding layer, thus simplifying the image processing model structure and training process.

[0063] Furthermore, after embedding the exposure time and obtaining the time tensor, pixel-by-pixel decoding is performed by the decoding unit to generate a clear GS frame. The decoding unit can implement pixel-by-pixel decoding using a multilayer perceptron, or at least one of an adversarial network, a spatial transformation network, or a deformable convolutional network.

[0064] In a scenario where pixel-by-pixel decoding is achieved using a multilayer perceptron, the decoding unit includes multiple multilayer perceptrons. In this scenario, the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame are decoded pixel-by-pixel using multiple multilayer perceptrons to obtain the GS frame corresponding to each exposure time of the RS frame.

[0065] As an example, Figure 4 This illustrates the framework for the decoding unit to implement pixel-by-pixel decoding, consisting of... Figure 4 It can be seen that the decoding unit combines STR data and time tensor T to directly decode clear GS frames through a 5-layer MLP architecture.

[0066] It should be noted that using a pixel-by-pixel decoding method can avoid the need for explicit location lookup, improve location lookup efficiency, and thus improve the generation efficiency of GS data.

[0067] In a scenario where pixel-by-pixel decoding is achieved using at least one of generative adversarial networks (GANs), spatial transformation networks (STRAs), and deformable convolutional networks (DCRs), the decoding unit includes at least one of these GANs. In this scenario, by using at least one of these GANs, the temporal tensor corresponding to the RS frame and the STR data corresponding to the RS frame are decoded pixel-by-pixel to obtain the GS frame corresponding to each exposure time of the RS frame.

[0068] It should be noted that using the generator in the generative adversarial network architecture as the decoder, adversarial training can improve the realism and detail of the recovered image, thereby improving the clarity of the GS frame; using spatial transformation networks or deformable convolutional networks for pixel-by-pixel decoding can enhance the image processing model's ability to handle image spatial transformations, and is especially suitable for handling complex spatial deformations caused by camera or object motion.

[0069] In one embodiment, the image processing model is trained as follows: First, RS frame sample data and the first RS frame corresponding to the RS frame sample data are acquired; then, the RS frame sample data is input into the initial image processing model to obtain the GS frame output by the initial image processing model; then, multiple consecutive GS frames are combined line by line to obtain the second RS frame, and the first RS frame and the second RS frame are compared to obtain feature difference data; next, the loss value of the loss function of the initial image processing model is calculated through the feature difference data, and the model parameters of the initial image processing model are adjusted according to the loss value until the loss function converges to obtain the image processing model.

[0070] In the above embodiments, the RS frame sample data includes at least RS frames and corresponding event data. The sharpness of the first RS frame is higher than that of the RS frame sample data. The first RS frame is the average of multiple RS frames generated by encoding exposure time embedding and pixel-by-pixel decoding of the RS frame sample data and event data. For example, Figure 5 The diagram illustrates the training of an image processing model, by... Figure 5 As shown, the exposure time corresponding to the RS frame sample data and the exposure time corresponding to the event data are embedded into the STR data through the time embedding unit. Then, the decoding unit decodes the data to obtain multiple clear RS frames corresponding to the RS frame sample data. The average value of the multiple clear RS frames is then calculated to obtain the first RS frame. The generation process of the GS frame is the same as that in steps S101 to S104, except that the weights of the two PPDs need to be shared. After obtaining multiple consecutive GS frames, the GS frames are combined row by row to obtain the second RS frame corresponding to the multiple consecutive GS frames. By comparing the first RS frame and the second RS frame, the restoration accuracy of the RS frame can be determined (characterized by feature difference data). Then, the parameters of the image processing model are adjusted according to the magnitude of the restoration accuracy to achieve the training of the image processing model.

[0071] It should be noted that the loss function in the training process of the image processing model can be determined by combining the integral loss and reconstruction loss guided by the blurred RS frame. The loss function can guide the accurate recovery of RS and GS frames during the learning process of the image processing model.

[0072] Furthermore, it should be noted that the training methods for image processing models are not limited to those mentioned above. In practical applications, adversarial training, transfer learning, and other methods can also be used to train image processing models. For adversarial training, adversarial training strategies can be introduced to further optimize the performance of the image processing model through competitive learning, especially achieving better results in the naturalness and detail restoration of generated images. For transfer learning, pre-trained models can be used as part of the encoder or decoder, adapting to specific image restoration tasks through transfer learning; this approach is suitable for applications with limited data.

[0073] Loss functions can include perceptual loss functions and edge-preserving loss functions. For perceptual loss functions, by introducing perceptual loss, the visual quality of the generated image can be improved by comparing differences at the feature level rather than pixel-level differences. For edge-preserving loss functions, by adding an edge-preserving loss term, the model's ability to recover edge and texture details is enhanced, improving image clarity and visual effect.

[0074] The method proposed in this application can be applied to fields such as video communication, media production, and digital entertainment. The following three scenarios are used as examples to illustrate the method proposed in this application.

[0075] In the first scenario, the goal is to recover clear video frames of athletes or vehicles captured during high-speed motion. In this scenario, the encoding unit employs a sparse learning-based encoder to handle blurred RS frames and event data; the temporal embedding unit uses a single-layer MLP to embed exposure times; the decoding unit performs pixel-by-pixel decoding using a 5-layer MLP to recover clear GS frames; and the loss function combines an integral loss guided by blurred RS frames and a reconstruction loss. In this scenario, a publicly available dataset of high-speed motion video captured by RS cameras is used as the dataset, and progressive learning is employed, gradually increasing the target frame rate from a low frame rate.

[0076] In the second scenario, clear frames of nighttime urban landscape videos are recovered under low-light conditions. In this scenario, the encoding unit employs depthwise separable convolutions to improve the model's computational efficiency and adapt to the feature extraction requirements in low-light conditions; the temporal embedding unit uses a dynamic temporal embedding network to dynamically adjust the embedding strategy based on different exposure times; the decoding unit employs an integrated attention mechanism to improve the detail quality of the restored image; and the loss function incorporates perceptual loss and edge-preserving loss terms to enhance the naturalness and clarity of the visual effects. In this scenario, a specially collected RS video dataset under low-light conditions is used as the dataset, and a pre-trained model is introduced, with transfer learning used to accelerate model convergence.

[0077] In the third scenario, used in filmmaking, high-quality video frames are recovered from shaky or fast-moving shots. In this scenario, the encoding unit employs a graph neural network-based encoder to better handle the unstructured features of the event data; the temporal embedding unit uses a periodic function to encode exposure time, capturing the periodicity and continuity of time; the decoding unit uses a generative adversarial network-based generator as a decoder to improve the realism of the recovered image; and the loss function combines adversarial loss to optimize the quality of the generated image. In this scenario, high-speed action scene videos acquired during filming are used as the dataset, and adversarial training is employed to optimize model performance and improve the realism and visual effects of the images.

[0078] In this application embodiment, the performance of the proposed method is also verified through verification experiments. The verification experiments mainly set experimental conditions from three aspects: data augmentation strategy, training strategy optimization, and loss function refinement.

[0079] To address the data augmentation strategy, high dynamic range event data was simulated in the validation experiments. By simulating high dynamic range event data, the training set was augmented, enabling the image processing model to better adapt to image restoration tasks under different lighting conditions and improving the model's generalization ability.

[0080] To optimize the training strategy, a progressive learning strategy is adopted, starting training from recovering low-frame-rate GS frames and gradually increasing the target frame rate, which helps stabilize the training process and improve the performance of the final model.

[0081] To refine the loss function, a loss term based on structural similarity is added to the total loss function and appropriately weighted to ensure the structural fidelity of the recovered image and further improve image quality.

[0082] During the experimental verification process, a cross-domain verification approach was adopted, conducting validation tests on datasets from multiple domains (e.g., street view, motion scenes, and indoor scenes) to ensure the robustness and effectiveness of the image processing model under different scenarios. Real-time testing was also performed under practical application conditions (e.g., embedded devices or mobile devices) to evaluate the processing speed and resource consumption of the image processing model, ensuring its suitability for real-time or near-real-time application scenarios.

[0083] In the validation experiments, experiments were conducted using the Fastec dataset, the GevRS dataset, and the GevRS-long Exposure dataset.

[0084] The experimental results on the Fastec dataset are shown in Table 1:

[0085] Table 1

[0086]

[0087]

[0088] In Table 1, UniNR represents the image processing method proposed in this application's embodiments. As shown in Table 1, the method proposed in this application uses fewer parameters than other methods, achieves higher PSNR (Peak Signal-to-Noise Ratio), higher SSIM (Structural Similarity), and lower LPIPS (Learned Perceptual Image Patch Similarity) than other methods. Therefore, the image processing method proposed in this application outperforms other methods in experimental results on the Fastec dataset.

[0089] The experimental results on the GevRS dataset are shown in Table 2:

[0090] Table 2

[0091] method Frames event Parameter (M) PSNR SSIM LPIPS DSUN

[24] 2 × 3.91 23.10 0.70 0.166 JCD

[58] 3 × 7.16 24.90 0.82 0.105 EvUnroll

[59] 1 √ 20.83 30.14 0.91 0.061 NIRE

[55] 1 √ - 29.86 0.91 - UniNR 1 √ 0.38 31.47 0.9327 0.038

[0092] As shown in Table 2, the method proposed in this application uses fewer parameters than other methods, has a higher PSNR, higher SSIM, and lower LPIPS. Therefore, the image processing method proposed in this application outperforms other methods in experimental results on the GevRS dataset.

[0093] The experimental results on the GevRS-long Exposure dataset are shown in Table 3:

[0094] Table 3

[0095]

[0096]

[0097] As shown in Table 3, the method proposed in this application uses fewer parameters than other methods, has a higher PSNR, higher SSIM, and lower LPIPS. Therefore, the image processing method proposed in this application outperforms other methods in experimental results on the GevRS-long Exposure dataset.

[0098] Figure 6 This shows the experimental visualization results on the GevRS-long Exposure dataset. Figure 7This shows the experimental visualization results on the GevRS-Real World dataset. Figure 8 The experimental visualization results on the Fastec dataset are shown. Figure 9 This shows the experimental visualization results on the GevRS dataset, by Figures 6 to 9 The visualization results show that the method proposed in this application embodiment can provide GS frames with higher clarity compared to other methods.

[0099] The above verification experiments and results demonstrate that the method proposed in this application has achieved significant progress and effects in multiple aspects, including technical, economic, and social aspects, as detailed below:

[0100] (1) Unified processing of complex tasks: The method proposed in this application can recover a clear global shutter frame at any frame rate from a single rolling shutter blurred frame, effectively integrating the three tasks of RS correction, deblurring and frame interpolation into one. This method breaks through the problem of accumulated error and human trace caused by processing these tasks separately in related technologies.

[0101] (2) Lightweight Model and High Efficiency: Compared with related technologies, the method proposed in this application uses a model with only 0.379M parameters, which significantly reduces the demand for computing resources. It achieves an inference speed of 2.83ms / frame with 31x frame interpolation, greatly improving processing speed and enabling real-time or near-real-time image processing applications.

[0102] (3) Superior performance: Extensive experimental verification shows that, compared with related technologies, the method proposed in this application has significant performance advantages in rolling shutter correction, deblurring, and frame interpolation. Evaluation results on standard datasets show that the method proposed in this application significantly improves both PSNR and SSIM image quality evaluation metrics.

[0103] (4) Economic Benefits: The method proposed in this application uses a lightweight model to reduce reliance on high-performance computing resources, thereby lowering the computational cost of related image processing tasks. This is particularly valuable for applications requiring large-scale deployment, such as surveillance and autonomous driving. Furthermore, the method proposed in this application has high processing efficiency, enabling it to process large amounts of image data in a short time, significantly improving production efficiency and accelerating product development cycles.

[0104] (5) Social Impact: The innovative application of the method proposed in this application in the field of image processing has promoted the development of rolling shutter correction and image enhancement technologies, and has a positive effect on promoting technological progress and innovation in related scientific and technological fields. In addition, the method proposed in this application can directly improve the visual experience of end users by providing higher quality image restoration results, especially in the fields of video communication, media production and digital entertainment.

[0105] Finally, on standard test datasets, compared with existing technologies, the method proposed in this application improves PSNR by an average of approximately 1.6 dB and SSIM by an average of approximately 0.02 dB, demonstrating its significant improvement in image quality. In the 31x frame interpolation task, the method proposed in this application improves inference speed by approximately 60 times compared to traditional methods while maintaining high image restoration quality.

[0106] In summary, the method proposed in this application represents a technological breakthrough, has a positive impact on the economy and society, and has significant application value and broad development prospects.

[0107] This application also provides an image processing apparatus, such as... Figure 10 As shown, the device 1000 includes: a data acquisition module 1001, an encoding module 1002, a data embedding module 1003, and a decoding module 1004.

[0108] The data acquisition module 1001 is used to acquire the rolling shutter RS ​​frames and the event data corresponding to the RS frames captured by the event camera;

[0109] The encoding module 1002 is used to encode RS frames and event data through the encoding unit of the image processing model to obtain the spatial-temporal implicit representation STR data corresponding to the RS frames;

[0110] The data embedding module 1003 is used to embed the exposure time of the global shutter GS frame corresponding to the RS frame into the STR data through the time embedding unit of the image processing model to obtain the time tensor corresponding to the RS frame.

[0111] The decoding module 1004 is used to decode the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame pixel by pixel through the decoding unit of the image processing model to generate the GS frame.

[0112] In one example, the encoding module includes a feature extraction module and a sparse learning module. The feature extraction module extracts features from the RS frame and event data using the encoding unit to obtain the spatial and temporal features corresponding to the RS frame. The sparse learning module performs sparse learning on the spatial and temporal features corresponding to the RS frame using the encoding unit to obtain the STR data corresponding to the RS frame.

[0113] In one example, the encoding unit includes a convolutional layer of a convolutional neural network, and the feature extraction module is specifically used to extract features from RS frames and event data through the convolutional layer to obtain the spatial and temporal features corresponding to the RS frames.

[0114] In one example, the encoding unit includes a graph neural network, and the feature extraction module is specifically used to extract features from RS frames and event data through the graph neural network to obtain the spatial and temporal features corresponding to the RS frames.

[0115] In one example, the data embedding module includes a time information acquisition module, a mapping module, and a time embedding module. The time information acquisition module acquires the exposure time corresponding to the GS frame and the timestamp corresponding to the RS frame, where the timestamp represents the start time of the RS frame's exposure. The mapping module constructs a mapping relationship between the exposure time of the GS frame and the timestamp of the RS frame using the time embedding unit. The time embedding module embeds the exposure time of the GS frame into the STR data according to the mapping relationship to obtain the time tensor corresponding to the RS frame.

[0116] In one example, the temporal embedding unit includes a multilayer perceptron. Specifically, the temporal embedding module is used to determine the embedding position corresponding to the exposure time of the GS frame through a mapping relationship; the exposure time of the GS frame is embedded into the embedding position through the multilayer perceptron to obtain the temporal tensor corresponding to the RS frame.

[0117] In one example, the time embedding module is specifically used to obtain the relationship between the timestamp and the exposure time embedding strategy, determine the target embedding strategy corresponding to the timestamp; determine the embedding position corresponding to the exposure time through the mapping relationship; and embed the exposure time into the embedding position of the STR data through the target embedding strategy to obtain the time tensor corresponding to the RS frame.

[0118] In one example, the decoding unit includes multiple multilayer perceptrons. The decoding module is specifically used to perform pixel-by-pixel decoding on the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame through the multiple multilayer perceptrons to obtain the GS frame corresponding to each exposure time of the RS frame.

[0119] In one example, the decoding unit includes at least one of generative adversarial network, spatial transformation network, and deformable convolutional network. Specifically, the decoding module is used to perform pixel-by-pixel decoding on the temporal tensor and STR data corresponding to the RS frame by using at least one of generative adversarial network, spatial transformation network, and deformable convolutional network to obtain the GS frame corresponding to each exposure time of the RS frame.

[0120] In one example, the image processing model is trained as follows: RS frame sample data and a first RS frame corresponding to the RS frame sample data are acquired, wherein the RS frame sample data includes at least the RS frame and the event data corresponding to the RS frame, and the first RS frame has higher resolution than the RS frame sample data; the RS frame sample data is input into the initial image processing model to obtain the GS frame output by the initial image processing model; multiple consecutive GS frames are combined line by line to obtain the second RS frame; feature comparison is performed on the first RS frame and the second RS frame to obtain feature difference data; the loss value of the loss function of the initial image processing model is calculated using the feature difference data; the model parameters of the initial image processing model are adjusted according to the loss value until the loss function converges, thus obtaining the image processing model.

[0121] The image processing apparatus provided in this application embodiment can implement the various processes implemented in the foregoing method embodiment, and will not be described again here to avoid repetition.

[0122] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0123] Figure 11 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.

[0124] The electronic device may include a processor 1101 and a memory 1102 storing computer program instructions.

[0125] Specifically, the processor 1101 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0126] Memory 1102 may include mass storage for data or instructions. For example, and not limitingly, memory 1102 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 1102 may include removable or non-removable (or fixed) media. Where appropriate, memory 1102 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 1102 is non-volatile solid-state memory.

[0127] Memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the methods according to one aspect of this disclosure.

[0128] The processor 1101 implements any of the image processing methods described in the above embodiments by reading and executing computer program instructions stored in the memory 1102.

[0129] In one example, the electronic device may also include a communication interface 1103 and a bus 1110. For example, Figure 11 As shown, the processor 1101, memory 1102, and communication interface 1103 are connected through bus 1110 and complete communication with each other.

[0130] The communication interface 1103 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0131] Bus 1110 includes hardware, software, or both, that couples components of an electronic device together. For example, and not limitingly, the bus may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 1110 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.

[0132] Furthermore, in conjunction with the image processing methods in the above embodiments, this application embodiment can provide a computer-readable storage medium for implementation. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the image processing methods in the above embodiments.

[0133] Furthermore, in conjunction with the image processing methods described in the above embodiments, this application can provide a computer program product for implementation. When the instructions in this computer program product are executed by the processor of an electronic device, the electronic device performs any of the image processing methods described in the above embodiments.

[0134] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0135] The functional modules shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. The programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0136] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0137] The foregoing flowcharts and / or block diagrams of image processing methods, apparatuses, devices, storage media, and computer program products according to embodiments of the present disclosure have described various aspects of the present disclosure. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor may be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0138] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. An image processing method, characterized in that, include: Acquire the rolling shutter RS ​​frames captured by the event camera and the event data corresponding to the RS frames; The RS frame and the event data are encoded by the encoding unit of the image processing model to obtain the spatial-temporal implicit representation STR data corresponding to the RS frame; The exposure time of the global shutter GS frame corresponding to the RS frame is embedded into the STR data through the temporal embedding unit of the image processing model to obtain the temporal tensor corresponding to the RS frame. The decoding unit of the image processing model performs pixel-by-pixel decoding on the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame to generate the GS frame; The time embedding unit of the image processing model embeds the exposure time of the global shutter (GS) frame corresponding to the RS frame into the STR data to obtain the time tensor corresponding to the RS frame. This includes: obtaining the exposure time corresponding to the GS frame and the timestamp corresponding to the RS frame, wherein the timestamp is used to characterize the start time of the RS frame exposure, and the exposure time corresponding to the GS frame is determined by the frame rate of the GS frame expected by the user; constructing a mapping relationship between the exposure time of the GS frame and the timestamp of the RS frame through the time embedding unit; and embedding the exposure time of the GS frame into the STR data according to the mapping relationship to obtain the time tensor corresponding to the RS frame.

2. The method according to claim 1, characterized in that, The RS frame and the event data are encoded by the encoding unit of the image processing model to obtain the spatial-temporal implicit representation (STR) data corresponding to the RS frame, including: The encoding unit extracts features from the RS frame and the event data to obtain the spatial and temporal features corresponding to the RS frame. The STR data corresponding to the RS frame is obtained by sparse learning of the spatial and temporal features corresponding to the RS frame through the encoding unit.

3. The method according to claim 2, characterized in that, The encoding unit includes a convolutional layer of a convolutional neural network. The encoding unit extracts features from the RS frame and the event data to obtain the spatial and temporal features corresponding to the RS frame, including: The convolutional layer extracts features from the RS frame and the event data to obtain the spatial and temporal features corresponding to the RS frame.

4. The method according to claim 2, characterized in that, The encoding unit includes a graph neural network. The encoding unit extracts features from the RS frame and the event data to obtain the spatial and temporal features corresponding to the RS frame, including: The RS frame and the event data are used to extract features through the graph neural network to obtain the spatial and temporal features corresponding to the RS frame.

5. The method according to claim 1, characterized in that, The temporal embedding unit includes a multilayer perceptron, wherein, according to the mapping relationship, the exposure time of the GS frame is embedded into the STR data to obtain the temporal tensor corresponding to the RS frame, including: The embedding position corresponding to the exposure time of the GS frame is determined by the mapping relationship; The exposure time of the GS frame is embedded into the embedding position using the multilayer perceptron to obtain the time tensor corresponding to the RS frame.

6. The method according to claim 1, characterized in that, Based on the mapping relationship, the exposure time of the GS frame is embedded into the STR data to obtain the time tensor corresponding to the RS frame, including: Obtain the relationship between the timestamp and the exposure time embedding strategy, and determine the target embedding strategy corresponding to the timestamp; The embedding position corresponding to the exposure time is determined by the mapping relationship; The exposure time is embedded into the embedding position of the STR data using the target embedding strategy to obtain the time tensor corresponding to the RS frame.

7. The method according to claim 1, characterized in that, The decoding unit includes multiple multilayer perceptrons, wherein the decoding unit of the image processing model performs pixel-by-pixel decoding on the temporal tensor corresponding to the RS frame and the STR data corresponding to the RS frame to generate the GS frame, including: The multiple multilayer perceptrons decode the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame pixel by pixel to obtain the GS frame corresponding to each exposure time of the RS frame.

8. The method according to claim 1, characterized in that, The decoding unit includes at least one of a generative adversarial network, a spatial transformation network, and a deformable convolutional network. The decoding unit of the image processing model performs pixel-by-pixel decoding on the temporal tensor corresponding to the RS frame and the STR data corresponding to the RS frame to generate the GS frame, including: By using at least one of the generative adversarial network, the spatial transformation network, and the deformable convolutional network, the temporal tensor corresponding to the RS frame and the STR data corresponding to the RS frame are decoded pixel by pixel to obtain the GS frame corresponding to each exposure time of the RS frame.

9. The method according to any one of claims 1 to 8, characterized in that, The image processing model is trained in the following manner: Obtain RS frame sample data and a first RS frame corresponding to the RS frame sample data, wherein the RS frame sample data includes at least the RS frame and the event data corresponding to the RS frame, and the clarity of the first RS frame is higher than that of the RS frame sample data. The RS frame sample data is input into the initial image processing model to obtain the GS frame output by the initial image processing model; The second RS frame is obtained by combining multiple consecutive GS frames line by line. Feature comparison is performed on the first RS frame and the second RS frame to obtain feature difference data; The loss value of the loss function of the initial image processing model is calculated using the feature difference data; The model parameters of the initial image processing model are adjusted based on the loss value until the loss function converges, thus obtaining the image processing model.

10. An image processing apparatus, characterized in that, include: The data acquisition module is used to acquire the rolling shutter RS ​​frames captured by the event camera and the event data corresponding to the RS frames; The encoding module is used to encode the RS frame and the event data through the encoding unit of the image processing model to obtain the spatial-temporal implicit representation STR data corresponding to the RS frame; The data embedding module is used to embed the exposure time of the global shutter GS frame corresponding to the RS frame into the STR data through the time embedding unit of the image processing model to obtain the time tensor corresponding to the RS frame. The decoding module is used to decode the time tensor corresponding to the RS frame and the STR data corresponding to the RS frame pixel by pixel through the decoding unit of the image processing model to generate the GS frame; The data embedding module includes: a time information acquisition module, used to acquire the exposure time corresponding to the GS frame and the timestamp corresponding to the RS frame, wherein the timestamp is used to characterize the start time of the exposure of the RS frame, and the exposure time corresponding to the GS frame is determined by the frame rate of the GS frame expected by the user; a mapping module, used to construct a mapping relationship between the exposure time of the GS frame and the timestamp of the RS frame through the time embedding unit; and a time embedding module, used to embed the exposure time of the GS frame into the STR data according to the mapping relationship to obtain the time tensor corresponding to the RS frame.

11. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the image processing method as described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the image processing method as described in any one of claims 1-9.

13. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the image processing method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Event-image deblurring method based on LSTM

    CN117058043A