Complex exposure scene online video super-resolution method based on multi-memory stream convergence
By adopting the multi-memory stream aggregation method in the super-resolution of online videos, adaptively detecting and correcting abnormal exposures, and decoupling and aligning long-term memory, the problem of memory stream interruption in complex exposure scenarios is solved, and the video super-resolution performance is significantly improved.
Patent Information
- Application Number
- CN202510136164.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-06
AI Technical Summary
Existing online video super-resolution methods are difficult to accurately gather long-term implicit memory in complex exposure scenarios, resulting in interruption and disappearance of memory streams, and unable to effectively improve video super-resolution performance.
A method based on multi-memory stream aggregation is proposed. Through adaptive detection and correction of abnormal exposure, dynamic and static decoupling accurately align past long-term memories, build multi-memory streams, and extract complementary information between memory streams through adaptive memory fusion module.
It effectively avoids memory stream interruption, accurately gathers long-term memory, significantly improves the super-resolution performance of online video in complex exposure scenarios, and realizes real-time, high-quality super-resolution reconstruction.
Smart Images

Figure CN119941512A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence and super-resolution technology, and in particular to an online video super-resolution method for complex exposure scenes based on multi-memory stream convergence. Background Art
[0002] The purpose of the online video super-resolution task is to reconstruct low-resolution videos into high-resolution videos when future frames cannot be obtained. It can be applied to fields such as autonomous driving, security monitoring, and live video broadcasting, greatly improving the detection accuracy of downstream visual models and the visual sensory experience.
[0003] In recent years, the field of super-resolution has made significant progress. However, unlike single-image super-resolution, the online video super-resolution task uses a continuous video stream as input. The key challenge is to accurately aggregate the implicit memory of past frames to mine complementary information to alleviate the ill-posed problem in video super-resolution. However, video shooting in actual scenes is often affected by hardware and environmental factors and contains abnormal exposure, resulting in drastic changes in the brightness of video frames. This damages the continuity of the video stream, making it impossible for existing work to accurately transfer the implicit memory of past frames, resulting in interrupted memory flow. In addition, long-term memory will be affected by factors such as noise during transmission, resulting in memory disappearance. Therefore, exploring how to accurately aggregate long-term implicit memory in complex exposure scenes has significant application value.
[0004] Existing online video super-resolution methods can be divided into two paradigms based on their memory aggregation methods - sliding window-based and loop-based. The sliding window-based method uses the window center reference frame and its adjacent frames as input, which uses optical flow estimation and deformable convolution to aggregate the information within the sliding window. "T.Xue, B.Chen, J.Wu, D.Wei, and W.T.Freeman.Video enhancement with task-oriented flow.[C] / / InternationalJournal of Computer Vision.2019:1106-1125" predicts the optical flow information between frames in the time window, and then warps the adjacent frames to align the reference frame to provide reference information. In addition, "S. Jin, M. Liu, Y. Guo, C. Yao, and MSO Baidat. Multi-frame correlated representation network for video super-resolution. [C] / / International Conference on Computer, Information and Telecommunication Systems. 2023: 1-7" proposed a related region selection strategy for searching similar patches in a sliding window to obtain reference information, which avoids the optical flow estimation error caused by large-scale displacement. However, the temporal receptive field of the sliding window-based method is limited, and it can only aggregate information within the sliding window. It cannot effectively utilize long-term memory and cannot solve the problem of memory interruption caused by abnormal exposure. In order to achieve long-term memory aggregation, a cycle-based video super-resolution method is proposed. The cycle-based video super-resolution method inputs video frames in time sequence, and uses optical flow information to distort the implicit memory of the previous moment to the next moment to provide complementary information. This makes the cycle-based video super-resolution method theoretically possible to achieve long-term memory aggregation. "D. Fuoli, M. Danelljan, R. Timofte, and L. Van Gool. Fast online video super-resolution with deformable attention pyramid. [C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2023: 1735-1744" proposes a deformable attention pyramid module to dynamically focus on salient locations in hidden memory, avoiding exhaustive operations and thus reducing computational costs."J. Xiao, X. Jiang, N. Zheng, H. Yang, Y. Yang, Y. Yang, D. Li, and K.-M. Lam. Online video super-resolution with convolutional kernel bypass grafts [J]. IEEE Transactions on Multimedia, 2023, 25: 8972-8987" uses reparameterization technology to greatly compress the parameters of the model and achieve real-time online video super-resolution. However, Shannon pointed out that information will be negatively affected by noise during transmission, resulting in loss or distortion. Therefore, the transmission of memory in online video super-resolution will be negatively affected as follows: 1) Abnormal exposure leads to a large optical flow estimation error, which interrupts the past memory flow; 2) The complementary information required by frames at different times is different, resulting in the gradual forgetting of long-term memory. Since the existing cycle-based online video super-resolution method only has a single memory flow transmission path, the above negative effects will lead to the problem of memory dissipation. This makes it impossible to effectively aggregate long-term memory to mine complementary information, thereby reducing the video super-resolution performance.
[0005] To solve the above problems, the present invention proposes an online video super-resolution paradigm based on multiple memory streams, which aims to adaptively detect and correct abnormal exposure, and at the same time create memory stream branches to achieve long-term memory aggregation, thereby solving the memory interruption and disappearance problems of online video super-resolution in complex exposure scenes and achieving real-time high-quality super-resolution reconstruction. Summary of the invention
[0006] In order to solve the above technical problems, the purpose of the present invention is to provide an online video super-resolution method for complex exposure scenes based on the convergence of multiple memory streams. This method can adaptively detect and correct abnormal exposure in complex scenes to avoid interruption of memory streams, and at the same time, based on the inter-frame pixel displacement, the corrected video frames are subjected to dynamic and static decoupling and precise alignment of past long-term memories to construct multiple memory streams, thereby effectively extracting complementary information between memory streams to solve the problem of memory dissipation, and significantly improving the online video super-resolution performance in complex exposure scenes.
[0007] The technical solution of the present invention is as follows: an online video super-resolution method for complex exposure scenes based on multi-memory stream convergence, the steps are as follows:
[0008] Step 1: Set the midpoint frame of each training sample as the abnormal exposure frame for subsequent training;
[0009] Step 2: The abnormal exposure detection module detects abnormal exposure frames of the training samples; the correction module corrects the detected abnormal exposure frames;
[0010] Step 3: static-dynamic decoupling alignment; In the static-dynamic decoupling stage, the exposure-corrected image is divided into a static area and a dynamic area according to the inter-frame pixel displacement; static area alignment and dynamic area alignment are performed separately to obtain the aligned static area long-term memory and the aligned dynamic area long-term memory, and the two are fused to obtain the complete aligned long-term memory;
[0011] Step 4: Adaptive memory fusion;
[0012] First, the recent memory that is not dynamic and static decoupling alignment Long-term memory after alignment with static-dynamic decoupling With the current frame The concatenation is performed in the channel dimension and input into the adaptive memory fusion module to extract the complementary attention map, where the current frame is used to guide the extraction of complementary information; and The complementary information between them is integrated to obtain the final complete integrated memory;
[0013] Step 5: High-resolution image reconstruction;
[0014] Step 6: Calculate the loss function; overexposure correction loss and super-resolution reconstruction loss are used for joint training. The trained network module is used for online video super-resolution reconstruction in complex exposure scenes.
[0015] Furthermore, in the detection stage, the abnormal exposure detection module uses the overfitting characteristics of optical flow estimation to detect abnormal exposure; defines an optical flow error threshold σ to detect abnormal exposure frames in the training set; and the calculation formula is as follows:
[0016]
[0017] Among them I t is the current frame, is the normal exposure frame at the previous moment, NE means normal exposure, F f is the pre-trained optical flow estimation network, Warp is the warping alignment operation, |.| means taking the absolute value, Mean means calculating the average value, Error is the optical flow alignment error, and σ is set to 0.08; S t is the predicted exposure state of the current frame, where 0 represents normal exposure and 1 represents abnormal exposure.
[0018] Furthermore, in the correction stage, the correction module uses the time series brightness information to perform exposure correction on the detected abnormal exposure frame; firstly, the brightness correction guide vector is calculated, and the current frame and the previous frame are simultaneously averaged and pooled; the pooled different frame images are subtracted to obtain the brightness correction guide vector; the current frame, the previous frame and the brightness correction guide vector are spliced in the channel dimension and input into the correction module to obtain the exposure corrected image; the calculation formula is:
[0019]
[0020] in is the current abnormal frame, AE indicates abnormal exposure, Avg_Pool indicates average pooling operation, LReLU is the activation function, Conv is the convolution function, Cat is the channel concatenation operation, R i Represents the residual convolution module; Indicates the image after exposure correction.
[0021] Furthermore, the static-dynamic decoupling alignment is specifically as follows: first, the previous frame is aligned with the current frame using a pre-trained optical flow prediction network, and then an average pooling operation is performed and then subtracted to calculate an alignment error map; an alignment error threshold δ is defined for static-dynamic decoupling; the calculation formula is:
[0022]
[0023] Among them, It-k is the past frame that needs to be memorized, k represents the time interval between the current frame and the past frame, and F f is a pre-trained optical flow estimation network, is the predicted inter-frame optical flow information, |.| represents the absolute value, Align_error is the dynamic-static decoupling alignment error, and δ is set to 0.08; DM t It is the dynamic and static decoupling mask map;
[0024] In the static region alignment stage, the static and dynamic decoupling mask map is used to extract the static region and the calculated optical flow information is used To align the past long-term memory; the calculation formula is:
[0025]
[0026] Among them, M t-k is the unaligned past long-term memory, DM t is the static-dynamic decoupling mask diagram, It is the long-term memory of the static region after alignment, and the superscript s stands for static;
[0027] In the dynamic region alignment stage, an optical flow-guided grid matching strategy is proposed for accurate dynamic region alignment. First, dynamic grid sampling is performed to decouple the dynamic and static mask map DM. t Split into several non-overlapping grids and count the number of 0 elements in each grid, 0 represents dynamic and 1 represents static; when the number of 0 elements is greater than or equal to 5, the current grid is marked as a dynamic grid; according to the coordinates (i, j) of the dynamic grid and the calculated optical flow information Locate the area to be matched, calculate the L1 distance between the grid of the tth frame and the grid of the tkth frame in the matching area, search for the optimal matching grid for dynamic alignment, and obtain the long-term memory of the aligned dynamic area; according to the dynamic-static decoupling mask map DM t The static area memory and the dynamic area memory are merged to obtain the complete aligned long-term memory; the calculation formula is:
[0028]
[0029] in It is the long-term memory of the aligned dynamic region. The superscript d represents dynamic. For long-term memory after complete alignment.
[0030] Furthermore, the adaptive memory fusion module obtains the final complete fusion memory calculation formula as follows:
[0031]
[0032] in represents the recent memory after optical flow alignment, a represents aligned, Cat represents the channel dimension splicing operation, AMF represents the adaptive memory fusion module, AM t-k represents the long-term complementary attention map, AM t-1 represents the recent complementary attention map, M t For the final complete fusion memory.
[0033] Furthermore, the step 5 is specifically as follows: sending the fused memory features to a reconstruction module to obtain a super-resolution result, wherein the reconstruction module includes 30 residual blocks, several convolutional layers and 2 upsampling layers; the calculation formula is:
[0034] SR t =R(M t ) (15)
[0035] Where R stands for reconstruction module, SR t This is the super-resolution reconstruction result.
[0036] Furthermore, the exposure correction loss and super-resolution reconstruction loss are calculated as follows:
[0037]
[0038] L Rec (SR t , HR t )=||SR t -HR t ||1 (17)
[0039] L overall =λLEc +L Rec (18)
[0040] Where L EC and L Rec are respectively exposure correction loss and super-resolution reconstruction loss, L overall For the overall loss, is the normal exposure frame corrected at the current time t, is the exposure frame of the label at the current moment, ||·|| represents the Euclidean norm; SR t is the super-resolution result at the current moment, HR t is the high-resolution frame labeled at the current moment; λ is the weight of the exposure correction loss function, which is set to 1.
[0041] The beneficial effects of the present invention are as follows: the present invention proposes an online video super-resolution method for complex exposure scenes based on the convergence of multiple memory streams, which can adaptively correct abnormal exposure and create multiple memory streams to accurately converge long-term memory. Specifically, firstly, an abnormal exposure detection-module and a correction module are proposed, which utilize the over-fitting characteristics of optical flow and time-series brightness information to detect and correct abnormal exposure to avoid memory stream interruption. In addition, a dynamic-static decoupling alignment strategy is proposed, which can adaptively select an alignment method based on pixel displacement, thereby accurately converging past long-term memories to create multiple memory streams. Furthermore, an adaptive memory fusion module is designed to mine the complementary information between multiple memory streams to solve the problem of memory dispersion, greatly improving the online video super-resolution performance in complex exposure scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a flow chart of the online video super-resolution method for complex exposure scenes based on multi-memory stream convergence of the present invention;
[0043] Figure 2 It is an overall framework diagram of a specific embodiment of the present invention;
[0044] Figure 3 It is a framework diagram of an embodiment of an exposure detection module and a correction module of the present invention;
[0045] Figure 4 It is a framework diagram of an embodiment of the dynamic-static decoupling alignment strategy of the present invention;
[0046] Figure 5 It is a framework diagram of an embodiment of the adaptive memory fusion module of the present invention;
[0047] Figure 6 The figure is a flow chart of the test steps of a specific embodiment of the present invention. DETAILED DESCRIPTION
[0048] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. The present invention includes but is not limited to the following embodiments.
[0049] like Figure 1 As shown, the present invention provides a specific embodiment of an online video super-resolution method for complex exposure scenes based on multi-memory stream convergence, and its specific implementation process is as follows:
[0050] 1. Video data preprocessing
[0051] The dataset is divided into a training set and a test set. Every 15 frames form a training sample. The midpoint frame is set as the abnormal exposure frame using gamma conversion, where the overexposure gamma value interval is [0.25, 0.5] and the underexposure gamma value interval is [2, 3]. The original video frame is used as the label exposure frame.
[0052] 2. Abnormal exposure detection and correction
[0053] like Figure 2 and Figure 3 As shown in the figure, in the detection stage, the overfitting characteristics of the pre-trained optical flow estimation network Spynet are used to detect abnormal exposure. Specifically, abnormal exposure damages the pixel correspondence between the current frame and the past frame, resulting in the inability of the pre-trained optical flow estimation network to generalize to abnormal exposure scenarios. Therefore, an optical flow error threshold σ is defined to detect abnormal exposure. The calculation formula is as follows:
[0054]
[0055] Among them I t is the current frame (t represents the current time), is the normal exposure frame of the previous moment (NE means normal exposure), F f is a pre-trained optical flow estimation network, Warp is a warp alignment operation, |.| means taking the absolute value, Mean represents the average value operation, Error is the optical flow alignment error, and σ is set to 0.08. t is the predicted exposure state of the current frame, where 0 represents normal exposure and 1 represents abnormal exposure.
[0056] In the correction stage, the detected abnormal exposure frames are corrected; exposure correction is performed using time-series brightness information to avoid interruption of memory flow caused by drastic changes in brightness. First, the brightness correction guidance vector is calculated, and the current frame and the previous frame are averaged and pooled at the same time to avoid the influence of pixel displacement. The pooled images are subtracted to obtain the brightness correction guidance vector. Then the current frame, the previous frame, and the brightness correction guidance vector are spliced in the channel dimension and input into the repair module to obtain the exposure-corrected image. The calculation formula is:
[0057]
[0058]
[0059] in is the current abnormal frame (AE means abnormal exposure), is the normal exposure frame of the previous moment (NE means normal exposure), Avg_Pool means the average pooling operation (the pooling kernel size is set to 11 and the step size is set to 1), LReLU is the activation function, Conv is the convolution function (the convolution kernel size is set to 3 and the step size is set to 1), Cat is the channel splicing operation, R i Represents a residual convolution module. Indicates the image after exposure correction.
[0060] 3. Dynamic and static decoupling alignment.
[0061] like Figure 2 and Figure 4 As shown in the figure, in the static-dynamic decoupling stage, the image is divided into static and dynamic areas according to the pixel displacement between frames. First, the pre-trained optical flow prediction network Spynet is used to align the past frame with the current frame, and then the average pooling operation is performed and then subtracted to calculate the alignment error map. Note that the pre-trained optical flow network cannot estimate large-scale pixel displacements, so areas with small pixel displacements (static) have lower alignment errors, while areas with large pixel displacements (dynamic) have larger alignment errors. Therefore, an alignment error threshold δ is defined for static-dynamic decoupling. The calculation formula is:
[0062]
[0063] Among them I t is the current frame, I t-k is the past frame that needs to be memorized (k represents the time interval with the past frame), F f is a pre-trained optical flow estimation network, is the predicted inter-frame optical flow information, Warp is the warping alignment operation, Mean represents the average value calculation operation, Avg_Pool represents the average pooling operation (the pooling kernel size is set to 5 and the step size is 1), |.| represents taking the absolute value, Align_error is the dynamic and static decoupling alignment error, and δ is set to 0.08. DM t It is the mask diagram for dynamic and static decoupling.
[0064] In the static region alignment stage, the static and dynamic decoupling mask map is used to extract the static region and the calculated optical flow information is used To align the past long-term memory. The calculation formula is:
[0065]
[0066] Warp is the warping alignment operation, M t-k is the unaligned past long-term memory (k represents the time interval with the past memory), DM t is the static-dynamic decoupling mask diagram, It is the long-term memory of the static region after alignment (the superscript s stands for static).
[0067] In the dynamic region alignment stage, an optical flow-guided grid matching strategy is proposed to perform accurate dynamic region alignment. First, dynamic grid sampling is performed to decouple the dynamic and static mask map DM. t Split into several non-overlapping grids (grid size is set to 10), and count the number of 0 elements (0 for dynamic, 1 for static) in each grid. When the number of 0 elements is greater than or equal to 5, the current grid is marked as a dynamic grid. Then, based on the coordinates (i, j) of the dynamic grid and the calculated optical flow information To locate the area to be matched, calculate the L1 distance between the grids of the tth frame and the tkth frame in the matching area to search for the optimal matching grid for dynamic alignment. Finally, based on the dynamic-static decoupling mask DM t The static area memory and the dynamic area memory are merged to obtain the complete aligned long-term memory. The calculation formula is:
[0068]
[0069] in is the long-term memory of the static region after alignment (the superscript s stands for static), is the long-term memory of the aligned dynamic region (the superscript d represents dynamic), For long-term memory after complete alignment.
[0070] 4. Adaptive memory fusion.
[0071] like Figure 2 and Figure 5 As shown, first, the recent memory that is not dynamic and static decoupling alignment is Long-term memory after alignment with static-dynamic decoupling With the current low-resolution frame The concatenation is done in the channel dimension and then fed into the adaptive memory fusion module to extract the complementary attention map, where the low-resolution frame is used to guide the extraction of complementary information. and The complementary information between them is integrated to obtain the final complete integrated memory. The calculation formula is:
[0072]
[0073] in and Represent recent memory and long-term memory respectively (k represents the time interval with past memory, a represents alignment), is the current low-resolution frame (NE means normal exposure), Cat represents the channel dimension stitching operation, AMF represents the adaptive memory fusion module, AM t-k and AM t-1 denote the recent and long-term complementary attention maps, M t For the final complete fusion memory.
[0074] 5. High-resolution image reconstruction
[0075] like Figure 2 As shown in Figure 1, the final complete fusion memory feature is sent to the reconstruction module to obtain the super-resolution result, where the reconstruction module contains 30 residual blocks, 1 convolution layer and 2 upsampling layers. The calculation formula is:
[0076] SR t =R(M t ) (33)
[0077] Where R stands for reconstruction module, M t represents the final complete fusion memory, SR t is the super-resolution reconstruction result (t represents the current time).
[0078] 6. Loss function calculation
[0079] Exposure correction loss and super-resolution reconstruction loss are used for joint training. The calculation formula is:
[0080]
[0081] L Rec (SR t , HR t )=||SR t -HR t ||1 (35)
[0082] L overall =λL EC +L Rec (36)
[0083] Where L EC and L Rec are respectively exposure correction loss and super-resolution reconstruction loss, L overall For the overall loss, is the normal exposure frame corrected at the current time t (NE represents normal exposure), is the exposure frame of the label at the current moment, and ||·|| represents the Euclidean norm. tis the super-resolution result at the current moment, HR t is the high-resolution frame labeled at the current moment. λ is the weight of the exposure correction loss function, which is set to 1.
[0084] 7. Test set video super-resolution
[0085] like Figure 6 As shown, the test video sequence is input into step 2 to detect whether there is abnormal exposure and make corrections, and then the obtained normal exposure data is input into step 3 for static-dynamic decoupling alignment, and then the aligned multi-memory streams are input into step 4 for adaptive memory fusion, and finally the fused memory features are input into step 5 for high-resolution image reconstruction to obtain super-resolution results.
[0086] In summary, the present invention discloses an online video super-resolution method for complex exposure scenes based on the convergence of multiple memory streams. The present invention can adaptively correct abnormal exposure in complex scenes to avoid memory stream interruption, while accurately aligning past long-term memories to construct multiple memory streams, and then effectively extract complementary information between memory streams to solve the problem of memory dissipation, greatly improving the online video super-resolution performance of complex exposure scenes.
[0087] First, the overfitting characteristics of optical flow and time-series brightness information are used to detect and correct abnormal exposure to avoid interruption of memory flow. Then, the video frame is decoupled into dynamic and static areas based on the pixel displacement between frames, and the alignment method is adaptively selected to accurately aggregate past long-term memories to create multiple memory streams. Finally, the complementary information between multiple memory streams is adaptively mined to solve the problem of memory diffusion, effectively improving the online video super-resolution performance in complex exposure scenes.
Claims
1. An online video super-resolution method for complex exposure scenes based on multi-memory stream convergence, characterized in that: Here are the steps: Step 1: Set the midpoint frame of each training sample as the abnormal exposure frame for subsequent training; Step 2: The abnormal exposure detection module detects abnormal exposure frames of the training samples; the correction module corrects the detected abnormal exposure frames; Step 3: Dynamic and static decoupling alignment; In the static-dynamic decoupling stage, the exposure-corrected image is divided into static and dynamic areas according to the inter-frame pixel displacement; Static region alignment and dynamic region alignment are performed respectively to obtain aligned static region long-term memory and aligned dynamic region long-term memory, and the two are fused to obtain a complete aligned long-term memory; Step 4: Adaptive memory fusion; First, the recent memory that is not dynamic and static decoupling alignment Long-term memory after alignment with static-dynamic decoupling With the current frame The concatenation is performed in the channel dimension and input into the adaptive memory fusion module to extract the complementary attention map, where the current frame is used to guide the extraction of complementary information; and The complementary information between them is integrated to obtain the final complete integrated memory; Step 5: High-resolution image reconstruction; Step 6: Calculate the loss function; Overexposure correction loss and super-resolution reconstruction loss are used for joint training. The trained network module is used for online video super-resolution reconstruction in complex exposure scenes.
2. The online video super-resolution method for complex exposure scenes based on multi-memory stream convergence according to claim 1 is characterized in that: The abnormal exposure detection module detects abnormal exposure by using the overfitting characteristics of optical flow estimation in the detection stage; defines an optical flow error threshold σ to detect abnormal exposure frames in the training set; The calculation formula is as follows: Among them I t is the current frame, is the normal exposure frame at the previous moment, NE means normal exposure, F f is the pre-trained optical flow estimation network, Warp is the warping alignment operation, |.| means taking the absolute value, Mean represents the average value operation, Error is the optical flow alignment error, and σ is set to 0.08; S t is the predicted exposure state of the current frame, where 0 represents normal exposure and 1 represents abnormal exposure.
3. The online video super-resolution method for complex exposure scenes based on multi-memory stream convergence according to claim 2 is characterized in that: In the correction stage, the correction module uses the time series brightness information to perform exposure correction on the detected abnormal exposure frame; firstly, the brightness correction guide vector is calculated, and the current frame and the previous frame are simultaneously averaged and pooled; Subtract the pooled frames to obtain the brightness correction guide vector; concatenate the current frame, the previous frame and the brightness correction guide vector in the channel dimension and input them into the correction module to obtain the exposure corrected image; the calculation formula is: in is the current abnormal frame, AE indicates abnormal exposure, Avg_Pool indicates average pooling operation, LReLU is the activation function, Conv is the convolution function, Cat is the channel concatenation operation, R i Represents the residual convolution module; Indicates the image after exposure correction.
4. The online video super-resolution method for complex exposure scenes based on multi-memory stream convergence according to claim 3 is characterized in that: The static-dynamic decoupling alignment is specifically as follows: firstly, the previous frame is aligned with the current frame using a pre-trained optical flow prediction network, and then an average pooling operation is performed and then subtracted to calculate an alignment error map; an alignment error threshold δ is defined for static-dynamic decoupling; The calculation formula is: Among them, I t-k is the past frame that needs to be memorized, k is the time interval between the current frame and the past frame, and F f is a pre-trained optical flow estimation network, is the predicted inter-frame optical flow information, |.| represents the absolute value, Align_error is the dynamic-static decoupling alignment error, and δ is set to 0.08; DM t It is the dynamic and static decoupling mask map; In the static region alignment stage, the static and dynamic decoupling mask map is used to extract the static region and the calculated optical flow information is used To align the past long-term memory; the calculation formula is: Among them, M t-k is the unaligned past long-term memory, DM t is the static-dynamic decoupling mask diagram, It is the long-term memory of the static region after alignment, and the superscript s stands for static; In the dynamic region alignment stage, an optical flow-guided grid matching strategy is proposed for accurate dynamic region alignment. First, dynamic grid sampling is performed to decouple the dynamic and static mask map DM. t Split into several non-overlapping grids and count the number of 0 elements in each grid, 0 represents dynamic and 1 represents static; when the number of 0 elements is greater than or equal to 5, the current grid is marked as a dynamic grid; according to the coordinates (i, j) of the dynamic grid and the calculated optical flow information Locate the area to be matched, calculate the L1 distance between the grid of the tth frame and the grid of the tkth frame in the matching area, search for the optimal matching grid for dynamic alignment, and obtain the long-term memory of the aligned dynamic area; according to the dynamic-static decoupling mask map DM t The static area memory and the dynamic area memory are merged to obtain the complete aligned long-term memory; the calculation formula is: in It is the long-term memory of the aligned dynamic region. The superscript d represents dynamic. For long-term memory after complete alignment.
5. The online video super-resolution method for complex exposure scenes based on multi-memory stream convergence according to claim 4, characterized in that: The adaptive memory fusion module obtains the final complete fusion memory calculation formula as follows: in represents the recent memory after optical flow alignment, a represents aligned, Cat represents the channel dimension splicing operation, AMF represents the adaptive memory fusion module, AM t-k represents the long-term complementary attention map, AM t-1 represents the recent complementary attention map, M t For the final complete fusion memory.
6. The online video super-resolution method for complex exposure scenes based on multi-memory stream convergence according to claim 1, characterized in that: The step 5 is specifically as follows: sending the fused memory features to a reconstruction module to obtain a super-resolution result, wherein the reconstruction module includes 30 residual blocks, several convolutional layers and 2 upsampling layers; the calculation formula is: SR t =R(M t ) (15) Where R stands for reconstruction module, SR t This is the super-resolution reconstruction result.
7. The online video super-resolution method for complex exposure scenes based on multi-memory stream convergence according to claim 1, characterized in that: The calculation formulas for the exposure correction loss and super-resolution reconstruction loss are: L Rec (SR t ,HR t )=||SR t -HR t ||1 (17) THE overall =λL EC +L Rec (18) Where L EC and L Rec are respectively exposure correction loss and super-resolution reconstruction loss, L overall For the overall loss, is the normal exposure frame corrected at the current time t, is the exposure frame of the label at the current moment, ||.|| represents the Euclidean norm; SR t is the super-resolution result at the current moment, HR t Label the high-resolution frame for the current moment; λ is the weight of the exposure correction loss function and is set to 1.
Citation Information
Patent Citations
Video super-resolution reconstruction method based on multi-memory and mixed loss
CN109118431A
Track irregularity dynamic and static detection data inversion method and system and storage medium
CN116307302A
Hardware-efficient neural frame prediction with low resolution optical flow
US20240311962A1
Video generation method, and server
WO2024228676A1