Video processing method and video processing device
Through the hidden motion compensation method, the multi-step feature filtering and queue update strategy are used to solve the problem of inaccurate optical flow estimation in the existing video super-resolution algorithm, and achieve higher quality and more efficient super-resolution image processing.
Patent Information
- Application Number
- CN202111275096.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-10-29
AI Technical Summary
In the existing video super-resolution algorithm, the method based on dominant motion compensation leads to poor super-resolution image quality due to inaccurate optical flow estimation and long time.
Using the method based on implicit motion compensation, the hidden layer features and super-segment features of multiple moments before the current moment are used to filter out relevant information through convolution and activation functions, and combined with the first-in-first-out queue update strategy, the timeliness of historical features are ensured and the accuracy of motion compensation is improved.
The sharpness of super-resolved images is improved, blurred, noise interference is avoided, image quality is ensured, and computational efficiency is improved.
Smart Images

Figure CN113989118B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and more particularly, to a video processing method and a video processing apparatus. Background Art
[0002] Transmission of low-resolution videos is a means of cost savings and bandwidth reduction. Video super-resolution algorithms can cooperate with this transmission strategy to restore these low-resolution videos to high-resolution videos and then distribute them to users. Therefore, it is very necessary to improve the performance of video super-resolution algorithms.
[0003] In related technologies, video super-resolution algorithms may include algorithms based on explicit motion compensation. Algorithms based on explicit motion compensation often use optical flow as the motion representation between video frames. However, the estimation of optical flow is often inaccurate and time-consuming. Summary of the Invention
[0004] The present disclosure provides a video processing method and a video processing apparatus to at least solve the problems existing in the above related technologies.
[0005] According to a first aspect of an embodiment of the present disclosure, a video processing method is provided, including: for each image frame at each moment in a video, based on the image frame at the current moment, the hidden layer features at m moments before the current moment, and the super-resolution features at n moments before the current moment, obtaining the input feature at the current moment; inputting the input feature at the current moment into a video super-resolution model to obtain the super-resolution feature at the current moment output by an output layer of the video super-resolution model and the hidden layer feature at the current moment output by a hidden layer of the video super-resolution model, where the hidden layer feature at the current moment is used to obtain the input features at m moments after the current moment; based on the super-resolution feature at the current moment and the image frame at the current moment, obtaining the super-resolution image at the current moment; where m is an integer greater than or equal to 1, n is an integer greater than or equal to 1, and m and n are not both 1 at the same time.
[0006] Optionally, the obtaining the input feature at the current moment based on the image frame at the current moment, the hidden layer features at m moments before the current moment, and the super-resolution features at n moments before the current moment includes: for the hidden layer feature at each of the m moments before the current moment, filtering out the information related to the image frame at the current moment from the hidden layer feature at that moment as the filtering result of the hidden layer feature at that moment; and / or, for the super-resolution feature at each of the n moments before the current moment, filtering out the information related to the image frame at the current moment from the super-resolution feature at that moment as the filtering result of the super-resolution feature at that moment; obtaining the input feature at the current moment based on the image frame at the current moment, the hidden layer features at m moments before the current moment or their filtering results, and the super-resolution features at n moments before the current moment or their filtering results.
[0007] Optionally, for each of the hidden layer features at m moments before the current moment, filtering out the information related to the image frame at the current moment from the hidden layer features at that moment as the filtering result of the hidden layer features at that moment includes: performing convolution processing on the image frame at the current moment and the hidden layer features at m moments before the current moment to obtain a convolution-processed hidden layer features, where a is equal to m, and the a convolution-processed hidden layer features correspond one by one to the m moments before the current moment; using an activation function to activate the part of the a convolution-processed hidden layer features whose correlation with the image frame at the current moment satisfies a first preset condition to obtain a activated hidden layer features; performing dot multiplication on the a activated hidden layer features and the hidden layer features at m moments before the current moment to obtain the filtering result of the hidden layer features at m moments before the current moment.
[0008] Optionally, for each of the super-resolution features at n moments before the current moment, filtering out the information related to the image frame at the current moment from the super-resolution features at that moment as the filtering result of the super-resolution features at that moment includes: performing convolution processing on the image frame at the current moment and the super-resolution features at n moments before the current moment to obtain b convolution-processed super-resolution features, where b is equal to n, and the b convolution-processed super-resolution features correspond one by one to the n moments before the current moment; using an activation function to activate the part of the b convolution-processed super-resolution features whose correlation with the image frame at the current moment satisfies a second preset condition to obtain b activated super-resolution features; performing dot multiplication on the b activated super-resolution features and the super-resolution features at n moments before the current moment to obtain the filtering result of the super-resolution features at n moments before the current moment.
[0009] Optionally, the hidden layer features at m moments before the current moment are obtained from a historical hidden layer feature queue for storing hidden layer features; wherein, the video processing method further includes: updating the historical hidden layer feature queue according to the hidden layer features at the current moment.
[0010] Optionally, the updating the historical hidden layer feature queue according to the hidden layer features at the current moment includes: deleting the historical hidden layer feature that was first stored in the historical hidden layer feature queue, and storing the hidden layer features at the current moment in the historical hidden layer feature queue; wherein, the historical hidden layer feature queue is a first-in first-out queue for storing m hidden layer features.
[0011] Optionally, the super-resolution features at n moments before the current moment are obtained from a historical super-resolution feature queue for storing super-resolution features; wherein, the video processing method further includes: updating the historical super-resolution feature queue according to the super-resolution features at the current moment.
[0012] Optionally, updating the historical super-resolution feature queue according to the super-resolution feature at the current moment includes: deleting the historical super-resolution feature that was first written into the historical super-resolution feature queue, and writing the super-resolution feature at the current moment into the historical super-resolution feature queue; wherein, the historical super-resolution feature queue is a first-in-first-out queue for storing n super-resolution features.
[0013] Optionally, the video super-resolution model includes: a first convolutional neural network, at least one residual module, and a second convolutional neural network; wherein, inputting the input feature at the current moment into the video super-resolution model to obtain the super-resolution feature at the current moment output by the output layer of the video super-resolution model and the hidden layer feature at the current moment output by the hidden layer of the video super-resolution model includes: inputting the input feature at the current moment into the first convolutional neural network to obtain a convolutional result; inputting the convolutional result into the at least one residual module to obtain the hidden layer feature at the current moment; and inputting the hidden layer feature at the current moment into the second convolutional neural network to obtain the super-resolution feature at the current moment.
[0014] According to a second aspect of the embodiments of the present disclosure, there is provided a video processing apparatus, including: a first acquisition module configured to obtain, for each image frame at each moment in a video, an input feature at the current moment based on the image frame at the current moment, the hidden layer features at m moments before the current moment, and the super-resolution features at n moments before the current moment; an input module configured to input the input feature at the current moment into a video super-resolution model to obtain the super-resolution feature at the current moment output by the output layer of the video super-resolution model and the hidden layer feature at the current moment output by the hidden layer of the video super-resolution model, wherein the hidden layer feature at the current moment is used to obtain the input features at m moments after the current moment; and a second acquisition module configured to obtain a super-resolution image at the current moment based on the super-resolution feature at the current moment and the image frame at the current moment; wherein m is an integer greater than or equal to 1, n is an integer greater than or equal to 1, and m and n are not both 1 at the same time.
[0015] Optionally, the first acquisition module is configured to: for the hidden layer feature at each of the m moments before the current moment, filter out the information related to the image frame at the current moment in the hidden layer feature at that moment as the filtering result of the hidden layer feature at that moment; and / or, for the super-resolution feature at each of the n moments before the current moment, filter out the information related to the image frame at the current moment in the super-resolution feature at that moment as the filtering result of the super-resolution feature at that moment; and obtain the input feature at the current moment based on the image frame at the current moment, the hidden layer features at m moments before the current moment or their filtering results, and the super-resolution features at n moments before the current moment or their filtering results.
[0016] Optionally, the first acquisition module is configured to: perform convolution processing on the image frame at the current moment and the hidden layer features at m moments before the current moment to obtain a convolution-processed hidden layer features, where a is equal to m, and the a convolution-processed hidden layer features correspond one-to-one to the m moments before the current moment; use an activation function to activate the part of the a convolution-processed hidden layer features whose correlation with the image frame at the current moment satisfies a first preset condition to obtain a activated hidden layer features; perform dot multiplication on the a activated hidden layer features and the hidden layer features at m moments before the current moment to obtain a filtering result of the hidden layer features at m moments before the current moment.
[0017] Optionally, the first acquisition module is configured to: perform convolution processing on the super-resolution features at n moments before the current moment and the image frame at the current moment to obtain b convolution-processed super-resolution features, where b is equal to n, and the b convolution-processed super-resolution features correspond one-to-one to the n moments before the current moment; use an activation function to activate the part of the b convolution-processed super-resolution features whose correlation with the image frame at the current moment satisfies a second preset condition to obtain b activated super-resolution features; perform dot multiplication on the b activated super-resolution features and the super-resolution features at n moments before the current moment to obtain a filtering result of the super-resolution features at n moments before the current moment.
[0018] Optionally, the hidden layer features at m moments before the current moment are obtained from a historical hidden layer feature queue for storing hidden layer features; wherein, the video processing device further includes: a first update module configured to update the historical hidden layer feature queue according to the hidden layer features at the current moment.
[0019] Optionally, the first update module is configured to: delete the historical hidden layer feature that was first stored in the historical hidden layer feature queue, and store the hidden layer features at the current moment in the historical hidden layer feature queue; wherein, the historical hidden layer feature queue is a first-in first-out queue for storing m hidden layer features.
[0020] Optionally, the super-resolution features at n moments before the current moment are obtained from a historical super-resolution feature queue for storing super-resolution features; wherein, the video processing device further includes: a second update module configured to update the historical super-resolution feature queue according to the super-resolution features at the current moment.
[0021] Optionally, the second update module is configured to: delete the historical super-resolution feature that was first written in the historical super-resolution feature queue, and write the super-resolution features at the current moment into the historical super-resolution feature queue; wherein, the historical super-resolution feature queue is a first-in first-out queue for storing n super-resolution features.
[0022] Optionally, the video super-resolution model includes: a first convolutional neural network, at least one residual module, and a second convolutional neural network; wherein, the input module is configured to: input the input features at the current moment into the first convolutional neural network to obtain a convolutional result; input the convolutional result into the at least one residual module to obtain the hidden layer features at the current moment; and input the hidden layer features at the current moment into the second convolutional neural network to obtain the super-resolution features at the current moment.
[0023] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to execute the instructions to implement the video processing method according to the present disclosure.
[0024] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the video processing method according to the present disclosure.
[0025] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which implements the video processing method according to the present disclosure when executed by a processor.
[0026] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0027] When obtaining a super-resolution image, the historical motion states at multiple moments before the current moment can be utilized. More temporal information of the video frames before the current moment is utilized, which can reduce the blurriness of the obtained super-resolution image and make it sharper.
[0028] Furthermore, the historical motion states can be filtered to select the historical motion states related to the image frame at the current moment, and the historical motion states irrelevant to the image frame at the current moment can be filtered out, which can avoid introducing noise and maximize the utilization of the information resources in the hidden layer and super-resolution features.
[0029] Further, a storage strategy for multi-step hidden layer features and / or multi-step super-resolution features is adopted. Further, by updating the historical hidden layer feature queue, it can be ensured that the historical hidden layer features in the historical hidden layer feature queue keep pace with the times. Furthermore, it can be ensured that each video frame in the video can use the historical hidden layer features at a similar time for motion compensation, avoiding the introduction of overly distant historical motion states unrelated to the current image frame, avoiding the introduction of noise, and ensuring the quality of the super-resolution image. Further, by updating the historical super-resolution feature queue, it can be ensured that the historical super-resolution features in the historical super-resolution feature queue keep pace with the times. Furthermore, it can be ensured that each video frame in the video can use the historical super-resolution features at a similar time for motion compensation, avoiding the introduction of overly distant historical motion states unrelated to the current image frame, avoiding the introduction of noise, and ensuring the quality of the super-resolution image.
[0030] Further, the historical hidden layer features in the historical hidden layer feature queue can be made to keep pace with the times by setting the historical hidden layer feature queue as a first-in, first-out queue. The implementation process is simple, convenient, and fast. Further, the historical super-resolution features in the historical super-resolution feature queue can be made to keep pace with the times by setting the historical super-resolution feature queue as a first-in, first-out queue. The implementation process is simple, convenient, and fast.
[0031] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings
[0032] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0033] Figure 1 is a flowchart of a video processing method shown according to an exemplary embodiment;
[0034] Figure 2 is a schematic diagram of obtaining the super-resolution features and hidden layer features at the current moment through a video super-resolution model shown according to an exemplary embodiment;
[0035] Figure 3 is a comparison schematic diagram of the results of implicit motion compensation between a video processing method and the RLSP method according to an exemplary embodiment;
[0036] Figure 4 is a block diagram of a video processing device shown according to an exemplary embodiment;
[0037] Figure 5It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation
[0038] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0039] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0040] It should be noted here that "at least one of several items" in the present disclosure all represents the inclusion of three types of parallel situations: "any one of the several items", "a combination of any multiple of the several items", and "the whole of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step one and step two", which means the following three parallel situations: (1) executing step one; (2) executing step two; (3) executing step one and step two.
[0041] The video super-resolution algorithm may include an algorithm based on implicit motion compensation. The algorithm based on implicit motion compensation regards the features stored in the hidden layer as motion states at different times, and can fuse the features stored in the hidden layer with the video frame. At this time, different motion states can be regarded as implicit motion compensation for the video frame. For example, the Recurrent Latent Space Propagation (RLSP) method can be used for implicit motion compensation as follows:
[0042] (1) Decompose a video sequence frame by frame.
[0043] (2) After decomposition, use the first frame and the second frame as the input of the RLSP module, and then obtain the feature h1 at the first moment and the super-resolved image y1.
[0044] (3) Input the hidden layer feature h1 at the first moment and the second and third video frames into the RLSP module. Then, within the RLSP module, fuse the hidden layer feature h1 at the first moment with the two video frames to obtain the hidden layer feature h2 at the second moment and the super-resolved image y2.
[0045] (4) Repeat step (2) until the entire video sequence is super-resolved.
[0046] This disclosure takes into account that the algorithm for implicit motion compensation only stores the motion state at the previous moment and ignores the motion states at more distant moments, using less temporal information of the video frames. This will result in a relatively blurred super-resolved image. Therefore, the video processing method proposed in this disclosure can utilize the historical motion states at multiple moments before the current moment when obtaining the super-resolved image, using more temporal information of the video frames before the current moment, which can reduce the blurriness of the obtained super-resolved image and make it sharper. Further, the historical motion states can be filtered to select the historical motion states relevant to the image frame at the current moment and filter out the historical motion states irrelevant to the image frame at the current moment, which can avoid introducing noise and maximize the utilization of the information resources in the hidden layer and super-resolution features. Further, a storage strategy of multi-step hidden layer features and / or multi-step super-resolution features is adopted. Further, by updating the historical hidden layer feature queue, it can be ensured that the historical hidden layer features in the historical hidden layer feature queue keep up with the times. Furthermore, for each video frame in the video, the historical hidden layer features at moments close to it can be used for motion compensation, which can avoid introducing overly distant historical motion states irrelevant to the image frame at the current moment, avoid introducing noise, and ensure the quality of the super-resolved image. Further, by updating the historical super-resolution feature queue, it can be ensured that the historical super-resolution features in the historical super-resolution feature queue keep up with the times. Furthermore, for each video frame in the video, the historical super-resolution features at moments close to it can be used for motion compensation, which can avoid introducing overly distant historical motion states irrelevant to the image frame at the current moment, avoid introducing noise, and ensure the quality of the super-resolved image. Further, the historical hidden layer feature queue can be set as a first-in-first-out queue to ensure that the historical hidden layer features in the historical hidden layer feature queue keep up with the times, and the implementation process is simple, convenient, and fast. Further, the historical super-resolution feature queue can be set as a first-in-first-out queue to ensure that the historical super-resolution features in the historical super-resolution feature queue keep up with the times, and the implementation process is simple, convenient, and fast.
[0047] Figure 1 It is a flowchart of a video processing method shown according to an exemplary embodiment.
[0048] Refer to Figure 1, in step 101, for each image frame at each moment in the video, the input feature at the current moment can be obtained based on the image frame X at the current moment t (i.e., taking the t-th moment as the current moment), the hidden layer features at m moments before the current moment, and the super-resolution features at n moments before the current moment. It should be understood that each image frame at each moment in the video can be sequentially used as the image frame at the current moment.
[0049] Among them, m is an integer greater than or equal to 1, n is an integer greater than or equal to 1, and m and n are not both 1 at the same time. It should be understood that the values of m and n can be determined according to aspects such as computing resources, requirements for computing efficiency and overhead, etc. For example, in the case of abundant computing resources, larger values can be taken for m and n.
[0050] For example, referring to Figure 2 , m = 3 and n = 3 can be set. At this time, for the image frame X at the current moment t , the hidden layer features h t-1 , h t-2 , h t-3 at the three nearest moments before the current moment, and the super-resolution features o t-1 , o t-2 , o t-3 at the three nearest moments before the current moment are concatenated to obtain the input feature at the current moment. It should be understood that m = 3 and n = 3 are only examples. For example, m = 3 and n = 1 can be set. At this time, for the image frame X at the current moment t , the hidden layer features h t-1 , h t-2 , h t-3 at the three moments before the current moment, and the super-resolution feature o t-1 at the previous moment of the current moment are concatenated to obtain the input feature at the current moment.
[0051] According to an exemplary embodiment of the present disclosure, it should be noted that as the motion state changes, the difference between frames will change. Therefore, if the motion state at the previous moment is less relevant to the current frame, it is difficult to play a role in motion compensation. For example, in diving, if the position of the athlete in the previous frame is far from the position of the athlete in the next frame, it is difficult to achieve the effect of motion compensation. Moreover, introducing historical information that is irrelevant to the current frame will instead interfere with the learning of the neural network and introduce inevitable noise. Therefore, for each hidden layer feature at each of the m moments before the current moment, information related to the image frame at the current moment can be filtered out as the filtering result of the hidden layer feature at that moment; and / or, for each super-resolution feature at each of the n moments before the current moment, information related to the image frame at the current moment can be filtered out as the filtering result of the super-resolution feature at that moment. Next, based on the image frame X at the current moment t and the hidden layer features or their filtering results at the m moments before the current moment, and the super-resolution features or their filtering results at the n moments before the current moment, the input feature at the current moment can be obtained. In this way, the historical motion state can be filtered, the historical motion state related to the image frame at the current moment can be selected, and the historical motion state irrelevant to the image frame at the current moment can be filtered out, which can avoid introducing noise and maximize the utilization of the information resources in the hidden layer and super-resolution features. For example, the image frame X at the current moment t and the filtering results of the hidden layer features at the m moments before the current moment, and the super-resolution features at the n moments before the current moment can be concatenated to obtain the input feature at the current moment. For example, the image frame X at the current moment t and the filtering results of the hidden layer features at the m moments before the current moment, and the filtering results of the super-resolution features at the n moments before the current moment can be concatenated to obtain the input feature at the current moment. For example, the image frame X at the current moment t and the hidden layer features at the m moments before the current moment, and the filtering results of the super-resolution features at the n moments before the current moment can be concatenated to obtain the input feature at the current moment.
[0052] According to an exemplary embodiment of the present disclosure, the image frame X at the current moment t and the hidden layer features at the m moments before the current moment can be subjected to convolution processing to obtain a hidden layer features after convolution processing. Among them, a is equal to m, and the a hidden layer features after convolution processing correspond one by one to the m moments before the current moment. Then, the obtained a hidden layer features after convolution processing can be concatenated together in the feature dimension to obtain a hidden layer feature concatenation result. Next, an activation function (for example, Softmax) can be used to process the hidden layer feature concatenation result in the above-mentioned feature dimension for the part related to the image frame X at the current moment tThe part whose relevance meets the first preset condition is activated to obtain a activated hidden layer features. Then, the a activated hidden layer features can be multiplied pointwise with the hidden layer features at m moments before the current moment to obtain the filtering result of the hidden layer features at m moments before the current moment. Specifically, each activated hidden layer feature is multiplied pointwise with the hidden layer feature at its corresponding moment. For example, the first preset condition can be: the correlation degree exceeds the first preset threshold.
[0053] According to an exemplary embodiment of the present disclosure, the image frame X at the current moment t and the super-resolution features at n moments before the current moment can be subjected to convolution processing to obtain b convolution-processed super-resolution features. Among them, b is equal to n, and the b convolution-processed super-resolution features correspond one by one to the n moments before the current moment. Next, the obtained b convolution-processed super-resolution features can be concatenated together in the feature dimension to obtain a super-resolution feature concatenation result. Next, the activation function (for example, Softmax) can be used to activate the part of the super-resolution feature concatenation result that meets the second preset condition in the above feature dimension with respect to the image frame X at the current moment t to obtain b activated super-resolution features. Then, the b activated super-resolution features can be multiplied pointwise with the super-resolution features at n moments before the current moment to obtain the filtering result of the super-resolution features at n moments before the current moment. Specifically, each activated super-resolution feature is multiplied pointwise with the super-resolution feature at its corresponding moment. For example, the second preset condition can be: the correlation degree exceeds the second preset threshold.
[0054] Return reference Figure 1 , in step 102, the input feature at the current moment can be input into the video super-resolution model to obtain the super-resolution feature o at the current moment output by the output layer of the video super-resolution model t and the hidden layer feature h at the current moment output by the hidden layer of the video super-resolution model t . Among them, the hidden layer feature h at the current moment t is used to obtain the input features at m moments after the current moment.
[0055] According to an exemplary embodiment of the present disclosure, the above video super-resolution model may include: a first convolutional neural network (Conv2D), at least one residual module (Res Block), and a second convolutional neural network. Figure 2 is a schematic diagram showing obtaining the super-resolution feature at the current moment and the hidden layer feature at the current moment through a video super-resolution model according to an exemplary embodiment.
[0056] Reference Figure 2, the input features at the current moment can be input into the first convolutional neural network, which can play a role in fusing the input, and the obtained convolutional result can be a feature of hxwx128. Then, the convolutional result output by the first convolutional neural network, that is, the feature of hxwx128, can be input into the at least one residual module to obtain the hidden layer feature h at the current moment. t . It should be noted that the above convolutional result can be input into the first residual module in the at least one residual module, and the output of the first residual module can be used as the input of the second residual module in the at least one residual module. And so on, the last residual module in the at least one residual module can output the above hidden layer feature h at the current moment. t . Next, the hidden layer feature h at the current moment t can be input into the second convolutional neural network to obtain the super-resolution feature o at the current moment output by the second convolutional neural network. t .
[0057] It should be understood that the hidden layer feature at each moment is: after inputting the input features at that moment into the video super-resolution model, the feature output by a specific hidden layer of the video super-resolution model. For example, in the above embodiment, the specific hidden layer is the at least one residual module.
[0058] Return for reference Figure 1 , in step 103, based on the super-resolution feature o at the current moment t and the image frame X at the current moment t , the super-resolution image at the current moment can be obtained.
[0059] It should be understood that the image frames at each moment in the video can be sequentially used as the image frame at the current moment to obtain its super-resolution image.
[0060] As an example, the image frame X at the current moment t can be first upsampled to obtain an upsampling result. Then, the obtained upsampling result can be superimposed with the super-resolution feature o at the current moment t to obtain the super-resolution image at the current moment.
[0061] According to an exemplary embodiment of the present disclosure, the hidden layer features at m moments before the current moment can be obtained from a historical hidden layer feature queue for storing hidden layer features.
[0062] The video processing method of the present disclosure can also be based on the hidden layer feature h at the current moment tUpdate the historical hidden layer feature queue. In this way, by updating the historical hidden layer feature queue, it can be ensured that the historical hidden layer features in the historical hidden layer feature queue keep pace with the times. Furthermore, it can be ensured that each video frame in the video can use the historical hidden layer features at a similar time for motion compensation, avoiding introducing overly distant historical motion states that are irrelevant to the current image frame, avoiding introducing noise, and ensuring the quality of the super-resolution image.
[0063] According to an exemplary embodiment of the present disclosure, the historical hidden layer feature that was first stored in the historical hidden layer feature queue can be deleted, and the hidden layer feature h at the current moment t is stored in the historical hidden layer feature queue. Among them, the historical hidden layer feature queue can be a first-in-first-out queue for storing m hidden layer features. For example, as described above, m can take the value of 3, and the hidden layer features at 3 moments before the current moment can be h t-3 、h t-2 、h t-1 . After obtaining the hidden layer feature h at the current moment t , the historical hidden layer feature that was first stored in the historical hidden layer feature queue, that is, the oldest historical hidden layer feature h t-3 , can be deleted, and the hidden layer feature h at the current moment t is stored in the historical hidden layer feature queue, obtaining the updated historical hidden layer feature queue: h t-2 、h t-1 、h t . In this way, it can be ensured that the historical hidden layer features in the historical hidden layer feature queue keep pace with the times by setting the historical hidden layer feature queue as a first-in-first-out queue, and the implementation process is simple, convenient, and fast.
[0064] According to an exemplary embodiment of the present disclosure, the super-resolution features at n moments before the current moment can be obtained from the historical super-resolution feature queue for storing super-resolution features.
[0065] The video processing method of the present disclosure can also update the historical super-resolution feature queue according to the super-resolution feature o at the current moment t . In this way, by updating the historical super-resolution feature queue, it can be ensured that the historical super-resolution features in the historical super-resolution feature queue keep pace with the times. Furthermore, it can be ensured that each video frame in the video can use the historical super-resolution features at a similar time for motion compensation, avoiding introducing overly distant historical motion states that are irrelevant to the current image frame, avoiding introducing noise, and ensuring the quality of the super-resolution image.
[0066] According to an exemplary embodiment of the present disclosure, the historical super-resolution feature that was first written into the historical super-resolution feature queue can be deleted, and the super-resolution feature o at the current moment tWrite into the historical super-resolution feature queue. Among them, the historical super-resolution feature queue can be a first-in-first-out queue for storing n super-resolution features. For example, as mentioned above, n can take the value of 3, and the super-resolution features at 3 moments before the current moment can be o t-3 、o t-2 、o t-1 。After obtaining the super-resolution feature o t at the current moment, the historical super-resolution feature that was first stored in the historical super-resolution feature queue, that is, the oldest historical super-resolution feature o t-3 , can be deleted, and the super-resolution feature o t at the current moment is stored in this historical super-resolution feature queue to obtain the updated historical super-resolution feature queue: o t-2 、o t-1 、o t 。In this way, it can be achieved that the historical super-resolution features in the historical super-resolution feature queue keep up with the times by setting the historical super-resolution feature queue as a first-in-first-out queue. The implementation process is simple, convenient and fast.
[0067] It should be noted that after obtaining the super-resolution image at time t, the next moment (t + 1) of time t can be used as the current moment. For the image frame X t+1 at time t + 1, the updated historical hidden layer feature queue: h t-2 、h t-1 、h t and the updated historical super-resolution feature queue: o t-2 、o t-1 、o t can be used to obtain its corresponding super-resolution image at time t + 1. And so on, until each image frame in the video obtains its corresponding super-resolution image.
[0068] It should be noted that since the hidden layer features can be 128 - dimensional, which is relatively large, while the super - resolution features can be 48 - dimensional, and the super - resolution features are much smaller than the 128 - dimensional hidden layer features. If the number m of historical hidden layer features in the hidden layer feature queue is made smaller, for example, making m = 1, and the number n of historical super - resolution features in the historical super - resolution feature queue is made larger, for example, making n>1, the computational efficiency will be relatively high. However, in this case, there may be a situation of performance degradation. For example, the obtained super - resolved image is relatively blurred and not sharp enough. If the number m of historical hidden layer features in the hidden layer feature queue is made larger, for example, making m>1, and the number n of historical super - resolution features in the historical super - resolution feature queue is made smaller, for example, making n = 1, the blurring degree of the obtained super - resolved image can be reduced and made sharper. However, since the hidden layer features are much larger than the super - resolution features, the computational efficiency at this time will be relatively low. Therefore, the lengths of the hidden layer feature queue and the historical super - resolution feature queue can be flexibly adjusted according to the actual situation and needs to achieve a better balance between the sharpness of the super - resolved image and the computational efficiency.
[0069] This disclosure was verified on the Video Super - Resolution Academic Set (Vid4). Vid4 can include four long - video test scenarios: the Foliage sequence, the Walk sequence, the City sequence, and the Calendar sequence. These videos have different resolutions and motion patterns. The Peak Signal to Noise Ratio (PSNR) and the Structural Similarity Index Measurement (SSIM) can be used to evaluate the advantages and disadvantages of the video processing method of this disclosure for implicit motion compensation compared to the RLSP method. The larger the PSNR and SSIM, the better the effect of implicit motion compensation. Figure 3 It is a comparison schematic diagram of the implicit motion compensation results between a video processing method according to an exemplary embodiment and the RLSP method.
[0070] Referring to Figure 3 , it can be seen that the implicit motion compensation effect of the video processing method of this disclosure is improved compared to the existing RLSP method in the above - mentioned four test scenarios. After further adding a screening and filtering strategy, the performance of the implicit motion compensation effect has been significantly improved. For example, in the City test scenario, the PSNR of the video processing method of this disclosure using the screening and filtering strategy is 28.39dB, and the PSNR of using the existing RLSP method is 27.89dB. It can be seen that the video processing method of this disclosure using the screening and filtering strategy has improved the PSNR by 0.5dB compared to the existing RLSP method.
[0071] Figure 4It is a block diagram of a video processing device 400 shown according to an exemplary embodiment.
[0072] Referring to Figure 4 , the video processing device 400 may include a first acquisition module 401, an input module 402, and a second acquisition module 403.
[0073] For each image frame at each moment in the video, the first acquisition module 401 can obtain the input feature at the current moment based on the image frame X at the current moment t (i.e., taking the t-th moment as the current moment), the hidden layer features at m moments before the current moment, and the super-resolution features at n moments before the current moment. It should be understood that each image frame in the video can be sequentially used as the image frame at the current moment.
[0074] Wherein, m is an integer greater than or equal to 1, n is an integer greater than or equal to 1, and m and n are not both 1 at the same time. It should be understood that the values of m and n can be determined according to aspects such as computing resources, requirements for computing efficiency and overhead. For example, when computing resources are abundant, larger values can be taken for m and n.
[0075] According to the exemplary embodiment of the present disclosure, it should be noted that as the motion state changes, the difference between frames will change. Therefore, if the motion state at the previous moment is less related to the current frame, it is difficult to play the role of motion compensation. For example, in diving, if the position of the athlete in the previous frame is far from the position of the athlete in the next frame, it is difficult to achieve the effect of motion compensation. Moreover, introducing historical information unrelated to the current frame will instead interfere with the learning of the neural network and introduce inevitable noise. Therefore, the first acquisition module 401 can also filter out the information related to the image frame at the current moment from the hidden layer features at each of the m moments before the current moment as the filtering result of the hidden layer features at that moment; and / or filter out the information related to the image frame at the current moment from the super-resolution features at each of the n moments before the current moment as the filtering result of the super-resolution features at that moment.
[0076] Next, the first acquisition module 401 can obtain the input feature at the current moment based on the image frame X at the current moment t , the hidden layer features at m moments before the current moment or their filtering results, and the super-resolution features at n moments before the current moment or their filtering results. In this way, the historical motion state can be filtered, the historical motion state related to the image frame at the current moment can be screened out, and the historical motion state unrelated to the image frame at the current moment can be filtered out, which can avoid introducing noise and make the most of the information resources in the hidden layer and super-resolution features. For example, the image frame X at the current moment t, the filtering results of the hidden layer features at m moments before the current moment and the super-resolution features at n moments before the current moment are concatenated to obtain the input features at the current moment. For example, the image frame X at the current moment can be used t , the filtering results of the hidden layer features at m moments before the current moment and the filtering results of the super-resolution features at n moments before the current moment are concatenated to obtain the input features at the current moment. For example, the image frame X at the current moment can be used t , the hidden layer features at m moments before the current moment and the filtering results of the super-resolution features at n moments before the current moment are concatenated to obtain the input features at the current moment.
[0077] According to an exemplary embodiment of the present disclosure, the first acquisition module 401 can perform convolution processing on the image frame X at the current moment t and the hidden layer features at m moments before the current moment to obtain a convolution-processed hidden layer features. Wherein, a is equal to m, and the a convolution-processed hidden layer features correspond to the m moments before the current moment one by one. Then, the first acquisition module 401 can concatenate the obtained a convolution-processed hidden layer features together in the feature dimension to obtain a hidden layer feature concatenation result. Next, the first acquisition module 401 can use an activation function (for example, Softmax) to activate the part of the hidden layer feature concatenation result that satisfies the first preset condition with respect to the image frame X at the current moment in the above feature dimension t to obtain a activated hidden layer features. Then, the first acquisition module 401 can perform dot product on the a activated hidden layer features and the hidden layer features at m moments before the current moment to obtain the filtering results of the hidden layer features at m moments before the current moment. Specifically, each activated hidden layer feature is dot-multiplied with the hidden layer feature at its corresponding moment. For example, the first preset condition can be: the correlation degree exceeds the first preset threshold.
[0078] According to an exemplary embodiment of the present disclosure, the first acquisition module 401 can also perform convolution processing on the image frame X at the current moment t and the super-resolution features at n moments before the current moment to obtain b convolution-processed super-resolution features. Wherein, b is equal to n, and the b convolution-processed super-resolution features correspond to the n moments before the current moment one by one. Next, the first acquisition module 401 can concatenate the obtained b convolution-processed super-resolution features together in the feature dimension to obtain a super-resolution feature concatenation result. Next, the first acquisition module 401 can use an activation function (for example, Softmax) to activate the part of the super-resolution feature concatenation result that satisfies the first preset condition with respect to the image frame X at the current moment in the above feature dimension tThe part whose relevance satisfies the second preset condition is activated to obtain b activated super-resolution features. Then, the first acquisition module 401 can perform dot multiplication on the b activated super-resolution features and the super-resolution features at n moments before the current moment to obtain the filtering result of the super-resolution features at n moments before the current moment. Specifically, each activated super-resolution feature is dot-multiplied with the super-resolution feature at its corresponding moment. For example, the second preset condition can be that the correlation degree exceeds the second preset threshold.
[0079] The input module 402 can input the input feature at the current moment into the video super-resolution model to obtain the super-resolution feature o at the current moment output by the output layer of the video super-resolution model t and the hidden layer feature h at the current moment output by the hidden layer of the video super-resolution model t . Among them, the hidden layer feature h at the current moment t is used to obtain the input features at m moments after the current moment.
[0080] According to an exemplary embodiment of the present disclosure, the above video super-resolution model may include: a first convolutional neural network (Conv2D), at least one residual module (Res Block), and a second convolutional neural network. The input module 402 can input the input feature at the current moment into the first convolutional neural network, and this first convolutional neural network can play a role in fusing the input, and the obtained convolutional result can be a feature of hxwx128. Then, the input module 402 can input the above convolutional result, that is, the feature of hxwx128, into the at least one residual module to obtain the hidden layer feature h at the current moment t . It should be noted that the above convolutional result can be input into the first residual module in the at least one residual module, and the output of this first residual module can be used as the input of the second residual module in the at least one residual module. And so on, the last residual module in the at least one residual module can output the hidden layer feature h at the current moment t .
[0081] Next, the input module 402 can input the hidden layer feature h at the current moment t into the second convolutional neural network to obtain the super-resolution feature o at the current moment output by the second convolutional neural network t .
[0082] The second acquisition module 403 can obtain the super-resolution image at the current moment based on the super-resolution feature o at the current moment t and the image frame X at the current moment t .
[0083] It should be understood that the image frame at each moment in the video can be used as the image frame at the current moment in turn to obtain its super-resolution image.
[0084] As an example, the second acquisition module 403 may first perform upsampling on the image frame X at the current moment t to obtain an upsampling result. Then, the obtained upsampling result may be superimposed with the super-resolution feature o at the current moment t to obtain the super-resolution image at the current moment.
[0085] According to an exemplary embodiment of the present disclosure, the hidden layer features at m moments before the current moment may be obtained from a historical hidden layer feature queue for storing hidden layer features. The video processing apparatus of the present disclosure may further include a first update module, and the first update module may update the historical hidden layer feature queue according to the hidden layer feature h at the current moment t In this way, by updating the historical hidden layer feature queue, it can be ensured that the historical hidden layer features in the historical hidden layer feature queue keep up with the times. Furthermore, it can be ensured that each video frame in the video can use the historical hidden layer features at moments close to it for motion compensation, and it is possible to avoid introducing historical motion states that are too far away and unrelated to the image frame at the current moment, avoid introducing noise, and ensure the quality of the super-resolution image.
[0086] According to an exemplary embodiment of the present disclosure, the first update module may delete the historical hidden layer feature that was first stored in the historical hidden layer feature queue and store the hidden layer feature h at the current moment t into the historical hidden layer feature queue. Among them, the historical hidden layer feature queue may be a first-in, first-out queue for storing m hidden layer features. For example, as described above, m may take the value of 3, and the hidden layer features at 3 moments before the current moment may be h t-3 , h t-2 , h t-1 . After obtaining the hidden layer feature h at the current moment t , the historical hidden layer feature that was first stored in the historical hidden layer feature queue, that is, the oldest historical hidden layer feature h t-3 , may be deleted, and the hidden layer feature h at the current moment t may be stored in the historical hidden layer feature queue to obtain an updated historical hidden layer feature queue:,, h t-2 , h t-1 , h t . In this way, it can be ensured that the historical hidden layer features in the historical hidden layer feature queue keep up with the times by setting the historical hidden layer feature queue as a first-in, first-out queue, and the implementation process is simple, convenient, and fast.
[0087] According to an exemplary embodiment of the present disclosure, the super-resolution features at n moments before the current moment may be obtained from a historical super-resolution feature queue for storing super-resolution features. Among them, the video processing apparatus of the present disclosure may further include a second update module, and the second update module may update the historical super-resolution feature queue according to the super-resolution feature o at the current momentt Update the historical super-resolution feature queue. In this way, by updating the historical super-resolution feature queue, it can be ensured that the historical super-resolution features in the historical super-resolution feature queue keep up with the times. Furthermore, it can be ensured that each video frame in the video can use the historical super-resolution features at a similar time for motion compensation, avoiding the introduction of overly distant historical motion states unrelated to the current image frame, avoiding the introduction of noise, and ensuring the quality of the super-resolution image.
[0088] According to an exemplary embodiment of the present disclosure, the second update module may delete the historical super-resolution feature that was first written into the historical super-resolution feature queue and write the super-resolution feature at the current moment o t into the historical super-resolution feature queue. Among them, the historical super-resolution feature queue may be a first-in, first-out queue for storing n super-resolution features. For example, as described above, n may take the value of 3, and the super-resolution features at the 3 moments before the current moment may be o t-3 、o t-2 、o t-1 . After obtaining the super-resolution feature o t at the current moment, the historical super-resolution feature that was first stored in the historical super-resolution feature queue, that is, the oldest historical super-resolution feature o t-3 , may be deleted, and the super-resolution feature o t at the current moment may be stored in this historical super-resolution feature queue to obtain the updated historical super-resolution feature queue: o t-2 、o t-1 、o t . In this way, it can be ensured that the historical super-resolution features in the historical super-resolution feature queue keep up with the times by setting the historical super-resolution feature queue as a first-in, first-out queue, and the implementation process is simple, convenient, and fast.
[0089] It should be noted that after obtaining the super-resolution image at time t, the next moment (t + 1) of time t may be used as the current moment. For the image frame X t+1 at time t + 1, the updated historical hidden layer feature queue: h t-2 、h t-1 、h t and the updated historical super-resolution feature queue: o t-2 、o t-1 、o t may be used to obtain the corresponding super-resolution image at time t + 1. And so on, until each image frame in the video obtains its corresponding super-resolution image.
[0090] It should be noted that since the hidden layer features can be 128-dimensional, which is relatively large, while the super-resolution features can be 48-dimensional, which is much smaller than the 128-dimensional hidden layer features. If the number m of historical hidden layer features in the hidden layer feature queue is made smaller, for example, making m = 1, and the number n of historical super-resolution features in the historical super-resolution feature queue is made larger, for example, making n > 1, the computational efficiency will be relatively high. However, in this case, there may be a situation of performance degradation. For example, the obtained super-resolution image is relatively blurred and not sharp enough. If the number m of historical hidden layer features in the hidden layer feature queue is made larger, for example, making m > 1, and the number n of historical super-resolution features in the historical super-resolution feature queue is made smaller, for example, making n = 1, the blurring degree of the obtained super-resolution image can be reduced and made sharper. However, since the hidden layer features are much larger than the super-resolution features, the computational efficiency at this time will be relatively low. Therefore, the lengths of the hidden layer feature queue and the historical super-resolution feature queue can be flexibly adjusted according to the actual situation and needs to achieve a better balance between the sharpness of the super-resolution image and the computational efficiency.
[0091] Figure 5 is a block diagram of an electronic device 500 shown according to an exemplary embodiment.
[0092] Referring to Figure 5 , the electronic device 500 includes at least one memory 501 and at least one processor 502. Instructions are stored in the at least one memory 501. When the instructions are executed by the at least one processor 502, a video processing method according to an exemplary embodiment of the present disclosure is executed.
[0093] As an example, the electronic device 500 may be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instructions. Here, the electronic device 500 does not have to be a single electronic device, but may also be any aggregate of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 500 may also be a part of an integrated control system or a system manager, or may be configured as a portable electronic device that can be interconnected locally or remotely (e.g., via wireless transmission).
[0094] In the electronic device 500, the processor 502 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.
[0095] The processor 502 can run instructions or code stored in the memory 501, where the memory 501 can also store data. The instructions and data can also be sent and received via the network interface device over the network, where the network interface device can employ any known transmission protocol.
[0096] The memory 501 can be integrated with the processor 502. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 501 can include separate devices such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 501 and the processor 502 can be operatively coupled or can communicate with each other, for example, via I / O ports, network connections, etc., such that the processor 502 can read files stored in the memory.
[0097] In addition, the electronic device 500 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device 500 can be connected to each other via a bus and / or network.
[0098] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium may also be provided. When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the above video processing method. Examples of the computer-readable storage medium herein include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as, multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium may run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files, and data structures are distributed on a networked computer system such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.
[0099] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided, including a computer program, which when executed by a processor, implements the video processing method according to the present disclosure.
[0100] According to the video processing method and video processing device of the present disclosure, when obtaining a super-resolution image, the historical motion states at multiple moments before the current moment can be utilized. Since more temporal information of the video frames before the current moment is utilized, the blurriness of the obtained super-resolution image can be reduced, making it sharper. Further, the historical motion states can be filtered to select the historical motion states relevant to the image frame at the current moment, while filtering out the historical motion states irrelevant to the image frame at the current moment, which can avoid introducing noise and maximize the utilization of the information resources in the hidden layer and super-resolution features. Further, a storage strategy of multi-step hidden layer features and / or multi-step super-resolution features is adopted. Further, by updating the historical hidden layer feature queue, it can be ensured that the historical hidden layer features in the historical hidden layer feature queue keep pace with the times. Furthermore, it can be ensured that each video frame in the video can use the historical hidden layer features at a moment close to it for motion compensation, which can avoid introducing overly old historical motion states irrelevant to the image frame at the current moment, avoid introducing noise, and ensure the quality of the super-resolution image. Further, by updating the historical super-resolution feature queue, it can be ensured that the historical super-resolution features in the historical super-resolution feature queue keep pace with the times. Furthermore, it can be ensured that each video frame in the video can use the historical super-resolution features at a moment close to it for motion compensation, which can avoid introducing overly old historical motion states irrelevant to the image frame at the current moment, avoid introducing noise, and ensure the quality of the super-resolution image. Further, the historical hidden layer features in the historical hidden layer feature queue can be made to keep pace with the times by setting the historical hidden layer feature queue as a first-in-first-out queue. The implementation process is simple, convenient, and fast. Further, the historical super-resolution features in the historical super-resolution feature queue can be made to keep pace with the times by setting the historical super-resolution feature queue as a first-in-first-out queue. The implementation process is simple, convenient, and fast.
[0101] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include the known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0102] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A video processing method, characterized in that, Including: For each image frame at each moment in the video, based on the image frame at the current moment, the hidden layer features at m moments before the current moment, and the super-resolution features at n moments before the current moment, obtain the input feature at the current moment; Input the input feature at the current moment into the video super-resolution model to obtain the super-resolution feature at the current moment output by the output layer of the video super-resolution model and the hidden layer feature at the current moment output by the hidden layer of the video super-resolution model, where the hidden layer feature at the current moment is used to obtain the input features at m moments after the current moment; Based on the super-resolution feature at the current moment and the image frame at the current moment, obtain the super-resolution image at the current moment; Where m is an integer greater than or equal to 1, n is an integer greater than or equal to 1, and m and n are not both 1 at the same time; Where the hidden layer features at m moments before the current moment are obtained from the historical hidden layer feature queue for storing hidden layer features, and the historical hidden layer feature queue is a first-in-first-out queue for storing m hidden layer features; Where the super-resolution features at n moments before the current moment are obtained from the historical super-resolution feature queue for storing super-resolution features, and the historical super-resolution feature queue is a first-in-first-out queue for storing n super-resolution features.
2. The method according to claim 1, characterized in that, The obtaining the input feature at the current moment based on the image frame at the current moment, the hidden layer features at m moments before the current moment, and the super-resolution features at n moments before the current moment includes: For the hidden layer feature at each of the m moments before the current moment, filter out the information related to the image frame at the current moment in the hidden layer feature at that moment as the filtering result of the hidden layer feature at that moment; and / or, for the super-resolution feature at each of the n moments before the current moment, filter out the information related to the image frame at the current moment in the super-resolution feature at that moment as the filtering result of the super-resolution feature at that moment; Based on the image frame at the current moment, the hidden layer features at m moments before the current moment or their filtering results, and the super-resolution features at n moments before the current moment or their filtering results, obtain the input feature at the current moment.
3. The method according to claim 2, characterized in that, The filtering out the information related to the image frame at the current moment in the hidden layer feature at each of the m moments before the current moment as the filtering result of the hidden layer feature at that moment includes: Perform convolution processing on the image frame at the current moment and the hidden layer features at m moments before the current moment to obtain a convolution-processed hidden layer features, where a is equal to m, and the a convolution-processed hidden layer features correspond one-to-one to the m moments before the current moment; Use an activation function to activate the part of the a convolution-processed hidden layer features whose correlation with the image frame at the current moment satisfies a first preset condition to obtain a activated hidden layer features; Perform a dot product on the a activated hidden layer features and the hidden layer features at m moments before the current moment to obtain the filtering result of the hidden layer features at m moments before the current moment.
4. The method according to claim 2, wherein For each of the n moments before the current moment, filter out the information related to the image frame at the current moment from the super-resolution features at that moment as the filtering result of the super-resolution features at that moment, including: Perform convolution processing on the image frame at the current moment and the super-resolution features at the n moments before the current moment to obtain b super-resolution features after convolution processing, where b is equal to n, and the b super-resolution features after convolution processing correspond one-to-one to the n moments before the current moment; Use an activation function to activate the part of the b super-resolution features after convolution processing that satisfies the second preset condition with respect to the image frame at the current moment to obtain b activated super-resolution features; Perform dot multiplication on the b activated super-resolution features and the super-resolution features at the n moments before the current moment to obtain the filtering result of the super-resolution features at the n moments before the current moment.
5. The method according to claim 1, wherein The video processing method further includes: Updating the historical hidden layer feature queue according to the hidden layer feature at the current moment.
6. The method according to claim 5, wherein The updating the historical hidden layer feature queue according to the hidden layer feature at the current moment includes: Deleting the historical hidden layer feature that was first stored in the historical hidden layer feature queue and storing the hidden layer feature at the current moment in the historical hidden layer feature queue.
7. The method according to claim 1, characterized in that, The video processing method further includes: Updating the historical super-resolution feature queue according to the super-resolution feature at the current moment.
8. The method according to claim 7, wherein The updating the historical super-resolution feature queue according to the super-resolution feature at the current moment includes: Deleting the historical super-resolution feature that was first written in the historical super-resolution feature queue and writing the super-resolution feature at the current moment into the historical super-resolution feature queue.
9. The method according to claim 1, wherein The video super-resolution model includes: a first convolutional neural network, at least one residual module, and a second convolutional neural network; Wherein, the inputting the input feature at the current moment into the video super-resolution model to obtain the super-resolution feature at the current moment output by the output layer of the video super-resolution model and the hidden layer feature at the current moment output by the hidden layer of the video super-resolution model includes: Inputting the input feature at the current moment into the first convolutional neural network to obtain a convolution result; Inputting the convolution result into the at least one residual module to obtain the hidden layer feature at the current moment; Inputting the hidden layer feature at the current moment into the second convolutional neural network to obtain the super-resolution feature at the current moment.
10. A video processing device, characterized in that, Includes: A first acquisition module configured to, for each moment's image frame in the video, obtain the input feature at the current moment based on the image frame at the current moment, the hidden layer features at the m moments before the current moment, and the super-resolution features at the n moments before the current moment; An input module configured to input the input feature at the current moment into the video super-resolution model to obtain the super-resolution feature at the current moment output by the output layer of the video super-resolution model and the hidden layer feature at the current moment output by the hidden layer of the video super-resolution model, wherein the hidden layer feature at the current moment is used to obtain the input features at the m moments after the current moment; A second acquisition module configured to obtain the super-resolution image at the current moment based on the super-resolution feature at the current moment and the image frame at the current moment. Wherein, m is an integer greater than or equal to 1, n is an integer greater than or equal to 1, and m and n are not both 1 at the same time; Wherein, the hidden layer features at m moments before the current moment are obtained from a historical hidden layer feature queue for storing hidden layer features, and the historical hidden layer feature queue is a first-in-first-out queue for storing m hidden layer features; Wherein, the super-resolution features at n moments before the current moment are obtained from a historical super-resolution feature queue for storing super-resolution features, and the historical super-resolution feature queue is a first-in-first-out queue for storing n super-resolution features.
11. The video processing device according to claim 10, wherein, The first obtaining module is configured to: For the hidden layer features at each of the m moments before the current moment, filter out the information related to the image frame at the current moment from the hidden layer features at that moment as the filtering result of the hidden layer features at that moment; and / or, for the super-resolution features at each of the n moments before the current moment, filter out the information related to the image frame at the current moment from the super-resolution features at that moment as the filtering result of the super-resolution features at that moment; Based on the image frame at the current moment, the hidden layer features at m moments before the current moment or their filtering results, and the super-resolution features at n moments before the current moment or their filtering results, obtain the input features at the current moment.
12. The video processing apparatus according to claim 11, wherein The first obtaining module is configured to: Perform convolution processing on the image frame at the current moment and the hidden layer features at m moments before the current moment to obtain a convolution-processed hidden layer features, where a is equal to m, and the a convolution-processed hidden layer features correspond one-to-one to the m moments before the current moment; Use an activation function to activate the part of the a convolution-processed hidden layer features whose correlation with the image frame at the current moment satisfies a first preset condition to obtain a activated hidden layer features; Perform dot multiplication on the a activated hidden layer features and the hidden layer features at m moments before the current moment to obtain the filtering result of the hidden layer features at m moments before the current moment.
13. The video processing apparatus according to claim 11, wherein The first obtaining module is configured to: Perform convolution processing on the image frame at the current moment and the super-resolution features at n moments before the current moment to obtain b convolution-processed super-resolution features, where b is equal to n, and the b convolution-processed super-resolution features correspond one-to-one to the n moments before the current moment; Use an activation function to activate the part of the b convolution-processed super-resolution features whose correlation with the image frame at the current moment satisfies a second preset condition to obtain b activated super-resolution features; Perform dot multiplication on the b activated super-resolution features and the super-resolution features at n moments before the current moment to obtain the filtering result of the super-resolution features at n moments before the current moment.
14. The video processing device according to claim 10, wherein The video processing device further includes: A first updating module, configured to update the historical hidden layer feature queue according to the hidden layer features at the current moment.
15. The video processing device according to claim 14, characterized in that, The first updating module is configured to: Delete the historical hidden layer feature that was first stored in the historical hidden layer feature queue, and store the hidden layer features at the current moment in the historical hidden layer feature queue.
16. The video processing device according to claim 10, wherein The video processing device further includes: A second update module, configured to update the historical super-resolution feature queue according to the super-resolution features at the current moment.
17. The video processing device according to claim 16, characterized in that, The second update module is configured to: Delete the historical super-resolution feature that was first written into the historical super-resolution feature queue, and write the super-resolution feature at the current moment into the historical super-resolution feature queue.
18. The video processing device according to claim 10, wherein The video super-resolution model includes: a first convolutional neural network, at least one residual module, and a second convolutional neural network; Wherein, the input module is configured to: Input the input features at the current moment into the first convolutional neural network to obtain a convolutional result; Input the convolutional result into the at least one residual module to obtain the hidden layer features at the current moment; Input the hidden layer features at the current moment into the second convolutional neural network to obtain the super-resolution features at the current moment.
19. An electronic device, characterized in that, It includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the video processing method according to any one of claims 1 to 9.
20. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the video processing method according to any one of claims 1 to 9.
21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the video processing method according to any one of claims 1-9.
Citation Information
Patent Citations
multispectral remote sensing image Pan-shift method based on a multilayer coupling convolutional neural network
CN109801218A