Video Processing Method, Apparatus, Electronic Device, and Storage Medium

The image fusion model handles interlaced video frames, which solves the problem of picture drawing and details loss in motion scenes, and achieves significant recovery and quality improvement of video images.

CN115633144BActive Publication Date: 2025-07-29DOUYIN VISION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211294643.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-21
Publication Date
2025-07-29
Estimated Expiration
2042-10-21

AI Technical Summary

Technical Problem

In the prior art, when processing interleaved videos, especially in moving scene videos, there are problems such as picture drawing and details loss, and the deinterleaving effect is poor, especially in moving object scenes.

Method used

The image fusion model is used to process interleaved frames. The image fusion model includes a feature processing sub-model and a motion-aware sub-model. Through feature extraction, fusion and motion perception, the target video frame is generated.

Benefits of technology

It significantly improves the recovery effect of video images, especially video images in sports scenes, solves the problems of picture drawing and details loss, and improves picture quality and clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115633144B_ABST
    Figure CN115633144B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a video processing method, apparatus, electronic device, and storage medium. Among them, the method includes: obtaining at least three to-be-processed interleaved frames; wherein, the to-be-processed interleaved frames are determined based on two adjacent to-be-processed video frames; inputting the at least three to-be-processed interleaved frames into an image fusion model pre-trained to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames; wherein, the image fusion model includes a feature processing sub-model and a motion perception sub-model; determining a target video based on the at least two target video frames. The technical solution of the embodiments of the present disclosure can effectively improve the restoration effect of video pictures. Especially for video pictures of motion scenes, a relatively significant restoration effect can also be achieved. At the same time, problems such as picture streaking and detail loss are solved, the picture quality and clarity of video pictures are improved, and the user experience is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of video processing technologies, and in particular, to a video processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the continuous development of network technologies, when scanning and displaying images, in order to improve the image display effect, more and more users adopt the progressive scanning method to scan and display images.

[0003] However, for early videos generated based on the interlaced scanning method before, due to the large time interval between the displays of two images, there will be image quality problems such as large image flickering, moiré patterns, and artifacts. Such a generated video is called an interlaced video. When displaying an interlaced video on an existing display interface, it needs to be deinterlaced before a complete video can be displayed.

[0004] Currently, deinterlacing is usually performed on an interlaced video to remove the combing effect in the interlaced video. However, the deinterlacing effect of this method is not good. Especially in a moving object scene, the combing areas are relatively blurred, and problems such as detail loss and combing in the picture are likely to occur. Summary of the Invention

[0005] The present disclosure provides a video processing method, apparatus, electronic device, and storage medium to achieve an effective restoration effect on video frames, especially for video frames in a motion scene, a relatively significant restoration effect can be achieved.

[0006] In a first aspect, embodiments of the present disclosure provide a video processing method, and the method includes:

[0007] Obtain at least three to-be-processed interlaced frames; wherein, the to-be-processed interlaced frames are determined based on two adjacent to-be-processed video frames;

[0008] Input the at least three to-be-processed interlaced frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interlaced frames; wherein, the image fusion model includes a feature processing sub-model and a motion perception sub-model;

[0009] Determine a target video based on the at least two target video frames.

[0010] In a second aspect, embodiments of the present disclosure further provide a video processing apparatus, and the apparatus includes:

[0011] A to-be-processed interlaced frame acquisition module, configured to obtain at least three to-be-processed interlaced frames; wherein, the to-be-processed interlaced frames are determined based on two adjacent to-be-processed video frames;

[0012] A target video frame determination module, configured to input the at least three to-be-processed interleaved frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames; wherein, the image fusion model includes a feature processing sub-model and a motion perception sub-model.

[0013] A target video determination module, configured to determine a target video based on the at least two target video frames.

[0014] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including:

[0015] One or more processors;

[0016] A storage device, configured to store one or more programs,

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method according to any one of the embodiments of the present disclosure.

[0018] In a fourth aspect, an embodiment of the present disclosure further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the video processing method according to any one of the embodiments of the present disclosure when executed by a computer processor.

[0019] The technical solution of the embodiment of the present disclosure, after obtaining at least three to-be-processed interleaved frames, can input the at least three to-be-processed interleaved frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames. Finally, based on the at least two target video frames, a target video is determined. When an interleaved video is displayed on an existing display device, the restoration effect of the video picture can be effectively improved. Especially for the video picture of a motion scene, a relatively significant restoration effect can also be achieved. At the same time, problems such as picture streaking and detail loss are solved, the picture quality and clarity of the video picture are improved, and the user experience is enhanced. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Combined with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.

[0021] Figure 1 is a schematic flowchart of a video processing method provided by an embodiment of the present disclosure;

[0022] Figure 2It is a schematic diagram of the image fusion model provided by an embodiment of the present disclosure;

[0023] Figure 3 It is a schematic diagram of the motion perception model provided by an embodiment of the present disclosure;

[0024] Figure 4 It is a schematic flowchart of a video processing method provided by an embodiment of the present disclosure;

[0025] Figure 5 It is a schematic diagram of a video frame to be processed provided by an embodiment of the present disclosure;

[0026] Figure 6 It is a schematic diagram of the structure of a video processing device provided by an embodiment of the present disclosure;

[0027] Figure 7 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0028] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0029] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.

[0030] As used herein, the term "including" and its variations are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.

[0031] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0032] It should be noted that the modifications of "one" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".

[0033] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0034] It can be understood that before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.

[0035] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operations of the technical solutions of this disclosure according to the prompt message.

[0036] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0037] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not constitute a limitation on the implementation manner of this disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of this disclosure.

[0038] It can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.

[0039] Before introducing the technical solution, an exemplary description of the application scenario can be given first. When generating a corresponding target video based on the original video, there are usually three common solutions. The first implementation can be to perform de-interlacing on the interlaced frames to be processed based on the de-interlacing algorithm (YADIF) to remove the combing effect of the original video. However, this method has poor restoration effects for moving object scenarios, with problems such as missing details and combing in the picture. The second implementation can be to input multiple interlaced frames to be processed into the ST-Deint deep learning neural network model that combines temporal and spatial information prediction, and process each interlaced frame to be processed based on the deep learning algorithm. Although this method is significantly better than the first implementation, it can only roughly restore the moving scene, and the detail restoration effect is not good, and there will still be the problem of combing in the picture. The third implementation can be to process the interlaced frames to be processed based on the deep learning model DIN, so that each interlaced frame to be processed first fills in the missing information and then fuses the inter-field content to obtain the processed video. Similar to the second implementation, although this method is better than the first implementation, it can only roughly restore the moving scene, and the detail restoration effect is not good, especially in the moving combing area, where blurring and detail loss are likely to occur. Based on the above, the existing video processing methods still have the problem of poor display effects of the output video. At this time, based on the technical solution of the embodiments of the present disclosure, the interlaced frames to be processed can be processed based on an image fusion model including multiple sub-models, thereby avoiding problems such as missing details, combing, and blurring in the output target video.

[0040] Figure 1 FIG. is a schematic flowchart of a video processing method provided by an embodiment of the present disclosure. The embodiments of the present disclosure are applicable to the situation of supplementing feature information for an original video using interlaced scanning, so that the obtained target video can be completely displayed on existing display devices. This method can be executed by a video processing device, which can be implemented in the form of software and / or hardware. Optionally, it is implemented by an electronic device, which can be a mobile terminal, a PC, or a server, etc. The technical solution provided by the embodiments of the present disclosure can be executed based on a client, or based on a server, or based on the cooperation of a client and a server.

[0041] As Figure 1 shown, the method includes:

[0042] S110. Obtain at least three interlaced frames to be processed.

[0043] Among them, the interlaced frames to be processed are determined based on two adjacent video frames to be processed.

[0044] First of all, it should be noted that the device for executing the video processing method provided in the embodiments of the present disclosure can be integrated in an application software that supports video processing functions, and the software can be installed in an electronic device. Optionally, the electronic device can be a mobile terminal or a PC, etc. The application software can be a type of software for image / video processing, and specific application software will not be elaborated here one by one, as long as it can implement image / video processing. It can also be a specially developed application program to implement video processing and display the output video, or it can be integrated in the corresponding page, and the user can implement the processing of special effect videos through the page integrated in the PC.

[0045] In this embodiment, the user can capture videos in real time based on the camera device of the mobile terminal, or actively upload videos based on the pre-developed controls in the application software. Therefore, it can be understood that the videos captured in real time or actively uploaded by the user obtained by the application are the videos to be processed. Further, by parsing the videos to be processed based on the pre-written program, multiple video frames to be processed can be obtained. Those skilled in the art should understand that due to the limitations of bandwidth and the processing speed of video devices, the early video display methods usually adopted the interlaced scanning method, that is, first scan the odd lines to obtain a video frame in which only the odd-line pixel points have rendering pixel values, and then scan the even lines to obtain a video frame in which only the even-line pixel points have rendering pixel values, and combine the two video frames to obtain a complete video frame. This display method will cause a relatively large time interval between the displays of two adjacent video frames, resulting in relatively large flicker, moiré, artifacts and other image quality problems in the video frame. At the same time, since the progressive scanning method is usually adopted when displaying videos currently, when displaying early videos based on existing video display devices, after deinterlacing processing, a complete video can be obtained. The video frames after interlaced scanning can be used as the video frames to be processed, that is, the video frames in which only the odd-line pixel points have rendering pixel values or the even-line pixel points have rendering pixel values, and the video frame combined by two adjacent video frames to be processed is the interlaced frame to be processed. Among them, deinterlacing processing refers to filling the missing half-field information of the odd and even fields of the interlaced frames of two adjacent images respectively to restore to the original frame size, and finally obtaining two frames of odd and even.

[0046] In this embodiment, since the video frames to be processed are the video frames in which only the odd-line pixel points have rendering pixel values or the even-line pixel points have rendering pixel values, when performing deinterlacing processing on the video frames to be processed, two adjacent video frames to be processed can be combined to obtain an interlaced frame to be processed, so that deinterlacing processing can be performed on the interlaced frame to be processed.

[0047] Exemplarily, I1, I2, I3, I4, I5, and I6 can be used as video frames to be processed. When performing fusion processing on two adjacent video frames to be processed, I1 and I2 can be combined to obtain a frame of interleaved frame D1 to be processed, I3 and I4 can be combined to obtain a frame of interleaved frame D2 to be processed, and I5 and I6 can be combined to obtain a frame of video frame D3 to be processed. The production principle of the interleaved frame to be processed can be expressed based on the following formula:

[0048]

[0049] It should be noted that the number of interleaved frames to be processed can be three, or more than three. The embodiments of the present disclosure do not make specific limitations on this.

[0050] That is to say, the number of interleaved frames to be processed corresponds to the number of video frames of the original video. The number of interleaved frames to be processed input into the model can be three frames or more than three frames.

[0051] S120. Input at least three interleaved frames to be processed into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interleaved frames to be processed.

[0052] In this embodiment, after obtaining at least three interleaved frames to be processed, each interleaved frame to be processed can be input into a pre-trained image fusion model. Among them, the image fusion model can be a deep learning neural network model including multiple sub-models. The image fusion model includes a feature processing sub-model and a motion perception sub-model.

[0053] It should also be noted that in the process of processing based on the image fusion model, the optical flow map between two adjacent interleaved frames to be processed needs to be combined to determine the final target video. Therefore, the number of interleaved frames to be processed is at least three frames.

[0054] Among them, the feature processing sub-model can be a neural network model including multiple convolutional modules. The feature processing sub-model can be used to extract, fuse, and perform other processing on the features in the interleaved frame to be processed. In this embodiment, the feature processing sub-model can include multiple 3D convolutional layers, so that the feature processing sub-model can not only process the time-domain feature information of multiple frames, but also process the spatial feature information, strengthening the information interaction between the interleaved frames to be processed.

[0055] Among them, the motion perception sub-model can be a neural network model for perceiving the inter-frame motion situation. The motion perception sub-model can be composed of at least one convolutional network, a network including a Backward Warping function, and a residual network. The Backward Warping function can realize the mapping between images. In practical applications, since there is a strong spatio-temporal correlation between the inter-frame contents of two adjacent frames, when processing the to-be-processed interleaved frames, the motion perception sub-model can be used to process the feature information between frames, so that the inter-frame contents are more continuous, and at the same time, the effect of complementing details with each other can be achieved.

[0056] In this embodiment, after inputting each to-be-processed interleaved frame into the image fusion model, each to-be-processed video frame can be processed based on each sub-model in the image fusion model, so that at least two target video frames corresponding to each to-be-processed interleaved frame can be obtained.

[0057] In practical applications, since the image fusion model includes multiple sub-models, when processing each to-be-processed interleaved frame based on the image fusion model, each to-be-processed video frame can be correspondingly processed through multiple sub-models in the model in sequence, so as to output at least two target video frames corresponding to each to-be-processed interleaved frame.

[0058] It should be noted that the image fusion model includes multiple sub-models, and their arrangement order can be arranged according to the data input and output order for each sub-model.

[0059] Optionally, the image fusion model includes a feature processing sub-model, a motion perception sub-model, and a 2D convolutional layer. Among them, the 2D convolutional layer can be a neural network layer that only performs feature processing on the height and width of the data.

[0060] In this embodiment, the advantage of determining the arrangement order of each sub-model in the image fusion model based on the data input and output order is that the image fusion model can not only process the feature information of the to-be-processed interleaved frames, but also perceive the motion situation between each to-be-processed interleaved frame, so that the inter-frame contents are more continuous, and at the same time, the effect of detail supplementation can be achieved.

[0061] It should be noted that when the to-be-processed interlaced frames are input into the image fusion model for processing, the solution adopted by the prior art is to split the to-be-processed interlaced frames into odd and even rows, that is, to halve the H dimension of the to-be-processed interlaced frames. Exemplarily, if the matrix of the to-be-processed interlaced frames is (H×W×C), then the matrix after splitting into odd and even rows is (2 / H×W×C). Such a solution may cause deformation of the objects in the to-be-processed interlaced frames in terms of structure, thus affecting the visual effect of the target video frames. The processing process of the embodiments of the present disclosure can be understood as that on the basis of splitting the to-be-processed interlaced frames into odd and even rows, the to-be-processed interlaced frames are also split into odd and even columns, that is, a dual-feature processing branch is adopted, so as to ensure that when the to-be-processed interlaced frames are processed based on the image fusion model, the overall structural feature information of the to-be-processed interlaced frames can be processed, and the high-frequency detail feature information of the to-be-processed interlaced frames can also be processed.

[0062] Based on this, on the basis of the above technical solutions, it further includes: the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch; the output of the first feature extraction branch is the input of the first motion perception sub-model in the motion perception sub-model, and the output of the second feature extraction branch is the input of the second motion perception sub-model in the motion perception sub-model; the output of the first motion perception sub-model and the output of the second motion perception sub-model are the input of the 2D convolutional layer, so that the 2D convolutional layer outputs the target video frame.

[0063] In this embodiment, the first feature extraction branch can be a neural network model for processing the structural features of the to-be-processed interlaced frames. Optionally, the first feature extraction branch includes a structural feature extraction network and a structural feature fusion network. Among them, the structural feature extraction network can be composed of at least one convolutional network, so that at least one convolutional network can process the to-be-processed interlaced frames according to a preset structural splitting ratio to obtain the structural features corresponding to the to-be-processed interlaced frames. The structural feature fusion network can be a neural network of a U-Net structure stacked by at least one 3D convolutional layer. It should be noted that the convolutional kernels of each 3D convolutional layer can be the same value or different values, and the embodiments of the present disclosure do not make specific limitations in this regard. Since the number of the to-be-processed interlaced frames is at least three, the structural feature fusion network can be used to strengthen the frame-interval information interaction, so that not only the spatial features of the to-be-processed interlaced frames can be processed, but also the temporal features between multiple frames can be strengthened.

[0064] In this embodiment, the second feature extraction branch may be a neural network model for processing the detailed features of the to-be-processed interleaved frames. Optionally, the second feature extraction branch includes a detailed feature extraction network and a detailed feature fusion network. Among them, the detailed feature extraction network may be composed of at least one convolutional layer, so that at least one convolutional layer can process the to-be-processed interleaved frames according to a preset detailed splitting ratio to obtain detailed features corresponding to the to-be-processed interleaved frames. The detailed feature fusion network may be a neural network of a U-Net structure stacked by at least one 3D convolutional layer. It should be noted that the convolutional kernels of the 3D convolutional layers may have the same value or different values, and the embodiments of the present disclosure do not make specific limitations in this regard. It should also be noted that the network structures of the detailed feature fusion network and the structural feature fusion network may be the same, and the effects achieved by these two networks are also the same, both of which are to enhance the inter-frame information of the to-be-processed interleaved frames. The following may be combined with Figure 2 to specifically describe the data input and output of each sub-model in the image fusion model.

[0065] Exemplarily, referring to Figure 2 , each to-be-processed interleaved frame is respectively input into the first feature extraction branch and the second feature extraction branch. After being processed by the structural feature extraction network and the structural feature fusion network in the first feature extraction branch for each to-be-processed interleaved frame, it can be input into the first motion perception sub-model. At the same time, after being processed by the detailed feature extraction network and the detailed feature fusion network in the second feature extraction branch for each to-be-processed interleaved frame, it can be input into the second motion perception sub-model. Further, after being processed by the first motion perception sub-model for the model input, it can be input into the 2D convolutional layer. At the same time, after being processed by the second motion perception sub-model for the model input, it is input into the 2D convolutional layer, so that the 2D convolutional layer can output the target video frame. The advantage of such a setting is that the image fusion model can not only process the feature information of the to-be-processed interleaved frames, but also perceive the motion situation between the to-be-processed interleaved frames, so that the inter-frame content is more continuous and the effect of detail supplementation is achieved at the same time.

[0066] In practical applications, when each to-be-processed interleaved frame is input into the image fusion model, it can be processed based on each sub-model in the model, so that the target video frame corresponding to the to-be-processed interleaved frame can be obtained. The following continues to combine Figure 2 to specifically describe the process of the image fusion model processing the to-be-processed interleaved frames.

[0067] Referring to Figure 2As shown, at least three to-be-processed interleaved frames are input into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames, including: performing equal-proportion feature extraction on the at least three to-be-processed interleaved frames based on a structural feature extraction network to obtain structural features corresponding to the to-be-processed interleaved frames; and performing odd-even field feature extraction on the at least three to-be-processed interleaved frames based on a detail feature extraction network to obtain detail features corresponding to the to-be-processed interleaved frames; processing the structural features based on a structural feature fusion network to obtain a first inter-frame feature map between two adjacent to-be-processed interleaved frames; and processing the detail features based on a detail feature fusion network to obtain a second inter-frame feature map between two adjacent to-be-processed interleaved frames; processing the first inter-frame feature map based on a first motion perception sub-model to obtain a first fusion feature map; and processing the second inter-frame feature map based on a second motion perception sub-model to obtain a second fusion feature map; processing the first fusion feature map and the second fusion feature map based on a 2D convolutional layer to obtain at least two target video frames.

[0068] In this embodiment, the structural feature may be a feature used to reflect the overall structural information of the to-be-processed interleaved frame. The detail feature may be a feature used to reflect the detail information of the to-be-processed interleaved frame. The detail feature may be a high-frequency feature, which is a higher-order feature than the structural feature.

[0069] In a specific implementation, each to-be-processed interleaved frame is input into an image fusion model. The structure feature extraction network can perform equal-proportion dimensionality reduction processing on each to-be-processed interleaved frame to obtain the corresponding structure features of each to-be-processed interleaved frame. At the same time, the detail feature extraction network can perform odd-even field splitting processing on each to-be-processed interleaved frame to obtain the corresponding detail features of each to-be-processed interleaved frame. Further, for the structure features, the structure feature fusion network can perform feature fusion processing on each structure feature, fusing the structure features of two adjacent to-be-processed interleaved frames, so as to obtain the fusion feature map between two adjacent to-be-processed interleaved frames, which is the first inter-frame feature map. At the same time, for the detail features, the detail feature fusion network can perform fusion processing on each detail feature, fusing the detail features of two adjacent to-be-processed interleaved frames, so as to obtain the fusion feature map between two adjacent to-be-processed interleaved frames, which is the second inter-frame feature map. Then, the first inter-frame feature map is input into the first motion perception sub-model, and the first motion perception sub-model performs feature filling processing on the first inter-frame feature map to obtain the first fusion feature map. At the same time, the second inter-frame feature map is input into the second motion perception sub-model, and the second motion perception sub-model performs feature filling processing on the second inter-frame feature map to obtain the second fusion feature map. Finally, the first fusion feature map and the second fusion feature map are input into the 2D convolutional layer, and the 2D convolutional layer processes each fusion feature map to obtain at least two target video frames corresponding to the to-be-processed interleaved frames. The advantage of this setting is that: adopting a dual-feature processing branch enables the image fusion model to process both the overall structural feature information and the detail feature information of the to-be-processed interleaved frames. Moreover, adopting a 3D convolutional layer can strengthen the information interaction between frames. The motion perception sub-model can perceive the motion situation between frames and perform feature alignment, making the content between frames more continuous, thereby improving the display effect of the target video frames.

[0070] It should be noted that since the first inter-frame feature map is obtained through feature fusion processing based on the structural features of two adjacent to-be-processed interleaved frames, when there are at least three to-be-processed interleaved frames, the first inter-frame feature map can include a first feature map and a second feature map. In practical applications, processing the first inter-frame feature map based on the first motion perception sub-model can be to process the first feature map and the second feature map respectively based on the first motion perception sub-model to obtain the first fusion feature map. The following will specifically describe Figure 3 the processing process of the first motion perception sub-model for the first inter-frame feature map.

[0071] See Figure 3As shown in the figure, processing the first inter-frame feature map based on the first motion perception sub-model to obtain the first fused feature map, including: processing the first feature map and the second feature map respectively based on the convolutional network in the first motion perception sub-model to obtain the first optical flow map and the second optical flow map; performing mapping processing on the first optical flow map and the second optical flow map based on the distortion network in the first motion perception sub-model to obtain the offset; determining the first fused feature map based on the first optical flow map, the second optical flow map, and the offset.

[0072] Those skilled in the art can understand that the optical flow map can represent the motion speed and motion direction of each pixel point in two adjacent frames of images. Optical flow is the instantaneous speed of the pixels of a spatial moving object in the observed imaging plane, and it is a method that uses the change of pixels in the time domain in an image sequence and the correlation between adjacent frames to find the corresponding relationship between the previous frame and the current frame, so as to calculate the motion information of the object between adjacent frames. The distortion network can be a network containing an inverse transformation (Backward Warping) function, which can realize the mapping between images. The offset can be obtained based on the mapping of the optical flow map and is used to represent the data of the feature displacement offset.

[0073] In a specific implementation, inputting the first inter-frame feature map into the first motion perception sub-model, the first feature map and the second feature map can be processed respectively based on the convolutional network, so as to obtain the first optical flow map for characterizing the motion speed and motion direction of each pixel point in two adjacent to-be-processed interleaved frames corresponding to the first feature map, and the second optical flow map for characterizing the motion speed and motion direction of each pixel point in two adjacent to-be-processed interleaved frames corresponding to the second feature map. Further, performing mapping processing on the first optical flow map and the second optical flow map based on the distortion network, the offset corresponding to the first optical flow map and the offset corresponding to the second optical flow map can be obtained. Finally, performing fusion processing on the first optical flow map, the second optical flow map, and the offsets corresponding to these two optical flow maps can obtain the first fused feature map. The advantage of such a setting is that it can perceive the motion of pixel points between two adjacent to-be-processed interleaved frames, and through feature alignment, make the feature content between frames more continuous, thereby improving the restoration effect of the moving object scene.

[0074] Continue to refer to Figure 3 As shown in the figure, determining the first fused feature map based on the first optical flow map, the second optical flow map, and the offset, including: performing residual processing on the first optical flow map and the offset to obtain the first feature map to be stitched; performing residual processing on the second optical flow map and the offset to obtain the second feature map to be stitched; obtaining the first fused feature map through the stitching processing of the first feature map to be stitched and the second feature map to be stitched.

[0075] In a specific embodiment, after obtaining the first optical flow map, the second optical flow map, and the offset, residual processing can be performed on the first optical flow map and the offset to align the optical flow features in the first optical flow map, thereby obtaining the first feature map to be stitched after feature alignment. At the same time, residual processing is performed on the second optical flow map and the offset to align the optical flow features in the second optical flow map, thereby obtaining the second feature map to be stitched after feature alignment. Further, the first feature map to be stitched and the second feature map to be stitched are subjected to stitching processing, so that a first fused feature map can be obtained. The advantage of such a setting is that it can achieve feature alignment of the offset features of the frames to be processed, making the frames to be processed more continuous with each other, and at the same time, it can also achieve the effect of complementing details with each other.

[0076] It should be noted that the processing process of the second inter-frame feature map based on the second motion perception sub-model is the same as the processing process of the first inter-frame feature map based on the first motion perception sub-model, and the embodiments of the present disclosure will not elaborate herein.

[0077] Exemplarily, taking three frames to be processed as an example, the processing process of the first motion perception sub-model for the first inter-frame feature map will be described exemplarily. D1, D2, and D3 can be used as the frames to be processed, and these three frames to be processed are input into the first feature extraction branch, and a first feature map and a second feature map can be obtained, which can be represented by F1 and F2. Further, F1 and F2 are input into the first motion perception sub-model, and based on the convolutional layer, F1 and F2 are processed respectively, and then the first optical flow map IF1 and the second optical flow map IF2 can be obtained. Then, based on the distortion network, mapping processing is performed on IF1 and IF2 respectively, and an offset can be obtained. Residual processing is performed on IF1 and the offset to obtain the first feature map to be stitched. Residual processing is performed on IF2 and the offset to obtain the second feature map to be stitched. Finally, for and Stitching processing is performed to obtain the first fused feature map F full .

[0078] S130. Determine a target video based on at least two target video frames.

[0079] In this embodiment, after obtaining each target video frame, each target video frame can be stitched, so that a target video composed of multiple consecutive target video frames can be obtained.

[0080] Optionally, determining a target video based on at least two target video frames includes: performing stitching processing on at least two target video frames in the time domain to obtain a target video.

[0081] In this embodiment, since each target video frame carries a corresponding timestamp, after processing each to-be-processed interleaved frame based on the image fusion model and outputting the corresponding target video frame, the application can splice multiple video frames according to the timestamps corresponding to the target video frames, so as to obtain the target video. It can be understood that by splicing multiple frames of pictures and generating the target video, the processed pictures can be displayed in a clear and coherent form.

[0082] Those skilled in the art should understand that after the application determines the target video, it can either directly play the video to display the processed video picture on the display interface, or store the target video in a specific space according to the preset path. The embodiments of the present disclosure do not make specific limitations on this.

[0083] The technical solution of the embodiments of the present disclosure can input at least three to-be-processed interleaved frames into a pre-trained image fusion model after obtaining at least three to-be-processed interleaved frames, to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames. Finally, based on the at least two target video frames, the target video is determined. When an interleaved video is displayed on an existing display device, the restoration effect of the video picture can be effectively improved. Especially for the video picture of a motion scene, a relatively significant restoration effect can also be achieved. At the same time, problems such as picture streaking and detail loss are solved, the picture quality and clarity of the video picture are improved, and the user experience is enhanced.

[0084] Figure 4 It is a schematic flowchart of a video processing method provided by the embodiments of the present disclosure. On the basis of the foregoing embodiments, the to-be-processed interleaved frames can be obtained by processing each to-be-processed video frame in the original video, and the specific implementation manner can refer to the technical solution of this embodiment. Wherein, the same or corresponding technical terms as those in the above embodiments will not be elaborated herein.

[0085] As Figure 4 shown, the method specifically includes the following steps:

[0086] S210. Obtain a plurality of to-be-processed video frames corresponding to the original video.

[0087] Among them, two to-be-processed video frames include an odd video frame and an even video frame, and the odd video frame and the even video frame are determined based on the order of the to-be-processed video frames in the original video.

[0088] In this embodiment, the original video may be a video formed by splicing interlaced video frames. The original video may be a video obtained by real-time shooting based on a terminal device, or a video pre-stored in a storage space by an application software, or a video uploaded by a user to a server or a client based on a pre-set video upload control, etc. The embodiments of the present disclosure do not make specific limitations thereto. Exemplarily, the original video may be an early imaging video. An odd video frame may be a video frame in which the number corresponding to the arrangement order in the original video is odd, and there are rendered pixel values for the odd-row pixel points, which can be rendered and displayed on the display interface. The pixel values of the even-row pixel points may be preset values and are displayed in the form of a black hole on the display interface. Correspondingly, an even video frame may be a video frame in which the number corresponding to the arrangement order in the original video is even, and there are rendered pixel values for the even-row pixel points, which can be rendered and displayed on the display interface. The pixel values of the odd-row pixel points may be preset values and are displayed in the form of a black hole on the display interface.

[0089] Exemplarily, as Figure 5 shown, where Figure 5 a may be an odd video frame, Figure 5 b may be an even video frame. The odd video frame is only scanned and sampled for odd rows. Therefore, Figure 5 only the pixel values of the odd-row pixel points in a are rendered pixel values, which may be blue and are rendered and displayed on the display interface, while the pixel values of the even-row pixel points may be preset values, which may be black. At this time, when the even-row pixel points are displayed on the display interface, they may be displayed in the form of a black hole on the display interface; similarly, the even video frame is only scanned and sampled for even rows. Therefore, Figure 5 only the pixel values of the even-row pixel points in b are rendered pixel values, which can be rendered and displayed on the display interface, while the pixel values of the odd-row pixel points may be preset values. At this time, when the odd-row pixel points are displayed on the display interface, they may be displayed in the form of a black hole on the display interface.

[0090] S220. Perform fusion processing on two adjacent video frames to be processed to obtain an interlaced frame to be processed.

[0091] In practical applications, after obtaining the original video, the original video can be parsed based on a pre-written program to obtain multiple video frames to be processed. Further, starting from the first video frame to be processed, two adjacent video frames to be processed are fused to obtain an interlaced frame to be processed.

[0092] It should be noted that since only the data of odd rows contain pixel points in odd video frames and only the data of even rows contain pixel points in even video frames, when fusing two adjacent video frames to be processed, the data with pixel points in odd video frames and the data with pixel points in even video frames can be extracted respectively, so as to obtain the interlaced frame to be processed that contains both pixel points of odd rows and pixel points of even rows.

[0093] Optionally, fusing two adjacent video frames to be processed to obtain an interlaced frame to be processed includes: extracting the odd-row data in the odd video frame and the even-row data in the even video frame; obtaining the interlaced frame to be processed by fusing the odd-row data and the even-row data.

[0094] In this embodiment, the odd-row data may be the pixel point information of the odd rows. The even-row data may be the pixel point information of the even rows. Those skilled in the art should understand that when the original video is displayed on the display interface based on the interlaced scanning method, the pixel point information of the odd rows can be sampled first to obtain the odd video frame, and then the pixel point information of the even rows can be sampled to obtain the even video frame. The pixel sampling information of the odd rows in the odd video frame can be used as the odd-row data, and the pixel sampling information of the even rows in the even video frame can be used as the even-row data.

[0095] In practical applications, after obtaining multiple video frames to be processed, for two adjacent video frames to be processed, the odd-row data of the odd video frame can be extracted, and the even-row data of the even video frame can be extracted. Further, the odd-row data and the even-row data are fused to obtain the interlaced frame to be processed. The advantage of such a setting is that: an interlaced frame to be processed that contains both the pixel information of the odd-row pixel points and the pixel information of the even-row pixel points can be obtained, so that after processing the interlaced frame to be processed, a target video frame that meets the user's requirements can be obtained.

[0096] S230. Obtain at least three interlaced frames to be processed.

[0097] S240. Input at least three interlaced frames to be processed into an image fusion model pre-trained to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed.

[0098] S250. Determine the target video based on at least two target video frames.

[0099] In the technical solution of the embodiment of the present disclosure, by obtaining a plurality of video frames to be processed corresponding to the original video, performing fusion processing on two adjacent video frames to be processed to obtain an interleaved frame to be processed, then, obtaining at least three interleaved frames to be processed, and inputting the at least three interleaved frames to be processed into an image fusion model pre-trained to obtain at least two target video frames corresponding to the at least three interleaved frames to be processed, and finally, determining a target video based on the at least two target video frames. When an interleaved video is displayed on an existing display device, the restoration effect of the video picture can be effectively improved. Especially for the video picture of a motion scene, a relatively significant restoration effect can also be achieved. At the same time, problems such as picture tearing and detail loss are solved, the picture quality and clarity of the video picture are improved, and the user experience is enhanced.

[0100] Figure 6 is a schematic structural diagram of a video processing device provided by an embodiment of the present disclosure, as Figure 6 shown, the device includes: an interleaved frame to be processed acquisition module 310, a target video frame determination module 320, and a target video determination module 330.

[0101] Among them, the interleaved frame to be processed acquisition module 310 is used to obtain at least three interleaved frames to be processed; among them, the interleaved frame to be processed is determined based on two adjacent video frames to be processed;

[0102] The target video frame determination module 320 is used to input the at least three interleaved frames to be processed into an image fusion model pre-trained to obtain at least two target video frames corresponding to the at least three interleaved frames to be processed; among them, the image fusion model includes a feature processing sub-model and a motion perception sub-model;

[0103] The target video determination module 330 is used to determine a target video based on the at least two target video frames.

[0104] Based on the above technical solutions, the device further includes: a video frame to be processed acquisition module and a video frame to be processed processing module.

[0105] The video frame to be processed acquisition module is used to obtain a plurality of video frames to be processed corresponding to the original video before obtaining the at least three interleaved frames to be processed;

[0106] The video frame to be processed processing module is used to perform fusion processing on two adjacent video frames to be processed to obtain the interleaved frame to be processed; among them, one of the two video frames to be processed includes an odd video frame and an even video frame, and the odd video frame and the even video frame are determined based on the order of the video frames to be processed in the original video.

[0107] Based on the above technical solutions, the video frame to be processed module includes: a data extraction unit and a data processing unit.

[0108] The data extraction unit is used to extract the odd-line data in the odd video frames and the even-line data in the even video frames;

[0109] The data processing unit is used to obtain the interlaced frame to be processed by performing fusion processing on the odd-line data and the even-line data.

[0110] Based on the above technical solutions, the image fusion model includes a feature processing sub-model, a motion perception sub-model, and a 2D convolutional layer.

[0111] Based on the above technical solutions, the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch;

[0112] The output of the first feature extraction branch is the input of the first motion perception sub-model in the motion perception sub-model, and the output of the second feature extraction branch is the input of the second motion perception sub-model in the motion perception sub-model;

[0113] The output of the first motion perception sub-model and the output of the second motion perception sub-model are the input of the 2D convolutional layer, so that the 2D convolutional layer outputs the target video frame.

[0114] Based on the above technical solutions, the first feature extraction branch includes a structural feature extraction network and a structural feature fusion network, and the second feature extraction branch includes a detail feature extraction network and a detail feature fusion network.

[0115] Based on the above technical solutions, the target video frame determination module 320 includes: a proportional feature extraction sub-module, an odd-even field feature extraction sub-module, a structural feature processing sub-module, a detail feature processing sub-module, a first fusion feature map determination sub-module, a second fusion feature map determination sub-module, and a target video frame determination sub-module.

[0116] The proportional feature extraction sub-module is used to perform proportional feature extraction on the at least three interlaced frames to be processed based on the structural feature extraction network, and obtain the structural features corresponding to the interlaced frames to be processed; and,

[0117] The odd-even field feature extraction sub-module is used to perform odd-even field feature extraction on the at least three interlaced frames to be processed based on the detail feature extraction network, and obtain the detail features corresponding to the interlaced frames to be processed;

[0118] A structural feature processing sub-module, configured to process the structural features based on the structural feature fusion network to obtain a first inter-frame feature map between two adjacent frames to be processed; and,

[0119] A detail feature processing sub-module, configured to process the detail features based on the detail feature fusion network to obtain a second inter-frame feature map between two adjacent frames to be processed;

[0120] A first fusion feature map determination sub-module, configured to process the first inter-frame feature map based on the first motion perception sub-model to obtain a first fusion feature map; and,

[0121] A second fusion feature map determination sub-module, configured to process the second inter-frame feature map based on the second motion perception sub-model to obtain a second fusion feature map;

[0122] A target video frame determination sub-module, configured to process the first fusion feature map and the second fusion feature map based on the 2D convolutional layer to obtain the at least two target video frames.

[0123] Based on the above technical solutions, the first fusion feature map determination sub-module includes: a feature map processing unit, an optical flow map mapping processing unit, and a first fusion feature map determination unit.

[0124] The feature map processing unit is configured to process the first feature map and the second feature map respectively based on the convolutional network in the first motion perception sub-model to obtain a first optical flow map and a second optical flow map;

[0125] The optical flow map mapping processing unit is configured to perform mapping processing on the first optical flow map and the second optical flow map based on the distortion network in the first motion perception sub-model to obtain an offset;

[0126] The first fusion feature map determination unit is configured to determine the first fusion feature map based on the first optical flow map, the second optical flow map, and the offset.

[0127] Based on the above technical solutions, the first fusion feature map determination unit is specifically configured to perform residual processing on the first optical flow map and the offset to obtain a first feature map to be stitched; perform residual processing on the second optical flow map and the offset to obtain a second feature map to be stitched; and obtain the first fusion feature map through stitching processing on the first feature map to be stitched and the second feature map to be stitched.

[0128] Based on the above technical solutions, the target video determination module 330 is specifically configured to perform stitching processing on the at least two target video frames in the time domain to obtain the target video.

[0129] In the technical solution of the embodiments of the present disclosure, after obtaining at least three to-be-processed interleaved frames, the at least three to-be-processed interleaved frames can be input into an image fusion model pre-trained to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames. Finally, based on the at least two target video frames, a target video is determined. When an interleaved video is displayed on an existing display device, the restoration effect of the video picture can be effectively improved. Especially for the video picture of a motion scene, a relatively significant restoration effect can also be achieved. At the same time, problems such as picture streaking and detail loss are solved, the picture quality and clarity of the video picture are improved, and the user experience is enhanced.

[0130] The video processing device provided by the embodiments of the present disclosure can execute the video processing method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the method.

[0131] It should be noted that the various units and modules included in the above device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present disclosure.

[0132] Figure 7 is a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. Next, refer to Figure 7 , which shows a schematic structural diagram of an electronic device 500 suitable for implementing the embodiments of the present disclosure (such as Figure 7 the terminal device or server in). The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0133] As Figure 7 shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The editing / output (I / O) interface 505 is also connected to the bus 504.

[0134] Typically, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 7 the electronic device 500 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.

[0135] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.

[0136] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0137] The electronic device provided by the embodiment of the present disclosure and the video processing method provided by the above embodiment belong to the same inventive concept. Technical details not described in detail in this embodiment can be seen in the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0138] The embodiment of the present disclosure provides a computer storage medium, on which a computer program is stored, and when the program is executed by a processor, the video processing method provided by the above embodiment is implemented.

[0139] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0140] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (for example, a communication network). Examples of the communication network include a local area network ("LAN"), a wide area network ("WAN"), the Internet (for example, the Internet), and a peer-to-peer network (for example, an ad hoc peer-to-peer network), as well as any currently known or future-developed network.

[0141] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0142] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to:

[0143] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:

[0144] Obtain at least three to-be-processed interleaved frames; wherein, the to-be-processed interleaved frames are determined based on two adjacent to-be-processed video frames;

[0145] Input the at least three to-be-processed interleaved frames into an image fusion model pre-trained to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames; wherein, the image fusion model includes a feature processing sub-model and a motion perception sub-model;

[0146] Determine a target video based on the at least two target video frames.

[0147] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet service provider via the Internet).

[0148] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0149] The units involved in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".

[0150] The functions described above in this article may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.

[0151] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0152] According to one or more embodiments of the present disclosure, [Example 1] provides a video processing method, which includes:

[0153] Obtain at least three to-be-processed interleaved frames; wherein, the to-be-processed interleaved frames are determined based on two adjacent to-be-processed video frames;

[0154] Input the at least three to-be-processed interleaved frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames; wherein, the image fusion model includes a feature processing sub-model and a motion perception sub-model;

[0155] Determine a target video based on the at least two target video frames.

[0156] According to one or more embodiments of the present disclosure, [Example 2] provides a video processing method, which further includes:

[0157] Optionally, obtain a plurality of to-be-processed video frames corresponding to the original video;

[0158] Fuse two adjacent video frames to be processed to obtain the interleaved frame to be processed;

[0159] Among them, the two video frames to be processed include an odd video frame and an even video frame, and the odd video frame and the even video frame are determined based on the order of the video frames to be processed in the original video.

[0160] According to one or more embodiments of the present disclosure, [Example Three] provides a video processing method, which further includes:

[0161] Optionally, extract the odd-line data in the odd video frame and the even-line data in the even video frame;

[0162] By fusing the odd-line data and the even-line data, the interleaved frame to be processed is obtained.

[0163] According to one or more embodiments of the present disclosure, [Example Four] provides a video processing method, which further includes:

[0164] Optionally, the image fusion model includes a feature processing sub-model, a motion perception sub-model, and a 2D convolutional layer.

[0165] According to one or more embodiments of the present disclosure, [Example Five] provides a video processing method, which further includes:

[0166] Optionally, the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch;

[0167] The output of the first feature extraction branch is the input of the first motion perception sub-model in the motion perception sub-model, and the output of the second feature extraction branch is the input of the second motion perception sub-model in the motion perception sub-model;

[0168] The output of the first motion perception sub-model and the output of the second motion perception sub-model are the inputs of the 2D convolutional layer, so that the 2D convolutional layer outputs the target video frame.

[0169] According to one or more embodiments of the present disclosure, [Example Six] provides a video processing method, which further includes:

[0170] Optionally, the first feature extraction branch includes a structural feature extraction network and a structural feature fusion network, and the second feature extraction branch includes a detail feature extraction network and a detail feature fusion network.

[0171] According to one or more embodiments of the present disclosure, [Example Seven] provides a video processing method, which further includes:

[0172] Optionally, based on the structural feature extraction network, perform equal-proportion feature extraction on the at least three to-be-processed interleaved frames to obtain structural features corresponding to the to-be-processed interleaved frames; and,

[0173] Based on the detail feature extraction network, perform odd-even field feature extraction on the at least three to-be-processed interleaved frames to obtain detail features corresponding to the to-be-processed interleaved frames;

[0174] Based on the structural feature fusion network, process the structural features to obtain a first inter-frame feature map between two adjacent to-be-processed interleaved frames; and,

[0175] Based on the detail feature fusion network, process the detail features to obtain a second inter-frame feature map between two adjacent to-be-processed interleaved frames;

[0176] Based on the first motion perception sub-model, process the first inter-frame feature map to obtain a first fused feature map; and,

[0177] Based on the second motion perception sub-model, process the second inter-frame feature map to obtain a second fused feature map;

[0178] Based on the 2D convolutional layer, process the first fused feature map and the second fused feature map to obtain the at least two target video frames.

[0179] According to one or more embodiments of the present disclosure, [Example VIII] provides a video processing method, and the method further includes:

[0180] Optionally, based on the convolutional network in the first motion perception sub-model, process the first feature map and the second feature map respectively to obtain a first optical flow map and a second optical flow map;

[0181] Based on the distortion network in the first motion perception sub-model, perform mapping processing on the first optical flow map and the second optical flow map to obtain an offset;

[0182] Based on the first optical flow map, the second optical flow map, and the offset, determine the first fused feature map.

[0183] According to one or more embodiments of the present disclosure, [Example IX] provides a video processing method, and the method further includes:

[0184] Optionally, perform residual processing on the first optical flow map and the offset to obtain a first to-be-spliced feature map;

[0185] Perform residual processing on the second optical flow map and the offset to obtain a second to-be-spliced feature map;

[0186] By performing splicing processing on the first feature map to be spliced and the second feature map to be spliced, the first fused feature map is obtained.

[0187] According to one or more embodiments of the present disclosure, [Example Ten] provides a video processing method, which further includes:

[0188] Optionally, perform splicing processing on the at least two target video frames in the time domain to obtain the target video.

[0189] According to one or more embodiments of the present disclosure, [Example Eleven] provides a video processing apparatus, which includes:

[0190] An interlaced frame to be processed acquisition module, configured to acquire at least three interlaced frames to be processed; wherein, the interlaced frames to be processed are determined based on two adjacent video frames to be processed;

[0191] A target video frame determination module, configured to input the at least three interlaced frames to be processed into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three interlaced frames to be processed; wherein, the image fusion model includes a feature processing sub-model and a motion perception sub-model;

[0192] A target video determination module, configured to determine a target video based on the at least two target video frames.

[0193] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present disclosure.

[0194] In addition, although the operations are depicted in a specific order, this should not be construed as requiring the operations to be performed in the specific order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0195] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims.

Claims

1. A video processing method, characterized in that, Including: Obtaining at least three to-be-processed interleaved frames; wherein, the to-be-processed interleaved frames are determined by fusing two adjacent to-be-processed video frames; the to-be-processed video frames are video frames obtained by interlaced scanning. Inputting the at least three to-be-processed interleaved frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames; wherein, the image fusion model includes a feature processing sub-model, a motion perception sub-model, and a 2D convolutional layer; the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch; the first feature extraction branch includes a structure feature extraction network and a structure feature fusion network, and the second feature extraction branch includes a detail feature extraction network and a detail feature fusion network. Determining a target video based on the at least two target video frames.

2. The method according to claim 1, wherein Before obtaining the at least three to-be-processed interleaved frames, it further includes: Obtaining a plurality of to-be-processed video frames corresponding to the original video. Performing fusion processing on two adjacent to-be-processed video frames to obtain the to-be-processed interleaved frames. Wherein, the two to-be-processed video frames include an odd video frame and an even video frame, and the odd video frame and the even video frame are determined based on the order of the to-be-processed video frames in the original video.

3. The method according to claim 2, wherein The performing fusion processing on two adjacent to-be-processed video frames to obtain to-be-processed interleaved frames includes: Extracting odd-line data in the odd video frame and even-line data in the even video frame. Obtaining the to-be-processed interleaved frames by performing fusion processing on the odd-line data and the even-line data.

4. The method according to claim 1, characterized in that, It further includes: The output of the first feature extraction branch is the input of the first motion perception sub-model in the motion perception sub-model, and the output of the second feature extraction branch is the input of the second motion perception sub-model in the motion perception sub-model. The output of the first motion perception sub-model and the output of the second motion perception sub-model are the input of the 2D convolutional layer, so that the 2D convolutional layer outputs target video frames.

5. The method according to claim 1, characterized in that, The inputting the at least three to-be-processed interleaved frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames includes: Performing equal-proportion feature extraction on the at least three to-be-processed interleaved frames based on the structure feature extraction network to obtain structure features corresponding to the to-be-processed interleaved frames; and Performing odd-even field feature extraction on the at least three to-be-processed interleaved frames based on the detail feature extraction network to obtain detail features corresponding to the to-be-processed interleaved frames. Processing the structure features based on the structure feature fusion network to obtain a first inter-frame feature map between two adjacent to-be-processed interleaved frames; and processing the detail features based on the detail feature fusion network to obtain a second inter-frame feature map between two adjacent to-be-processed interleaved frames. Processing the first inter-frame feature map based on the first motion perception sub-model to obtain a first fusion feature map; and processing the second inter-frame feature map based on the second motion perception sub-model to obtain a second fusion feature map. Processing the first fused feature map and the second fused feature map based on the 2D convolutional layer to obtain the at least two target video frames.

6. The method according to claim 5, wherein The first inter-frame feature map includes a first feature map and a second feature map. Processing the first inter-frame feature map based on the first motion perception sub-model to obtain a first fused feature map includes: Processing the first feature map and the second feature map respectively based on the convolutional network in the first motion perception sub-model to obtain a first optical flow map and a second optical flow map; Performing mapping processing on the first optical flow map and the second optical flow map based on the distortion network in the first motion perception sub-model to obtain an offset; Determining the first fused feature map based on the first optical flow map, the second optical flow map, and the offset.

7. The method according to claim 6, wherein The determining the first fused feature map based on the first optical flow map, the second optical flow map, and the offset includes: Performing residual processing on the first optical flow map and the offset to obtain a first feature map to be stitched; Performing residual processing on the second optical flow map and the offset to obtain a second feature map to be stitched; Obtaining the first fused feature map by performing stitching processing on the first feature map to be stitched and the second feature map to be stitched.

8. The method according to claim 1, wherein The determining the target video based on the at least two target video frames includes: Performing stitching processing on the at least two target video frames in the time domain to obtain the target video.

9. A video processing apparatus, characterized in that, Including: A to-be-processed interleaved frame acquisition module, configured to acquire at least three to-be-processed interleaved frames; wherein, the to-be-processed interleaved frames are determined by fusing two adjacent to-be-processed video frames; the to-be-processed video frames are video frames obtained by interlaced scanning; A target video frame determination module, configured to input the at least three to-be-processed interleaved frames into a pre-trained image fusion model to obtain at least two target video frames corresponding to the at least three to-be-processed interleaved frames; wherein, the image fusion model includes a feature processing sub-model, a motion perception sub-model, and a 2D convolutional layer; the feature processing sub-model includes a first feature extraction branch and a second feature extraction branch; the first feature extraction branch includes a structural feature extraction network and a structural feature fusion network, and the second feature extraction branch includes a detail feature extraction network and a detail feature fusion network; A target video determination module, configured to determine a target video based on the at least two target video frames.

10. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method according to any one of claims 1-8.

11. A storage medium containing computer-executable instructions, the computer-executable instructions being used to execute the video processing method according to any one of claims 1-8 when executed by a computer processor.

Citation Information

Patent Citations

  • Deinterlacing via deep learning

    US20220014708A1