Video Repair Processing Method, Device, Computer Equipment and Storage Medium
By performing field replacement processing, interleaving detection and residual feature extraction and superposition on the original video, the problems of low video repair efficiency and interleaving field patterns in traditional technology are solved, and efficient and automatic video repair processing is achieved.
Patent Information
- Application Number
- CN202211356009.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-01
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-11-01
AI Technical Summary
When processing videos processed by film over tape, the prior art has problems such as interlaced field patterns and picture lags, and the traditional reverse film over tape method requires manual intervention and low efficiency.
By performing field replacement processing on the original image in the original video, interleaved field patterns are detected and filtered, residual feature data is extracted, and superimposed on the corresponding image data to repair the interleaved field patterns and generate the target image after deinterleaving repair.
It realizes efficient repair of interlaced field patterns in videos without manual intervention, and improves the efficiency and quality of video repair processing.
Smart Images

Figure CN115801979B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a video repair processing method, device, computer equipment and storage medium. Background Art
[0002] Many old videos were processed by telecine when they were produced in order to adapt to different TV formats and digital video discs. However, with the development of video processing technology, when videos that have been processed by telecine are played on modern progressive scan monitors, there will be image quality problems such as interlaced field patterns and image freezes.
[0003] In traditional technology, the reverse film advance method is used to deal with the film advance problem. Since each frame in the video needs to be analyzed and then restored in reverse using different methods, the repair process must be manually intervened, which will cause the problem of low efficiency of video repair. Summary of the invention
[0004] Based on this, it is necessary to provide a video repair processing method, device, computer equipment, computer-readable storage medium and computer program product that can improve efficiency in order to solve the above technical problems.
[0005] In a first aspect, the present application provides a video repair processing method. The method comprises:
[0006] Performing field replacement processing on an original image in an original video to obtain a plurality of first images;
[0007] Performing interlacing detection on the first image to obtain an interlacing detection result; the interlacing detection result is used to reflect the number of interlaced field patterns in the first image;
[0008] Screening the plurality of first images according to the interlaced detection result to obtain a second image;
[0009] Perform residual feature extraction based on the second image to obtain residual feature data corresponding to the field data in the second image;
[0010] The residual feature data is superimposed on the corresponding field data in the second image to obtain a target image after de-interlacing and restoration; the target image is used as a video frame to constitute a restored target video.
[0011] In one embodiment, the process of performing field replacement on the original images in the original video to obtain a plurality of first images includes: determining an image group based on the time sequence corresponding to the original images in the original video; the image group includes a plurality of original images with consecutive time sequences; respectively using the even-field data in the plurality of original images in the image group to replace the even-field data in the median original image, and retaining the odd-field data in the median original image, to obtain a plurality of first images corresponding to the image group; the median original image is the original image with the median time sequence in the image group.
[0012] In one embodiment, the interlacing detection result includes an interlacing score; the process of performing interlacing detection on the first images to obtain an interlacing detection result includes: inputting the first images into an interlacing detection model, and detecting the interlaced field patterns in the first images through the interlacing detection model to predict the interlacing score corresponding to the first images.
[0013] In one embodiment, the process of performing residual feature extraction on the second image to obtain the residual feature data corresponding to the field data in the second image includes: performing residual feature extraction on the even-field data and odd-field data in the second image to obtain the even-field residual data corresponding to the even-field data in the second image, and the odd-field residual data corresponding to the odd-field data in the second image; the process of superimposing the residual feature data onto the corresponding field data in the second image to obtain the de-interlaced and repaired target image includes: obtaining the de-interlaced even-field image by superimposing the even-field residual data and the even-field data in the second image; obtaining the de-interlaced odd-field image by superimposing the odd-field residual data and the odd-field data in the second image; screening out the target image from the even-field image and the odd-field image.
[0014] In one embodiment, the process of performing residual feature extraction on the even-field data and odd-field data in the second image to obtain the even-field residual data corresponding to the even-field data in the second image, and the odd-field residual data corresponding to the odd-field data in the second image includes: inputting the second image into the interlacing segmentation layer in the de-interlacing model, to determine the odd-field data and even-field data from the second image through the interlacing segmentation layer; the de-interlacing model further includes an attention feature extraction layer; merging and inputting the odd-field data and even-field data of the second image into the attention feature extraction layer to perform residual feature extraction, to obtain the even-field residual data corresponding to the even-field data in the second image, and the odd-field residual data corresponding to the odd-field data in the second image.
[0015] In one embodiment, there are multiple target images obtained by deinterlacing and repairing the original video; the method further includes: segmenting the multiple target images to obtain interlaced video segments; the interlaced video segments have corresponding repeated frame numbers; for the multiple target images in the interlaced video segments, determining the similarity between adjacent images among the multiple target images; according to the similarity, determining repeated frame images that meet the repeated frame number from the interlaced video segments; extracting the repeated frame images from the interlaced video segments; the remaining target images after extracting the repeated frame images are used as video frames to form the repaired target video.
[0016] In one embodiment, the method further includes: extracting texture feature information from the target images; performing super-resolution enhancement on the target images based on the texture feature information to obtain super-resolution enhanced target images.
[0017] In a second aspect, the present application also provides a video repair processing device. The device includes:
[0018] A deinterlacing module, configured to perform field replacement processing on the original images in the original video to obtain multiple first images; performing interlacing detection on the first images to obtain an interlacing detection result; the interlacing detection result is used to reflect the number of interlaced field patterns in the first images; screening the multiple first images according to the interlacing detection result to obtain second images;
[0019] A repair module, configured to extract residual feature data corresponding to the field data in the second images based on the second images; superimposing the residual feature data on the corresponding field data in the second images to obtain deinterlaced and repaired target images; the target images are used as video frames to form the repaired target video.
[0020] In one embodiment, the deinterlacing module is configured to determine an image group based on the timing corresponding to the original images in the original video; the image group includes multiple sequentially timed original images; respectively replacing the even-field data in the median original image with the even-field data in the multiple original images in the image group, and retaining the odd-field data in the median original image to obtain multiple first images corresponding to the image group; the median original image is the original image corresponding to the median timing in the image group.
[0021] In one embodiment, the interlacing detection result includes an interlacing score; the deinterlacing module is configured to input the first images into an interlacing detection model, and detect the interlaced field patterns in the first images through the interlacing detection model to predict the interlacing score corresponding to the first images.
[0022] In one embodiment, the repair module is configured to extract residual features from the even-field data and odd-field data in the second image to obtain even-field residual data corresponding to the even-field data in the second image and odd-field residual data corresponding to the odd-field data in the second image; and is configured to obtain an anti-interlaced even-field image by superimposing the even-field residual data on the even-field data in the second image; obtain an anti-interlaced odd-field image by superimposing the odd-field residual data on the odd-field data in the second image; and screen a target image from the even-field image and the odd-field image.
[0023] In one embodiment, the repair module is configured to input the second image into an interlace segmentation layer in an anti-interlacing model to determine odd-field data and even-field data from the second image through the interlace segmentation layer; the anti-interlacing model further includes an attention feature extraction layer; and the odd-field data and even-field data of the second image are combined and input into the attention feature extraction layer to perform residual feature extraction, so as to obtain even-field residual data corresponding to the even-field data in the second image and odd-field residual data corresponding to the odd-field data in the second image.
[0024] In one embodiment, there are multiple target images obtained by anti-interlacing repair for the original video; the repair module is configured to segment the multiple target images to obtain interlaced video segments; the interlaced video segments have corresponding repeated frame numbers; for the multiple target images in the interlaced video segments, determine the similarity between adjacent images among the multiple target images; according to the similarity, determine repeated frame images that meet the repeated frame numbers from the interlaced video segments; extract the repeated frame images from the interlaced video segments; and the remaining target images after extracting the repeated frame images are used as video frames to constitute a repaired target video.
[0025] In one embodiment, the repair module is configured to extract texture feature information from the target image; and perform super-resolution enhancement on the target image based on the texture feature information to obtain a super-resolution enhanced target image.
[0026] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps in each embodiment of the method of the present application are implemented.
[0027] In a fourth aspect, the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in each embodiment of the method of the present application are implemented.
[0028] Fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the steps in the embodiments of the method of the present application.
[0029] For the above video restoration processing method, device, computer device, storage medium and computer program product, perform field replacement processing on the original images in the original video to obtain a plurality of first images; perform interlace detection on the first images to obtain an interlace detection result; the interlace detection result is used to reflect the number of interlace field patterns in the first images; screen the plurality of first images according to the interlace detection result to obtain second images; perform residual feature extraction based on the second images to obtain residual feature data corresponding to the field data in the second images; superimpose the residual feature data on the corresponding field data in the second images to obtain a de-interlaced and restored target image; the target image is used as a video frame to form a restored target video. By performing interlace detection on the first images obtained by performing field replacement processing on the original images, screening out the second images, and after extracting the residual feature data corresponding to the field data in the second images, obtaining the target image by superimposing the residual feature data and the corresponding field data, the video restoration processing of the original video can be completed. Compared with traditional methods, it does not require manual intervention and improves the efficiency of video restoration processing. Description of the Drawings
[0030] Figure 1 It is an application environment diagram of the video restoration processing method in one embodiment;
[0031] Figure 2 It is a schematic flowchart of the video restoration processing method in one embodiment;
[0032] Figure 3 It is a schematic diagram of performing field replacement processing on the original images in one embodiment;
[0033] Figure 4A It is a schematic diagram of a de-interlacing model in one embodiment;
[0034] Figure 4B It is a schematic diagram of a focus feature unit in the attention feature extraction layer in one embodiment;
[0035] Figure 4C It is a schematic diagram of the feature amplification layer of the de-interlacing model in one embodiment;
[0036] Figure 5 It is a schematic diagram of extracting repeated frame images from an interlaced video segment in one embodiment;
[0037] Figure 6 It is a schematic diagram of a super-resolution enhancement model in one embodiment;
[0038] Figure 7It is a simple flowchart of a video restoration processing method in an embodiment;
[0039] Figure 8 It is a structural block diagram of a video restoration processing device in an embodiment;
[0040] Figure 9 It is an internal structure diagram of a computer device in an embodiment;
[0041] Figure 10 It is an internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0042] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0043] The video restoration processing method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. The server 104 can perform field replacement processing on the original images in the original video to obtain a plurality of first images; the server 104 can perform interlace detection on the first images to obtain an interlace detection result; the server 104 can screen the plurality of first images according to the interlace detection result to obtain second images; the server 104 can perform residual feature extraction based on the second images to obtain residual feature data corresponding to the field data in the second images; the server 104 can superimpose the residual feature data on the field data in the corresponding second images to obtain a de-interlaced and restored target image; it can be understood that the server 104 can determine the target image as the target video composed of video frames and send the target video to the terminal 102. The terminal 102 can display the target video.
[0044] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0045] In one embodiment, as Figure 2 shown, a video restoration processing method is provided. The method is applied to Figure 1Taking the server in [the relevant context] as an example for illustration, it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0046] Step 202, perform field replacement processing on the original images in the original video to obtain multiple first images.
[0047] Among them, the original video is the video to be subjected to video repair processing. The original image is a video frame in the original video. The first image is obtained by performing field replacement processing on the original image. It can be understood that the video frames in the original video include interlaced odd and even fields. Such an original video cannot meet modern video playback standards in terms of picture quality, specifications, etc., so video repair processing is required. When obtaining an image, two scanning fields that are exchanged for display in the vertical direction constitute each complete frame of the picture. Each scanning field only contains half of the total number of rows of the scanned image. One scanning field that is all odd-numbered rows is called the odd field, and the other scanning field that is all even-numbered rows is called the even field. Therefore, the original images in the original video include odd-field data and even-field data.
[0048] Specifically, the server can determine the field data in the original image, and use the field data of other original images except this original image to perform field replacement processing on the field data in the original image to obtain multiple first images. It can be understood that the first image is obtained after the field data in the original image is replaced.
[0049] In one embodiment, after determining the current original image to be subjected to field replacement processing, the server can determine the reference original image corresponding to the current original image based on the timing of each original image in the original video. The server can use the field data in the reference original image to replace the field data in the current original image, and retain the other field data in the current original image. It can be understood that the other field data is the field data except the replaced field data. For example, if the odd-field data in the current original image is replaced, the server will retain the even-field data in the current original image, so that the even-field data of the current original image and the odd-field data of the reference original image constitute the first image.
[0050] In one embodiment, the field data of the current original image to be replaced should be consistent with the field data in the reference original image to be replaced. For example, the server can use the odd-field data in the reference original image to replace the odd-field data in the current original image. The server can also use the even-field data in the reference original image to replace the even-field data in the current original image.
[0051] In one embodiment, the server may determine a current original image to be subjected to field replacement processing, at least one previous original image whose timing is before the current original image, and at least one next original image whose timing is after the current original image. The server may use at least one of the previous original image and the next original image as a reference original image, and use the field data in the reference original image to replace the field data in the current original image to obtain a plurality of first images.
[0052] In one embodiment, the original video may be an old video. For example, anime videos, movie videos, and TV videos, etc.
[0053] Step 204: Perform interlace detection on the first images to obtain an interlace detection result; screen the plurality of first images according to the interlace detection result to obtain second images.
[0054] Among them, the interlace detection result is used to reflect the number of interlace field patterns in the first images.
[0055] Specifically, the server may perform interlace detection on the first images to obtain an interlace detection result. The server may determine the interlace field pattern data reflected by the interlace detection result, so as to screen out the first image with the fewest interlace field patterns from the plurality of first images to obtain second images.
[0056] Step 206: Extract residual features based on the second images to obtain residual feature data corresponding to the field data in the second images; superimpose the residual feature data on the corresponding field data in the second images to obtain a de-interlaced and restored target image.
[0057] Among them, the target image is used as a video frame to constitute a restored target video.
[0058] Specifically, the server may jointly extract residual features from the odd-field data and even-field data in the second images to obtain residual feature data corresponding to the odd-field data and even-field data in the second images. The server may obtain a de-interlaced and restored target image by superimposing the residual feature data on the corresponding odd-field data and even-field data in the second images.
[0059] In the above video restoration processing method, the original images in the original video are subjected to field replacement processing to obtain a plurality of first images; the first images are subjected to interlace detection to obtain an interlace detection result; the interlace detection result is used to reflect the number of interlace field patterns of the first images; the plurality of first images are screened according to the interlace detection result to obtain second images; residual feature extraction is performed based on the second images to obtain residual feature data corresponding to the field data in the second images; the residual feature data is superimposed on the corresponding field data in the second images to obtain a deinterlaced restored target image; the target image is used as a video frame to form a restored target video. By performing interlace detection on the first images obtained by subjecting the original images to field replacement processing, screening out the second images, and after extracting the residual feature data corresponding to the field data in the second images, obtaining the target image by superimposing the residual feature data and the corresponding field data, the video restoration processing of the original video can be completed. Compared with the traditional method, it does not require manual intervention and improves the efficiency of video restoration processing.
[0060] Moreover, by performing residual feature extraction on the second images and superimposing the residual feature data on the corresponding field data in the second images to obtain the target image, the quality of the target image after residual optimization is higher, improving the instruction of video restoration processing.
[0061] In one embodiment, subjecting the original images in the original video to field replacement processing to obtain a plurality of first images includes: determining an image group based on the timing corresponding to the original images in the original video; the image group includes a plurality of original images with consecutive timings; respectively using the even-field data in the plurality of original images in the image group to replace the even-field data in the median original image, and retaining the odd-field data in the median original image, to obtain a plurality of first images corresponding to the image group; the median original image is the original image with the median timing in the image group.
[0062] Specifically, the server can determine the pre-reference original image corresponding to the previous timing of each original image and the post-reference original image corresponding to the next timing according to the timing corresponding to the original images in the original video, to obtain an image group composed of each original image, the pre-reference image, and the post-reference image. It can be understood that each original image can be used as the median original image in an image group. For the original image corresponding to the first timing, there is no pre-reference original image. At this time, the original image corresponding to the first timing is equivalent to the median original image of the first image group. For the original image corresponding to the last timing, there is no post-reference original image. At this time, the original image corresponding to the last timing is equivalent to the median original image of the last image group. The server can respectively use the even-field data in the pre-reference original image and the post-reference original image in the image group to replace the even-field data in the median original image, and retain the odd-field data in the median original image, to obtain a plurality of first images corresponding to the image group.
[0063] In one embodiment, as Figure 3 shown, a schematic diagram of performing field replacement processing on the original image is provided. The pre-reference original image includes odd field data (Odd(t-1)) and even field data (Even(t-1)). The post-reference original image includes odd field data (Odd(t+1)) and even field data (Even(t+1)). The median original image includes odd field data (Odd(t)) and even field data (Even(t)). The image group includes 3 frames of original images, denoted as P(t-1), P(t), and P(t+1) respectively. Among them, P(t) is the median original image, P(t-1) is the pre-reference image, and P(t+1) is the post-reference image. The server can combine the even field data Even(t-1) and Even(t+1) of P(t-1) and P(t+1) with the odd field data Odd(t) of P(t) to form two first images. It can be understood that including P(t), after the field replacement processing, the server can obtain 3 first images.
[0064] In this embodiment, by determining the image group based on the timing corresponding to the original images in the original video; respectively using the even field data in multiple original images in the image group to replace the even field data in the median original image and retaining the odd field data in the median original image, multiple first images corresponding to the image group are obtained. Subsequently, the second image can be screened out based on the first image, so that the residual features of the second image can be extracted, thereby improving the quality of video restoration.
[0065] In one embodiment, the interleaving detection result includes an interleaving score; performing interleaving detection on the first image to obtain the interleaving detection result includes: inputting the first image into the interleaving detection model, and detecting the interleaving field patterns in the first image through the interleaving detection model to predict the interleaving score corresponding to the first image.
[0066] Among them, the interleaving score can evaluate the number of interleaving field patterns in the first image. It can be understood that the interleaving score and the interleaving field patterns can be positively correlated or negatively correlated. For example, the higher the interleaving score of the first image, the fewer the number of interleaving field patterns in it. When performing field replacement on the original image, if the field data used for replacement does not match the field data in the original image well, the number of interleaving field patterns in the first image obtained after the field replacement will be large, that is, the number of interleaving field patterns and the matching degree are negatively correlated. It can be understood that the pixels in the image are highly correlated with the neighboring pixels. If the field data in the image do not match, interleaving field patterns will appear.
[0067] Specifically, the server can input the first image into the interleaving detection model, detect the interleaving field patterns in the first image through the interleaving detection model to determine the number of interleaving field patterns in the first image, and obtain the interleaving score corresponding to the first image predicted by the interleaving detection model. It can be understood that the interleaving detection model is used to detect the interleaving field patterns in the image.
[0068] In one embodiment, the server can screen out the first image with the highest interleaving score from multiple first images corresponding to the image group to obtain the second image. It can be understood that the image group is the unit for performing field replacement processing on the current original image. The matching degree of odd-field data and even-field data in the second image is the highest among the multiple first images. Among them, the current original image to be subjected to field replacement processing is the median original image in the image group.
[0069] In one embodiment, it further includes the training step of the interleaving detection model. The interleaving detection model can be a deep convolutional model. After preprocessing the interleaving data set, use the preprocessed interleaving data set to train the deep convolutional model. After the training is completed, obtain the model file of the interleaving detection model. When performing interleaving detection, the server can call the model file of the interleaving detection model and input the obtained multiple first images into the interleaving detection model to obtain the interleaving score corresponding to each first image. For example, the server can input the obtained 3 first images into the interleaving detection model respectively to obtain the interleaving scores corresponding to these 3 first images respectively. It can be understood that the computer device used to train the model can include at least one of a terminal and a server.
[0070] In one embodiment, the interleaving data set includes training interleaving images with different numbers of interleaving field patterns. When preprocessing the interleaving data set, a person can set labels for each image through a computer device. It can be understood that a clean image, that is, an image with very few or no interleaving field patterns, is marked as 1, while an image with many interleaving field patterns is marked as 0. Then, normalize the interleaving data set after setting the labels to obtain a normalized interleaving data set. During the training of the interleaving detection model, the computer device can extract the interleaving field pattern features of the images in the normalized interleaving data set through multiple residual networks, and then use the AdaptiveAvgPool2d function to downsample the interleaving field pattern features to reduce the number of parameters. Finally, output the interleaving score through a linear activation layer. Among them, the activation function of the linear activation layer can be the softmax function.
[0071] In this embodiment, by inputting the first image into the interleaving detection model and detecting the interleaving field patterns in the first image through the interleaving detection model to predict the interleaving score corresponding to the first image, the second image with the best field replacement effect can be determined based on the interleaving score, thereby improving the quality of video restoration.
[0072] In one embodiment, residual feature extraction is performed based on the second image, and the residual feature data corresponding to the field data in the second image is obtained, including: performing residual feature extraction on the even-field data and odd-field data in the second image to obtain the even-field residual data corresponding to the even-field data in the second image and the odd-field residual data corresponding to the odd-field data in the second image; superimposing the residual feature data on the corresponding field data in the second image to obtain the target image after deinterlacing repair, including: obtaining the deinterlaced even-field image by superimposing the even-field residual data on the even-field data in the second image; obtaining the deinterlaced odd-field image by superimposing the odd-field residual data on the odd-field data in the second image; and screening out the target image from the even-field image and the odd-field image.
[0073] Among them, the odd-field residual data is used to supplement the odd-field data to obtain the image data of the complete odd-field image. The even-field residual data is used to supplement the even-field data to obtain the image data of the complete even-field image.
[0074] Specifically, after the field replacement process, more or less interlaced field patterns will appear. There may also be a small amount of interlaced field patterns in the second image. The pixels in the image are highly correlated with the neighboring pixels. By supplementing the information between the two field data, the odd-field data is extended to obtain the deinterlaced odd-field image, and the even-field data is extended to obtain the deinterlaced even-field image, further reducing the interlaced field patterns. The server can perform residual feature extraction on the even-field data and odd-field data in the second image to obtain the even-field residual data corresponding to the even-field data in the second image and the odd-field residual data corresponding to the odd-field data in the second image. The server can obtain the deinterlaced even-field image by superimposing the even-field residual data on the even-field data in the second image; obtain the deinterlaced odd-field image by superimposing the odd-field residual data on the odd-field data in the second image; and screen out the target image from the even-field image and the odd-field image.
[0075] In one embodiment, the server can input the second image into the deinterlacing model to obtain the odd-field image and the even-field image output by the deinterlacing model. Among them, the deinterlacing model is used to extend the odd-field data and even-field data in the second image to obtain the deinterlaced odd-field image and even-field image.
[0076] In one embodiment, the server can fixedly select one of the even-field image or the odd-field image to obtain the target image. It can be understood that since the original video is interlaced, only the odd-field image extended from the odd-field data can ensure that the selected target image is continuous, and only the even-field image extended from the even-field data can also ensure that the selected target image is continuous.
[0077] In this embodiment, by performing residual feature extraction and adding the extracted odd-field residual data and even-field residual data to the odd-field data and even-field data respectively, the de-interlaced odd-field image and even-field image are obtained. Compared with the second image obtained by directly replacing the field data, the target image can further avoid the influence of interlaced field patterns, thereby further improving the quality of video restoration.
[0078] In one embodiment, performing residual feature extraction on the even-field data and odd-field data in the second image to obtain the even-field residual data corresponding to the even-field data in the second image and the odd-field residual data corresponding to the odd-field data in the second image includes: inputting the second image into the interlace segmentation layer in the de-interlacing model to determine the odd-field data and even-field data from the second image through the interlace segmentation layer; the de-interlacing model further includes an attention feature extraction layer; combining and inputting the odd-field data and even-field data of the second image into the attention feature extraction layer to perform residual feature extraction, obtaining the even-field residual data corresponding to the even-field data in the second image, and the odd-field residual data corresponding to the odd-field data in the second image.
[0079] Specifically, the server can input the second image into the interlace segmentation layer in the de-interlacing model to determine the odd-field data and even-field data from the second image through the interlace segmentation layer. It can be understood that the interlace segmentation layer is used to split the odd-field data and even-field data from the second image. The server can combine and input the odd-field data and even-field data of the second image into the attention feature extraction layer to perform residual feature extraction, obtaining the even-field residual data corresponding to the even-field data in the second image, and the odd-field residual data corresponding to the odd-field data in the second image.
[0080] In one embodiment, the video restoration processing method further includes the training step of the de-interlacing model. The de-interlacing model can be a deep convolutional model. During the training process, first preprocess the de-interlacing dataset and use the preprocessed de-interlacing dataset to train the deep convolutional model. After training is completed, the model file of the de-interlacing model is obtained. Among them, the de-interlacing dataset includes multiple groups of training pairs composed of interlaced images and intact de-interlaced images.
[0081] In one embodiment, the video restoration processing method further includes a step of obtaining an anti-interlaced data set. The computer device can determine an image data set, split the odd-field data and the even-field data from two consecutive frames in the image data set, and respectively take the odd-field data and the even-field data corresponding to different images to form an interlaced image. To avoid the anti-interlaced model being unable to adapt to the aliasing problem, the computer device can use the nearest neighbor downsampling method to downsample and then restore the above interlaced image to simulate the aliasing effect, and finally normalize the anti-interlaced data set. For example, internally, 90,000 segments of standard-definition image data sets each containing 7-frame sequences can be collected. The computer device can generate an interlaced image based on the above standard-definition image data set, and downsample the produced interlaced image by a factor of 2 to 4 using the nearest neighbor downsampling method and then restore it.
[0082] In one embodiment, during the training process of the anti-interlaced model, the computer device can split the odd-field data and the even-field data from the images in the anti-interlaced data set through an interlaced split layer, and then divide them into two channels. One channel is used to learn residual features. The computer device can merge the odd-field data and the even-field data, and then learn the residual features in the odd-field data and the even-field data through an attentive feature block2 to obtain odd-field residual data and even-field residual data. It can be understood that the dimensions of the odd-field residual data and the even-field residual data at this time are equivalent to half of the complete image data dimension of an image. The computer device can amplify the odd-field residual data and the even-field residual data through a vertical pixel shuffle layer in the anti-interlaced model. The other channel is used to retain the odd-field data and the even-field data. The computer device can amplify the odd-field data and the even-field data through a vertical pixel shuffle layer in the anti-interlaced model. Then, the computer device can stack the amplified odd-field data and the amplified odd-field residual data through an Add layer, and obtain the anti-interlaced odd-field image through an odd_output layer. The computer device can stack the amplified even-field data and the amplified even-field residual data through an Add layer, and obtain the anti-interlaced even-field image through an even_output layer.
[0083] In one embodiment, as Figure 4AThe following provides a schematic diagram of the deinterlacing model. The deinterlacing model includes an input layer, an interlaced segmentation layer, an attention feature extraction layer, a feature amplification layer, a superposition layer, an odd-field output layer, and an even-field output layer. Among them, the odd-field output layer and the even-field output layer respectively include a 1*1 convolutional layer. It can be understood that the 1*1 convolutional layer is used to fuse the amplified odd-field data and the amplified odd-field residual data after superposition, and to fuse the amplified even-field data and the amplified even-field residual data after superposition. The feature amplification layer includes a 3*3 convolutional layer, which is used to amplify the odd-field data, even-field data, odd-field residual data, and even-field residual data to match the dimension of a complete frame of image. The attention feature extraction layer includes a 3*3 convolutional layer and multiple attention feature units. For example, the attention feature extraction layer may include 1 3*3 convolutional layer and 9 attention feature units.
[0084] In one embodiment, as Figure 4B The following provides a schematic diagram of the attention feature unit in the attention feature extraction layer. The attention feature unit has two branches. For the first branch, the residual feature data output by the previous i-1 attention feature units will be merged through the first-branch input layer and input into the i-th attention feature unit. For the second branch, the residual feature data output by the i-1-th attention feature unit is input into the current i-th attention feature unit through the second-branch input layer. The output of the i-th attention feature unit provides input for the subsequent i+1-th attention feature unit. The first parameter (λ0), the second parameter (λ1), and the third parameter (λ2) are all learnable parameters, which assign different attention weights to the input residual feature data for residual feature extraction. It can be understood that the attention weight is used to indicate the importance degree of the residual feature data.
[0085] The feature mapping layer of the first branch performs projection mapping on the residual feature data output by the i-th to i-1-th attention feature units. When performing feature fusion, the projection mappings of different channels have different importance degrees. Therefore, the subsequent exponential activation layer is used to learn the importance degrees of the projection mappings of different channels. Among them, the feature mapping layer is a 1*1 convolutional layer. The activation function of the exponential activation layer is the sigmoid function, which is a 1*1 convolutional layer. The first convolutional layer, the piecewise linear activation layer, and the second convolutional layer of the second branch learn the residual feature data output by the previous i-1 attention feature unit. Among them, the first convolutional layer and the second convolutional layer are 3*3 convolutional layers. The activation function of the piecewise linear activation layer is the rectified linear unit (relu).
[0086] In one embodiment, as Figure 4CThe principle schematic diagram of the feature amplification layer of the deinterlacing model is provided as shown. The feature amplification layer essentially enlarges the field data dimension to match that of a complete frame of image based on the vertical pixel feature data corresponding to the field data, obtaining the amplified field data. The feature amplification layer includes a 3×3 convolutional layer and a vertical pixel combination layer. The server can input the odd-field data and the even-field data into the 3×3 convolutional layer for joint feature extraction, obtaining the odd-field vertical pixel feature data corresponding to the odd-field data and the even-field vertical pixel feature data corresponding to the even-field data. The server can use the odd-field vertical pixel feature data to amplify the odd-field data through the vertical pixel combination layer, and use the even-field vertical pixel feature data to amplify the even-field data through the vertical pixel combination layer.
[0087] In this embodiment, residual feature extraction is performed through the deinterlacing model, and the extracted odd-field residual data and even-field residual data are respectively superimposed on the odd-field data and the even-field data to obtain the deinterlaced odd-field image and even-field image. Compared with the second image obtained by directly replacing the field data, the target image can further avoid the influence of interlaced field patterns, thereby further improving the quality of video restoration.
[0088] In one embodiment, there are multiple target images obtained by deinterlacing and restoring the original video; the method further includes: segmenting the multiple target images to obtain interlaced video segments; the interlaced video segments have corresponding repeated frame numbers; for the multiple target images in the interlaced video segments, determining the similarity between adjacent images among the multiple target images; according to the similarity, determining the repeated frame images that meet the repeated frame numbers from the interlaced video segments; extracting the repeated frame images from the interlaced video segments; the remaining target images after extracting the repeated frame images are used as video frames to form the restored target video.
[0089] Among them, the interlaced video segment corresponds to the video segment with interlaced images in the original video. The repeated frame number is the number of frames with repeated images in the interlaced video segment. For example, if the original video is a video that is converted from 24fps (frames per second, Frames Per Second) to 30fps format after film rewind, then 2 out of every 5 frames will be interlaced. Therefore, after deinterlacing, the target images corresponding to two original images will be repeated, that is, there will be two repeated target images. At this time, there are 5 target images in an interlaced video segment, and the repeated frame number is 1.
[0090] Specifically, the server can segment multiple target images according to the film threading method corresponding to the original video and the timing of the original image corresponding to the target image in the original video, so as to obtain interleaved video segments. The server can determine the similarity between adjacent images among the multiple target images in the interleaved video segments. It can be understood that the timing of the target images after de-interlacing repair is consistent with that of the original images. Each target image in the interleaved video segments has a corresponding timing. Adjacent images are target images adjacent in timing, that is, the current target image, the target image corresponding to the previous timing, and the target image corresponding to the next timing are adjacent. The server can determine the similarity corresponding to each target image and, based on the maximum similarity value corresponding to each target image, determine a group of repeated frame images that meet the number of repeated frames from the interleaved video segments, and determine one repeated frame image from each determined group of repeated frame images. It can be understood that the greater the similarity corresponding to the target image, the more similar the target image is to the adjacent image, the higher the degree of repetition, and the greater the probability that the target image is a repeated frame image. Among them, the group of repeated frame images includes two adjacent target images with the same maximum similarity value. The server can extract the repeated frame images from the interleaved video segments.
[0091] In one embodiment, the server can segment multiple target images according to the film threading method of the original video to obtain interleaved video segments and determine the number of repeated frames in the interleaved video segments. Suppose each interleaved video segment has n frame images, and the number of repeated frames to be extracted is m. The server can calculate the similarity between each of the n target images in each interleaved video segment and its adjacent images before and after. Suppose the similarities are k(n, -) and k(n, +) respectively. The server can determine the maximum similarity value from the two similarities, that is, k(n) = max(k(n, -), k(n, +)), and use the maximum similarity value to represent the similarity score of the nth frame. The server can determine the m groups of adjacent target images with the highest similarity scores among the n frame images in each interleaved video segment, and extract one target image from each group of adjacent target images respectively. It can be understood that the extracted target images are the repeated frame images.
[0092] In this embodiment, for multiple target images in the interleaved video segments, the similarity between adjacent images among the multiple target images is determined; according to the similarity, repeated frame images that meet the number of repeated frames are determined from the interleaved video segments; the repeated frame images are extracted from the interleaved video segments to obtain the target video after frame extraction processing. Without manual intervention, by calculating the similarity between target pictures, the frame extraction processing of repeated images in the target video can be completed, improving the efficiency of video repair.
[0093] In one embodiment, the method further includes: extracting texture feature information from the target image; and performing super-resolution enhancement on the target image based on the texture feature information to obtain the super-resolution enhanced target image.
[0094] Specifically, the server can input the target image into the super-resolution enhancement model, extract the texture feature information from the target image through the texture feature extraction layer in the super-resolution enhancement model, and then fuse the texture feature information through the feature fusion layer of the super-resolution enhancement model to obtain the fused texture feature information. It can be understood that the pixels in the image are highly correlated with the neighboring pixels. By fusing the existing texture feature information in the image, the new texture feature information obtained can well supplement the information lacking in the texture details of the low-resolution image. The server can use the fused texture feature information to perform super-resolution enhancement on the target image to obtain the super-resolution enhanced target image output by the super-resolution enhancement model. It can be understood that the super-resolution enhanced target image has a higher resolution than the non-super-resolution enhanced target image.
[0095] In one embodiment, the server can perform super-resolution enhancement on the target image in the target video after frame extraction. It can be understood that the super-resolution enhanced target image is used as a video frame to constitute the finally restored target video.
[0096] In one embodiment, such as Figure 6The following provides a schematic diagram of the super-resolution enhancement model. The super-resolution enhancement model includes a texture feature extraction layer (feature extractor), a feature fusion layer (feature fuse), a residual network layer (resblock), a first upsampling convolutional layer (up_conv_1), a second upsampling convolutional layer (up_conv_x2), and a super-resolution enhancement output layer (out_conv). The server can input the target image into the super-resolution enhancement model, extract the texture feature information in the target image through the texture feature extraction layer, then fuse the texture features of multiple channels included in the texture feature information through the feature fusion layer to obtain the fused texture feature information, then perform further feature extraction on the fused texture feature information through the residual network layer, then double the texture feature information extracted by the further feature extraction through the first upsampling convolutional layer, and then further double it through the second upsampling convolutional layer, and finally output the super-resolution enhanced target image through the super-resolution enhancement output layer. It can be understood that when the resolution of the target image is 128*128, the texture feature extraction layer outputs texture feature information of 32 channels and a size of 64*64, the feature fusion layer outputs fused texture feature information of 32 channels and a size of 64*64, the residual network layer outputs texture feature information of 32 channels and a size of 64*64 extracted by the further feature extraction, the first upsampling convolutional layer outputs doubled texture feature information of 8 channels and a size of 128*128, the second upsampling convolutional layer outputs further doubled texture feature information of 8 channels and a size of 256*256, and the super-resolution enhancement output layer outputs the super-resolution enhanced target image with a resolution of 256*256.
[0097] In one embodiment, the video restoration processing method further includes a training step for the super-resolution enhancement model. The super-resolution enhancement model can be a deep convolutional model. The computer device can preprocess the low-resolution image dataset and use the preprocessed low-resolution image dataset to train the deep convolutional model, and obtain the model file of the super-resolution enhancement model after the training is completed.
[0098] In one embodiment, in order to make the super-resolution result have more details and sharpness, during the iterative training of the super-resolution enhancement model, only when the preset loss function combination converges can the iteration be stopped. The loss function combination can include at least two of the absolute value loss function (L1), the original loss function (gan), and the perceptual loss function (perception).
[0099] In one embodiment, the super-resolution enhancement model can include multiple fully connected network layers.
[0100] In one embodiment, the video restoration processing method further includes a step of obtaining a low-resolution image data set. In reality, there are many cases where images are damaged. A single degradation method cannot meet the requirement of restoring high definition of images and is also likely to magnify defects. A computer device can simulate a large number of degradation methods, using high-definition images as samples to simulate and produce low-resolution images. The degradation methods may include one or more of different levels of Gaussian blur, additive noise, compression, nearest neighbor downsampling, and bilinear downsampling, etc. It can be understood that when the computer device uses multiple degradation methods to produce low-resolution images, the order of using the degradation methods may not be fixed. After normalizing the produced low-resolution images, the computer device obtains a low-resolution image data set.
[0101] In this embodiment, by extracting texture feature information from the target image; and performing super-resolution enhancement on the target image based on the texture feature information to obtain the super-resolution enhanced target image, the quality of video restoration can be improved.
[0102] In one embodiment, as Figure 7 shown, a simple flowchart of the video restoration processing method is provided. The server can perform video decoding on the original video to obtain the original image. The server can perform deinterlacing restoration on the original image to obtain the target image after deinterlacing restoration. It can be understood that the original video may be a video that has undergone film threading processing. For the target image obtained by performing deinterlacing restoration on each original image in the original video, there will inevitably be duplicates. After obtaining the target image after deinterlacing restoration, the server can determine the duplicate frame images. After extracting the duplicate frame images, the target image after frame extraction processing is obtained. It can be understood that the original video has a relatively large time span compared to the present, and the resolution of the original images therein is low. Therefore, the server can perform super-resolution enhancement on the target images in the target video after frame extraction processing to obtain the super-resolution enhanced target images, so that the super-resolution enhanced target images can be used as video frames for video encoding to obtain the restored target video.
[0103] The server can send the target video to the terminal for manual acceptance of the target video. During the acceptance process, the human can screen out the target video segments that do not meet the acceptance criteria from the target video. After manually repairing the video segment, the manually repaired target video segment is sent to the server through the terminal. The server can perform frame extraction processing on the manually repaired target video segment. After frame extraction processing, super-resolution enhancement is performed on the target images in the manually repaired target video segment. Finally, the manually repaired target video segment is used to replace the original target video segment for video encoding to obtain the target video. When the target video meets the acceptance criteria, the entire video restoration processing flow ends, and the restored target video is obtained.
[0104] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0105] Based on the same inventive concept, an embodiment of the present application further provides a video restoration processing device for implementing the above-mentioned video restoration processing method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the video restoration processing device provided below can refer to the limitations on the video restoration processing method in the above text, and will not be repeated here.
[0106] In one embodiment, as Figure 8 shown, a video restoration processing device 800 is provided, including: an anti-interlacing module 802 and a restoration module 804, where:
[0107] The anti-interlacing module 802 is configured to perform field replacement processing on the original images in the original video to obtain a plurality of first images; perform interlace detection on the first images to obtain an interlace detection result; the interlace detection result is used to reflect the number of interlace field patterns in the first images; screen the plurality of first images according to the interlace detection result to obtain second images;
[0108] The restoration module 804 is configured to extract residual feature data corresponding to the field data in the second images based on the second images; superimpose the residual feature data on the corresponding field data in the second images to obtain a target image after anti-interlacing restoration; the target image is used as a video frame to form a restored target video.
[0109] In one of the embodiments, the anti-interlacing module 802 is configured to determine an image group based on the timing corresponding to the original images in the original video; the image group includes a plurality of original images with consecutive timings; respectively use the even-field data in the plurality of original images in the image group to replace the even-field data in the median original image, and retain the odd-field data in the median original image, to obtain a plurality of first images corresponding to the image group; the median original image is the original image corresponding to the median timing in the image group.
[0110] In one embodiment, the interleaving detection result includes an interleaving score; the deinterleaving module 802 is configured to input a first image into an interleaving detection model, and detect the interleaved field pattern in the first image through the interleaving detection model to predict the interleaving score corresponding to the first image.
[0111] In one embodiment, the restoration module 804 is configured to extract residual features from the even-field data and odd-field data in a second image to obtain even-field residual data corresponding to the even-field data in the second image and odd-field residual data corresponding to the odd-field data in the second image; and is configured to obtain a deinterleaved even-field image by superimposing the even-field residual data on the even-field data in the second image; obtain a deinterleaved odd-field image by superimposing the odd-field residual data on the odd-field data in the second image; and screen out a target image from the even-field image and the odd-field image.
[0112] In one embodiment, the restoration module 804 is configured to input the second image into an interleaving segmentation layer in a deinterleaving model to determine odd-field data and even-field data from the second image through the interleaving segmentation layer; the deinterleaving model further includes an attention feature extraction layer; the odd-field data and even-field data of the second image are combined and input into the attention feature extraction layer to perform residual feature extraction, so as to obtain even-field residual data corresponding to the even-field data in the second image and odd-field residual data corresponding to the odd-field data in the second image.
[0113] In one embodiment, there are multiple target images obtained by deinterleaving and restoring the original video; the restoration module 804 is configured to segment the multiple target images to obtain interleaved video segments; the interleaved video segments have corresponding repeated frame numbers; for the multiple target images in the interleaved video segments, determine the similarity between adjacent images among the multiple target images; according to the similarity, determine repeated frame images that meet the repeated frame numbers from the interleaved video segments; extract the repeated frame images from the interleaved video segments; the remaining target images after extracting the repeated frame images are used as video frames to form a restored target video.
[0114] In one embodiment, the restoration module 804 is configured to extract texture feature information from a target image; and perform super-resolution enhancement on the target image based on the texture feature information to obtain a super-resolution enhanced target image.
[0115] Each module in the above video restoration processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of it, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0116] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in Figure 9 . The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the model file of the deinterlacing model. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements a video restoration processing method.
[0117] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in Figure 10 . The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a video restoration processing method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad set on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0118] Those skilled in the art can understand that Figure 9 and Figure 10The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0119] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0120] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0121] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0122] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0123] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0124] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0125] The above-described embodiments only represent several implementation manners of the present application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A video restoration processing method, characterized in that, The method includes: Performing field replacement processing on the original images in the original video to obtain a plurality of first images; The performing field replacement processing on the original images in the original video to obtain a plurality of first images includes: Determining the field data in the original image, and using the field data of other original images in the original video except the original image to perform field replacement processing on the field data in the original image to obtain a plurality of the first images; Performing interlace detection on the first images to obtain an interlace detection result; the interlace detection result is used to reflect the number of interlace field patterns in the first images; Screening the plurality of first images according to the interlace detection result to obtain second images; Performing residual feature extraction based on the second images to obtain residual feature data corresponding to the field data in the second images; Superimposing the residual feature data on the corresponding field data in the second images to obtain a de-interlaced and repaired target image; the target image is used as a video frame to form a repaired target video.
2. The method according to claim 1, wherein The performing field replacement processing on the original images in the original video to obtain a plurality of first images includes: Determining an image group based on the time sequence corresponding to the original images in the original video; the image group includes a plurality of original images with consecutive time sequences; Respectively using the even-field data in the plurality of original images in the image group to replace the even-field data in the median original image, and retaining the odd-field data in the median original image to obtain a plurality of first images corresponding to the image group; the median original image is the original image with the median time sequence in the image group.
3. The method according to claim 1, characterized in that, The interlace detection result includes an interlace score; the performing interlace detection on the first images to obtain an interlace detection result includes: Inputting the first image into an interlace detection model, and detecting the interlace field patterns in the first image through the interlace detection model to predict the interlace score corresponding to the first image.
4. The method according to claim 1, wherein The performing residual feature extraction based on the second images to obtain residual feature data corresponding to the field data in the second images includes: Performing residual feature extraction on the even-field data and odd-field data in the second images to obtain even-field residual data corresponding to the even-field data in the second images and odd-field residual data corresponding to the odd-field data in the second images; The superimposing the residual feature data on the corresponding field data in the second images to obtain a de-interlaced and repaired target image includes: Obtaining a de-interlaced even-field image by superimposing the even-field residual data on the even-field data in the second image; Obtaining a de-interlaced odd-field image by superimposing the odd-field residual data on the odd-field data in the second image; Selecting a target image from the even-field image and the odd-field image.
5. The method according to claim 4, wherein The performing residual feature extraction on the even-field data and odd-field data in the second images to obtain even-field residual data corresponding to the even-field data in the second images and odd-field residual data corresponding to the odd-field data in the second images includes: Input the second image into the interlaced segmentation layer in the de-interlacing model to determine odd-field data and even-field data from the second image through the interlaced segmentation layer; the de-interlacing model further includes an attention feature extraction layer; Input the combined odd-field data and even-field data of the second image into the attention feature extraction layer to perform residual feature extraction, obtaining even-field residual data corresponding to the even-field data of the second image and odd-field residual data corresponding to the odd-field data of the second image.
6. The method according to claim 1, wherein There are multiple target images obtained after de-interlacing repair for the original video processing; the method further includes: Segment the multiple target images to obtain interlaced video segments; the interlaced video segments have corresponding repeated frame numbers; For the multiple target images in the interlaced video segment, determine the similarity between adjacent images among the multiple target images; According to the similarity, determine repeated frame images that meet the repeated frame number from the interlaced video segment; Extract the repeated frame images from the interlaced video segment; the remaining target images after extracting the repeated frame images are used as video frames to form the repaired target video.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: Extract texture feature information from the target image; Based on the texture feature information, perform super-resolution enhancement on the target image to obtain a super-resolution enhanced target image.
8. A video restoration processing device, characterized in that, The device includes: A de-interlacing module, configured to perform field replacement processing on the original images in the original video to obtain multiple first images; perform interlacing detection on the first images to obtain an interlacing detection result; the interlacing detection result is used to reflect the number of interlaced field patterns of the first images; screen the multiple first images according to the interlacing detection result to obtain second images; The performing field replacement processing on the original images in the original video to obtain multiple first images includes: Determine the field data in the original image, and use the field data of other original images in the original video except the original image to perform field replacement processing on the field data in the original image to obtain multiple of the first images; A repair module, configured to perform residual feature extraction based on the second image to obtain residual feature data corresponding to the field data in the second image; superimpose the residual feature data on the corresponding field data in the second image to obtain a de-interlaced repaired target image; the target image is used as a video frame to form the repaired target video.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video image de-interlacing method and device, electronic equipment and storage medium
CN112218081A
Video image coding method and device, electronic equipment and storage medium
CN112218096A