Video processing method, device, electronic device and computer-readable storage medium
The time and spatial characteristics of video images are processed through the time domain branches and airspace branches of the target network model, and the problem of inconsistent image timing in video repair is solved and the quality of video repair is improved.
Patent Information
- Application Number
- CN202210971512.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-08-12
AI Technical Summary
The prior art fails to effectively process the changes in time and spatial information between images during video repair, resulting in poor repair quality.
The time-domain branches and airspace branches of the target network model process the temporal correlation characteristics between images and the spatial characteristics of each frame of images respectively, and the transformer-perceptual repair network model is used for feature fusion and repair.
The quality of video repair is improved, and the problem of poor video repair quality caused by the lack of refinement of temporal and spatial information characteristics is solved, and the accumulation of deviations during frame-by-frame repair is avoided.
Smart Images

Figure CN115239598B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a video processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Before the digital age, the carriers of video films were basically films. Due to technical and material limitations, the preservation of films had a time limit. Among them, these video films not only carried the memories of many people, but also contained precious historical records. Although manual restoration could achieve relatively good results, it was time-consuming and laborious, and the promotion effect was extremely poor. Therefore, in recent years, the problem of video restoration has attracted the attention of the industry.
[0003] In the prior art, the method of video restoration is usually to extract the optical flow information between video frame sequences and the mask information of the missing area, so as to obtain a relatively rough optical flow estimate, and use a cascaded optimization strategy to repair the optical flow, and then complete each video frame frame by frame in a way guided by the optical flow, so as to repair the video. However, for video images, when the image scene changes, the temporal information and spatial information of the image will change. In the prior art, when repairing multiple frames of images of a video, the change of temporal information and spatial information between different images is not considered. Therefore, the quality of the finally restored video is poor due to the temporal inconsistency between multiple frames of images.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a video processing method, apparatus, electronic device, and computer-readable storage medium, so as to at least solve the technical problem of poor video restoration quality caused by the lack of feature refinement of temporal information and spatial information when the prior art repairs a video.
[0006] According to one aspect of the embodiments of the present application, a video processing method is provided, including: obtaining a plurality of image sets from a video to be processed, where each image set includes K to-be-repaired images and K edge images, and each to-be-repaired image corresponds to one edge image; determining a first input eigenvalue and a second input eigenvalue according to the image sets, where the first input eigenvalue is an eigenvalue obtained by fusing K first eigenvalues, the second input eigenvalue is an eigenvalue obtained by fusing one first eigenvalue and one second eigenvalue, the first eigenvalue is an eigenvalue obtained by performing feature extraction on the to-be-repaired image, and the second eigenvalue is an eigenvalue obtained by performing feature extraction on the edge image; inputting the first input eigenvalue into the temporal branch of the target network model to obtain a first target eigenvalue, where the first target eigenvalue is used to represent the temporal correlation feature between the K to-be-repaired images in the image set; inputting the second input eigenvalue into the spatial branch of the target network model to obtain a second target eigenvalue, where the second target eigenvalue is used to represent the spatial feature of each to-be-repaired image in the image set; and repairing each to-be-repaired image in the video to be processed according to the first target eigenvalue and the second target eigenvalue to obtain a repaired video.
[0007] Further, the video processing method further includes: obtaining a video to be processed, where the video to be processed includes N to-be-repaired images, and the N to-be-repaired images include K to-be-repaired images; preprocessing each to-be-repaired image in the video to be processed to obtain N first images, where each first image corresponds to one to-be-repaired image; performing edge extraction processing on each first image to obtain N edge images, where each edge image corresponds to one first image; and determining a plurality of image sets according to the N edge images and the N to-be-repaired images.
[0008] Further, the target network model is trained and generated by the following method: obtaining training data, where the training data at least includes M historical to-be-repaired images and M historical target images, and the historical target images are images obtained by repairing the historical to-be-repaired images; preprocessing and performing edge extraction processing on the M historical to-be-repaired images to obtain M historical edge images, where each historical edge image corresponds to one historical to-be-repaired image; determining a plurality of historical image sets according to the M historical to-be-repaired images and the M historical edge images, where each historical image set includes K historical to-be-repaired images and K historical edge images; and training the target network model according to the plurality of historical image sets.
[0009] Further, the training process of the target network model further includes: extracting features from K frames of historical images to be repaired in each historical image set to obtain first feature values corresponding to each frame of historical image to be repaired; extracting features from K frames of historical edge images in each historical image set to obtain second feature values corresponding to each frame of historical edge image; performing feature fusion on the first feature values corresponding to the K frames of historical images to be repaired to obtain first training feature values; performing feature fusion on the second feature values corresponding to the K frames of historical edge images to obtain second training feature values; training the target network model according to the first training feature values, the second training feature values, and M frames of historical target images.
[0010] Further, the training process of the target network model further includes: inputting the first training feature values and the second training feature values into a neural network to obtain an initial network model; determining a target loss function according to the M frames of historical target images and the M frames of historical images to be repaired, where the target loss function includes a perceptual loss function, an adversarial loss function, and a reconstruction loss function; determining a loss value corresponding to the initial network model according to the target loss function; adjusting the initial network model according to the loss value to obtain the target network model.
[0011] Further, the video processing method further includes: performing a summation process on the first target feature values and the second target feature values to obtain third target feature values corresponding to each frame of image to be repaired in the image set; superimposing the third target feature values on each frame of image to be repaired in the image set to obtain target images, where each image set corresponds to K frames of target images; determining a repaired video according to the target images.
[0012] Further, the video processing method further includes: obtaining all target images corresponding to multiple image sets; forming a target image sequence from all the target images according to the third target feature values corresponding to each target image; converting the target image sequence into a repaired video according to a preset video format and storing the repaired video in a preset area.
[0013] According to another aspect of the embodiments of the present application, there is also provided a video processing device, including: an acquisition module, configured to acquire a plurality of image sets from a video to be processed, where each image set includes K to-be-repaired images and K edge images, and each to-be-repaired image corresponds to one edge image; a determination module, configured to determine a first input eigenvalue and a second input eigenvalue according to the image set, where the first input eigenvalue is an eigenvalue obtained by performing feature fusion on K first eigenvalues, the second input eigenvalue is an eigenvalue obtained by performing feature fusion on one first eigenvalue and one second eigenvalue, the first eigenvalue is an eigenvalue obtained by performing feature extraction on the to-be-repaired image, and the second eigenvalue is an eigenvalue obtained by performing feature extraction on the edge image; a first input module, configured to input the first input eigenvalue into the temporal branch of the target network model to obtain a first target eigenvalue, where the first target eigenvalue is used to characterize the temporal correlation feature between the K to-be-repaired images in the image set; a second input module, configured to input the second input eigenvalue into the spatial branch of the target network model to obtain a second target eigenvalue, where the second target eigenvalue is used to characterize the spatial feature of each to-be-repaired image in the image set; a repair module, configured to repair each to-be-repaired image in the video to be processed according to the first target eigenvalue and the second target eigenvalue to obtain a repaired video.
[0014] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, in which a computer program is stored, where the computer program is configured to execute the above-mentioned video processing method when running.
[0015] According to another aspect of the embodiments of the present application, there is also provided an electronic device, which includes one or more processors; a memory, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement a program for running, where the program is configured to execute the above-mentioned video processing method when running.
[0016] In the present application, a method of determining a first target feature value through a time domain branch of a target network model and a second target feature value through a spatial domain branch of the target network model is adopted. First, multiple image sets are obtained from the video to be processed, and then a first input feature value and a second input feature value are determined according to the image sets. The first input feature value is input into the time domain branch of the target network model to obtain the first target feature value, and the second input feature value is input into the spatial domain branch of the target network model to obtain the second target feature value. Finally, each frame of the image to be repaired in the video to be processed is repaired according to the first target feature value and the second target feature value to obtain a repaired video. Among them, each image set contains K frames of images to be repaired and K frames of edge images, and each frame of the image to be repaired corresponds to a frame of edge image; the first input eigenvalue is a eigenvalue obtained by feature fusion of K first eigenvalues, the second input eigenvalue is a eigenvalue obtained by feature fusion of a first eigenvalue and a second eigenvalue, the first eigenvalue is a eigenvalue obtained by feature extraction of the image to be repaired, and the second eigenvalue is a eigenvalue obtained by feature extraction of the edge image; the first target eigenvalue is used to characterize the temporal correlation characteristics between the K frames of images to be repaired in the image set; wherein the second target eigenvalue is used to characterize the spatial characteristics of each frame of the image to be repaired in the image set.
[0017] From the above content, it can be seen that the present application uses the time domain branch of the target network model to determine the time correlation features between multiple frames of images to be repaired, and uses the spatial domain branch to determine the spatial features of each frame of the image to be repaired, so that the time features and spatial features are also used as supplementary features of the repaired image. Compared with the prior art, the present application overcomes the problem of image time sequence inconsistency when the DFC-Net network repairs the image, achieves the purpose of improving the repair quality of each frame of the image to be repaired, and then realizes the effect of the overall video repair quality.
[0018] It can be seen that through the technical solution of the present application, the purpose of establishing a global temporal dependency relationship between multiple frame images in a video is achieved, thereby achieving the effect of improving the quality of video image restoration, and further solving the technical problem of poor video restoration quality caused by the lack of feature refinement of temporal information and spatial information when repairing the video in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 is a flowchart of an optional video processing method according to an embodiment of the present application;
[0021] Figure 2 It is a schematic diagram of an optional video restoration system according to an embodiment of the present application;
[0022] Figure 3 It is a flowchart of an optional method for obtaining multiple image sets according to an embodiment of the present application;
[0023] Figure 4 It is a flowchart of an optional feature processing according to an embodiment of the present application;
[0024] Figure 5 It is a flowchart of a model training according to an embodiment of the present application;
[0025] Figure 6 It is a flowchart of an optional image restoration according to an embodiment of the present application;
[0026] Figure 7 It is a flowchart of an optional method for generating video data according to an embodiment of the present application;
[0027] Figure 8 It is a schematic diagram of an optional video processing device according to an embodiment of the present application;
[0028] Figure 9 It is a schematic diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0029] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0031] Embodiment 1
[0032] According to an embodiment of the present application, an embodiment of a video processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0033] Figure 1 is a flowchart of an optional video processing method according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:
[0034] Step S101, obtain a plurality of image sets from the video to be processed.
[0035] In step S101, each image set contains K frames of images to be repaired and K frames of edge images, and each frame of image to be repaired corresponds to one frame of edge image.
[0036] Specifically, Figure 2 shows a schematic diagram of an optional video repair system according to an embodiment of the present application. As Figure 2 shown, the video repair system includes a negative film scanning module, a video sequence repair module, and a video assembly and storage module. Among them, the negative film scanning module is used to obtain a plurality of image sets from the video to be processed.
[0037] Specifically, the video to be processed consists of N frames of images to be repaired, where N is greater than K. Figure 3 shows a flowchart of an optional method for obtaining a plurality of image sets according to an embodiment of the present application. As Figure 3 shown, assuming that the video to be processed is a video in the form of a film, by scanning the film content with a negative film scanner, a video to be processed consisting of N frames of images to be repaired can be obtained. Subsequently, the negative film scanner will send the video to be processed to the negative film scanning module, and the negative film scanning module will arrange the N frames of images to be repaired in sequence to obtain a complete image sequence.
[0038] Furthermore, the negative film scanning module will perform preprocessing and edge extraction processing on the N frames of images to be repaired in the image sequence to obtain N frames of edge images. Among them, the edge images and the images to be repaired are in a one-to-one correspondence relationship, that is, each frame of image to be repaired will correspond to one frame of edge image.
[0039] Finally, in order to improve the efficiency of video repair, the negative film scanning module will divide a plurality of image sets according to the N frames of images to be repaired and the N frames of edge images. Among them, each image set contains K frames of images to be repaired and K frames of edge images, and the images to be repaired and the edge images in the image set are also in a one-to-one correspondence.
[0040] Step S102: Determine the first input eigenvalue and the second input eigenvalue according to the image set.
[0041] In step S102, the first input eigenvalue is the eigenvalue obtained by performing feature fusion on K first eigenvalues, the second input eigenvalue is the eigenvalue obtained by performing feature fusion on one first eigenvalue and one second eigenvalue, the first eigenvalue is the eigenvalue obtained by performing feature extraction on the image to be repaired, and the second eigenvalue is the eigenvalue obtained by performing feature extraction on the edge image.
[0042] In an alternative embodiment, Figure 4 shows a flowchart of an alternative feature processing according to an embodiment of the present application. As Figure 4 shown, after the negative film scanning module obtains a plurality of image sets, it will send the plurality of image sets to the video sequence repair module. The video sequence repair module will perform a series of feature processing on the images in the plurality of image sets simultaneously in a parallel manner. Specifically, taking the image set A as an example, first, the video sequence repair module will perform feature extraction on the K frames of images to be repaired in the image set A to obtain the first eigenvalue corresponding to each frame of the image to be repaired. At the same time, the video sequence repair module will also perform feature extraction on the K frames of edge images in the image set A to obtain the second eigenvalue corresponding to each frame of the edge image.
[0043] Further, the video sequence repair module performs feature fusion on the first eigenvalues corresponding to the K frames of images to be repaired to obtain the first input eigenvalue, and performs feature fusion on the second eigenvalues corresponding to the K frames of edge images to obtain the second input eigenvalue.
[0044] Step S103: Input the first input eigenvalue into the time domain branch of the target network model to obtain the first target eigenvalue.
[0045] In step S103, the first target eigenvalue is used to characterize the time correlation feature between the K frames of images to be repaired in the image set.
[0046] Step S104: Input the second input eigenvalue into the spatial domain branch of the target network model to obtain the second target eigenvalue.
[0047] In step S104, the second target eigenvalue is used to characterize the spatial feature of each frame of the image to be repaired in the image set.
[0048] Optionally, the above target network model is a transformer perceptual repair network model, which is a model that uses an attention mechanism for deep learning and includes a multi-head attention layer, a normalization layer, and a fully connected layer.
[0049] AsFigure 4 As shown in the figure, the video sequence repair module uses the target network model to optimize the features of the first input eigenvalue and the second input eigenvalue. Specifically, the target network model includes a temporal branch and a spatial branch. Among them, the temporal branch is used to extract the temporal feature information between multiple frames of images in the video, and the spatial branch is used to extract the spatial feature information of each frame of image.
[0050] Optionally, assuming that the value of K is 7, the first input eigenvalue corresponding to the above-mentioned image set A can be used as a group of inputs, and the second input eigenvalue can be used as 7 groups of inputs. The video sequence repair module will input the first input eigenvalue into a temporal branch in the target network model. Then, the temporal branch will construct the association information between 7 frames of images to be repaired based on the first input eigenvalue, and further determine the temporal global dependence relationship between the 7 frames of images to be repaired. In addition, the video sequence repair module will also input 7 groups of second input eigenvalues into 7 spatial branches in the target network model. Then, each spatial branch will determine the spatial feature information of the corresponding image to be repaired according to the input second input eigenvalue, and further determine the spatial global dependence relationship of each frame of image to be repaired.
[0051] It should be noted that the above-mentioned spatial branch can be understood as a transformer spatial network with an encoder-decoder structure, and the above-mentioned temporal branch can be understood as a transformer temporal network with an encoder-decoder structure.
[0052] Step S105: Repair each frame of the image to be repaired in the video to be processed according to the first target eigenvalue and the second target eigenvalue, and obtain the repaired video.
[0053] In an alternative embodiment, as Figure 4 shown, still taking the value of K as 7 as an example, the video sequence repair module will add the first target eigenvalue output by the above-mentioned temporal branch to the 7 second target eigenvalues output by the 7 spatial branches respectively, so as to obtain 7 third target eigenvalues. Finally, the video sequence repair module superimposes the 7 third target eigenvalues on the corresponding images to be repaired in the image set A respectively, and the repaired images can be obtained. Since multiple image sets are processed in parallel, the image repair process of other image sets is the same as that of the image set A, and the present application will not elaborate here.
[0054] In addition, after repairing the images to be repaired in all image sets, the video sequence repair module will send all the repaired images to the video assembly and storage module. The video assembly and storage module will reassemble all the repaired images according to the order of the images, so as to obtain the repaired video.
[0055] In an alternative embodiment, to improve the video restoration efficiency, the video restoration system disassembles the video to be processed into multiple image sets, so as to perform parallel processing on the multiple image sets, which can not only improve the video restoration efficiency, but also avoid the problem of deviation accumulation that is prone to occur when repairing images frame by frame. Specifically, the negative film scanning module first obtains the video to be processed, then preprocesses each frame of the image to be restored in the video to be processed to obtain N frames of first images, performs edge extraction processing on each frame of the first images to obtain N frames of edge images, and finally the negative film scanning module determines multiple image sets according to the N frames of edge images and the N frames of images to be restored. Among them, the video to be processed contains N frames of images to be restored, and the N frames of images to be restored contain K frames of images to be restored; each frame of the first image corresponds to one frame of the image to be restored; each frame of the edge image corresponds to one frame of the first image.
[0056] Optionally, assuming N is 100 and K is 7, it means that the video to be processed is a video composed of 100 frames of images to be restored. The negative film scanning module will uniformly preprocess the 100 frames of images to be restored to obtain the corresponding 100 frames of first images. Among them, the preprocessing includes data cleaning processes such as removing image burrs and removing image background noise. In addition, the negative film scanning module will also perform edge extraction processing on the 100 frames of first images to obtain the corresponding 100 frames of edge images. Finally, the negative film scanning module divides the 100 frames of images to be restored and the 100 frames of edge images into 14 image sets. Among them, according to the principle that if the number of frames in the last image set is less than 7, it will be filled in the previous image set, each of the first 13 image sets contains 7 frames of images to be restored and the corresponding 7 frames of edge images, and the 14th image set contains 9 frames of images to be restored and the corresponding 9 frames of edge images.
[0057] It should be noted that in the prior art, each frame of the image to be restored is repaired frame by frame. However, in this repair method, if there is a repair deviation in the repair process of a certain frame of image in the front, this repair deviation will continue to affect the subsequent image repair, and the repair deviation will continuously accumulate as the number of frames increases, and finally the repair effect of the entire video is poor. In this application, by dividing the video to be processed into multiple image sets and independently processing each image set, the problem that the previous repair errors will accumulate frame by frame when using frame-by-frame repair is avoided, thereby improving the video restoration quality.
[0058] In an alternative embodiment, the target network model is trained and generated by the following method: The video sequence restoration module first obtains training data, where the training data includes at least M frames of historical images to be restored and M frames of historical target images, and the historical target images are the images obtained after restoring the historical images to be restored. Then the video sequence restoration module preprocesses and performs edge extraction on the M frames of historical images to be restored to obtain M frames of historical edge images, where each frame of historical edge image corresponds to one frame of historical image to be restored. Finally, the video sequence restoration module determines multiple historical image sets based on the M frames of historical images to be restored and the M frames of historical edge images, and trains the target network model based on the multiple historical image sets. Each historical image set contains K frames of historical images to be restored and K frames of historical edge images.
[0059] Optionally, Figure 5 shows a flowchart of a model training according to an embodiment of the present application, as Figure 5 shown, first the video sequence restoration module will collect M frames of historical images to be restored and M frames of historical target images to form training data, then the video sequence restoration module will preprocess the M frames of historical images to be restored to obtain M frames of historical first images, and on this basis, the video sequence restoration module continues to perform edge extraction on the M frames of historical first images to obtain M frames of historical edge images.
[0060] Furthermore, as Figure 5 shown, the video sequence restoration module will divide multiple historical image sets based on the M frames of historical edge images and the M frames of historical images to be restored, and each historical image set also contains K frames of historical edge images and K frames of historical images to be restored. Finally, the video sequence restoration module trains the target network model based on the multiple historical image sets.
[0061] In an alternative embodiment, image features are required when training the target network model. In the present application, the following process is required when extracting image features. First, the video sequence restoration module extracts features from the K frames of historical images to be restored in each historical image set to obtain the first feature value corresponding to each frame of historical image to be restored, and extracts features from the K frames of historical edge images in each historical image set to obtain the second feature value corresponding to each frame of historical edge image. Then, the video sequence restoration module fuses the first feature values corresponding to the K frames of historical images to be restored to obtain the first training feature value, and fuses the second feature values corresponding to the K frames of historical edge images to obtain the second training feature value. Finally, the video sequence restoration module trains the target network model based on the first training feature value, the second training feature value, and the M frames of historical target images.
[0062] Optionally, the video sequence repair module inputs the first training eigenvalue and the second training eigenvalue into a neural network to obtain an initial network model, and determines a target loss function based on M frames of historical target images and M frames of historical images to be repaired. The target loss function includes a perceptual loss function, an adversarial loss function, and a reconstruction loss function. Finally, the video sequence repair module determines the loss value corresponding to the initial network model according to the target loss function, and adjusts the initial network model according to the loss value to obtain a target network model.
[0063] Specifically, as Figure 5 shown, after obtaining the first training eigenvalue and the second training eigenvalue, the video sequence repair module inputs the first training eigenvalue and the second training eigenvalue into the neural network respectively. After the neural network is trained, an initial network model is obtained. To improve the accuracy and robustness of the initial network model, the present application uses the target loss function to adjust the initial network model until the loss value of the model reaches a preset threshold, and then determines the adjusted initial network model as the target network model.
[0064] It should be noted that the target loss function in the present application includes a perceptual loss function, an adversarial loss function, and a reconstruction loss function. By adjusting the initial network model with the three loss functions, it is possible to ensure the structural integrity of the damaged model while improving the accuracy of the model for perceptual information.
[0065] In an optional embodiment, after the target network model outputs the first target eigenvalue and the second target eigenvalue, the video sequence repair module performs a summation process on the first target eigenvalue and the second target eigenvalue to obtain a third target eigenvalue corresponding to each image to be repaired in the image set, and superimposes the third target eigenvalue on each image to be repaired in the image set to obtain a target image, where each image set corresponds to K frames of target images. Finally, the video sequence repair module determines the repaired video according to the target image.
[0066] Optionally, Figure 6 shows an optional image repair flowchart according to an embodiment of the present application. As Figure 6 shown, assuming that there are 7 images to be repaired and 7 corresponding edge images in an image set, first, the video sequence repair module extracts features from all the images in the image set to obtain 7 first eigenvalues and 7 second eigenvalues. Among them, the first eigenvalue is the eigenvalue of the image to be repaired, and the second eigenvalue is the eigenvalue of the edge image. Then, the video sequence repair module fuses the 7 first eigenvalues to obtain a first input eigenvalue, and fuses each first eigenvalue with its corresponding second eigenvalue to obtain a second input eigenvalue. For example, Figure 6It shows the feature fusion of the first eigenvalue corresponding to the image to be repaired 1 and the second eigenvalue corresponding to the edge image 1, the feature fusion of the first eigenvalue corresponding to the image to be repaired 2 and the second eigenvalue corresponding to the edge image 2, and the feature fusion of the first eigenvalue corresponding to the image to be repaired 7 and the second eigenvalue corresponding to the edge image 7.
[0067] It should be noted that the first input eigenvalue and the second input eigenvalue obtained by feature fusion are input into the attention module in the target network model. Among them, the attention module contains a time domain branch and a spatial domain branch. The first input eigenvalue is input into the time domain branch and processed by the time domain branch to obtain the first target eigenvalue. The second input eigenvalue is input into the spatial domain branch and processed by the spatial domain branch to obtain the second target eigenvalue. Further, the video sequence repair module also calls the reconstruction module of the target network model to perform a summation process on the first target eigenvalue and the second target eigenvalue to obtain the third target eigenvalue, and the reconstruction module also superimposes the third target eigenvalue on the original image to be repaired to obtain the repaired target image, such as Figure 6 the target images 1, 2, and 7 shown in it.
[0068] In an optional embodiment, the video assembly and storage module obtains all the target images corresponding to multiple image sets, forms all the target images into a target image sequence according to the third target eigenvalue corresponding to each target image, and finally converts the target image sequence into a repaired video according to a preset video format and stores the repaired video in a preset area.
[0069] Optionally, as Figure 7 shown, the video assembly and storage module obtains all the target images corresponding to multiple image sets from the video sequence repair module. Since the third target eigenvalue represents the temporal correlation relationship between multiple images to be repaired, the video assembly and storage module can form all the target images into a target image sequence in chronological order according to the third target eigenvalue corresponding to each target image. Finally, the video assembly and storage module converts the target image sequence into video data according to a preset video format to obtain the repaired video, and the video assembly and storage module also stores the repaired video in a preset area. The preset video format can be the avi format or the flv format; in addition to storing the repaired video, the preset area can also store the target image sequence to implement the backup of the target image sequence.
[0070] It should be noted that in this application, the spatial and temporal information of video images are refined by using a Transformer structure with a global attention mechanism, thereby improving the global dependence relationship of video images, and a convolutional discriminator is used to enhance the local detail perception features of video images. At the same time, in this application, video images are repaired in the form of an image set, which also avoids the problem that the pre-repair errors in frame-by-frame repair will accumulate frame by frame, thereby further improving the video repair quality.
[0071] Embodiment 2
[0072] According to an embodiment of the present application, there is also provided a video processing device, wherein, Figure 8 is a schematic diagram of an optional video processing device according to an embodiment of the present application, as Figure 8 shown, the device includes: an acquisition module 801, configured to acquire a plurality of image sets from a video to be processed, wherein each image set includes K to-be-repaired images and K edge images, and each to-be-repaired image corresponds to one edge image; a determination module 802, configured to determine a first input feature value and a second input feature value according to the image set, wherein the first input feature value is a feature value obtained by fusing K first feature values, and the second input feature value is a feature value obtained by fusing one first feature value and one second feature value, the first feature value is a feature value obtained by performing feature extraction on the to-be-repaired image, and the second feature value is a feature value obtained by performing feature extraction on the edge image; a first input module 803, configured to input the first input feature value into the time domain branch of the target network model to obtain a first target feature value, wherein the first target feature value is used to represent the temporal correlation feature between the K to-be-repaired images in the image set; a second input module 804, configured to input the second input feature value into the spatial domain branch of the target network model to obtain a second target feature value, wherein the second target feature value is used to represent the spatial feature of each to-be-repaired image in the image set; a repair module 805, configured to repair each to-be-repaired image in the video to be processed according to the first target feature value and the second target feature value to obtain a repaired video.
[0073] It should be noted that the above acquisition module 801, determination module 802, first input module 803, second input module 804, and repair module 805 correspond to steps S101 to S105 in the above embodiment 1. The examples and application scenarios implemented by the five modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiment 1.
[0074] Optionally, the above-mentioned acquisition module further includes: a first acquisition unit, a first preprocessing unit, a first edge extraction processing unit, and a first determination unit. Among them, the first acquisition unit is configured to acquire the video to be processed, where the video to be processed includes N frames of images to be repaired, and the N frames of images to be repaired include the K frames of images to be repaired; the first preprocessing unit is configured to preprocess each frame of the image to be repaired in the video to be processed to obtain N frames of first images, where each frame of the first image corresponds to a frame of the image to be repaired; the first edge extraction processing unit is configured to perform edge extraction processing on each frame of the first image to obtain N frames of edge images, where each frame of the edge image corresponds to a frame of the first image; the first determination unit is configured to determine the multiple image sets according to the N frames of edge images and the N frames of images to be repaired.
[0075] Optionally, the video processing device further includes: a first acquisition module, a first processing module, a first determination module, and a first training module. Among them, the first acquisition module is configured to acquire training data, where the training data at least includes M frames of historical images to be repaired and M frames of historical target images, and the historical target images are the images obtained after repairing the historical images to be repaired; the first processing module is configured to perform the preprocessing and the edge extraction processing on the M frames of historical images to be repaired to obtain M frames of historical edge images, where each frame of the historical edge image corresponds to a frame of the historical image to be repaired; the first determination module is configured to determine multiple historical image sets according to the M frames of historical images to be repaired and the M frames of historical edge images, where each historical image set includes K frames of historical images to be repaired and K frames of historical edge images; the first training module is configured to train the target network model according to the multiple historical image sets.
[0076] Optionally, the above-mentioned first training module further includes: a first feature extraction unit, a second feature extraction unit, a first feature fusion unit, a second feature fusion unit, and a first training unit. Among them, the first feature extraction unit is configured to perform feature extraction on the K frames of historical images to be repaired in each historical image set to obtain a first feature value corresponding to each frame of the historical image to be repaired; the second feature extraction unit is configured to perform feature extraction on the K frames of historical edge images in each historical image set to obtain a second feature value corresponding to each frame of the historical edge image; the first feature fusion unit is configured to perform feature fusion on the first feature values corresponding to the K frames of historical images to be repaired to obtain a first training feature value; the second feature fusion unit is configured to perform feature fusion on the second feature values corresponding to the K frames of historical edge images to obtain a second training feature value; the first training unit is configured to train the target network model according to the first training feature value, the second training feature value, and the M frames of historical target images.
[0077] Optionally, the first training unit includes: an input subunit, a first determination subunit, a second determination subunit, and an adjustment unit. Among them, the input subunit is configured to input the first training eigenvalue and the second training eigenvalue into a neural network to obtain an initial network model; the first determination subunit is configured to determine a target loss function according to the M-frame historical target images and the M-frame historical images to be repaired, where the target loss function includes a perceptual loss function, an adversarial loss function, and a reconstruction loss function; the second determination subunit is configured to determine a loss value corresponding to the initial network model according to the target loss function; the adjustment unit is configured to adjust the initial network model according to the loss value to obtain the target network model.
[0078] Optionally, the above-mentioned repair module further includes: a summation processing unit, a superposition unit, and a first repair unit. Among them, the summation processing unit is configured to perform summation processing on the first target eigenvalue and the second target eigenvalue to obtain a third target eigenvalue corresponding to each image to be repaired in the image set; the superposition unit is configured to superimpose the third target eigenvalue on each image to be repaired in the image set to obtain a target image, where each image set corresponds to K-frame target images; the first repair unit is configured to determine the repaired video according to the target image.
[0079] Optionally, the first repair unit further includes: a first acquisition subunit, an assembly subunit, and a conversion subunit. Among them, the first acquisition subunit is configured to acquire all target images corresponding to the multiple image sets; the assembly subunit is configured to form the target images into a target image sequence according to the third target eigenvalue corresponding to each target image; the conversion subunit is configured to convert the target image sequence into the repaired video according to a preset video format and store the repaired video in a preset area.
[0080] Embodiment 3
[0081] According to an embodiment of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the video processing method in the above Embodiment 1 when running.
[0082] Embodiment 4
[0083] According to an embodiment of the present application, there is also provided an embodiment of an electronic device, where Figure 9 is a schematic diagram of an optional electronic device according to an embodiment of the present application, as Figure 9 shown, the electronic device includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the following steps are implemented:
[0084] Obtain multiple image sets from the video to be processed. Each image set contains K frames of images to be repaired and K frames of edge images, and each frame of the image to be repaired corresponds to one frame of the edge image. Determine the first input eigenvalue and the second input eigenvalue according to the image set. The first input eigenvalue is the eigenvalue obtained by fusing K first eigenvalues, and the second input eigenvalue is the eigenvalue obtained by fusing one first eigenvalue and one second eigenvalue. The first eigenvalue is the eigenvalue obtained by extracting features from the image to be repaired, and the second eigenvalue is the eigenvalue obtained by extracting features from the edge image. Input the first input eigenvalue into the temporal branch of the target network model to obtain the first target eigenvalue, where the first target eigenvalue is used to represent the temporal correlation features between the K frames of images to be repaired in the image set. Input the second input eigenvalue into the spatial branch of the target network model to obtain the second target eigenvalue, where the second target eigenvalue is used to represent the spatial features of each frame of the image to be repaired in the image set. Repair each frame of the image to be repaired in the video to be processed according to the first target eigenvalue and the second target eigenvalue to obtain the repaired video.
[0085] Optionally, when the processor executes the program, the following steps are also implemented: Obtain the video to be processed, where the video to be processed contains N frames of images to be repaired, and the N frames of images to be repaired include K frames of images to be repaired. Preprocess each frame of the image to be repaired in the video to be processed to obtain N frames of first images, where each frame of the first image corresponds to one frame of the image to be repaired. Perform edge extraction processing on each frame of the first image to obtain N frames of edge images, where each frame of the edge image corresponds to one frame of the first image. Determine multiple image sets according to the N frames of edge images and the N frames of images to be repaired.
[0086] Optionally, when the processor executes the program, the following steps are also implemented: Obtain training data, where the training data at least includes M frames of historical images to be repaired and M frames of historical target images, and the historical target image is the image obtained by repairing the historical image to be repaired. Preprocess and perform edge extraction processing on the M frames of historical images to be repaired to obtain M frames of historical edge images, where each frame of the historical edge image corresponds to one frame of the historical image to be repaired. Determine multiple historical image sets according to the M frames of historical images to be repaired and the M frames of historical edge images, where each historical image set contains K frames of historical images to be repaired and K frames of historical edge images. Train the target network model according to the multiple historical image sets.
[0087] Optionally, when the processor executes the program, the following steps are further implemented: extracting features from K frames of historical images to be repaired in each historical image set to obtain a first feature value corresponding to each frame of historical image to be repaired; extracting features from K frames of historical edge images in each historical image set to obtain a second feature value corresponding to each frame of historical edge image; performing feature fusion on the first feature values corresponding to the K frames of historical images to be repaired to obtain a first training feature value; performing feature fusion on the second feature values corresponding to the K frames of historical edge images to obtain a second training feature value; training a target network model according to the first training feature value, the second training feature value, and M frames of historical target images.
[0088] Optionally, when the processor executes the program, the following steps are further implemented: inputting the first training feature value and the second training feature value into a neural network to obtain an initial network model; determining a target loss function according to the M frames of historical target images and the M frames of historical images to be repaired, where the target loss function includes a perceptual loss function, an adversarial loss function, and a reconstruction loss function; determining a loss value corresponding to the initial network model according to the target loss function; adjusting the initial network model according to the loss value to obtain a target network model.
[0089] Optionally, when the processor executes the program, the following steps are further implemented: performing a summation process on the first target feature value and the second target feature value to obtain a third target feature value corresponding to each frame of the image to be repaired in the image set; superimposing the third target feature value on each frame of the image to be repaired in the image set to obtain a target image, where each image set corresponds to K frames of target images; determining a repaired video according to the target image.
[0090] Optionally, when the processor executes the program, the following steps are further implemented: obtaining all target images corresponding to multiple image sets; forming a target image sequence from all the target images according to the third target feature value corresponding to each target image; converting the target image sequence into a repaired video according to a preset video format and storing the repaired video in a preset area.
[0091] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0092] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0093] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0094] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0095] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0096] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0097] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A video processing method, characterized in that, Including: Obtaining a plurality of image sets from the video to be processed, where each image set includes K frames of images to be repaired and K frames of edge images, and each frame of the image to be repaired corresponds to one frame of the edge image; Determining a first input eigenvalue and a second input eigenvalue according to the image set, where the first input eigenvalue is an eigenvalue obtained by fusing K first eigenvalues, and the second input eigenvalue is an eigenvalue obtained by fusing one first eigenvalue and one second eigenvalue. The first eigenvalue is an eigenvalue obtained by performing feature extraction on the image to be repaired, and the second eigenvalue is an eigenvalue obtained by performing feature extraction on the edge image; Inputting the first input eigenvalue into the time domain branch of the target network model to obtain a first target eigenvalue, where the first target eigenvalue is used to characterize the temporal correlation feature between K frames of images to be repaired in the image set; Inputting the second input eigenvalue into the spatial domain branch of the target network model to obtain a second target eigenvalue, where the second target eigenvalue is used to characterize the spatial feature of each frame of the image to be repaired in the image set; Repairing each frame of the image to be repaired in the video to be processed according to the first target eigenvalue and the second target eigenvalue to obtain a repaired video; Wherein, repairing each frame of the image to be repaired according to the first target eigenvalue and the second target eigenvalue to obtain a repaired video includes: Performing a summation process on the first target eigenvalue and the second target eigenvalue to obtain a third target eigenvalue corresponding to each frame of the image to be repaired in the image set; Superimposing the third target eigenvalue on each frame of the image to be repaired in the image set to obtain a target image, where each image set corresponds to K frames of target images; Determining the repaired video according to the target image; Wherein, the method further includes processing the plurality of image sets in parallel.
2. The method according to claim 1, wherein Obtaining a plurality of image sets from the video to be processed includes: Obtaining the video to be processed, where the video to be processed includes N frames of images to be repaired, and the N frames of images to be repaired include the K frames of images to be repaired; Performing preprocessing on each frame of the image to be repaired in the video to be processed to obtain N frames of first images, where each frame of the first image corresponds to one frame of the image to be repaired; Performing edge extraction processing on each frame of the first image to obtain N frames of edge images, where each frame of the edge image corresponds to one frame of the first image; Determining the plurality of image sets according to the N frames of edge images and the N frames of images to be repaired.
3. The method according to claim 2, wherein The target network model is generated by training through the following method: Obtaining training data, where the training data at least includes M frames of historical images to be repaired and M frames of historical target images, and the historical target image is an image obtained by repairing the historical image to be repaired; Performing the preprocessing and the edge extraction processing on the M frames of historical images to be repaired to obtain M frames of historical edge images, where each frame of the historical edge image corresponds to one frame of the historical image to be repaired; Determine a plurality of historical image sets according to the M-frame historical images to be repaired and the M-frame historical edge images, where each historical image set includes K frames of historical images to be repaired and K frames of historical edge images; Train the target network model according to the plurality of historical image sets.
4. The method according to claim 3, characterized in that Training the target network model according to the plurality of historical image sets includes: Extract features from the K frames of historical images to be repaired in each historical image set to obtain a first eigenvalue corresponding to each frame of historical image to be repaired; Extract features from the K frames of historical edge images in each historical image set to obtain a second eigenvalue corresponding to each frame of historical edge image; Fuse the first eigenvalues corresponding to the K frames of historical images to be repaired to obtain a first training eigenvalue; Fuse the second eigenvalues corresponding to the K frames of historical edge images to obtain a second training eigenvalue; Train the target network model according to the first training eigenvalue, the second training eigenvalue, and the M-frame historical target images.
5. The method according to claim 4, wherein Training the target network model according to the first training eigenvalue, the second training eigenvalue, and the M-frame historical target images includes: Input the first training eigenvalue and the second training eigenvalue into a neural network to obtain an initial network model; Determine a target loss function according to the M-frame historical target images and the M-frame historical images to be repaired, where the target loss function includes a perceptual loss function, an adversarial loss function, and a reconstruction loss function; Determine the loss value corresponding to the initial network model according to the target loss function; Adjust the initial network model according to the loss value to obtain the target network model.
6. The method according to claim 1, wherein Determine the repaired video according to the target image, including: Obtain all target images corresponding to the plurality of image sets; According to the third target eigenvalue corresponding to each target image, form the target images into a target image sequence; Convert the target image sequence into the repaired video according to a preset video format and store the repaired video in a preset area.
7. A video processing device, characterized in that, Including: An acquisition module, configured to acquire a plurality of image sets from a video to be processed, where each image set includes K frames of images to be repaired and K frames of edge images, and each frame of image to be repaired corresponds to a frame of edge image; A determination module, configured to determine a first input eigenvalue and a second input eigenvalue according to the image set, where the first input eigenvalue is an eigenvalue obtained by fusing K first eigenvalues, and the second input eigenvalue is an eigenvalue obtained by fusing a first eigenvalue and a second eigenvalue, the first eigenvalue is an eigenvalue obtained by extracting features from the image to be repaired, and the second eigenvalue is an eigenvalue obtained by extracting features from the edge image; A first input module, configured to input the first input eigenvalue into the time domain branch of the target network model to obtain a first target eigenvalue, where the first target eigenvalue is used to characterize the temporal correlation feature between K to-be-restored images in the image set; A second input module, configured to input the second input eigenvalue into the spatial domain branch of the target network model to obtain a second target eigenvalue, where the second target eigenvalue is used to characterize the spatial feature of each to-be-restored image in the image set; A restoration module, configured to restore each to-be-restored image in the to-be-processed video according to the first target eigenvalue and the second target eigenvalue to obtain a restored video; Wherein, the restoration module is further configured to perform a summation process on the first target eigenvalue and the second target eigenvalue to obtain a third target eigenvalue corresponding to each to-be-restored image in the image set; superimpose the third target eigenvalue on each to-be-restored image in the image set to obtain a target image, where each image set corresponds to K target images; determine the restored video according to the target image; Wherein, the apparatus is further configured to perform parallel processing on the multiple image sets.
8. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program is configured to execute the video processing method described in any one of claims 1 to 6 when running.
9. An electronic device, characterized in that, The electronic device includes one or more processors; A memory, configured to store one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement a program for running, where the program is configured to execute the video processing method described in any one of claims 1 to 6 when running.
Citation Information
Patent Citations
Method, system and terminal for video restoration by using deep convolutional neural network
CN111787187A
Image processing method and device, computer equipment and storage medium
CN114299105A